ML Dashboard for Hyper-Converged Root Cause Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In hyper-converged infrastructure environments, existing resource utilization monitoring systems lack the ability to efficiently identify the root causes of user issues and failures due to inadequate real-time monitoring and analysis of resource utilization parameters across nodes.
Innovation Solution
A machine-learning based self-populating dashboard that continuously monitors select resource utilization parameters across nodes, assessing cause-and-effect relationships to dynamically provide administrators with a user interface that identifies potential root causes of issues and failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional monitoring systems are used to track resource utilization parameters, then basic data collection is achieved, but the ability to identify root causes of issues and failures is insufficient
Solution Approach 1:
The system implements feedback mechanisms by continuously monitoring resource utilization parameters and using machine learning models to analyze correlations between parameters and issues. The dashboard provides feedback to administrators about potential root causes, enabling them to make informed decisions. This feedback loop transforms raw monitoring data into actionable insights, directly addressing the insufficient root cause identification capability while maintaining efficient troubleshooting.
Solution Approach 2:
The patent replaces traditional mechanical monitoring approaches with machine learning-based analysis. Instead of relying on simple threshold alerts and manual analysis, the system uses ML models to automatically correlate resource utilization parameters with issues and failures. This substitution enables precise root cause identification without compromising troubleshooting efficiency, as the ML models process data automatically and provide insights rapidly.
2Loss of information
If comprehensive resource utilization parameters are monitored across all nodes, then complete visibility is achieved, but system complexity and computational requirements increase
Solution Approach 1:
The system extracts only the most relevant resource utilization parameters and correlations using machine learning analysis. Instead of presenting all monitored parameters equally, the ML models identify and extract the specific parameters that are most strongly correlated with issues and failures. This extraction approach maintains complete monitoring information while reducing system complexity by focusing computational resources on the most critical parameters and relationships.
Solution Approach 2:
The dashboard implements local quality by providing different levels of monitoring detail to different users or for different situations. The system monitors comprehensive parameters across all nodes but presents customized views based on user roles, specific issues being investigated, and node criticality. This approach ensures complete information is captured while managing system complexity through selective presentation and analysis focus.
3Reliability
If real-time analysis of resource utilization parameters is performed, then immediate issue detection is achieved, but computational resources and processing time are consumed
Solution Approach 1:
The system implements periodic action by analyzing resource utilization parameters at strategically determined intervals rather than continuously processing every data point in real-time. The machine learning models are trained to identify patterns and correlations that persist across time periods, allowing the system to perform comprehensive analysis at regular intervals while maintaining reliable issue detection. This periodic approach reduces computational resource consumption while preserving the ability to detect issues timely.
Solution Approach 2:
The system performs preliminary action by pre-training machine learning models on historical resource utilization data and issue patterns before deployment. This preliminary training enables the models to quickly analyze new data with high accuracy without requiring extensive real-time computational resources. The pre-processed knowledge from historical data allows immediate issue detection while minimizing ongoing computational resource consumption during operational monitoring.
Data Source
AI summary
A machine-learning based self-populating dashboard for resource utilization monitoring in hyper-converged information technology (IT) environments. Specifically, the method and system disclosed herein entail the continuous monitoring of select resource utilization parameters across various nodes in an environment. Cause and effect relationships between these select resource utilization parameters and other utilization parameters are periodically assessed to dynamically populate a user interface (e.g., a web-based or non-web-based dashboard). The user interface provides environment administrators with a tool that intelligently identifies which resource utilization parameters may be strong contenders as root causes of user issues and/or failures occurring in the environment.


