Cluster-Wide Correlation Analysis for Distributed Storage Troubleshooting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional troubleshooting techniques are inefficient and time-consuming for diagnosing issues in large-scale distributed storage systems, particularly in complex computing environments, often requiring significant effort and time to identify root causes and provide effective remedies, which can lead to poor user experience and high costs.

Innovation Solution

Implementing cluster-wide correlation analysis using health monitoring agents, a correlation engine, a cluster-wide health analyzer, and a troubleshooting workflow engine to identify potential causes of issues, assess their impact, and provide guided remediation workflows, reducing reliance on support engineering staff and enabling proactive issue resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional troubleshooting techniques are used in large-scale distributed storage systems, then troubleshooting can be performed with simple methods, but the troubleshooting process becomes time-consuming and inefficient

Engineering Contradiction:
Improvetroubleshooting efficiencyVSAvoidtime to identify root cause
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The troubleshooting system is segmented into specialized components: health monitoring agents deployed across nodes, a correlation engine for pattern recognition, a cluster-wide health analyzer for aggregate assessment, and a troubleshooting workflow engine for guided remediation. This segmentation allows parallel processing of monitoring, analysis, and resolution tasks, dramatically improving troubleshooting efficiency while reducing time to identify root causes in large-scale distributed storage systems.

Inventive Principle:
Principle #1Segmentation

2Reliability

If manual troubleshooting by support engineering staff is used, then complex issues can be diagnosed, but significant effort and costs are required

Engineering Contradiction:
Improveaccuracy of issue diagnosisVSAvoidcomplexity of troubleshooting system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements self-service troubleshooting through automated health monitoring agents that continuously collect metrics, a correlation engine that automatically identifies patterns and potential root causes, and a workflow engine that provides guided remediation steps. This automation maintains high diagnostic accuracy for complex issues while eliminating the need for extensive manual support engineering effort, thereby reducing operational costs without sacrificing reliability.

Inventive Principle:
Principle #25Self-service

3Productivity

If cluster-wide correlation analysis is implemented, then quick and efficient diagnosis is achieved, but the system requires multiple components including health monitoring agents, correlation engine, and troubleshooting workflow engine

Engineering Contradiction:
Improvespeed of issue resolutionVSAvoidnumber of system components
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Each component in the cluster-wide correlation analysis system is designed with multi-functionality to justify its inclusion. Health monitoring agents not only collect metrics but also validate data quality and correlate local events. The correlation engine performs both pattern recognition and root cause identification, while the workflow engine provides both automated remediation and guidance for manual intervention. This universality ensures that each component contributes multiple values, maintaining high issue resolution speed while managing system complexity through consolidated, multi-purpose elements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11714701B2Troubleshooting for a distributed storage system by cluster wide correlation analysis
Publication Date: 2023.08.01 VMWARE INC
  • US11714701B2 patent drawing
  • US11714701B2 patent drawing
  • US11714701B2 patent drawing

AI summary

A troubleshooting technique provides faster and more efficient troubleshooting of issues in a distributed system, such as a distributed storage system provided by a virtualized computing environment. The distributed system includes a plurality of hosts arranged in a cluster. The troubleshooting technique uses cluster-wide correlation analysis to identify potential causes of a particular issue in the distributed system, and executes workflows to remedy the particular issue.