Storage Failure Snapshot Logging for Rapid Root Cause Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing storage device failures is challenging due to the vast amount of information scattered throughout general system logs, making it difficult to identify the cause of a failure, especially as the number of storage devices increases, leading to a time-consuming and potentially incomplete analysis process.
Innovation Solution
Collecting and analyzing relevant information about the storage device and its environment in a timely manner, including input/output data and adjacent storage device status, to determine the reason for failure, and presenting this information to storage system administrators in a concise manner.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If general system logs are used to store all system events, then comprehensive system monitoring is achieved, but finding storage device failure information becomes difficult and time-consuming
Solution Approach 1:
The patent segments the general system log into specific storage device failure logs. When a storage device failure is detected, the system creates a dedicated log entry that consolidates all relevant failure information from multiple sources (system logs, storage device logs, environmental data) into a single organized record, eliminating the need to search through entire system logs.
Solution Approach 2:
The patent introduces an intermediary component (storage system/logger) that captures and processes storage device failure information before it reaches the general system log. This intermediary creates specialized failure logs that serve as a bridge between detailed storage device data and general system records, making failure information easily accessible.
2Quantity of substance
If the number of storage devices in the system increases, then system capacity increases, but the general system log becomes larger and more difficult to review
Solution Approach 1:
The patent segments log information by storage device identifier, creating separate organized sections for each device's failure events. This segmentation allows the system to handle an increasing number of storage devices without proportionally increasing log review complexity, as each device's information is self-contained and easily locatable.
Solution Approach 2:
The patent performs preliminary organization of log data by creating structured failure records with predefined fields and categories before the actual log review process. This preliminary structuring of storage device information makes subsequent analysis efficient regardless of the total number of devices in the system.
3Measurement precision
If manual log review is performed to analyze storage device failures, then detailed analysis is possible, but the process is time-consuming and may miss important information
Solution Approach 1:
The patent introduces an automated intermediary system that processes and analyzes storage device failure logs. This intermediary automatically correlates information from multiple sources, identifies failure patterns, and generates analysis results, maintaining high accuracy while dramatically reducing the time required compared to manual review.
Solution Approach 2:
The patent implements feedback mechanisms where the system automatically processes failure log entries, correlates them with environmental data and device history, and generates refined failure analysis. This automated feedback loop ensures comprehensive analysis of all relevant information without human oversight limitations.
4Measurement precision
If detailed environmental information is collected at the time of failure, then failure cause determination is improved, but information collection and processing complexity increases
Solution Approach 1:
The patent collects and stores environmental information and device status data in advance of actual failures using preliminary monitoring. This preliminary data collection creates a ready repository of contextual information that can be immediately accessed when failures occur, improving cause determination without requiring complex real-time data gathering at the moment of failure.
Solution Approach 2:
The patent segments environmental and device information into organized categories (system environment, storage device status, error history) that are independently collected and stored. This segmentation simplifies the overall collection system while enabling comprehensive failure analysis, as each segment can be processed and stored independently.
Data Source
AI summary
A storage device failure in a computer storage system can be analyzed by the storage system by examining relevant information about the storage device and its environment. Information about the storage device is collected in real-time and stored; this is an on-going process such that some information is continuously available. The information can include information relating to the storage device, such as input/output related information, and information relating to a storage shelf where the storage device is located, such as a status of adjacent storage devices on the shelf. All of the relevant information is analyzed to determine a reason for the storage device failure. Optionally, additional information may be collected and analyzed by the storage system to help determine the reason for the storage device failure. The analysis and supporting information can be stored in a log and/or presented to a storage system administrator to view.


