Storage Failure Snapshot Logging for Rapid Root Cause Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Analyzing storage device failures is challenging due to the vast amount of information scattered throughout general system logs, making it difficult to identify the cause of a failure, especially as the number of storage devices increases, leading to a time-consuming and potentially incomplete analysis process.

Innovation Solution

Collecting and analyzing relevant information about the storage device and its environment in a timely manner, including input/output data and adjacent storage device status, to determine the reason for failure, and presenting this information to storage system administrators in a concise manner.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If general system logs are used to store all system events, then comprehensive system monitoring is achieved, but finding storage device failure information becomes difficult and time-consuming

Engineering Contradiction:
Improvestorage device failure informationVSAvoidtime to review logs
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent segments the general system log into specific storage device failure logs. When a storage device failure is detected, the system creates a dedicated log entry that consolidates all relevant failure information from multiple sources (system logs, storage device logs, environmental data) into a single organized record, eliminating the need to search through entire system logs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary component (storage system/logger) that captures and processes storage device failure information before it reaches the general system log. This intermediary creates specialized failure logs that serve as a bridge between detailed storage device data and general system records, making failure information easily accessible.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If the number of storage devices in the system increases, then system capacity increases, but the general system log becomes larger and more difficult to review

Engineering Contradiction:
Improvenumber of storage devicesVSAvoidlog review complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments log information by storage device identifier, creating separate organized sections for each device's failure events. This segmentation allows the system to handle an increasing number of storage devices without proportionally increasing log review complexity, as each device's information is self-contained and easily locatable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary organization of log data by creating structured failure records with predefined fields and categories before the actual log review process. This preliminary structuring of storage device information makes subsequent analysis efficient regardless of the total number of devices in the system.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual log review is performed to analyze storage device failures, then detailed analysis is possible, but the process is time-consuming and may miss important information

Engineering Contradiction:
Improvefailure analysis accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an automated intermediary system that processes and analyzes storage device failure logs. This intermediary automatically correlates information from multiple sources, identifies failure patterns, and generates analysis results, maintaining high accuracy while dramatically reducing the time required compared to manual review.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback mechanisms where the system automatically processes failure log entries, correlates them with environmental data and device history, and generates refined failure analysis. This automated feedback loop ensures comprehensive analysis of all relevant information without human oversight limitations.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If detailed environmental information is collected at the time of failure, then failure cause determination is improved, but information collection and processing complexity increases

Engineering Contradiction:
Improvefailure cause determinationVSAvoidinformation collection system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent collects and stores environmental information and device status data in advance of actual failures using preliminary monitoring. This preliminary data collection creates a ready repository of contextual information that can be immediately accessed when failures occur, improving cause determination without requiring complex real-time data gathering at the moment of failure.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments environmental and device information into organized categories (system environment, storage device status, error history) that are independently collected and stored. This segmentation simplifies the overall collection system while enabling comprehensive failure analysis, as each segment can be processed and stored independently.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9354966B2Creating environmental snapshots of storage device failure events
Publication Date: 2016.05.31 NETAPP INC
  • US9354966B2 patent drawing
  • US9354966B2 patent drawing
  • US9354966B2 patent drawing

AI summary

A storage device failure in a computer storage system can be analyzed by the storage system by examining relevant information about the storage device and its environment. Information about the storage device is collected in real-time and stored; this is an on-going process such that some information is continuously available. The information can include information relating to the storage device, such as input/output related information, and information relating to a storage shelf where the storage device is located, such as a status of adjacent storage devices on the shelf. All of the relevant information is analyzed to determine a reason for the storage device failure. Optionally, additional information may be collected and analyzed by the storage system to help determine the reason for the storage device failure. The analysis and supporting information can be stored in a log and/or presented to a storage system administrator to view.