Storage Device Error Log Analysis for Drive Health Recommendations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data storage array utilities are limited in their ability to handle drive errors, often resorting to either killing or sparing storage drives without considering alternative solutions like firmware updates or detailed diagnostics, and lack the capability to detect complex drive issues or provide tailored recommendations.

Innovation Solution

A tool that analyzes storage device error logs to determine the health of storage devices and provides recommendations for activities such as upgrading firmware, reseating, or replacing drives, using configurable text-based factors to map error events to individual weights and recommended actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the conventional local utility simply kills or spares storage drives when error weights exceed a threshold, then drive availability is maintained, but diagnostic precision and flexibility are lost

Engineering Contradiction:
Improvedrive availabilityVSAvoiddiagnostic precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the monolithic kill/spare decision into multiple independent diagnostic components: error pattern recognition, firmware version analysis, manufacturer-specific error grouping, and activity recommendation generation. Each component processes specific aspects of drive health independently, then combines results to form comprehensive recommendations, thereby maintaining reliability while dramatically improving diagnostic precision

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary diagnostic layer between the error detection utility and the drive management action. This intermediary analyzes error patterns, firmware versions, and manufacturer data to generate recommended activities (firmware updates, reseating, replacement) before executing any drive management action. This intermediary prevents premature kill/spare decisions while maintaining system reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Stability of the object's composition

If the local utility is updated infrequently during quarterly operating system upgrades, then system stability is maintained, but detection of new drive errors and updated diagnostics is delayed

Engineering Contradiction:
Improvesystem stabilityVSAvoidlatency in detecting new errors
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent extracts the diagnostic evaluation logic from the quarterly operating system upgrade cycle and places it in a separate, independently updatable utility component. This extracted diagnostic module can be updated separately and more frequently than the full operating system, reducing latency in detecting new error types while maintaining overall system stability through controlled update independence

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If generic error weights are used by the conventional utility, then simplicity is maintained, but manufacturing precision and manufacturer-specific diagnostics are lost

Engineering Contradiction:
Improveutility simplicityVSAvoidmanufacturer-specific diagnostics
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent implements dynamic error weights that adapt based on manufacturer, drive type, and error pattern rather than using static generic weights. The utility automatically adjusts diagnostic thresholds and evaluation criteria based on the specific manufacturer and drive characteristics being analyzed, achieving manufacturing precision while maintaining utility simplicity through automated adaptation

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by using manufacturer-specific error groups and customized evaluation criteria for different drive manufacturers and types. Each manufacturer's drives are evaluated with tailored error weights and patterns specific to that manufacturer's hardware characteristics, achieving precision diagnostics while maintaining overall system simplicity through localized customization

Inventive Principle:
Principle #3Local quality

4Productivity

If the conventional utility focuses primarily on high availability, then drive management speed is improved, but versatility of diagnostic capabilities deteriorates

Engineering Contradiction:
Improvedrive management speedVSAvoiddiagnostic versatility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary diagnostic actions by analyzing error patterns, firmware versions, and manufacturer data before executing drive management decisions. This preliminary analysis phase identifies the most appropriate recommended activities (firmware updates, reseating, replacement) upfront, enabling fast execution of pre-determined actions while providing versatile diagnostic capabilities during the analysis phase

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9311176B1Evaluating a set of storage devices and providing recommended activities
Publication Date: 2016.04.12 EMC IP HLDG CO LLC
  • US9311176B1 patent drawing
  • US9311176B1 patent drawing
  • US9311176B1 patent drawing

AI summary

A technique evaluates a set of storage devices (e.g., magnetic disk drives, solid state drives, etc.). The technique involves receiving, by processing circuitry (e.g., a storage processor, a standalone computer, etc.), storage device evaluation factors which (i) map possible storage device error events to individual weights and (ii) map cumulative weights to recommended activities. The technique further involves receiving, by the processing circuitry, a storage device error log containing storage device error entries identifying actual storage device error events which were encountered by the set of storage devices while performing data storage operations over a period of time. The technique further involves analyzing, by the processing circuitry, the storage device error entries based on the storage device evaluation factors to produce a set of evaluation results identifying a set of recommended activities to be performed on the set of storage devices.