Disk Drive Media Failure Analysis for Virtual Tape Servers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Virtual tape servers face challenges in identifying and addressing disk drive media failures, leading to increased activation time and costs due to widespread firmware updates, as defective disk drives with faulty microcode or manufacturing issues are difficult to detect and manage.

Innovation Solution

A system and method for managing disk drive media failures by receiving and analyzing failure data across multiple virtual tape servers, using a data repository to identify problems and determine the need for firmware or hardware updates, thereby minimizing unnecessary updates and optimizing maintenance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all DDM firmware changes/updates are automatically installed during activation, then reliability of DDM is improved, but activation time is greatly increased

Engineering Contradiction:
ImproveDDM reliabilityVSAvoidactivation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of DDM failure data before activation to identify problematic firmware versions. By pre-processing failure information and creating a database of known issues, the system can quickly determine which DDM require firmware updates without performing comprehensive checks during the activation process itself, thus reducing activation time while maintaining reliability improvements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where DDM failure data from the field is continuously collected, analyzed, and used to generate firmware update recommendations. This feedback loop allows the system to learn from actual failures and improve firmware deployment strategies, ensuring that updates are applied based on real-world performance data rather than blanket updates to all DDM.

Inventive Principle:
Principle #23Feedback

2Reliability

If widespread firmware updates are performed to address defective DDM, then reliability is improved, but costs and time consumption increase

Engineering Contradiction:
ImproveDDM reliabilityVSAvoidmaintenance resources
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

Instead of applying firmware updates uniformly across all DDM, the system analyzes failure data to identify specific DDM models, manufacturers, or firmware versions that are problematic. Updates are then targeted only to the specific DDM that require them, rather than performing widespread updates across the entire installed base. This localized approach maintains reliability improvements while significantly reducing the time and resources consumed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs partial firmware updates only on the subset of DDM that are identified as problematic through failure data analysis, rather than performing excessive blanket updates on all DDM. This partial action approach applies the principle of doing just enough maintenance to address actual problems without over-maintaining systems that are already functioning correctly.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If DDM are installed without pre-checking for defects, then ease of installation is improved, but identification of defective DDM becomes difficult

Engineering Contradiction:
Improveinstallation easeVSAvoiddefect detection difficulty
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary analysis of DDM before installation by checking the identified DDM against a database of known failures and problematic firmware versions. This pre-check process occurs automatically in the background without requiring manual intervention from installers, thus maintaining ease of installation while enabling early detection of potential defects through automated comparison with known issue patterns.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary layer between DDM installation and operation - a data repository and analysis system that acts as a mediator to evaluate incoming DDM against known failure patterns. This intermediary automatically assesses whether installed DDM have potential defects by comparing them with historical failure data, providing defect detection capability without adding complexity to the installation process itself.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8751858B2System, method, and computer program product for physical drive failure identification, prevention, and minimization of firmware revisions
Publication Date: 2014.06.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8751858B2 patent drawing
  • US8751858B2 patent drawing
  • US8751858B2 patent drawing

AI summary

In one embodiment, a system includes logic adapted for receiving information relating to disk drive media (DDM) failures in an installed base of DDM across multiple virtual tape servers, a storage device adapted for storing the information relating to the DDM failures in a data repository, and a processor adapted for analyzing the information stored in the data repository to identify problems in an installed base of DDM, the analysis comprising analyzing comparative DDM failure data comprising vectors. In another embodiment, a method for managing DDM failures includes receiving DDM failure information in virtual tape servers, storing the DDM failure information in a data repository, and analyzing the information to identify problems in an installed base of DDM. Other systems, methods, and computer program products are also described according to more embodiments.