Storage Media Soft Failure Detection and Sector Rewriting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing information handling systems face challenges in proactively preventing and predicting storage media failures, particularly due to single bit errors evolving into multi-bit errors on hard disk drives and data decay in solid state drives, which can lead to uncorrectable errors and susceptibility in critical boot sectors.

Innovation Solution

A method and system that includes a processor and a failure analysis module to detect soft failures during boot, rewrite affected sectors, and predict storage media failures by using SMART commands to compare error counts and relocate or recharge sectors, thereby correcting errors and anticipating potential catastrophic failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If error correction codes are used to correct single bit errors, then error correction capability is improved, but multi-bit errors become uncorrectable

Engineering Contradiction:
Improveerror correction capabilityVSAvoidmulti-bit errors
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary actions by continuously monitoring storage media health parameters and proactively rewriting sectors before multi-bit errors occur. The failure analysis module detects early signs of degradation and initiates preventive rewriting operations to maintain data integrity before errors become uncorrectable.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms through SMART commands that continuously monitor storage media parameters including error counts, reallocation sectors, and pending sectors. This feedback loop enables the failure analysis module to adjust rewriting strategies based on actual storage media condition and predict future failures.

Inventive Principle:
Principle #23Feedback

2Reliability

If boot sectors are made read-only to prevent wear, then boot sector integrity is improved, but susceptibility to data decay increases

Engineering Contradiction:
Improveboot sector integrityVSAvoiddata decay
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary rewriting of boot sectors at scheduled intervals during system boot-up. The failure analysis module proactively rewrites boot sectors before data decay becomes significant, maintaining their integrity without requiring them to be read-only. This preventive action balances boot sector protection with decay prevention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically changes the operational parameters of boot sectors by monitoring their health status and adjusting rewriting frequency and intensity accordingly. Based on SMART data and decay rates, the system adapts the rewriting strategy to optimize both integrity and decay prevention.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If continuous monitoring of storage media is implemented, then failure prediction accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvefailure prediction accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The failure analysis module is integrated into existing system components such as the BIOS or operating system boot process, making it multi-functional. It combines monitoring, analysis, and preventive rewriting operations within a unified module that leverages existing hardware and software resources, thereby reducing overall system complexity while maintaining high prediction accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs self-monitoring and self-diagnosis using SMART commands during normal boot operations. The failure analysis module automatically detects storage media health issues and initiates corrective actions without requiring external intervention or complex separate monitoring infrastructure, enabling the system to serve itself in maintaining data integrity.

Inventive Principle:
Principle #25Self-service

4Reliability

If frequent rewriting of storage sectors is performed, then data integrity is improved, but storage media wear increases

Engineering Contradiction:
Improvedata integrityVSAvoidstorage media lifespan
Core Design Contradiction:
ReliabilityVSDuration of action of stationary object

Solution Approach 1:

The system dynamically adjusts rewriting parameters including frequency, timing, and intensity based on real-time storage media health status. By monitoring SMART parameters and decay rates, the system performs rewriting only when necessary and at optimized intervals, maintaining data integrity while minimizing wear on storage media.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs preliminary rewriting actions at scheduled intervals during boot-up rather than continuously. This proactive approach addresses decay before it becomes problematic, reducing the need for reactive rewriting operations that would increase wear. The failure analysis module plans and executes rewriting operations in advance based on predicted failure timelines.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11126502B2Systems and methods for proactively preventing and predicting storage media failures
Publication Date: 2021.09.21 DELL PROD LP
  • US11126502B2 patent drawing
  • US11126502B2 patent drawing
  • US11126502B2 patent drawing

AI summary

A method may include, during a boot of an information handling system, detecting a soft failure associated with a read request to storage media of the information handling system wherein the soft failure is not visible to an operating system of the information handling system and in response to detecting the soft failure, rewriting a sector of the storage media affected by the soft failure to correct the soft failure.