Disk Array Controller Error Counting and Warning Mode

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing disk array systems with redundancy face inefficiencies in data restoration due to high error rates in disk drives, leading to prolonged processing times and potential failure acceleration when a disk drive is recognized as likely to fail, which can disrupt data rebuilding and access operations.

Innovation Solution

An array controller with error counters, a failure estimation unit, and a mode-setting unit that sets a disk drive in a warning mode to reduce access frequency and maintain it as a member of the array, making it less accessible than other drives, thereby elongating its lifespan and preventing premature failure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a disk drive recognized as likely to fail is accessed in the same manner as normal disk drives during rebuild process, then data restoration can be performed, but the time when the disk drive actually fails is accelerated

Engineering Contradiction:
Improvedata restoration speedVSAvoiddisk drive lifespan
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies local quality by treating the recognized-at-risk disk drive differently from normal disk drives. Specifically, the controller restricts access to the at-risk drive only to rebuild operations while allowing normal drives to be accessed freely for both rebuild and host I/O operations. This differentiated access policy preserves the at-risk drive longer while still enabling necessary data restoration.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements preliminary action by proactively identifying disk drives at risk of failure through error counting before actual failure occurs. The controller monitors error rates and preemptively restricts access to drives exceeding error thresholds, preventing further degradation and potential failure while maintaining system functionality through alternative access paths.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If a disk drive recognized as likely to fail is accessed frequently, then data restoration process can proceed, but the disk drive may fail before rebuild is completed

Engineering Contradiction:
Improverebuild completion speedVSAvoidrebuild completion success
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The controller proactively identifies drives at risk through error monitoring before failure occurs, enabling preemptive protection measures. By restricting access to at-risk drives only for rebuild operations and preventing other access patterns, the system ensures rebuild completion while avoiding premature drive failure that would interrupt the process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses the controller as an intermediary that manages access to the at-risk disk drive. The controller acts as a gatekeeper, allowing only necessary rebuild operations to access the vulnerable drive while blocking other access patterns that could accelerate failure. This intermediary control enables safe completion of rebuild processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If a disk drive is made less accessible to elongate its lifespan, then premature failure is prevented, but data access speed may be reduced

Engineering Contradiction:
Improvedisk drive lifespanVSAvoiddata access speed
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The patent applies local quality by implementing selective access restrictions only on identified at-risk disk drives while leaving normal drives fully accessible. The controller differentiates between safe and unsafe drives, applying access limitations only where necessary. This targeted approach preserves overall system performance while protecting vulnerable components.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses redundancy and copying mechanisms to mitigate the impact of restricted access. Data from at-risk drives is restored to safe drives or spare drives, creating copies that can serve future access requests. This copying approach maintains data availability and access speed while protecting the original at-risk drive from further stress.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS7779202B2Apparatus and method for controlling disk array with redundancy and error counting
Publication Date: 2010.08.17 KK TOSHIBA
  • US7779202B2 patent drawing
  • US7779202B2 patent drawing
  • US7779202B2 patent drawing

AI summary

According to one embodiment, a read/write control unit controls read/write access to at least two disk drives that provide a disk array. Error counters are provided for the respective disk drives for counting respective numbers of errors if the errors occur when the disk drives are accessed. A failure estimation unit detects, as a disk drive which is very likely to fail, a disk drive included in the disk array and having a high error occurrence degree, based on the numbers of errors counted by the error counters. A mode-setting unit sets the detected disk drive in a particular mode in which the detected disk drive is maintained as a member of the disk array and is made more inaccessible than the remaining disk drive of the disk array.