Storage Controller Slowdown Detection via Response Time Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In storage systems, disks in a slowdown state are not detected effectively, leading to performance degradation and potential two-point failures, as existing failure detection methods rely on timeout thresholds that can misclassify response delays, causing unnecessary log analysis and disk detachment issues.

Innovation Solution

A storage controlling apparatus that collects and analyzes response time period data to calculate average values and standard deviations, allowing for the detection of performance degradation in disks before they reach critical failure points, enabling early intervention and maintaining RAID group performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If timeout threshold-based failure detection is used, then critical failures can be detected, but disks in slowdown state are not detected effectively and false positives occur

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidslowdown state detection precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent changes the detection parameter from a fixed timeout threshold to a dynamic statistical model using response achievement time periods, average values, and standard deviations. This allows the system to adapt to varying disk performance characteristics and detect slowdown states without false positives from normal performance variations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical timeout threshold mechanism with a statistical analysis system that calculates average response times and standard deviations. This substitution enables more nuanced detection of performance degradation by considering variability in response times rather than using a single fixed threshold.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If timeout threshold detection is used, then failure detection is simple, but unnecessary log analysis and disk detachment occur

Engineering Contradiction:
Improvedetection method complexityVSAvoidsystem stability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the system continuously monitors response achievement time periods, calculates statistical parameters, and uses this information to make informed decisions about disk health. This feedback loop prevents unnecessary log analysis and disk detachment by providing accurate real-time assessment of disk performance.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary statistical analysis by calculating average values and standard deviations from multiple response time measurements before making failure detection decisions. This preliminary action reduces false positives by establishing a baseline of normal performance variation before detecting actual failures.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If statistical point addition is used for failure detection, then critical failures are detected, but performance degradation in slowdown state is not detected

Engineering Contradiction:
Improvefailure detection capabilityVSAvoidperformance degradation detection precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent introduces dynamics to the detection system by continuously updating statistical parameters (average response time and standard deviation) as new measurements are collected. This dynamic approach allows the system to detect gradual performance degradation in slowdown states while maintaining the ability to detect critical failures.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent adds another dimension to failure detection by incorporating response achievement time periods and their statistical variations into the detection criteria. This dimensional expansion allows the system to distinguish between temporary performance fluctuations and actual performance degradation, enabling detection of slowdown states that were previously invisible to timeout-based methods.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10725665B2Storage controlling apparatus, recording medium for recording storage control program and storage controlling method
Publication Date: 2020.07.28 FUJITSU LTD
  • US10725665B2 patent drawing
  • US10725665B2 patent drawing
  • US10725665B2 patent drawing

AI summary

A storage controlling apparatus, includes: a memory configured to store a program; and a processor configured to control a plurality of storage devices based on the program, wherein the processor: collects information relating to a data access performed for the plurality of storage devices; and decides performance degradation of a first storage device from among the plurality of storage devices based on a response achievement time period for a first data access request performed for the first storage device, and a response time period average value and a response time period standard deviation which are calculated based on response achievement time periods with respect to a plurality of data access requests performed for the first storage device before the first data access request.