Storage Device Service Takeover via Resource-Aware Arbitration Delay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In two-center storage systems, communication faults between data centers disrupt data synchronization, leading to inefficient service takeover, as existing methods do not consider the current system resource usage of storage devices before sending arbitration requests to a quorum server.
Innovation Solution
A method where storage devices delay sending arbitration requests to a quorum server based on their current system resource usage, allowing the quorum server to select a storage device in a better running status to take over host services, by determining a delay duration based on the running status value of processor, hard disk, cache, and host bandwidth resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If storage devices immediately send arbitration requests to the quorum server when a communication fault occurs, then the service takeover response time is reduced, but the service takeover efficiency deteriorates because devices with poor resource availability may be selected
Solution Approach 1:
The storage device performs preliminary assessment of its own system resource usage before sending the arbitration request. By evaluating processor usage, hard disk usage, cache usage, and host bandwidth usage in advance, the device determines whether it is in a suitable state to take over services, preventing premature arbitration participation
Solution Approach 2:
The arbitration request sending mechanism is made dynamic by introducing a delay timer that is adjusted based on real-time system resource usage. When resources are heavily utilized, the delay is extended; when resources are available, the delay is reduced or eliminated, making the arbitration response adaptive to current system state
2Device complexity
If storage devices send arbitration requests without considering system resource usage, then the arbitration process is simplified, but the reliability of service takeover deteriorates due to selection of poorly loaded devices
Solution Approach 1:
The storage device autonomously evaluates its own system resource status and determines its suitability for service takeover without requiring external assessment. The device self-monitors processor usage, hard disk usage, cache usage, and host bandwidth usage, and makes its own decision on whether to participate in arbitration, simplifying the overall system architecture while improving reliability
Data Source
AI summary
The present disclosure describes example service takeover methods, storage devices, and service takeover apparatuses. In one example method, when a communication fault occurs between two storage devices in a storage system, the two storage devices respectively obtain running statuses of the two storage devices. A running status can reflect current usage of one or more system resources of a particular storage device. Then, a delay duration is determined according to the running statuses, where the delay duration is a duration for which the storage device waits before sending an arbitration request to a quorum server. The two storage devices respectively send, after the delay duration, arbitration requests to the quorum server to request to take over a service. The quorum server then can select a storage device in a relatively better running status to take over a host service.


