Storage system fault handling method and electronic equipment

By detecting storage system load information and classifying light and heavy load modes, and implementing differentiated fault handling strategies, the lack of flexibility in fault tolerance mechanisms in existing technologies is solved, thereby improving the fault tolerance and stability of the storage system.

CN120704619BActive Publication Date: 2025-10-28INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511214609.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-10-28
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing storage systems lack flexible and dynamic adaptability in their fault tolerance mechanisms, resulting in sharp performance drops under high loads and low recovery efficiency under low loads, which affects the overall availability of the system.

Method used

By detecting the load information of the storage system, light load mode and heavy load mode are dynamically divided, and differentiated fault handling strategies are implemented according to the load information and fault level to avoid blind recovery operations and optimize fault recovery speed.

Benefits of technology

It improves the fault tolerance and stability of the storage system, reduces service interruption time caused by failures, and ensures that the system does not crash under high load and recovers quickly under low load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704619B_ABST
    Figure CN120704619B_ABST
Patent Text Reader

Abstract

This application discloses a storage system fault handling method and electronic device, which relates to the field of computer technology. Based on the load information of the storage system, this application determines whether its load mode is a light load mode or a heavy load mode, and implements differentiated fault handling strategies to handle load faults by combining the lightness and heaviness of the system load information and the current fault level. This avoids system crashes caused by blindly performing recovery operations under high load, while accelerating the recovery process under low load, thereby comprehensively improving the fault tolerance and stability of the storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a storage system fault handling method and electronic device. Background Technology

[0002] In large-scale storage systems, storage arrays composed of numerous disks are the core of data storage. To ensure data security and service continuity, storage systems typically possess a certain degree of fault recovery capability. To address disk failures, related technologies often employ Redundant Array of Independent Disks (RAID) technology. When a disk failure is detected, the system automatically reconstructs the data on a spare disk using redundant information to ensure data integrity. Simultaneously, some systems set fixed detection cycles to periodically check the disk health status. However, current storage system fault tolerance mechanisms often employ preset, fixed strategies, such as automatically activating a spare disk for data reconstruction when a disk error occurs. These fault tolerance strategies lack flexibility and dynamic adaptability. On one hand, regardless of the current system load, initiating a fixed reconstruction process whenever a disk failure is detected can consume significant system resources under high load scenarios, leading to a substantial increase in front-end service response latency. On the other hand, using the same handling method for faults of varying severity (such as minor sector errors and complete disk failure) may result in resource waste or untimely recovery, impacting the overall system availability. However, this approach lacks the ability to finely perceive and dynamically adjust the real-time status of the system. Under high load, forcing data reconstruction may cause a sharp decline in system performance or even trigger a chain of failures. Under low load, the recovery strategy may be too conservative, increasing the risk of data loss. Summary of the Invention

[0003] This application provides a storage system fault handling method and electronic device to at least solve the problem that in the related art, when a disk fault is detected, the current storage system's fault tolerance mechanism mostly adopts a preset fixed strategy, using the same processing method for high load and low load, resulting in a sharp drop in performance under high load and low recovery efficiency under low load, causing the fault tolerance strategy to lack flexibility and dynamic adaptability, thereby improving the fault tolerance and stability of the storage system.

[0004] This application provides a storage system fault handling method, including:

[0005] Detect the load information of the storage system, and determine the load mode of the storage system based on the load information, wherein the load mode includes a light load mode and a heavy load mode;

[0006] In response to the storage system being in a light load mode, the light load fault level is obtained under the light load mode, the current fault level of the storage system in the light load fault level is monitored in real time, and when the load information is less than a first threshold, the first utilization rate under the current fault level is obtained, and a preset light load fault handling strategy is selected to handle the load fault based on the first utilization rate and the current fault level.

[0007] When the first utilization rate is greater than the first ratio value and the current fault level is greater than the preset light load fault level, prompt whether to switch the storage system to heavy load mode;

[0008] In response to the storage system being in a heavy load mode, the heavy load fault level is obtained under the heavy load mode, the current fault level of the storage system in the heavy load fault level is monitored in real time, and when the load information is less than a second threshold, the second utilization rate under the current fault level is obtained, and a preset heavy load fault handling strategy is selected to handle the load fault based on the second utilization rate and the current fault level.

[0009] This application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described memory system fault handling methods when executing the computer program.

[0010] Detect the load information of the storage system, and determine the load mode of the storage system based on the load information, wherein the load mode includes a light load mode and a heavy load mode;

[0011] In response to the storage system being in a light load mode, the light load fault level is obtained under the light load mode, the current fault level of the storage system in the light load fault level is monitored in real time, and when the load information is less than a first threshold, the first utilization rate under the current fault level is obtained, and a preset light load fault handling strategy is selected to handle the load fault based on the first utilization rate and the current fault level.

[0012] When the first utilization rate is greater than the first ratio value and the current fault level is greater than the preset light load fault level, prompt whether to switch the storage system to heavy load mode;

[0013] In response to the storage system being in a heavy load mode, the heavy load fault level is obtained under the heavy load mode, the current fault level of the storage system in the heavy load fault level is monitored in real time, and when the load information is less than a second threshold, the second utilization rate under the current fault level is obtained, and a preset heavy load fault handling strategy is selected to handle the load fault based on the second utilization rate and the current fault level.

[0014] This application determines whether the storage system's load mode is light or heavy based on its load information. It then implements differentiated fault handling strategies to address load-related faults by combining the severity of the system load information with the current fault level. This avoids system crashes caused by blindly performing recovery operations under high loads, while accelerating the recovery process under low loads, thereby comprehensively improving the fault tolerance and stability of the storage system. Attached Figure Description

[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a diagram illustrating the application environment of a storage system fault handling method in one embodiment of this application.

[0017] Figure 2 This is a flowchart illustrating a storage system fault handling method in one embodiment of this application;

[0018] Figure 3 This is a structural block diagram of a storage system fault handling device in one embodiment of this application;

[0019] Figure 4 This is an internal structural diagram of a computer device in one embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0021] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] As mentioned in the background section, current storage systems often employ pre-defined, fixed fault tolerance mechanisms, which have significant limitations. For example, once a disk failure is detected, the system immediately initiates a standardized rebuild process. This "one-size-fits-all" approach lacks necessary flexibility. Specifically, regardless of the current system load, the same rebuild operation is performed. In high-load scenarios, this consumes a large amount of system resources, leading to a significant increase in front-end service response latency. Furthermore, using the same handling method for faults of varying severity, such as minor sector errors versus complete disk failure, can waste resources and result in delayed recovery from critical faults, ultimately impacting overall system availability.

[0024] The main drawback of existing technologies lies in the lack of fine-grained perception and dynamic adjustment capabilities for the real-time status of the system. Under high load conditions, forced data reconstruction may cause a precipitous drop in system performance, or even trigger a chain of failures; while during low load periods, overly conservative recovery strategies may inadvertently increase the risk of data loss. This static fault-tolerance mechanism is no longer adequate for the resilience and intelligence requirements of modern storage systems.

[0025] The storage system fault handling method provided in this application can be applied to, for example... Figure 1 In the application environment shown, the business host and the server's storage system communicate via a network. The business host can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be a standalone server or a server cluster consisting of multiple servers. The server's storage system includes multiple storage nodes. The business host connects to the server via Ethernet or a switch. The server's storage nodes connect to the business host via a storage controller and a front-end interface module. The storage controller includes an Intelligent Surveillance Module (ISM), a caching module, an alarm module, and a management module. The business host monitors the load information of multiple storage nodes and manages data buffering, alarm information dissemination, and fault handling within the storage controller.

[0026] The business host is a server that provides data processing and access services to users, connected to the storage system via a network. The front-end interface module, located at the front end of the storage system, is responsible for data interaction with the business host, supporting multiple interface types such as Ethernet and Fibre Channel. The intelligent monitoring module can obtain business access pressure through this module. Storage nodes are storage units composed of multiple disks, carrying the actual data storage tasks and monitored and managed by the intelligent monitoring module. The caching module is used for temporary storage of frequently accessed data, improving the read and write performance of the storage system. The intelligent monitoring module can monitor its cache hit rate and utilization rate in real time. The alarm module, controlled by the intelligent monitoring module, issues alarm information via sound, light, SMS, email, etc., when the system malfunctions or experiences anomalies. The management module provides an interactive interface between the user and the intelligent monitoring module, allowing for parameter configuration, status query, and function activation.

[0027] The intelligent monitoring module is deployed on the storage controller and utilizes high-performance programmable chips such as FPGAs. This module dynamically manages storage nodes, disk arrays, and cache modules under both light and heavy load modes. Business hosts connect to the storage system's front-end interface module via Ethernet or Fibre Channel to access storage services. The intelligent monitoring module collects system performance data in real time through the storage nodes, including CPU utilization, memory usage, disk I / O throughput, and response time. It also embeds maximum system load data for different fault levels. Fault levels are categorized into multiple levels; lower numbers indicate less severe faults. For example, a level 1 fault is less severe than a level 2 fault, meaning its impact on the system is less than that of a level 2 fault. Intelligent monitoring modules interconnect via InfiniBand or 10 Gigabit Ethernet to synchronize fault and load information. Furthermore, the intelligent monitoring module can obtain the current access pressure of the business hosts through the front-end interface module.

[0028] like Figure 2 As shown, an embodiment of this application provides a storage system fault handling method, including the following steps:

[0029] Step S1: Detect the load information of the storage system and determine the load mode of the storage system based on the load information, wherein the load mode includes light load mode and heavy load mode.

[0030] Step S2: In response to the storage system being in light load mode, obtain the light load fault level divided in light load mode, monitor the current fault level of the storage system in the light load fault level in real time, obtain the first utilization rate under the current fault level when the load information is less than the first threshold, and select a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level.

[0031] Step S3: In response to the first utilization rate being greater than the first ratio value and the current fault level being greater than the preset light load fault level, prompt whether to switch the storage system to heavy load mode.

[0032] Step S4: In response to the storage system being in heavy load mode, obtain the heavy load fault level divided in heavy load mode, monitor the current fault level of the storage system in the heavy load fault level in real time, obtain the second utilization rate under the current fault level when the load information is less than the second threshold, and select the preset heavy load fault handling strategy to handle the load fault according to the second utilization rate and the current fault level.

[0033] This embodiment combines the severity of system load information with the current fault level to implement differentiated fault handling strategies for load-related faults. This avoids system crashes caused by blindly performing recovery operations under high load, while accelerating the recovery process under low load, thereby comprehensively improving the fault tolerance and stability of the storage system. Under light load conditions, it can quickly respond to faults and perform recovery; under heavy load conditions, it can flexibly adjust the handling method according to the severity of the fault, avoiding excessive impact on critical business operations due to recovery operations. This effectively improves the fault tolerance and stability of the storage system and reduces service interruption time caused by faults.

[0034] In this embodiment, the load information of the storage system is detected, and the load mode of the storage system is determined based on the load information. The load mode includes a light load mode and a heavy load mode.

[0035] Real-time monitoring of storage nodes, disk arrays, and cache modules in the storage system is performed to obtain real-time load information of storage nodes, disk arrays, and cache modules respectively, and the load information of the storage system is determined based on the real-time load information.

[0036] If the load information of the storage system is less than the preset load threshold, the storage system is determined to be in a light load mode.

[0037] If the load information of the storage system is greater than or equal to a preset load threshold, the storage system is determined to be in heavy load mode.

[0038] In this embodiment, when the load information is less than a first threshold, the first utilization rate under the current fault level is obtained, and a preset light load fault handling strategy is selected to handle the load fault based on the first utilization rate and the current fault level, including:

[0039] When the load information is less than the first threshold, obtain the system performance data and maximum load data under the current fault level, and determine the first utilization rate under the current fault level;

[0040] When the first utilization rate is less than or equal to the first ratio value and the current fault level is less than or equal to the preset light load fault level, a fault recovery operation is performed.

[0041] When the first utilization rate is greater than the first proportion value, or the current fault level is greater than the preset light load fault level, a first-level fault tolerance warning is issued and a fault recovery operation is performed.

[0042] When the first utilization rate is greater than the first proportion value and the current fault level is greater than the preset light load fault level, a level 1 fault tolerance alarm is issued and no fault recovery operation is performed.

[0043] In this embodiment, when the load information is less than a first threshold, the system performance data and maximum load data under the current fault level are obtained, and the first utilization rate under the current fault level is determined, including:

[0044] When the load information is less than the first threshold, a fault tolerance test is performed and the system performance data in the test results is obtained. The system performance data includes processor utilization, memory usage, disk I / O throughput, and response time.

[0045] Obtain the maximum load data under the current fault level;

[0046] The first utilization rate is obtained by dividing the system performance data by the maximum capacity data.

[0047] In this embodiment, prompting whether to switch the storage system to heavy load mode in response to a first utilization rate greater than a first proportion value and a current fault level greater than a preset light-load fault level further includes:

[0048] If no response is received from the user to the prompt within a preset time, the storage system is switched to heavy load mode; and / or, if the load information of the storage system remains below a preset load threshold after running in heavy load mode for a period of time, the storage system is switched to light load mode.

[0049] This includes the addition of automated mode switching logic, including automatic switching after timeout and mechanisms for automatically switching back from heavy load mode to light load mode.

[0050] In this embodiment, when the load information is less than a second threshold, the second utilization rate under the current fault level is obtained. The load fault is then processed using a preset heavy load fault handling strategy based on the second utilization rate and the current fault level, including:

[0051] When the load information is less than the second threshold, obtain the system performance data and maximum load data under the current fault level, and determine the second utilization rate under the current fault level;

[0052] When the second utilization rate is less than or equal to the second ratio value and the current fault level is less than or equal to the preset light load fault level, a fault recovery operation is performed.

[0053] When the second utilization rate is greater than the second proportion value, or the current fault level is greater than the first preset heavy load fault level, a second-level fault tolerance warning is issued and a fault recovery operation is performed.

[0054] When the second utilization rate is greater than the second ratio value and the current fault level is greater than the preset light load fault level, a level 2 fault tolerance alarm is issued and a fault recovery operation is performed.

[0055] When the current fault level is greater than the second preset heavy load fault level, the non-critical front-end services are suspended and fault recovery operations are performed.

[0056] Specifically, after adjusting the fault recovery execution speed, the system continuously monitors the storage system's load information, utilization, and fault recovery progress. Based on the monitoring results, the fault recovery execution speed is dynamically adjusted to complete the recovery as quickly as possible without affecting system stability and critical business operations. A dynamic feedback and readjustment mechanism for the recovery speed has been added.

[0057] Specifically, when the load information is less than the second threshold, the system performance data and maximum load data under the current fault level are obtained, and the second utilization rate under the current fault level is determined, including:

[0058] When the load information is less than the second threshold, a fault tolerance test is performed and the system performance data in the test results is obtained. The system performance data includes processor utilization, memory usage, disk I / O throughput, and response time.

[0059] Obtain the maximum load data under the current fault level;

[0060] The second utilization rate is obtained by dividing the system performance data by the maximum capacity data.

[0061] In this embodiment, when the load information is less than the second threshold, the second utilization rate under the current fault level is obtained. The process of selecting a preset heavy load fault handling strategy based on the second utilization rate and the current fault level to handle the load fault further includes:

[0062] When the current fault level is greater than the second preset heavy load fault level, check whether the fault recovery operation has been completed.

[0063] In response to the completion of the fault recovery operation, obtain the third utilization rate under the current fault level;

[0064] In response to the third utilization rate being less than or equal to the third ratio value, resume non-critical front-end business operations.

[0065] Among them, obtaining the third utilization rate under the current fault level after the fault recovery operation is completed includes:

[0066] After the fault recovery operation is completed, a fault tolerance test is performed and the system performance data in the test results is obtained. The system performance data includes processor utilization, memory usage, disk I / O throughput, and response time.

[0067] After completing the fault recovery operation, re-monitor the current fault level of the storage system under heavy load fault level, and obtain the maximum data capacity under the current fault level;

[0068] The third utilization rate is obtained by dividing the system performance data by the maximum capacity data.

[0069] In this embodiment, the storage system fault handling method further includes:

[0070] When responding to a storage system failure recovery operation, the failure recovery execution speed is determined based on load information, the current failure level, and the system utilization rate under the current failure level, and the failure recovery execution speed in heavy load mode is controlled to be lower than that in light load mode.

[0071] The fault recovery operations include, but are not limited to: data reconstruction (e.g., RAID reconstruction), hot spare disk switching, data migration, snapshot recovery, component restart, system degradation, load balancing adjustment, or fault isolation.

[0072] In this embodiment, the storage system fault handling method further includes:

[0073] Real-time monitoring of the operating status and performance metrics of storage nodes, disk arrays, and cache modules in the storage system;

[0074] Based on operational status and performance indicators, combined with historical data, predict whether there are potential failures in the storage system;

[0075] In response to the prediction of a potential failure in the storage system, the severity of the potential failure is obtained. Based on load information, the severity of the potential failure, and / or the current system utilization, a preset preventive maintenance strategy is selected and executed before the failure occurs. The preventive maintenance strategy includes, but is not limited to: data pre-migration, adding redundancy, adjusting load information, replacing components, or issuing preventive warnings.

[0076] Integrating predictive maintenance into existing storage systems transforms the system from a passive response to an active prevention approach, further improving fault tolerance.

[0077] In this embodiment, when issuing a Level 1 fault tolerance warning, a Level 1 fault tolerance alarm, a Level 2 fault tolerance warning, or a Level 2 fault tolerance alarm, the following is also included:

[0078] Send notifications to system administrators or relevant operations and maintenance personnel via email, SMS, system management interface, API interface, or log records, along with fault details, suggested handling measures, and / or current system status reports.

[0079] The update includes more specific notification methods and additional information for early warnings and alerts, improving their usability.

[0080] Specifically, in light-load mode, ISM classifies system fault levels into P levels, with level 1 being the least severe and level P the most severe. Within time period S1, ISM monitors the system's fault levels in real time and performs a fault tolerance test (FT) when the system load information falls below the percentage E1. Let Y represent the maximum load data for each fault level. If the system performance data X in a single FT result does not exceed Y×F1, and the fault level is not higher than P1, the system can automatically perform fault recovery. If the system performance data X in the FT result exceeds Y×F1, or the fault level is higher than P1, a level 1 fault tolerance warning is issued. If the system performance data X in the FT result exceeds Y×F1, and the fault level is higher than P1, a level 1 fault tolerance alarm is issued, prompting whether to switch to heavy-load mode. In this case, the system temporarily does not perform automatic fault recovery.

[0081] In heavy load mode, ISM classifies system fault levels into Q levels, with level 1 being the least severe and level Q the most severe. Within time period S2, ISM monitors the system's fault level in real time and performs a fault assessment (FT) when the system load information is below the percentage E2. If the system performance data X in the FT result does not exceed Y×F2 and the fault level is not higher than level Q1, the system can automatically perform fault recovery. If the system performance data X in the FT result exceeds Y×F2, or the fault level is higher than level Q1, a level 2 fault tolerance warning is issued. If the system performance data X in the FT result exceeds Y×F2 and the fault level is higher than level Q1, a level 2 fault tolerance alarm is issued. At this time, if the fault level is not higher than level Q2, the system can automatically perform fault recovery; if the fault level is higher than level Q2, the system will first suspend non-critical front-end services, prioritize fault recovery, and after recovery, perform another FT within time period S3. If the system performance data X does not exceed Y×F3, the non-critical front-end services will be restored.

[0082] Among them, Y, E1, E2, F1, F2, F3, P, P1, Q, Q1, Q2, S1, S2, and S3 are all system preset parameters that can be adjusted through the system management interface or command-line tools. E1 is the first threshold, preferably E1=40%; X / Y is the first or second utilization rate; F1 is the first proportional value, preferably F1=70%; E2 is the second threshold, preferably E2=20%; F2 is the second proportional value, preferably F2=50%; F3 is the third proportional value, preferably F3=60%; P is the light load fault level, preferably P=6; and Q is the heavy load fault level, preferably Q=10.

[0083] For example, in light load mode, ISM classifies system fault levels into 6 levels, with level 1 being the least severe and level 6 the most severe. Within a 60-second time period, ISM monitors the system's fault level in real time and performs a fault assessment (FT) when the system load falls below 40%. If the system performance data X in the FT result does not exceed 70% of Y × 70%, and the fault level is not higher than level 3, the system can automatically perform fault recovery. If the system performance data X in the FT result exceeds 70% of Y × 70%, or the fault level is higher than level 3, a level 1 fault tolerance warning is issued. If the system performance data X in the FT result exceeds 70% of Y × 70%, and the fault level is higher than level 3, a level 1 fault tolerance alarm is issued, prompting the user to switch to heavy load mode. In this mode, the system temporarily suspends automatic fault recovery.

[0084] In heavy load mode, ISM classifies system fault levels into 10 levels, with level 1 being the least severe and level 10 the most severe. Within a 30-second time period, ISM monitors the system's fault level in real time and performs a fault assessment (FT) when the system load falls below 20%. If the system performance data X in the FT result does not exceed 50% of Y, and the fault level is not higher than level 5, the system can automatically perform fault recovery. If the system performance data X in the FT result exceeds 50% of Y, or the fault level is higher than level 5, a level 2 fault tolerance warning is issued. If the system performance data X in the FT result exceeds 50% of Y, and the fault level is higher than level 5, a level 2 fault tolerance alarm is issued. In this case, if the fault level is not higher than level 7, the system can automatically perform fault recovery; if the fault level is higher than level 7, the system will first suspend non-critical front-end services, prioritize fault recovery, and after recovery, perform another FT within a 45-second time period. If X does not exceed 60% of Y, then the non-critical front-end services will resume.

[0085] The aforementioned storage system fault handling method determines the storage system's load mode (light or heavy) based on its load information. It then implements differentiated fault handling strategies based on the severity of the load and the current fault level. This avoids system crashes caused by blindly performing recovery operations under high load, while accelerating the recovery process under low load, thereby comprehensively improving the storage system's fault tolerance and stability. Under light load conditions, it can quickly respond to faults and perform recovery; under heavy load conditions, it can flexibly adjust the handling method according to the severity of the fault, avoiding excessive impact on critical business operations due to recovery operations. This effectively improves the storage system's fault tolerance and stability, and reduces service interruption time caused by faults.

[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0087] In one embodiment, such as Figure 3 As shown, a storage system fault handling device 10 is provided, including: a load information detection module 1, a light load fault handling module 2, a mode switching module 3, a heavy load fault handling module 4, a fault recovery speed control module 5, and a predictive maintenance module 6.

[0088] The load information detection module 1 is used to detect the load information of the storage system and determine the load mode of the storage system based on the load information. The load mode includes light load mode and heavy load mode.

[0089] The light load fault handling module 2 is used to respond to the storage system being in light load mode, obtain the light load fault level divided in light load mode, monitor the current fault level of the storage system in light load fault level in real time, obtain the first utilization rate under the current fault level when the load information is less than the first threshold, and select a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level.

[0090] The mode switching module 3 is used to prompt whether to switch the storage system to heavy load mode when the first utilization rate is greater than the first ratio value and the current fault level is greater than the preset light load fault level.

[0091] The heavy load fault handling module 4 is used to respond to the storage system being in heavy load mode, obtain the heavy load fault level divided in heavy load mode, monitor the current fault level of the storage system in the heavy load fault level in real time, obtain the second utilization rate under the current fault level when the load information is less than the second threshold, and select a preset heavy load fault handling strategy to handle the load fault based on the second utilization rate and the current fault level.

[0092] In this embodiment, the load information of the storage system is detected, and the load mode of the storage system is determined based on the load information. The load mode includes a light load mode and a heavy load mode.

[0093] Real-time monitoring of storage nodes, disk arrays, and cache modules in the storage system is performed to obtain real-time load information of storage nodes, disk arrays, and cache modules respectively, and the load information of the storage system is determined based on the real-time load information.

[0094] If the load information of the storage system is less than the preset load threshold, the storage system is determined to be in a light load mode.

[0095] If the load information of the storage system is greater than or equal to a preset load threshold, the storage system is determined to be in heavy load mode.

[0096] In this embodiment, when the load information is less than a first threshold, the first utilization rate under the current fault level is obtained, and a preset light load fault handling strategy is selected to handle the load fault based on the first utilization rate and the current fault level, including:

[0097] When the load information is less than the first threshold, obtain the system performance data and maximum load data under the current fault level, and determine the first utilization rate under the current fault level;

[0098] When the first utilization rate is less than or equal to the first ratio value and the current fault level is less than or equal to the preset light load fault level, a fault recovery operation is performed.

[0099] When the first utilization rate is greater than the first proportion value, or the current fault level is greater than the preset light load fault level, a first-level fault tolerance warning is issued and a fault recovery operation is performed.

[0100] When the first utilization rate is greater than the first proportion value and the current fault level is greater than the preset light load fault level, a level 1 fault tolerance alarm is issued and no fault recovery operation is performed.

[0101] In this embodiment, when the load information is less than a first threshold, the system performance data and maximum load data under the current fault level are obtained, and the first utilization rate under the current fault level is determined, including:

[0102] When the load information is less than the first threshold, a fault tolerance test is performed and the system performance data in the test results is obtained. The system performance data includes processor utilization, memory usage, disk I / O throughput, and response time.

[0103] Obtain the maximum load data under the current fault level;

[0104] The first utilization rate is obtained by dividing the system performance data by the maximum capacity data.

[0105] In this embodiment, prompting whether to switch the storage system to heavy load mode in response to a first utilization rate greater than a first proportion value and a current fault level greater than a preset light-load fault level further includes:

[0106] If no response is received from the user to the prompt within a preset time, the storage system is switched to heavy load mode; and / or, if the load information of the storage system remains below a preset load threshold after running in heavy load mode for a period of time, the storage system is switched to light load mode.

[0107] In this embodiment, when the load information is less than a second threshold, the second utilization rate under the current fault level is obtained. The load fault is then processed using a preset heavy load fault handling strategy based on the second utilization rate and the current fault level, including:

[0108] When the load information is less than the second threshold, obtain the system performance data and maximum load data under the current fault level, and determine the second utilization rate under the current fault level;

[0109] When the second utilization rate is less than or equal to the second ratio value and the current fault level is less than or equal to the preset light load fault level, a fault recovery operation is performed.

[0110] When the second utilization rate is greater than the second proportion value, or the current fault level is greater than the first preset heavy load fault level, a second-level fault tolerance warning is issued and a fault recovery operation is performed.

[0111] When the second utilization rate is greater than the second ratio value and the current fault level is greater than the preset light load fault level, a level 2 fault tolerance alarm is issued and a fault recovery operation is performed.

[0112] When the current fault level is greater than the second preset heavy load fault level, the non-critical front-end services are suspended and fault recovery operations are performed.

[0113] In this embodiment, when the load information is less than the second threshold, the second utilization rate under the current fault level is obtained. The process of selecting a preset heavy load fault handling strategy based on the second utilization rate and the current fault level to handle the load fault further includes:

[0114] When the current fault level is greater than the second preset heavy load fault level, check whether the fault recovery operation has been completed.

[0115] In response to the completion of the fault recovery operation, obtain the third utilization rate under the current fault level;

[0116] In response to the third utilization rate being less than or equal to the third ratio value, resume non-critical front-end business operations.

[0117] In this embodiment, the fault recovery speed control module 5 is used for:

[0118] When responding to a storage system failure recovery operation, the failure recovery execution speed is determined based on load information, the current failure level, and the system utilization rate under the current failure level, and the failure recovery execution speed in heavy load mode is controlled to be lower than that in light load mode.

[0119] In this embodiment, the predictive maintenance module 6 is used for:

[0120] Real-time monitoring of the operating status and performance metrics of storage nodes, disk arrays, and cache modules in the storage system;

[0121] Based on operational status and performance indicators, combined with historical data, predict whether there are potential failures in the storage system;

[0122] In response to the prediction of a potential failure in the storage system, the severity of the potential failure is obtained. Based on load information, the severity of the potential failure, and / or the current system utilization, a preset preventive maintenance strategy is selected and executed before the failure occurs. The preventive maintenance strategy includes, but is not limited to: data pre-migration, adding redundancy, adjusting load information, replacing components, or issuing preventive warnings.

[0123] In the aforementioned storage system fault handling device, the load mode of the storage system is determined to be either light load mode or heavy load mode based on the load information of the storage system. The device combines the severity of the system load information with the current fault level to implement differentiated fault handling strategies to handle load faults. This avoids system crashes caused by blindly performing recovery operations under high load, while accelerating the recovery process under low load, thereby comprehensively improving the fault tolerance and stability of the storage system.

[0124] For a description of the features in the embodiment corresponding to the storage system fault handling device, please refer to the relevant description of the embodiment corresponding to the storage system fault handling method, which will not be repeated here.

[0125] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the storage system fault handling method.

[0126] In one embodiment, the electronic device may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the electronic device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores storage system fault handling data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a storage system fault handling method.

[0127] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the storage system fault handling method when running:

[0128] Detect the load information of the storage system and determine the load mode of the storage system based on the load information, including light load mode and heavy load mode.

[0129] In response to the storage system being in a light-load mode, the light-load fault level is obtained under the light-load mode. The current fault level of the storage system in the light-load fault level is monitored in real time. When the load information is less than the first threshold, the first utilization rate under the current fault level is obtained. Based on the first utilization rate and the current fault level, a preset light-load fault handling strategy is selected to handle the load fault.

[0130] When the first utilization rate is greater than the first proportion value and the current fault level is greater than the preset light load fault level, a prompt will be made asking whether to switch the storage system to heavy load mode.

[0131] In response to the storage system being in heavy load mode, the heavy load fault level is obtained under heavy load mode. The current fault level of the storage system in the heavy load fault level is monitored in real time. When the load information is less than the second threshold, the second utilization rate under the current fault level is obtained. Based on the second utilization rate and the current fault level, a preset heavy load fault handling strategy is selected to handle the load fault.

[0132] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0133] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the storage system fault handling method:

[0134] Detect the load information of the storage system and determine the load mode of the storage system based on the load information, including light load mode and heavy load mode.

[0135] In response to the storage system being in a light-load mode, the light-load fault level is obtained under the light-load mode. The current fault level of the storage system in the light-load fault level is monitored in real time. When the load information is less than the first threshold, the first utilization rate under the current fault level is obtained. Based on the first utilization rate and the current fault level, a preset light-load fault handling strategy is selected to handle the load fault.

[0136] When the first utilization rate is greater than the first proportion value and the current fault level is greater than the preset light load fault level, a prompt will be made asking whether to switch the storage system to heavy load mode.

[0137] In response to the storage system being in heavy load mode, the heavy load fault level is obtained under heavy load mode. The current fault level of the storage system in the heavy load fault level is monitored in real time. When the load information is less than the second threshold, the second utilization rate under the current fault level is obtained. Based on the second utilization rate and the current fault level, a preset heavy load fault handling strategy is selected to handle the load fault.

[0138] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the storage system fault handling method:

[0139] Detect the load information of the storage system and determine the load mode of the storage system based on the load information, including light load mode and heavy load mode.

[0140] In response to the storage system being in a light-load mode, the light-load fault level is obtained under the light-load mode. The current fault level of the storage system in the light-load fault level is monitored in real time. When the load information is less than the first threshold, the first utilization rate under the current fault level is obtained. Based on the first utilization rate and the current fault level, a preset light-load fault handling strategy is selected to handle the load fault.

[0141] When the first utilization rate is greater than the first proportion value and the current fault level is greater than the preset light load fault level, a prompt will be made asking whether to switch the storage system to heavy load mode.

[0142] In response to the storage system being in heavy load mode, the heavy load fault level is obtained under heavy load mode. The current fault level of the storage system in the heavy load fault level is monitored in real time. When the load information is less than the second threshold, the second utilization rate under the current fault level is obtained. Based on the second utilization rate and the current fault level, a preset heavy load fault handling strategy is selected to handle the load fault.

[0143] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0144] The foregoing has provided a detailed description of a storage system fault handling method and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for handling storage system faults, characterized in that, include: Detect the load information of the storage system, and determine the load mode of the storage system based on the load information, wherein the load mode includes a light load mode and a heavy load mode; In response to the storage system being in a light load mode, the light load fault level is obtained under the light load mode, the current fault level of the storage system in the light load fault level is monitored in real time, and when the load information is less than a first threshold, the first utilization rate under the current fault level is obtained, and a preset light load fault handling strategy is selected to handle the load fault based on the first utilization rate and the current fault level. When the first utilization rate is greater than the first ratio value and the current fault level is greater than the preset light load fault level, prompt whether to switch the storage system to heavy load mode; In response to the storage system being in a heavy load mode, the heavy load fault level is obtained under the heavy load mode, the current fault level of the storage system in the heavy load fault level is monitored in real time, and when the load information is less than a second threshold, the second utilization rate under the current fault level is obtained, and a preset heavy load fault handling strategy is selected to handle the load fault based on the second utilization rate and the current fault level.

2. The storage system fault handling method according to claim 1, characterized in that, The detection of the storage system's load information, and the determination of the storage system's load mode based on the load information, wherein the load mode includes a light load mode and a heavy load mode, including: The storage nodes, disk arrays, and cache modules in the storage system are monitored in real time, and the real-time load information of the storage nodes, disk arrays, and cache modules is obtained respectively. The load information of the storage system is determined based on the real-time load information. If the load information of the storage system is less than a preset load threshold, the storage system is determined to be in a light load mode. If the load information of the storage system is greater than or equal to a preset load threshold, the storage system is determined to be in a heavy load mode.

3. The storage system fault handling method according to claim 1, characterized in that, The step of obtaining a first utilization rate under the current fault level when the load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level includes: When the load information is less than a first threshold, obtain the system performance data and maximum load data under the current fault level, and determine the first utilization rate under the current fault level; When the first utilization rate is less than or equal to the first ratio value and the current fault level is less than or equal to the preset light load fault level, a fault recovery operation is performed. In response to the first utilization rate being greater than the first ratio value, or the current fault level being greater than the preset light load fault level, a first-level fault tolerance warning is issued and a fault recovery operation is performed. When the first utilization rate is greater than the first ratio value and the current fault level is greater than the preset light load fault level, a level 1 fault tolerance alarm is issued and no fault recovery operation is performed.

4. The storage system fault handling method according to claim 3, characterized in that, When the load information is less than a first threshold, acquiring system performance data and maximum load data under the current fault level, and determining the first utilization rate under the current fault level includes: When the load information is less than a first threshold, a fault tolerance test is performed and system performance data in the test results is obtained, wherein the system performance data includes processor utilization, memory usage, disk I / O throughput, and response time; Obtain the maximum load data under the current fault level; The first utilization rate is obtained by dividing the system performance data by the maximum capacity data.

5. The storage system fault handling method according to claim 1, characterized in that, The step of prompting whether to switch the storage system to heavy load mode in response to the first utilization rate being greater than a first proportion value and the current fault level being greater than a preset light load fault level further includes: If no response is received from the user to the prompt within a preset time, the storage system is switched to heavy load mode; and / or If, after the storage system has been running in heavy load mode for a period of time, its load information remains below a preset load threshold, the storage system is switched to light load mode.

6. The storage system fault handling method according to claim 1, characterized in that, The step of obtaining a second utilization rate under the current fault level when the load information is less than a second threshold, and selecting a preset heavy load fault handling strategy to handle the load fault based on the second utilization rate and the current fault level includes: When the load information is less than the second threshold, obtain the system performance data and maximum load data under the current fault level, and determine the second utilization rate under the current fault level; When the second utilization rate is less than or equal to the second ratio value and the current fault level is less than or equal to the preset light load fault level, a fault recovery operation is performed. In response to the second utilization rate being greater than the second ratio value, or the current fault level being greater than the first preset heavy load fault level, a level two fault tolerance warning is issued, and a fault recovery operation is performed. When the second utilization rate is greater than the second ratio value and the current fault level is greater than the preset light load fault level, a level 2 fault tolerance alarm is issued and a fault recovery operation is performed. When the current fault level is greater than the second preset heavy load fault level, the non-critical front-end services are suspended and a fault recovery operation is performed.

7. The storage system fault handling method according to claim 6, characterized in that, The step of obtaining the second utilization rate under the current fault level when the load information is less than the second threshold, and selecting a preset heavy load fault handling strategy to handle the load fault based on the second utilization rate and the current fault level, further includes: When the current fault level is greater than the second preset heavy load fault level, it is detected whether the fault recovery operation has been completed. In response to the completion of the fault recovery operation, the third utilization rate under the current fault level is obtained; In response to the third utilization rate being less than or equal to the third ratio value, non-critical front-end services are restored.

8. The storage system fault handling method according to claim 7, characterized in that, The storage system fault handling method further includes: When the storage system performs a fault recovery operation, the fault recovery execution speed is determined based on the load information, the current fault level, and the system utilization rate under the current fault level, and the fault recovery execution speed in the heavy load mode is controlled to be lower than the fault recovery execution speed in the light load mode.

9. The storage system fault handling method according to claim 1, characterized in that, The storage system fault handling method further includes: Real-time monitoring of the operating status and performance indicators of storage nodes, disk arrays, and cache modules in the storage system; Based on the operating status and performance indicators, combined with historical data, predict whether the storage system has potential failures; In response to the prediction of a potential failure in the storage system, the severity of the potential failure is obtained. Based on the load information, the severity of the potential failure, and / or the current system utilization, a preset preventive maintenance strategy is selected and executed before the failure occurs. The preventive maintenance strategy includes, but is not limited to: data pre-migration, adding redundancy, adjusting load information, replacing components, or issuing preventive warnings.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the storage system fault handling method as described in any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Fault processing method and device, computer equipment, readable storage medium and program product

    CN119201536A

  • Disk scheduling method

    CN120428922A