Storage system fault processing method and electronic equipment
By dynamically dividing load patterns and implementing differentiated fault handling strategies, the lack of flexibility in the fault tolerance mechanism in the storage system is resolved, the system's fault tolerance and stability are improved, and service interruption time is reduced.
Patent Information
- Application Number
- CN202511214609.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-28
AI Technical Summary
The fault tolerance mechanism of existing storage systems lacks flexibility and dynamic adaptability, resulting in a sudden drop in performance under high load and low recovery efficiency under low load, affecting the overall availability and stability of the system.
By detecting the load information of the storage system, dynamically dividing it into light load mode and heavy load mode, and implementing differentiated fault handling strategies based on the load information and fault level, blind recovery operations are avoided and the fault recovery process is optimized.
Improves the fault tolerance and stability of the storage system, reduces service interruption time caused by failures, and ensures that the system does not crash under high load and recovers quickly under low load.
Smart Images

Figure CN120704619A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a storage system fault handling method and electronic equipment. Background Art
[0002] In large-scale storage systems, storage arrays comprised of numerous disks are the core of data storage. To ensure data security and service continuity, storage systems typically possess a certain degree of fault recovery capability. To address disk failures, related technologies typically employ redundant array of arrays (RAID) technology. Upon detecting a disk failure, data is automatically rebuilt on a spare disk using redundant information to ensure data integrity. Furthermore, some systems implement fixed detection cycles to regularly check disk health. However, current storage system fault tolerance mechanisms often rely on pre-set fixed policies. For example, when a disk error occurs, a spare disk is automatically activated for data reconstruction. These fault tolerance policies lack flexibility and dynamic adaptability. First, regardless of the system's current load, a fixed reconstruction process is initiated upon detecting a disk failure. This can consume significant system resources in high-load scenarios, significantly increasing front-end service response latency. Second, applying the same approach to faults of varying severity (e.g., minor sector errors and complete disk failures) can waste resources or delay recovery, impacting overall system availability. However, this approach lacks refined perception of the system's real-time status and the ability to dynamically adjust. Under high load conditions, forced data reconstruction may cause a sharp decline in system performance or even trigger a chain reaction of failures. Under low load conditions, the recovery strategy may be too conservative, increasing the risk of data loss. Summary of the Invention
[0003] The present application provides a storage system fault handling method and electronic device to at least solve the problem in the related art that when a disk failure is detected, the fault tolerance mechanism of the current storage system mostly adopts a preset fixed strategy, and adopts the same processing method for high load and low load, resulting in a sudden drop in performance under high load and low recovery efficiency under low load, resulting in a lack of flexibility and dynamic adaptability of the fault tolerance strategy, thereby improving the fault tolerance and stability of the storage system.
[0004] This application provides a storage system fault handling method, including: Detecting load information of a storage system, and determining a load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode; In response to the storage system being in a light load mode, obtaining a light load fault level classified under the light load mode, monitoring in real time a current fault level of the storage system within the light load fault level, obtaining a first utilization rate at the current fault level when the load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level; In response to the first utilization being greater than a first proportional value and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to a heavy load mode; In response to the storage system being in heavy load mode, a heavy load fault level divided under the heavy load mode is obtained, the current fault level of the storage system in the heavy load fault level is monitored in real time, a second utilization rate at the current fault level is obtained when the load information is less than a second threshold, and a preset heavy load fault handling strategy is selected according to the second utilization rate and the current fault level to handle the load fault.
[0005] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned storage system fault handling methods when executing the computer program: Detecting load information of a storage system, and determining a load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode; In response to the storage system being in a light load mode, obtaining a light load fault level classified under the light load mode, monitoring in real time a current fault level of the storage system within the light load fault level, obtaining a first utilization rate at the current fault level when the load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level; In response to the first utilization being greater than a first proportional value and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to a heavy load mode; In response to the storage system being in heavy load mode, a heavy load fault level divided under the heavy load mode is obtained, the current fault level of the storage system in the heavy load fault level is monitored in real time, a second utilization rate at the current fault level is obtained when the load information is less than a second threshold, and a preset heavy load fault handling strategy is selected according to the second utilization rate and the current fault level to handle the load fault.
[0006] Through this application, the load mode of the storage system is determined to be a light load mode or a heavy load mode based on the load information of the storage system, and differentiated fault handling strategies are implemented to handle load failures in combination with the severity of the system load information and the current fault level, thereby avoiding system crashes caused by blindly executing recovery operations under high load, and accelerating the recovery process under low load, thereby comprehensively improving the fault tolerance and stability of the storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0008] Figure 1 This is an application environment diagram of a storage system fault handling method in one embodiment of the present application; Figure 2 This is a flowchart of a method for handling a storage system failure in one embodiment of the present application; Figure 3 This is a structural block diagram of a storage system fault processing device in one embodiment of the present application; Figure 4 This is a diagram of the internal structure of a computer device in one embodiment of the present application. DETAILED DESCRIPTION
[0009] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of this application.
[0010] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0011] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0012] As mentioned in the background, the fault tolerance mechanisms of current storage systems mostly use preset fixed strategies, which have obvious limitations. For example, once a disk failure is detected, the system will immediately initiate a standardized reconstruction process. This "one-size-fits-all" approach lacks the necessary flexibility. Specifically, regardless of the current load information status of the system, the same reconstruction operation will be performed. For example, in high-load scenarios, this will occupy a large amount of system resources, resulting in a significant increase in front-end business response delays; for failures of different severity, such as minor sector errors and complete disk failure, the same processing method is used, which may not only waste resources, but also lead to delayed recovery of critical failures, ultimately affecting the overall availability of the system.
[0013] The main drawback of existing technologies is their lack of refined awareness of the system's real-time status and the ability to dynamically adjust. Under high-load conditions, forced data reconstruction can cause a drastic drop in system performance or even trigger a cascading failure. During low-load periods, overly conservative recovery strategies can inadvertently increase the risk of data loss. This static fault-tolerance mechanism is no longer able to meet the elastic and intelligent requirements of modern storage systems.
[0014] The storage system fault handling method provided in this application can be applied to Figure 1 In the application environment shown. The service host communicates with the storage system of the server through a network. The service host can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers. The storage system of the server includes multiple storage nodes. The service host is connected to the server via Ethernet or a switch. The storage nodes of the server are connected to the service host via a storage controller and a front-end interface module. The storage controller includes an intelligent monitoring module (ISM), a cache module, an alarm module, and a management module. The service host detects load information on multiple storage nodes and performs data buffering, alarm information release, and fault handling management of the storage nodes in the storage controller.
[0015] The business host is a server that provides data processing and access services to users and is connected to the storage system through a network. The front-end interface module is located at the front end of the storage system and is responsible for data interaction with the business host. It supports multiple interface types such as Ethernet and Fibre Channel. The intelligent monitoring module can obtain business access pressure through this module. The storage node is a storage unit composed of multiple disks, which carries the actual data storage tasks and is monitored and managed by the intelligent monitoring module. The cache module is used to temporarily store frequently accessed data to improve the read and write performance of the storage system. The intelligent monitoring module can monitor its cache hit rate and usage rate in real time. The alarm module is controlled by the intelligent monitoring module. When the system fails or anomalies, it will issue alarm information through sound, light, SMS, email, etc. The management module provides an interactive interface between the user and the intelligent monitoring module, which can perform operations such as parameter configuration, status query, and function activation.
[0016] The intelligent monitoring module, deployed on the storage controller, utilizes high-performance programmable chips such as FPGAs. It dynamically manages storage nodes, disk arrays, and cache modules in both light-load and heavy-load modes. Service hosts connect to the storage system's front-end interface module via Ethernet or Fibre Channel to access storage services. The intelligent monitoring module collects real-time system performance data from the storage nodes, including CPU utilization, memory usage, disk I / O throughput, and response time. It also embeds data on the system's maximum load capacity for different fault levels. Fault levels are categorized into multiple levels, with lower levels indicating less severe failures. For example, a level 1 failure is less severe than a level 2 failure, meaning that a level 1 failure has less impact on the system than a level 2 failure. Intelligent monitoring modules are interconnected via InfiniBand or 10 Gigabit Ethernet to synchronize fault and load information. Furthermore, the intelligent monitoring module obtains access pressure information from the front-end interface module.
[0017] like Figure 2 As shown, an embodiment of the present application provides a storage system fault handling method, comprising the following steps: Step S1, detecting load information of a storage system, and determining a load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode; Step S2: In response to the storage system being in light load mode, obtaining a light load fault level classified in the light load mode, monitoring the current fault level of the storage system in the light load fault level in real time, obtaining a first utilization rate at the current fault level when load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level; Step S3, in response to the first utilization being greater than the first ratio value and the current fault level being greater than the preset light load fault level, prompting whether to switch the storage system to a heavy load mode; Step S4, in response to the storage system being in heavy load mode, obtain the heavy load fault level divided in the heavy load mode, monitor the current fault level of the storage system in the heavy load fault level in real time, obtain the second utilization at the current fault level when the load information is less than the second threshold, and select the preset heavy load fault handling strategy to handle the load fault according to the second utilization and the current fault level.
[0018] This embodiment implements differentiated fault handling strategies to handle load faults based on the severity of system load information and the current fault level. This prevents system crashes caused by blindly executing recovery operations under high load, while accelerating the recovery process under low load, thereby comprehensively improving the fault tolerance and stability of the storage system. Under light load conditions, it can quickly respond to faults and perform recovery. Under heavy load conditions, it can flexibly adjust the handling method based on the severity of the fault, avoiding excessive impact of recovery operations on critical services. This effectively improves the fault tolerance and stability of the storage system and reduces service interruption time caused by faults.
[0019] In this embodiment, load information of the storage system is detected, and a load mode of the storage system is determined according to the load information, where the load mode includes a light load mode and a heavy load mode. Monitor the storage nodes, disk arrays, and cache modules in the storage system in real time, obtain real-time load information of the storage nodes, disk arrays, and cache modules respectively, and determine the load information of the storage system based on the real-time load information; In response to the load information of the storage system being less than a preset load threshold, determining that the storage system is in a light load mode; In response to the load information of the storage system being greater than or equal to a preset load threshold, it is determined that the storage system is in a heavy load mode.
[0020] In this embodiment, when the load information is less than the first threshold, obtaining a first utilization ratio at the current fault level, and selecting a preset light load fault handling strategy to handle the load fault according to the first utilization ratio and the current fault level include: When the load information is less than a first threshold, obtaining system performance data and maximum load data at a current fault level, and determining a first utilization rate at the current fault level; In response to the first utilization being less than or equal to the first proportional value and the current fault level being less than or equal to the preset light load fault level, performing a fault recovery operation; In response to the first utilization being greater than the first proportional value, or the current fault level being greater than a preset light load fault level, issuing a first-level fault tolerance warning and performing a fault recovery operation; In response to the first utilization being greater than the first proportional value and the current fault level being greater than a preset light load fault level, a first-level fault tolerance alarm is issued and no fault recovery operation is performed.
[0021] In this embodiment, when the load information is less than the first threshold, obtaining system performance data and maximum load data at the current fault level, and determining a first utilization rate at the current fault level includes: When the load information is less than a first threshold, a fault tolerance test is performed and system performance data in the test result is obtained, wherein the system performance data includes processor utilization, memory occupancy, disk input and output throughput, and response time; Get the maximum load data under the current fault level; The first utilization is obtained by dividing the system performance data by the maximum load data.
[0022] In this embodiment, in response to the first utilization being greater than the first ratio and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to the heavy load mode further includes: In response to not receiving a user response to the prompt within a preset time, the storage system is switched to heavy load mode; and / or, in response to the storage system running in heavy load mode for a period of time and its load information continues to be lower than a preset load threshold, the storage system is switched to light load mode.
[0023] Among them, the automatic logic of mode switching has been added, including the mechanism of automatic switching upon timeout and automatic switching from heavy load mode to light load mode.
[0024] In this embodiment, when the load information is less than the second threshold, obtaining a second utilization ratio at the current fault level, and selecting a preset heavy load fault handling strategy to handle the load fault according to the second utilization ratio and the current fault level include: When the load information is less than a second threshold, obtaining system performance data and maximum load data at a current fault level, and determining a second utilization rate at the current fault level; In response to the second utilization being less than or equal to the second proportional value, and the current fault level being less than or equal to the preset light load fault level, performing a fault recovery operation; In response to the second utilization being greater than the second proportional value, or the current fault level being greater than the first preset heavy load fault level, issuing a second-level fault tolerance warning and performing a fault recovery operation; In response to the second utilization being greater than the second proportional value and the current fault level being greater than a preset light load fault level, issuing a second-level fault tolerance alarm and performing a fault recovery operation; In response to the current fault level being greater than a second preset heavy load fault level, the front-end non-critical service is suspended and a fault recovery operation is performed.
[0025] After adjusting the recovery speed, the system continuously monitors the storage system's load, utilization, and recovery progress. Based on these results, the system dynamically adjusts the recovery speed to ensure the fastest possible recovery without impacting system stability or critical business operations. A dynamic feedback and readjustment mechanism for recovery speed has been added.
[0026] When the load information is less than the second threshold, obtaining system performance data and maximum load data at the current fault level, and determining a second utilization rate at the current fault level includes: When the load information is less than a second threshold, a fault tolerance test is performed and system performance data in the test result is obtained, wherein the system performance data includes processor utilization, memory occupancy, disk input and output throughput, and response time; Get the maximum load data under the current fault level; The second utilization rate is obtained by dividing the system performance data by the maximum load data.
[0027] In this embodiment, when the load information is less than the second threshold, obtaining a second utilization ratio at the current fault level, and selecting a preset heavy load fault handling strategy to handle the load fault according to the second utilization ratio and the current fault level further includes: In response to the current fault level being greater than a second preset heavy load fault level, detecting whether a fault recovery operation has been performed; In response to the completion of the fault recovery operation, obtaining a third utilization rate at the current fault level; In response to the third utilization being less than or equal to the third ratio value, the front-end non-critical service is restored.
[0028] In response to the completion of the fault recovery operation, obtaining the third utilization rate at the current fault level includes: After the failure recovery operation is completed, a fault tolerance test is performed and system performance data is obtained from the test results, where the system performance data includes processor utilization, memory usage, disk input and output throughput, and response time; After the failure recovery operation is completed, the current fault level of the storage system in the heavy load fault level is re-monitored to obtain the maximum load data under the current fault level; The third utilization is obtained by dividing the system performance data by the maximum load data.
[0029] In this embodiment, the storage system fault handling method further includes: In response to a storage system fault recovery operation, the fault recovery execution speed is determined based on load information, current fault level, and system utilization at the current fault level, and the fault recovery execution speed in heavy load mode is controlled to be lower than that in light load mode.
[0030] Among them, fault recovery operations include but are not limited to: data reconstruction (such as RAID reconstruction), hot spare disk switching, data migration, snapshot recovery, component restart, system degradation, load balancing adjustment or fault isolation.
[0031] In this embodiment, the storage system fault handling method further includes: Real-time monitoring of the operating status and performance indicators of storage nodes, disk arrays, and cache modules in the storage system; Based on operating status and performance indicators combined with historical data, predict whether there are potential failures in the storage system; In response to predicting a potential failure of the storage system, the severity of the potential failure is obtained, and based on the load information, the severity of the potential failure and / or the current system utilization, a preset preventive maintenance strategy is selected and executed before the failure occurs, where the preventive maintenance strategy includes but is not limited to: data pre-migration, increasing redundancy, adjusting load information, replacing components or issuing preventive warnings.
[0032] Among them, integrating predictive maintenance into the existing storage system enables the system to change from passive response to active prevention, further improving fault tolerance.
[0033] In this embodiment, when a level one fault tolerance warning, a level one fault tolerance alarm, a level two fault tolerance warning, or a level two fault tolerance alarm is issued, the following is also included: Notifications are sent to system administrators or relevant operations and maintenance personnel via email, SMS, system management interface, API interface, or log records, along with fault details, recommended solutions, and / or current system status reports.
[0034] Among them, specific notification methods and additional information for early warnings and alarms have been added to improve practicality.
[0035] Specifically, in light load mode, the ISM categorizes system fault levels into P levels, with Level 1 being the least severe and Level P being the most severe. During time period S1, the ISM monitors the system fault level in real time and performs a fault tolerance test (FT) when the system load falls below a percentage E1. The maximum load data for each fault level is represented by Y. If the system performance data X in the FT result does not exceed Y × F1 and the fault level is not higher than P1, the system automatically performs fault recovery. If the system performance data X in the FT result exceeds Y × F1, or the fault level is higher than P1, a Level 1 fault tolerance warning is issued. If the system performance data X in the FT result exceeds Y × F1 and the fault level is higher than P1, a Level 1 fault tolerance warning is issued and a prompt is given to switch to heavy load mode. Automatic fault recovery is not performed at this time.
[0036] In heavy load mode, the ISM categorizes system fault levels into Q levels, with Level 1 being the least severe and Level Q being the most severe. During time period S2, the ISM monitors the system fault level in real time and performs a fault recovery (FT) when the system load falls below a percentage of E2. If the system performance data X in the FT results does not exceed Y × F2 and the fault level is not higher than Q1, the system automatically initiates fault recovery. If the system performance data X in the FT results exceeds Y × F2, or the fault level is higher than Q1, a Level 2 fault tolerance warning is issued. If the system performance data X in the FT results exceeds Y × F2 and the fault level is higher than Q1, a Level 2 fault tolerance alarm is issued. At this point, if the fault level is not higher than Q2, the system automatically initiates fault recovery. If the fault level is higher than Q2, the system suspends non-critical front-end services and prioritizes fault recovery. After recovery is complete, the FT is performed again during time period S3. If the system performance data X does not exceed Y × F3, non-critical front-end services are restored.
[0037] Among them, Y, E1, E2, F1, F2, F3, P, P1, Q, Q1, Q2, S1, S2, and S3 are system preset parameters and can be adjusted through the system management interface or command line tools. E1 is the first threshold, preferably E1 = 40%, X / Y is the first utilization or the second utilization, F1 is the first ratio value, preferably F1 = 70%, E2 is the second threshold, preferably E2 = 20%, F2 is the second ratio value, preferably F2 = 50%, F3 is the third ratio value, preferably F3 = 60%, P is the light load fault level, preferably P = 6, and Q is the heavy load fault level, preferably Q = 10.
[0038] For example, in light load mode, the ISM categorizes system fault levels into six levels, with Level 1 being the least severe and Level 6 being the most severe. Within a 60-second period, the ISM monitors the system fault level in real time and executes a fault recovery (FT) when the system load falls below 40%. If the system performance data X in the FT results does not exceed Y × 70% and the fault level is no higher than Level 3, the system automatically performs fault recovery. If the system performance data X in the FT results exceeds Y × 70%, or the fault level is higher than Level 3, a Level 1 fault tolerance warning is issued. If the system performance data X in the FT results exceeds Y × 70%, and the fault level is higher than Level 3, a Level 1 fault tolerance warning is issued, prompting whether to switch to heavy load mode. The system does not perform automatic fault recovery at this time.
[0039] In heavy load mode, the ISM categorizes system fault severity into 10 levels, with Level 1 being the least severe and Level 10 being the most severe. Within a 30-second period, the ISM monitors the system fault severity in real time and performs a failover (FT) when the system load falls below 20%. If the system performance data (X) in the FT results does not exceed Y × 50% and the fault severity is not higher than Level 5, the system automatically initiates fault recovery. If the system performance data (X) in the FT results exceeds Y × 50% or the fault severity is higher than Level 5, a Level 2 fault tolerance warning is issued. If the system performance data (X) in the FT results exceeds Y × 50% and the fault severity is higher than Level 5, a Level 2 fault tolerance alarm is issued. At this point, if the fault severity is not higher than Level 7, the system automatically initiates fault recovery. If the fault severity is higher than Level 7, the system suspends non-critical front-end services and prioritizes failover. After recovery is complete, the FT is performed again within a 45-second period. If X does not exceed Y × 60%, non-critical front-end services are restored.
[0040] In the above-mentioned storage system fault handling method, the load mode of the storage system is determined to be light load mode or heavy load mode based on its load information. A differentiated fault handling strategy is implemented to handle load faults in combination with the severity of the system load information and the current fault level. This avoids system crashes caused by blindly executing recovery operations under high load, while accelerating the recovery process under low load, thereby comprehensively improving the fault tolerance and stability of the storage system. Under light load information, the system can quickly respond to faults and perform recovery. Under heavy load information, the handling method can be flexibly adjusted according to the severity of the fault, avoiding excessive impact of recovery operations on key services. This effectively improves the fault tolerance and stability of the storage system and reduces the service interruption time caused by faults.
[0041] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0042] In one embodiment, Figure 3 As shown, a storage system fault processing device 10 is provided, including: a load information detection module 1, a light load fault processing module 2, a mode switching module 3, a heavy load fault processing module 4, a fault recovery speed control module 5 and a predictive maintenance module 6.
[0043] The load information detection module 1 is used to detect the load information of the storage system and determine the load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode.
[0044] The light load fault processing module 2 is used to respond to the storage system being in light load mode, obtain the light load fault level divided in the light load mode, monitor the current fault level of the storage system in the light load fault level in real time, obtain the first utilization rate at the current fault level when the load information is less than the first threshold, and select a preset light load fault processing strategy to handle the load fault according to the first utilization rate and the current fault level.
[0045] The mode switching module 3 is configured to prompt whether to switch the storage system to the heavy load mode in response to the first utilization being greater than the first proportional value and the current fault level being greater than the preset light load fault level.
[0046] The heavy load fault processing module 4 is used to respond to the storage system being in heavy load mode, obtain the heavy load fault level divided in the heavy load mode, monitor the current fault level of the storage system in the heavy load fault level in real time, obtain the second utilization rate at the current fault level when the load information is less than the second threshold, and select a preset heavy load fault processing strategy to handle the load fault according to the second utilization rate and the current fault level.
[0047] In this embodiment, load information of the storage system is detected, and a load mode of the storage system is determined according to the load information, where the load mode includes a light load mode and a heavy load mode. Monitor the storage nodes, disk arrays, and cache modules in the storage system in real time, obtain real-time load information of the storage nodes, disk arrays, and cache modules respectively, and determine the load information of the storage system based on the real-time load information; In response to the load information of the storage system being less than a preset load threshold, determining that the storage system is in a light load mode; In response to the load information of the storage system being greater than or equal to a preset load threshold, it is determined that the storage system is in a heavy load mode.
[0048] In this embodiment, when the load information is less than the first threshold, obtaining a first utilization ratio at the current fault level, and selecting a preset light load fault handling strategy to handle the load fault according to the first utilization ratio and the current fault level include: When the load information is less than a first threshold, obtaining system performance data and maximum load data at a current fault level, and determining a first utilization rate at the current fault level; In response to the first utilization being less than or equal to the first proportional value and the current fault level being less than or equal to the preset light load fault level, performing a fault recovery operation; In response to the first utilization being greater than the first proportional value, or the current fault level being greater than a preset light load fault level, issuing a first-level fault tolerance warning and performing a fault recovery operation; In response to the first utilization being greater than the first proportional value and the current fault level being greater than a preset light load fault level, a first-level fault tolerance alarm is issued and no fault recovery operation is performed.
[0049] In this embodiment, when the load information is less than the first threshold, obtaining system performance data and maximum load data at the current fault level, and determining a first utilization rate at the current fault level includes: When the load information is less than a first threshold, a fault tolerance test is performed and system performance data in the test result is obtained, wherein the system performance data includes processor utilization, memory occupancy, disk input and output throughput, and response time; Get the maximum load data under the current fault level; The first utilization is obtained by dividing the system performance data by the maximum load data.
[0050] In this embodiment, in response to the first utilization being greater than the first ratio and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to the heavy load mode further includes: In response to not receiving a user response to the prompt within a preset time, the storage system is switched to heavy load mode; and / or, in response to the storage system running in heavy load mode for a period of time and its load information continues to be lower than a preset load threshold, the storage system is switched to light load mode.
[0051] In this embodiment, when the load information is less than the second threshold, obtaining a second utilization ratio at the current fault level, and selecting a preset heavy load fault handling strategy to handle the load fault according to the second utilization ratio and the current fault level include: When the load information is less than a second threshold, obtaining system performance data and maximum load data at a current fault level, and determining a second utilization rate at the current fault level; In response to the second utilization being less than or equal to the second proportional value, and the current fault level being less than or equal to the preset light load fault level, performing a fault recovery operation; In response to the second utilization being greater than the second proportional value, or the current fault level being greater than the first preset heavy load fault level, issuing a second-level fault tolerance warning and performing a fault recovery operation; In response to the second utilization being greater than the second proportional value and the current fault level being greater than a preset light load fault level, issuing a second-level fault tolerance alarm and performing a fault recovery operation; In response to the current fault level being greater than a second preset heavy load fault level, the front-end non-critical service is suspended and a fault recovery operation is performed.
[0052] In this embodiment, when the load information is less than the second threshold, obtaining a second utilization ratio at the current fault level, and selecting a preset heavy load fault handling strategy to handle the load fault according to the second utilization ratio and the current fault level further includes: In response to the current fault level being greater than a second preset heavy load fault level, detecting whether a fault recovery operation has been performed; In response to the completion of the fault recovery operation, obtaining a third utilization rate at the current fault level; In response to the third utilization being less than or equal to the third ratio value, the front-end non-critical service is restored.
[0053] In this embodiment, the fault recovery speed control module 5 is used to: In response to a storage system fault recovery operation, the fault recovery execution speed is determined based on load information, current fault level, and system utilization at the current fault level, and the fault recovery execution speed in heavy load mode is controlled to be lower than that in light load mode.
[0054] In this embodiment, the predictive maintenance module 6 is used to: Real-time monitoring of the operating status and performance indicators of storage nodes, disk arrays, and cache modules in the storage system; Based on operating status and performance indicators combined with historical data, predict whether there are potential failures in the storage system; In response to predicting a potential failure of the storage system, the severity of the potential failure is obtained, and based on the load information, the severity of the potential failure and / or the current system utilization, a preset preventive maintenance strategy is selected and executed before the failure occurs, where the preventive maintenance strategy includes but is not limited to: data pre-migration, increasing redundancy, adjusting load information, replacing components or issuing preventive warnings.
[0055] In the above-mentioned storage system fault handling device, the load mode of the storage system is determined to be a light load mode or a heavy load mode based on the load information of the storage system, and differentiated fault handling strategies are implemented to handle load failures in combination with the severity of the system load information and the current fault level, thereby avoiding system crashes caused by blindly executing recovery operations under high loads, and accelerating the recovery process under low loads, thereby comprehensively improving the fault tolerance and stability of the storage system.
[0056] For descriptions of features in the embodiments corresponding to the storage system fault handling apparatus, reference can be made to the descriptions of the embodiments corresponding to the storage system fault handling method, which will not be detailed here.
[0057] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned storage system fault handling method embodiments.
[0058] In one embodiment, the electronic device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The electronic device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store storage system fault processing data. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a storage system fault processing method is implemented.
[0059] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned storage system fault handling method embodiments when executed: Detecting load information of the storage system and determining a load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode; In response to the storage system being in a light load mode, obtaining a light load fault level classified in the light load mode, monitoring in real time a current fault level of the storage system in the light load fault level, obtaining a first utilization rate at the current fault level when load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level; In response to the first utilization being greater than the first ratio value and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to a heavy load mode; In response to the storage system being in heavy load mode, a heavy load fault level divided in the heavy load mode is obtained, the current fault level of the storage system in the heavy load fault level is monitored in real time, a second utilization rate at the current fault level is obtained when the load information is less than a second threshold, and a preset heavy load fault handling strategy is selected to handle the load fault based on the second utilization rate and the current fault level.
[0060] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0061] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned storage system fault handling method embodiments are implemented: Detecting load information of the storage system and determining a load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode; In response to the storage system being in a light load mode, obtaining a light load fault level classified in the light load mode, monitoring in real time a current fault level of the storage system in the light load fault level, obtaining a first utilization rate at the current fault level when load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level; In response to the first utilization being greater than the first ratio value and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to a heavy load mode; In response to the storage system being in heavy load mode, a heavy load fault level divided in the heavy load mode is obtained, the current fault level of the storage system in the heavy load fault level is monitored in real time, a second utilization rate at the current fault level is obtained when the load information is less than a second threshold, and a preset heavy load fault handling strategy is selected to handle the load fault based on the second utilization rate and the current fault level.
[0062] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned storage system fault handling method embodiments are implemented: Detecting load information of the storage system and determining a load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode; In response to the storage system being in a light load mode, obtaining a light load fault level classified in the light load mode, monitoring in real time a current fault level of the storage system in the light load fault level, obtaining a first utilization rate at the current fault level when load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level; In response to the first utilization being greater than the first ratio value and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to a heavy load mode; In response to the storage system being in heavy load mode, a heavy load fault level divided in the heavy load mode is obtained, the current fault level of the storage system in the heavy load fault level is monitored in real time, a second utilization rate at the current fault level is obtained when the load information is less than a second threshold, and a preset heavy load fault handling strategy is selected to handle the load fault based on the second utilization rate and the current fault level.
[0063] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0064] The above is a detailed introduction to a storage system fault handling method and electronic device provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A storage system fault handling method, characterized in that: include: Detecting load information of a storage system, and determining a load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode; In response to the storage system being in a light load mode, obtaining a light load fault level classified under the light load mode, monitoring in real time a current fault level of the storage system within the light load fault level, obtaining a first utilization rate at the current fault level when the load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault based on the first utilization rate and the current fault level; In response to the first utilization being greater than a first proportional value and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to a heavy load mode; In response to the storage system being in heavy load mode, a heavy load fault level divided under the heavy load mode is obtained, the current fault level of the storage system in the heavy load fault level is monitored in real time, a second utilization rate at the current fault level is obtained when the load information is less than a second threshold, and a preset heavy load fault handling strategy is selected according to the second utilization rate and the current fault level to handle the load fault.
2. The storage system fault handling method according to claim 1, wherein: The detecting load information of the storage system and determining the load mode of the storage system according to the load information, wherein the load mode includes a light load mode and a heavy load mode, comprises: Performing real-time monitoring on storage nodes, disk arrays, and cache modules in the storage system, obtaining real-time load information of the storage nodes, the disk arrays, and the cache modules, respectively, and determining load information of the storage system based on the real-time load information; In response to the load information of the storage system being less than a preset load threshold, determining that the storage system is in a light load mode; In response to the load information of the storage system being greater than or equal to a preset load threshold, it is determined that the storage system is in a heavy load mode.
3. The storage system fault handling method according to claim 1, characterized in that: The acquiring of a first utilization ratio at the current fault level when the load information is less than a first threshold, and selecting a preset light load fault handling strategy to handle the load fault according to the first utilization ratio and the current fault level includes: When the load information is less than a first threshold, obtaining system performance data and maximum load data at the current fault level, and determining a first utilization rate at the current fault level; In response to the first utilization being less than or equal to a first proportional value, and the current fault level being less than or equal to a preset light load fault level, performing a fault recovery operation; In response to the first utilization being greater than a first proportional value, or the current fault level being greater than a preset light load fault level, issuing a first-level fault tolerance warning and performing a fault recovery operation; In response to the first utilization being greater than a first proportional value and the current fault level being greater than a preset light load fault level, a first-level fault tolerance alarm is issued, and no fault recovery operation is performed.
4. The storage system fault handling method according to claim 3, characterized in that: When the load information is less than a first threshold, acquiring system performance data and maximum load data at the current fault level, and determining a first utilization rate at the current fault level includes: When the load information is less than a first threshold, a fault tolerance test is performed and system performance data in the test result is obtained, wherein the system performance data includes processor utilization, memory occupancy, disk input and output throughput, and response time; Obtaining maximum load data under the current fault level; A first utilization rate is obtained by dividing the system performance data by the maximum load data.
5. The storage system fault handling method according to claim 1, wherein: In response to the first utilization being greater than a first proportional value and the current fault level being greater than a preset light load fault level, prompting whether to switch the storage system to a heavy load mode further includes: In response to not receiving a response to the prompt from the user within a preset time, switching the storage system to a heavy load mode; and / or In response to the load information of the storage system being continuously lower than a preset load threshold after the storage system has been running in the heavy load mode for a period of time, the storage system is switched to the light load mode.
6. The storage system fault handling method according to claim 1, wherein: The acquiring of a second utilization rate at the current fault level when the load information is less than a second threshold, and selecting a preset heavy load fault handling strategy to handle the load fault according to the second utilization rate and the current fault level includes: When the load information is less than a second threshold, obtaining system performance data and maximum load data at the current fault level, and determining a second utilization rate at the current fault level; In response to the second utilization being less than or equal to a second proportional value, and the current fault level being less than or equal to a preset light load fault level, performing a fault recovery operation; In response to the second utilization being greater than a second proportional value, or the current fault level being greater than a first preset heavy load fault level, issuing a second-level fault tolerance warning and performing a fault recovery operation; In response to the second utilization being greater than a second proportional value and the current fault level being greater than a preset light load fault level, issuing a level 2 fault tolerance alarm and performing a fault recovery operation; In response to the current fault level being greater than a second preset heavy load fault level, the front-end non-critical service is suspended and a fault recovery operation is performed.
7. The storage system fault handling method according to claim 6, characterized in that: The step of acquiring a second utilization rate at the current fault level when the load information is less than a second threshold, and selecting a preset heavy load fault handling strategy to handle the load fault according to the second utilization rate and the current fault level further includes: In response to the current fault level being greater than a second preset heavy load fault level, detecting whether a fault recovery operation is performed; In response to the completion of the fault recovery operation, obtaining a third utilization rate at the current fault level; In response to the third utilization being less than or equal to a third ratio value, the front-end non-critical service is restored.
8. The storage system fault handling method according to claim 7, characterized in that: The storage system fault handling method further includes: In response to the storage system performing a fault recovery operation, the fault recovery execution speed is determined based on the load information, the current fault level, and the system utilization under the current fault level, and the fault recovery execution speed in the heavy load mode is controlled to be lower than the fault recovery execution speed in the light load mode.
9. The storage system fault handling method according to claim 1, wherein: The storage system fault handling method further includes: Real-time monitoring of the operating status and performance indicators of storage nodes, disk arrays, and cache modules in the storage system; Based on the operating status and performance indicators and combined with historical data, predict whether the storage system has a potential failure; In response to predicting a potential failure of the storage system, the severity of the potential failure is obtained, and based on the load information, the severity of the potential failure and / or the current system utilization, a preset preventive maintenance strategy is selected and executed before the failure occurs, wherein the preventive maintenance strategy includes but is not limited to: data pre-migration, increasing redundancy, adjusting load information, replacing components or issuing preventive warnings.
10. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the storage system fault handling method according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Fault processing method and device, computer equipment, readable storage medium and program product
CN119201536A
Disk scheduling method
CN120428922A
Disk array device and failure handling method in disk array device
JP2020119233A
Intelligent stress testing and raid rebuild to prevent data loss
US20170147437A1
Storage system, load rebalancing method thereof and access control method thereof
US20190026039A1
Cited By
Redundant array of independent disks (RAID) reconstruction method and electronic equipment
CN120892264A