Storage system controller fault recovery method and electronic equipment
By maintaining the operational status of services and monitoring the initialization progress when the controller in the storage system recovers from a failure, the problem of long service interruption time is solved, and high availability and system stability are achieved.
Patent Information
- Application Number
- CN202511204157.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, the service interruption time is long when the storage system controller fails and recovers, which cannot meet the requirements of high availability.
In the event that a controller to be recovered exists in the target cluster, other controllers are controlled to maintain the business operation status, and the initialization progress of the controller to be recovered is monitored in real time. Initialization is performed in stages until the business process is stopped and the status is updated after completion.
It effectively shortened business downtime, improved the high availability of the storage system, and ensured the continuity and stability of the system.
Smart Images

Figure CN121029501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cluster fault recovery technology, and in particular to a storage system controller fault recovery method and electronic device. Background Technology
[0002] High-end storage systems, as a core component of enterprise-level applications, are widely used in data centers, cloud computing platforms, and large enterprise information systems. Centralized storage systems are widely adopted due to their simple management and superior performance. In current centralized storage systems, controllers typically employ a redundant architecture, i.e., a dual-controller or multi-controller architecture, to improve the system's fault tolerance. This also places extremely high demands on the reliability and fault recovery capabilities of the controllers.
[0003] In related technologies, when a controller fails and recovers before joining the cluster, the cluster module notifies the business modules to suspend current services and wait for the new controller to complete initialization before gradually restoring services. However, this method results in long service interruptions, failing to meet current high availability requirements for storage systems and urgently needing a solution. Summary of the Invention
[0004] This invention provides a storage system controller fault recovery method and electronic device to at least solve the problems of long service interruption time and failure to meet the current high availability requirements of storage systems in the prior art, thereby improving the high availability of storage systems.
[0005] This invention provides a method for recovering from a storage system controller failure, comprising the following steps: Determine if the target cluster has a controller that needs to be recovered; If the controller to be recovered exists in the target cluster, control other controllers in the target cluster to maintain the corresponding service operation state, initialize the controller to be recovered, and monitor whether the controller to be recovered has completed the initialization action; When the controller to be recovered completes its initialization action, the target business process of the business module in the target cluster is stopped, and the target information of the controller to be recovered is distributed to the business module, so that the business module performs a status update based on the target information to restore the controller to be recovered to the target cluster.
[0006] The present invention provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described storage system controller fault recovery methods.
[0007] This invention addresses the issue of long service interruptions and failure to meet high availability requirements in existing storage systems when a controller to be recovered exists within the target cluster. It controls other controllers within the cluster to maintain their corresponding service operation states, initializes the controller to be recovered, and monitors whether the controller has completed its initialization process. Upon successful initialization, the invention stops the target service processes of the business modules within the target cluster and distributes the target information of the controller to the business modules. This allows the business modules to update their states based on the target information, thereby restoring the controller to the target cluster. Thus, by maintaining service operation while simultaneously initializing newly added controllers, the invention solves the problems of long service interruption times and failure to meet current high availability requirements for storage systems in existing technologies, thereby improving the high availability of the storage system. Attached Figure Description
[0008] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A flowchart of a storage system controller fault recovery method provided in an embodiment of the present invention; Figure 2 A flowchart of another storage system controller fault recovery method provided in an embodiment of the present invention; Figure 3 A block diagram of a storage system controller fault recovery device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0011] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0012] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0013] The present invention provides a storage system controller fault recovery method, and the execution flow of the storage system controller fault recovery method is described in detail below.
[0014] Figure 1 This is a flowchart of a storage system controller fault recovery method according to an embodiment of the present invention.
[0015] For example, such as Figure 1 As shown, the storage system controller fault recovery method includes the following steps: In step S101, it is determined whether there is a controller to be recovered in the target cluster.
[0016] Understandably, a multi-controller redundant architecture storage system (i.e., a centralized high-end storage system) is a high-availability design aimed at improving the system's fault tolerance and data processing capabilities through the collaborative work of multiple controllers. The controller is the hardware component of the storage system, typically containing a CPU (Central Processing Unit), memory, cache, network interfaces, etc., used to perform actual data read / write operations and manage hardware resources. In a multi-controller redundant architecture storage system, under normal operating conditions, the system can distribute business requests evenly across controllers through business modules to fully utilize the processing power of each controller to perform actual data read / write operations, thereby improving overall performance. Data is redundantly stored across multiple controllers, typically using technologies such as mirroring or RAID (Redundant Array of Independent Disks) to ensure data availability even if a controller fails. Furthermore, the controllers maintain state synchronization, ensuring that each controller can take over the business of other controllers when needed. When one or more controllers fail, business modules can stop sending business requests to the failed controller and switch business requests to other healthy controllers, isolating the failed controller and preventing the failure from spreading. Once the faulty controller recovers, it will be restored to the cluster.
[0017] In this embodiment of the invention, a controller to be recovered refers to a controller that has recovered from a failure and needs to be rejoined to the cluster after recovery. Determining whether a controller to be recovered exists allows for timely triggering of subsequent fault recovery procedures. If a controller to be recovered is determined to exist, the process proceeds to step S102.
[0018] In step S102, if there is a controller to be recovered in the target cluster, control other controllers in the target cluster to maintain the corresponding business operation state, initialize the controller to be recovered, and monitor whether the controller to be recovered has completed the initialization action.
[0019] Specifically, in related technologies, when a new controller (i.e., the controller to be restored) is added, the original controller's services are usually paused, and services are only restored after the new controller completes initialization, resulting in a long service interruption time. The embodiments of the present invention can detect that a controller to be restored is to be added to the target cluster, without immediately interrupting the service operations of the currently running original controller (i.e., other controllers in the target cluster). Instead, while ensuring the continued stable operation of the original controller's services, the addition and initialization process of the controller to be restored can be carried out simultaneously, including loading configuration information, clearing old data, synchronizing status, etc., and the initialization process of the controller to be restored is monitored in real time to determine whether the controller to be restored has completed the initialization action.
[0020] This parallel processing approach effectively avoids interruptions to the original controller's services caused by the addition of the controller to be restored, thereby minimizing the potential service interruption time during controller recovery and ensuring the continuity and efficiency of the entire system.
[0021] Next, we will explain in detail how to monitor whether the controller to be restored has completed the initialization process.
[0022] As one possible implementation method, in some embodiments, monitoring whether the controller to be restored has completed the initialization action includes: adding a status flag bit corresponding to the controller to be restored to the state machine; using the status flag bit to monitor the status of the controller to be restored, and determining whether the controller to be restored has completed the initialization action based on the monitoring results.
[0023] Specifically, in storage systems with multi-controller redundancy architectures, due to the complexity of multiple controller failures and recovery scenarios, including but not limited to simultaneous failures of multiple controllers, cascading effects of controller failures, and data consistency issues during fault recovery, this embodiment of the invention can employ bitwise operations to add state flag bits corresponding to each controller to be recovered to the state machine. These flag bits are then used to monitor the controller's state, allowing the system to understand the controller's initialization progress and status in real time. Based on the monitoring results, the system can determine whether the controller has completed its initialization actions and take appropriate measures.
[0024] For example, two flags can be used to indicate the status of the controller to be recovered: Flag 1 indicates whether the controller is in the process of initialization (0 indicates uninitialized or initialization completed, 1 indicates initialization in progress); Flag 2 indicates whether the controller is faulty (0 indicates normal, 1 indicates fault). When a controller to be recovered joins the target cluster and initializes, status flag 1 is set to 1, indicating "initializing," until the controller's initialization is fully completed, at which point flag 1 is cleared to 0. If a controller fails during initialization, status flag 2 can be set to 1; otherwise, it is set to 0. In the runtime code, bitwise OR operations can be used to set specific flags, for example, setting the controller to the "initializing" state; bitwise AND operations can be used to check if a specific flag is 1, for example, checking if the controller is in the "initializing" state; and bitwise AND and bitwise NOT operations can be used to clear specific flags, for example, clearing the controller's "initializing" state.
[0025] Therefore, by setting, checking, and clearing specific flag bits through bit operations, the system can monitor the status of controllers to be restored in real time and take corresponding measures based on these statuses, thereby achieving efficient management of the status of multiple controllers to be restored.
[0026] The following details how to determine whether the controller to be restored has completed the initialization process based on the monitoring results.
[0027] Optionally, in some embodiments, determining whether the controller to be restored has completed the initialization action based on the monitoring results includes: determining whether the controller to be restored has finished initialization based on the monitoring results, and whether the controller to be restored is in a preset fault state; if the controller to be restored has finished initialization, and the controller to be restored is not in a preset fault state, then it is determined that the controller to be restored has completed the initialization action, and the controller to be restored is marked as a first preset state.
[0028] Specifically, based on the monitoring results, if the status flag 1 of a controller to be restored is 0 and the status flag 2 is 0, it means that the controller to be restored has completed initialization and is not in the preset fault state (i.e., no fault). That is, it can be determined that the controller to be restored has completed the initialization action, and the controller to be restored that has completed the initialization action is marked as the first preset state, indicating that the controller to be restored has completed the initialization action normally, so as to prompt the business module to perform further operations on the controller to be restored.
[0029] Therefore, by using status flags to monitor the status of the controller to be restored, the system can more flexibly respond to and handle the fault recovery and initialization process of multiple controllers.
[0030] Optionally, in other embodiments, determining whether the controller to be restored has completed the initialization action based on the monitoring results includes: determining whether the controller to be restored is in a preset fault state based on the monitoring results; if the controller to be restored is in a preset fault state, determining that the controller to be restored has not completed the initialization action, marking the controller to be restored as a second preset state, and sending a fault notification to the service module.
[0031] Specifically, based on the monitoring results, if the status flag 2 of a controller to be recovered is 1, regardless of the value of status flag 1, it indicates that the controller to be recovered is in a preset fault state. This further determines that the controller to be recovered has not completed the initialization process, and the controller that has not completed initialization is marked as being in the second preset state, indicating that it is a controller that has not completed initialization. Simultaneously, the target cluster can generate a fault notification event and send it to all business modules to prompt them to take appropriate action regarding the controller to be recovered, such as clearing the status flag of the faulty controller and controlling its exit from the target cluster.
[0032] Therefore, the above process can effectively handle controllers that fail again during initialization, thereby ensuring system stability and business continuity.
[0033] The fault notification includes at least one of the following: identification information of the controller to be restored that is in a preset fault state and the cause of the fault.
[0034] Understandably, a fault notification is an event or message generated when a fault is detected in a controller awaiting recovery. It is used to inform business modules, cluster management modules, etc., that a controller awaiting recovery has failed again. The information it contains is crucial for quickly locating and handling faulty controllers awaiting recovery.
[0035] Specifically, a fault notification typically includes the following information: the identification information of the controller to be recovered, which is used to uniquely identify the controller that has failed and can include the controller ID (a unique number or string), IP (Internet Protocol) address, network port number, etc.; the cause of the failure of the controller to be recovered, which can include hardware failure (such as disk failure, memory failure, network interface failure, etc.) and software failure (such as software crash, configuration error, communication timeout, etc.).
[0036] Therefore, by including this key information, fault notifications can help the system respond quickly to faults, ensuring business continuity and system stability.
[0037] In step S103, after the controller to be recovered completes the initialization action, the target business process of the business module in the target cluster is stopped, and the target information of the controller to be recovered is distributed to the business module, so that the business module updates its status based on the target information to restore the controller to be recovered to the target cluster.
[0038] In some embodiments, the target information includes at least one of the following: identification information, configuration information, and status information of the controller to be restored.
[0039] Specifically, once the controller to be recovered completes its initialization process, the business module can formally begin the process of adding the controller to the target cluster. During this stage, a state machine process is used to control the cessation of the target business process (i.e., the current business process) of the business module in the target cluster. This ensures that the system state remains consistent before switching to the more controllers mode. Then, the cluster module can distribute the target information of the controller to be recovered to the business module through event notifications or direct communication. This information includes the controller's identification information, configuration information (such as IP address, port number, storage resource allocation, etc.), and status information (such as the data synchronization status and progress between the controller to be recovered and other controllers (existing controllers) in the target cluster). After receiving this target information, the business module can update its internal controller status information (i.e., update the status of the controller to be recovered from "initializing" to "initialization complete") and record the configuration information of the controller to be recovered. After completing this series of operations, the system can be officially switched to a redundant control mode with more controllers (other controllers in the target cluster and the controller to be recovered that has completed initialization). At the same time, it can start the data synchronization process and the status synchronization process to synchronize the data of other controllers in the target cluster to the newly added controller to be recovered, and synchronize the status information of each controller to ensure data consistency and status consistency of all controllers. After completing the data synchronization and status synchronization, the business module resumes the business processing flow and begins to gradually switch business requests to the newly added controller to be recovered, thereby completing the entire process of restoring the controller to be recovered to the target cluster.
[0040] Furthermore, in some embodiments, after initializing the controller to be restored, the method further includes: monitoring the initialization progress of the controller to be restored; and adjusting the business processing strategy of the business module according to the initialization progress.
[0041] In other words, after the controller to be recovered is initialized (i.e., during the initialization process of the controller to be recovered), the system can monitor the initialization progress of the controller to be recovered in real time. The initialization process can be divided into multiple stages, for example: 0% progress: initialization has started; 25% progress: configuration information loading is complete; 50% progress: old data cleanup is complete; 75% progress: state synchronization is complete; 100% progress: initialization is complete. Based on the real-time initialization progress, business modules can dynamically adjust their business processing strategies to ensure that business modules can continue to process business requests during the initialization process of the controller to be recovered, while preparing for the addition of the controller to be recovered.
[0042] As one possible implementation, in some embodiments, the business processing strategy of the business module is adjusted according to the initialization progress, including: maintaining the business operation status of other controllers in the target cluster when the initialization progress meets a first preset threshold; maintaining the business operation status of other controllers in the target cluster when the initialization progress meets a second preset threshold, and preparing for communication between the business module and the controller to be restored; maintaining the business operation status of other controllers in the target cluster when the initialization progress meets a third preset threshold, and adjusting the communication link between the business module and the controller to be restored; maintaining the business operation status of other controllers in the target cluster when the initialization progress meets a fourth preset threshold, and sending preset non-critical business requests to the controller to be restored; and stopping the target business process of the business module when the initialization progress meets a fifth preset threshold.
[0043] Specifically, the following are examples of specific adjustment strategies: Initialization Start (0%): Business modules can continue to maintain the operational status of other controllers (existing controllers) in the target cluster, without sending any business requests to the controller to be recovered. Configuration Information Loading Complete (25%): Business modules can continue to maintain the operational status of other controllers in the target cluster, but begin preparing for communication with the controller to be recovered. For example, business modules can synchronize the configuration information of the controller to be recovered (such as IP address, port number, storage resource allocation, etc.) to their own configuration management module to ensure that the business modules can correctly interact with the controller to be recovered in subsequent communications; business modules update their internal routing tables to ensure that business requests can be correctly sent to the controller to be recovered; business modules begin establishing a communication link with the controller to be recovered, including steps such as initializing network connections, setting communication protocols, and verifying connections. Old Data Cleanup Complete (50%): Business modules continue to maintain the operational status of other controllers in the target cluster, but can begin warming up the communication link with the controller to be recovered. For example, a business module can send lightweight test requests to the controller to be recovered to verify the connectivity and stability of the communication link. The controller to be recovered receives and responds to these test requests, and the business module adjusts the parameters of the communication link based on the response results to ensure the efficiency and reliability of communication. State synchronization complete (75%): The business module continues to maintain the business operation status of other controllers in the target cluster, but can begin to gradually send preset non-critical business requests (non-critical business requests have characteristics such as low priority, retryability, and low latency requirements) to the controller to be recovered. By processing some non-critical business requests, the load on the controller to be recovered is gradually increased, allowing it to smoothly transition to normal operating status. Initialization complete (100%): The business module stops its current business processing flow, completes synchronization with the controller to be recovered, and then resumes its business processing flow.
[0044] Therefore, by monitoring the initialization progress in stages, the system can promptly detect and address potential problems during the initialization process, thereby improving the system's reliability.
[0045] Furthermore, in some embodiments, after stopping the target business process of the business module in the target cluster, the method further includes: controlling the business module to synchronize the status information of the controller to be restored to other controllers in the target cluster; and determining the business processing logic of other controllers in the target cluster based on the status information of the controller to be restored, so that other controllers in the target cluster can perform business processing according to the corresponding business processing logic.
[0046] Specifically, after the controller to be recovered completes initialization and is ready to join the cluster, the cluster module can synchronize the status information of the newly joined controller to other controllers in the target cluster via broadcast or point-to-point communication. Upon receiving the status information of the newly joined controller, other controllers can adjust their own business processing logic based on this information, such as adjusting the allocation of business requests according to load balancing strategies, to ensure high performance and high availability of the system. This ensures that the business processing logic of the entire cluster remains consistent during controller failure recovery, thereby achieving business continuity and data consistency.
[0047] Furthermore, in some embodiments, before stopping the target business process of the business module in the target cluster, the method further includes: sending a preset stage event notification to the business module in the target cluster based on a preset broadcast method.
[0048] Understandably, the default broadcast method is a communication mechanism that sends messages to all nodes in the cluster. This method ensures that all nodes receive the same message, thus maintaining state synchronization. During controller failure recovery, multiple stages may occur, each with specific tasks and states. By sending these notifications to all business modules in the target cluster through the default broadcast method, it can be ensured that all business modules receive the same notification, thereby maintaining state consistency.
[0049] Specifically, such as Figure 2 As shown, after the controller fault recovery process begins, the target cluster can notify each business module through three phases of events. After the target cluster initiates the Phase 1 event (i.e., after the first phase event notification), the business modules do not interrupt their operations. Instead, they keep the other controllers within the target cluster (i.e., the original controllers) operational, while simultaneously identifying the status of the newly added controller to be recovered. Although the controller to be recovered has been added to the system, because it has not yet completed the controller initialization process, the target business processes and related notifications will not be sent to the controller to be recovered. After the target cluster completes the Phase 1 event notification, it means that each business module has received the notification of the controller to be recovered and obtained the relevant controller information. At this point, the target cluster can issue the Phase 2 controller recovery event, which can be considered the main stage of the relevant business processing.
[0050] After the target cluster initiates the Phase 2 event (i.e., after the second phase event notification), the controller to be recovered formally starts the initialization process, clearing its own retained old data. During this process, other controllers within the target cluster maintain normal business operations and identify the status and processing stage of the target controller to be recovered by adding status flags to the state machine. If the Phase 2 event processing is not completed, the controller to be recovered remains unaffected by normal business operations during this stage.
[0051] After the target cluster completes the Phase 2 event processing, the controller to be recovered has completed the relevant initialization and preparation processes for joining the target cluster. After the target cluster initiates the Phase 3 event (i.e., after the third phase event notification), the business module can formally complete the relevant processes for joining the controller to be recovered into the target cluster. In this phase, the business module can first trigger the cessation of its internal business processes via a state machine process; then, it obtains the specific information of the controller to be recovered through the cluster module, modifies the control object state information within its own business module, formally switches to the more controllers redundant mode, and initiates the synchronization process of relevant business processes within the module. After the relevant business processes are completed, the module's business is restored, thus completing the entire controller fault recovery process.
[0052] The storage system controller fault recovery method proposed in this embodiment of the invention, when a controller to be recovered exists in the target cluster, controls other controllers in the target cluster to maintain their corresponding business operation states, initializes the controller to be recovered, and monitors whether the controller to be recovered has completed its initialization action. If the controller to be recovered completes its initialization action, the target business process of the business modules in the target cluster can be stopped, and the target information of the controller to be recovered can be distributed to the business modules. This allows the business modules to update their states based on the target information, thereby restoring the controller to be recovered to the target cluster. Therefore, by maintaining the business operation state while initializing newly added controllers, the method solves the problems of long business interruption times and failure to meet current high availability requirements of storage systems in existing technologies, thus improving the high availability of the storage system.
[0053] Through the above description of the embodiments, those skilled in the art can clearly understand that the system according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0054] Secondly, embodiments of the present invention also provide a storage system controller fault recovery device.
[0055] Figure 3 This is a block diagram of a storage system controller fault recovery device according to an embodiment of the present invention.
[0056] like Figure 3As shown, the storage system controller fault recovery device 10 includes: a judgment module 100, a first processing module 200, and a second processing module 300.
[0057] Among them, the judgment module 100 is used to determine whether there is a controller to be recovered in the target cluster; The first processing module 200 is used to control other controllers in the target cluster to maintain the corresponding business operation state when there is a controller to be recovered in the target cluster, and to initialize the controller to be recovered, and to monitor whether the controller to be recovered has completed the initialization action. The second processing module 300 is used to stop the target business process of the business module in the target cluster when the controller to be recovered completes the initialization action, and to distribute the target information of the controller to be recovered to the business module, so that the business module can update its status based on the target information to restore the controller to be recovered to the target cluster.
[0058] Optionally, in some embodiments, the first processing module 200 includes: The addition unit is used to add the status flag bit corresponding to the controller to be recovered to the state machine. The judgment unit is used to monitor the status of the controller to be restored by using the status flag bit, and to determine whether the controller to be restored has completed the initialization action based on the monitoring result.
[0059] Optionally, in some embodiments, the determining unit is specifically used for: Based on the monitoring results, determine whether the controller to be restored has finished initialization and whether the controller to be restored is in a preset fault state; If the controller to be restored has completed initialization and is not in a preset fault state, then the controller to be restored is determined to have completed the initialization action and is marked as the first preset state.
[0060] Optionally, in some embodiments, the determining unit is specifically used for: Based on the monitoring results, determine whether the controller to be restored is in a preset fault state; If the controller to be restored is in a preset fault state, it is determined that the controller to be restored has not completed the initialization action, and the controller to be restored is marked as the second preset state, and a fault notification is sent to the business module.
[0061] Optionally, in some embodiments, after initializing the controller to be restored, the first processing module 200 further includes: The monitoring unit is used to monitor the initialization progress of the controller to be restored; The adjustment unit is used to adjust the business processing strategy of the business module according to the initialization progress.
[0062] Optionally, in some embodiments, the adjustment unit is specifically used for: If the initialization progress meets the first preset threshold, maintain the business operation status of other controllers in the target cluster; If the initialization progress meets the second preset threshold, maintain the business operation status of other controllers in the target cluster, and prepare for communication between the business modules and the controllers to be restored. If the initialization progress meets the third preset threshold, maintain the business operation status of other controllers in the target cluster and adjust the communication link between the business module and the controller to be restored. If the initialization progress meets the fourth preset threshold, maintain the business operation status of other controllers in the target cluster, and send preset non-critical business requests to the controller to be recovered; If the initialization progress meets the fifth preset threshold, stop the target business process of the business module.
[0063] Optionally, in some embodiments, after stopping the target business process of the business module in the target cluster, the second processing module 300 is further configured to: The control module synchronizes the status information of the controller to be restored to other controllers in the target cluster; Based on the status information of the controller to be recovered, the business processing logic of other controllers in the target cluster is determined, so that the other controllers in the target cluster can perform business processing according to the corresponding business processing logic.
[0064] Optionally, in some embodiments, the target information includes at least one of the identification information, configuration information, and status information of the controller to be restored.
[0065] Optionally, in some embodiments, before stopping the target business process of the business module in the target cluster, the second processing module 300 is further configured to: Based on a preset broadcast method, preset stage event notifications are sent to the business modules in the target cluster.
[0066] The storage system controller fault recovery device proposed in this embodiment of the invention, when a controller to be recovered exists in the target cluster, controls other controllers in the target cluster to maintain their corresponding business operation states, initializes the controller to be recovered, and monitors whether the controller to be recovered has completed its initialization action. If the controller to be recovered completes its initialization action, the device can stop the target business process of the business modules in the target cluster and distribute the target information of the controller to be recovered to the business modules, enabling the business modules to update their states based on the target information to restore the controller to be recovered to the target cluster. Therefore, by maintaining the business operation state while initializing newly added controllers, the device solves the problems of long business interruption times and failure to meet current high availability requirements of storage systems in existing technologies, thereby improving the high availability of the storage system.
[0067] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device may include: The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0068] When processor 402 executes the program, it implements the steps in any of the above-described embodiments of the storage system controller fault recovery method.
[0069] Furthermore, electronic devices also include: Communication interface 403 is used for communication between memory 401 and processor 402.
[0070] The memory 401 is used to store computer programs that can run on the processor 402.
[0071] The memory 401 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0072] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0073] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0074] Processor 402 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement embodiments of the present invention.
[0075] Embodiments of the present invention also provide a non-volatile computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the storage system controller fault recovery method when it is run.
[0076] In one exemplary embodiment, the aforementioned non-volatile computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory, portable hard drives, magnetic disks, or optical disks.
[0077] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the storage system controller fault recovery method.
[0078] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the storage system controller fault recovery method.
[0079] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0080] The above provides a detailed description of a storage system controller fault recovery method provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of the above embodiments are only intended to help understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for recovering from a storage system controller failure, characterized in that, Includes the following steps: Determine if the target cluster has a controller that needs to be recovered; If the controller to be recovered exists in the target cluster, control other controllers in the target cluster to maintain the corresponding service operation state, initialize the controller to be recovered, and monitor whether the controller to be recovered has completed the initialization action; When the controller to be recovered completes its initialization action, the target business process of the business module in the target cluster is stopped, and the target information of the controller to be recovered is distributed to the business module, so that the business module performs a status update based on the target information to restore the controller to be recovered to the target cluster.
2. The storage system controller fault recovery method according to claim 1, characterized in that, The monitoring of whether the controller to be restored has completed the initialization action includes: Add the status flag bit corresponding to the controller to be restored to the state machine; The status flag is used to monitor the status of the controller to be restored, and the monitoring results are used to determine whether the controller to be restored has completed the initialization action.
3. The storage system controller fault recovery method according to claim 2, characterized in that, The step of determining whether the controller to be restored has completed the initialization action based on the monitoring results includes: Based on the monitoring results, it is determined whether the controller to be restored has finished initialization and whether the controller to be restored is in a preset fault state. If the controller to be restored has finished initialization and is not in the preset fault state, then it is determined that the controller to be restored has completed the initialization action and is marked as the first preset state.
4. The storage system controller fault recovery method according to claim 2, characterized in that, The step of determining whether the controller to be restored has completed the initialization action based on the monitoring results includes: Based on the monitoring results, it is determined whether the controller to be restored is in a preset fault state; If the controller to be restored is in the preset fault state, it is determined that the controller to be restored has not completed the initialization action, and the controller to be restored is marked as the second preset state, and a fault notification is sent to the service module.
5. The storage system controller fault recovery method according to claim 1, characterized in that, After initializing the controller to be restored, the process also includes: Monitor the initialization progress of the controller to be restored; Adjust the business processing strategy of the business module according to the initialization progress.
6. The storage system controller fault recovery method according to claim 5, characterized in that, The step of adjusting the business processing strategy of the business module according to the initialization progress includes: If the initialization progress meets the first preset threshold, maintain the service operation status of other controllers in the target cluster; If the initialization progress meets the second preset threshold, maintain the service operation status of other controllers in the target cluster, and prepare for communication between the service module and the controller to be restored. If the initialization progress meets the third preset threshold, maintain the service operation status of other controllers in the target cluster, and adjust the communication link between the service module and the controller to be restored. If the initialization progress meets the fourth preset threshold, maintain the business operation status of other controllers in the target cluster, and send preset non-critical business requests to the controller to be recovered. If the initialization progress meets the fifth preset threshold, the target business process of the business module is stopped.
7. The storage system controller fault recovery method according to claim 1, characterized in that, After stopping the target business process of the business module in the target cluster, the following is also included: The control module synchronizes the status information of the controller to be recovered to other controllers within the target cluster. Based on the status information of the controller to be recovered, the business processing logic of other controllers in the target cluster is determined, so that the other controllers in the target cluster perform business processing according to the corresponding business processing logic.
8. The storage system controller fault recovery method according to claim 1, characterized in that, The target information includes at least one of the identification information, configuration information, and status information of the controller to be restored.
9. The storage system controller fault recovery method according to claim 1, characterized in that, Before stopping the target business process of the business module in the target cluster, the following steps are also included: Based on a preset broadcast method, preset stage event notifications are sent to the service modules in the target cluster.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the storage system controller fault recovery method as described in any one of claims 1 to 9 when executing the computer program.