Fault processing method of multi-control storage system and electronic equipment
By identifying a mirror pair failure in a multi-controller storage system and immediately executing the second fault handling process of the cache state machine, business processing can continue using the surviving or recovered controllers. This solves the problem of long-term business interruption in multi-controller storage systems and improves the business continuity and stability of the system.
Patent Information
- Application Number
- CN202511463779.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-14
AI Technical Summary
In multi-controller storage systems, when two controllers in the same mirror pair fail simultaneously, existing technologies can cause prolonged business interruptions, affecting system response speed and performance.
After a mirror pair failure is detected, the control cache state machine executes the first failure handling process, and immediately executes the second failure handling process before the first failure handling is completed. The surviving or recovered controller continues business processing, avoiding waiting for the first failure handling process to complete.
It enables rapid recovery of business processes when complex faults overlap, reduces business interruption time and data processing latency, and enhances business continuity and system stability.
Smart Images

Figure CN120929313A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data storage technology, and in particular to fault handling methods and electronic devices for multi-controller storage systems. Background Technology
[0002] In related technologies, multi-controller storage systems employ mirror pairs to improve data redundancy and system availability. The two controllers in a mirror pair back up each other's cached data. When one controller fails, the other can take over, maintaining continuous data access. However, when both controllers in the same mirror pair fail simultaneously, business processing is suspended until the failed controller recovers. But in a multi-controller storage system, even if other controllers remain healthy, business operations cease, leading to prolonged service interruptions and impacting system responsiveness and performance.
[0003] Therefore, the relevant technologies suffer from the problem of prolonged business interruption in multi-controller storage systems under fault scenarios. Summary of the Invention
[0004] This application provides a fault handling method and electronic device for a multi-controller storage system, in order to at least solve the problem of long-term business interruption in multi-controller storage systems under fault scenarios in related technologies.
[0005] This application provides a fault handling method for a multi-controller storage system, the multi-controller storage system including multiple controllers and a cache state machine, the multiple controllers forming at least one mirror pair, and one mirror pair in the at least one mirror pair including two controllers from the multiple controllers; the method includes:
[0006] If a specified mirror pair is identified in the at least one mirror pair, the cache state machine is controlled to execute a first fault handling procedure, wherein the specified mirror pair is a mirror pair in which both controllers are faulty;
[0007] If the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution phase of the first fault handling process to control the multi-controller storage system to continue business processing. The designated controller is either the controller that triggered the recovery among the multiple controllers, or the controller that does not belong to the designated mirror pair and triggered the fault among the multiple controllers.
[0008] This application also provides a fault handling device for a multi-controller storage system, the multi-controller storage system including multiple controllers and a cache state machine, the multiple controllers forming at least one mirror pair, and one mirror pair in the at least one mirror pair including two controllers from the multiple controllers; the device includes:
[0009] The first execution module is configured to control the cache state machine to execute a first fault handling process when a specified mirror pair is detected in the at least one mirror pair, wherein the specified mirror pair is a mirror pair in which both controllers are faulty;
[0010] The second execution module is used to execute the second fault handling process of the cache state machine based on the execution stage of the first fault handling process when the first fault handling process has not been completed and a designated controller is identified, so as to control the multi-controller storage system to continue business processing. The designated controller is either the controller that triggered recovery among the multiple controllers, or the controller that does not belong to the designated mirror pair and triggered a fault among the multiple controllers.
[0011] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the fault handling method of any of the above-described multi-controller storage systems when executing the computer program.
[0012] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the fault handling method of any of the above-described multi-controller storage systems.
[0013] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described fault handling methods for a multi-controller storage system.
[0014] This application describes a method for controlling a cache state machine to execute a first fault handling process when a specified mirror pair is identified among at least one mirror pair. The specified mirror pair is defined as a pair containing two controllers that have both failed. If the first fault handling process is not completed and a specified controller is identified, a second fault handling process is executed based on the execution phase of the first fault handling process. This ensures the multi-controller storage system continues business processing. The specified controller is either the controller that triggered recovery among multiple controllers, or a controller that does not belong to the specified mirror pair and has triggered a failure. The cache state machine can immediately respond to fault handling processes in cases of recovery of a surviving controller or failure of a new controller during the execution of the first fault handling process. This enables rapid recovery of business processing when faced with complex overlapping faults, reducing business interruption time and data processing latency. It solves the problem of prolonged business interruption in multi-controller storage systems under fault scenarios in related technologies, significantly enhancing business continuity and system stability. Attached Figure Description
[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a hardware structure block diagram of a fault handling method for a multi-controller storage system according to an embodiment of this application.
[0017] Figure 2 This is a flowchart of a fault handling method for a multi-controller storage system according to an embodiment of this application.
[0018] Figure 3 This is a schematic diagram of a fault handling process according to an embodiment of this application.
[0019] Figure 4 This is a schematic diagram of a multi-controller storage system according to an embodiment of this application.
[0020] Figure 5 This is a schematic diagram of a fault handling method for a multi-controller storage system according to an embodiment of this application.
[0021] Figure 6 This is a schematic diagram of another fault handling method for a multi-controller storage system according to an embodiment of this application.
[0022] Figure 7 This is a structural block diagram of a fault handling device for a multi-controller storage system according to an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0024] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0025] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] The specific application environment architecture or specific hardware architecture on which the fault handling method of the multi-controller storage system depends is described here.
[0027] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a fault handling method for a multi-controller storage system according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a central processing unit (CPU), microprocessor (MCU), or programmable logic device (FPGA), etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0028] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the fault handling method of the multi-controller storage system in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0029] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0030] The embodiments of this application provide a fault handling method for a multi-controller storage system. The method is described in detail below in conjunction with the execution flow of the fault handling method for a multi-controller storage system.
[0031] This embodiment provides a fault handling method for a multi-controller storage system. Figure 2 This is a flowchart of a fault handling method for a multi-controller storage system according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:
[0032] Step S202: If a specified mirror pair is detected in at least one mirror pair, the cache state machine is controlled to execute the first fault handling process, wherein the specified mirror pair is a mirror pair in which both controllers are faulty.
[0033] Step S204: If the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution phase of the first fault handling process to control the multi-controller storage system to continue business processing. The designated controller is either the controller that triggered recovery among multiple controllers, or the controller that does not belong to the designated mirror pair and triggered a fault among multiple controllers.
[0034] Through the above steps, since the cache state machine can immediately respond to the recovery of the live controller or the failure of the new controller during the execution of the first fault handling process, it can quickly restore business processing when faced with complex fault superposition, reduce business interruption time and data processing latency, solve the problem of long-term business interruption in multi-controller storage systems under fault scenarios in related technologies, and significantly enhance business continuity and system stability.
[0035] It should be noted that a multi-controller storage system can include multiple controllers, a cache state machine, and cache modules. The multiple controllers are the core components of the multi-controller storage system, working together to improve its performance and availability. These controllers can form a cluster. Each controller can handle a portion of the (input / output) operations and data management tasks. The cache state machine manages the state and processes of the cache modules, changing their behavior based on system events and controller states. The cache modules improve data access speed and efficiency within the multi-controller storage system. Multiple controllers in a multi-controller storage system can form at least one mirror pair. A mirror pair can contain two controllers.
[0036] In the embodiment provided in step S202, the designated mirror pair can be a mirror pair where both controllers are faulty, and the first fault handling process can be a preset fault handling process in the multi-controller storage system. If the designated mirror pair exists in at least one mirror pair, the control cache state machine is triggered to execute the first fault handling process.
[0037] Optionally, if it is determined that a specified mirror pair exists in at least one mirror pair, a corresponding fault handling event is generated. The cache state machine receives and responds to the fault handling event, triggering the execution of a first fault handling process. The first fault handling process includes at least a pause process to control the suspension of the fault controller's business processing.
[0038] Optionally, each controller sends a heartbeat message to other controllers in the system within each specified heartbeat cycle, and returns an acknowledgment response upon receiving the heartbeat message. This acknowledgment response proves that the sending controller was active in the previous specified heartbeat cycle, and also indicates that the receiving controller itself has the ability to respond to external communication. If a controller fails to receive a heartbeat acknowledgment from its mirror pair or other controllers in the cluster within a specified heartbeat cycle, the controller is marked as potentially faulty. To reduce false alarms, a fault tolerance threshold can be set, meaning that a controller is officially marked as faulty only after failing to receive acknowledgments for several consecutive heartbeat cycles. In a multi-controller storage system, the existence of a specified mirror pair can be determined using heartbeat messages and the fault tolerance threshold. If the existence of a specified mirror pair is determined in at least one mirror pair, the first fault handling procedure of the cache state machine is triggered to implement a silent caching module, thereby stopping the business processing related to the specified mirror pair and preventing data inconsistency.
[0039] In the embodiment provided in step S204, during the process of controlling the cache state machine to execute the first fault handling process, if a designated controller is detected, the cache state machine will not simply wait for the first process to complete, but will immediately start the second fault handling process based on the execution stage of the first fault handling process to ensure that the system can continue to perform business processing.
[0040] The specified controller can be the controller that triggers the recovery event or the controller that triggers the fault (not the controller that belongs to the same mirror pair as the two controllers that caused the fault).
[0041] Optionally, after the faulty controller is repaired and comes back online, it can respond to heartbeat detection from other controllers to confirm that the faulty controller has recovered. Alternatively, after the faulty controller is repaired and comes back online, it can actively send heartbeat signals to inform other components in the system (such as the cluster manager, cache state machine, etc.) that its state has changed from faulty to active. Or, after the faulty controller is repaired and comes back online, it can report its status information to the cache state machine, which then determines that the faulty controller has recovered based on the status information.
[0042] In one example, assuming both controllers A and B in a specified mirror pair fail, the control cache state machine executes the first fault handling procedure. During this procedure, a fault in controller C is detected. At this point, controller C becomes the "specified controller" because it does not belong to the "specified mirror pair" consisting of A and B. The cache state machine executes the second fault handling procedure based on the current processing stage, instead of waiting for the first procedure to complete. This ensures that the system can quickly resume business processing even in complex fault scenarios, avoiding unnecessary business interruptions and data delays.
[0043] According to the embodiments of this application, when a specified mirror pair is identified among at least one mirror pair, the cache state machine is controlled to execute a first fault handling process, wherein the specified mirror pair is a mirror pair in which both controllers have failed; if the first fault handling process has not been completed and the specified controller is identified, a second fault handling process of the cache state machine is executed based on the execution stage of the first fault handling process to control the multi-controller storage system to continue business processing, wherein the specified controller is the controller that triggered recovery among multiple controllers, or the controller that does not belong to the specified mirror pair and has triggered a failure among multiple controllers. The cache state machine can immediately respond to the fault handling process in the case of recovery of the surviving controller or failure of a new controller during the execution of the first fault handling process, enabling rapid recovery of business processing when faced with complex overlapping faults, reducing business interruption time and data processing latency, solving the problem of long-term business interruption in multi-controller storage systems under fault scenarios in related technologies, and significantly enhancing business continuity and system stability.
[0044] As an optional implementation, the execution order of the processes in the first fault handling process is the interruption process, the first update process, and the recovery process, wherein the interruption process is used to indicate the interruption of the business processing of the multi-controller storage system, and the recovery process is used to indicate the recovery of the business processing of the multi-controller storage system. Step S204 includes: determining the cluster state of the multi-controller storage system when the designated controller is identified as the controller that triggered recovery among multiple controllers, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed; when the cluster state of the multi-controller storage system is in a stable state, executing a mirror pair construction process, wherein the mirror pair construction process is used to indicate the determination of a first controller from a group of candidate controllers in the multi-controller storage system to form a first designated mirror pair with the designated controller, wherein the first designated mirror pair is used to execute the services of a prior mirror pair including the designated controller and having a controller fault, and the candidate controllers in the group of candidate controllers are controllers in the multi-controller storage system that are not in a fault state; when the mirror pair construction process is completed, executing a second update process, wherein the second update process is used to indicate the updating of the cluster state information of the multi-controller storage system, and the cluster state information of the multi-controller storage system is used to indicate the online information of multiple controllers and the mirror pair state information of multiple controllers; when the second update process is completed, executing a recovery process.
[0045] It should be noted that the first fault handling process may include an interruption process, a first update process, and a recovery process. The interruption process can be used to instruct the interruption of business processing of the multi-controller storage system, and the recovery process can be used to instruct the resumption of business processing of the multi-controller storage system. Specifically, the first fault handling process can be found in [reference needed]. Figure 3 Upon receiving an event notification (such as a fault event), the cache state machine first interrupts the multi-controller storage system's business processing by initiating an interruption request (such as a queuing request) to ensure the safety of the current operation. Subsequently, after the state machine receives the interruption request, it switches the business mode to send an update request (such as an acknowledgment request) to the business processing end to realize the business mode switch. Finally, after the state machine receives the update request, it initiates a recovery request (such as a resume request) to instruct the business processing end to resume the multi-controller storage system's business processing. After the business processing end resumes the multi-controller storage system's business processing, it reports back to the state machine that the business recovery is complete.
[0046] Optionally, if it is determined that at least one mirror pair contains a specified mirror pair (a mirror pair in which both controllers are faulty), a cluster fault notification event is triggered. The cache state machine receives the cluster fault notification event, confirms that a mirror pair is faulty, and executes an interruption process to suspend the business processing of the cache module. After the interruption process, the cache state machine executes the first update process (acknowledgement phase) to update the cluster information and realize the synchronization of cluster information.
[0047] Optionally, if the execution phase of the first fault handling process instructs the completion of the first update process but does not execute the recovery process, and if it is determined that a specified controller (such as the previously faulty controller has recovered) exists, the cluster status of the multi-controller storage system is evaluated, for example, whether there are any new faults or unstable factors in the multi-controller storage system.
[0048] Optionally, if the specified controller is identified as the controller that triggered recovery among multiple controllers, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, the second fault handling process may include the image pair construction process, the second update process, and the recovery process. That is, if the specified controller is identified as the controller that triggered recovery among multiple controllers, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, the order of the entire fault handling process is: interruption process, first update process, image pair construction process, second update process, and recovery process.
[0049] Once it is determined that the cluster state of the multi-controller storage system is stable, a mirror pair construction process is executed. This involves selecting a controller from a group of candidate controllers that are not in a faulty state and constructing a new mirror pair with the specified controller, namely the first specified mirror pair. The first specified mirror pair can be used to execute services that include the specified controller and where a controller failure exists in the prior mirror pair. Optionally, the selection of a controller from multiple candidate controllers that are not in a faulty state can be determined based on load balancing, performance evaluation, and mirror pair formation rules, ensuring data redundancy and service processing capabilities.
[0050] Optionally, the online information of multiple controllers can be used to indicate whether the controllers in the multi-controller storage system are online. If a controller malfunctions, the recorded online information indicates an offline status; if the controller is functioning correctly, the recorded online information indicates an online status. The mirror pair status information of multiple controllers can be used to indicate the mirror pair to which each controller belongs. A controller can belong to at least one mirror pair; that is, a controller can belong to multiple mirror pairs. Taking a four-controller storage system as an example... Figure 4 As shown, there are controllers 1, 2, 3, and 4. Controller 1 can form a mirror pair 1 with controller 2, controller 2 can form a mirror pair 2 with controller 3, controller 3 can form a mirror pair 3 with controller 4, and controller 4 can form a mirror pair 4 with controller 1. Different services can be executed in different mirror pairs belonging to a single controller.
[0051] Optionally, a set of candidate controllers can be controllers in the multi-controller storage system that are not in a fault state and whose performance conditions meet specified performance conditions. Optionally, a first pre-trained model can be deployed in each controller of the multi-controller storage system. The performance condition of the controller is determined by the response time between the predicted data output by the first pre-trained model after the training data in the training set is input into the first pre-trained model. The response time is negatively correlated with the controller performance; the longer the response time, the worse the controller performance, and the shorter the response time, the better the controller performance and the faster the response. In this case, the specified performance condition can be a specified response time. Controllers in the multi-controller storage system that are not in a fault state and whose response time is less than the specified response time can be selected as candidate controllers. This ensures that the selected controllers can meet the performance requirements for building new image pairs and avoids business processing delays or instability caused by low-performance controllers.
[0052] Optionally, the first pre-trained model can be a performance identification model, or it can be something other than a performance identification model. In a multi-controller storage system, the first pre-trained model can be at least one of the following functional models: a fault prediction model, a workload prediction model, and a security prediction model. The fault prediction model can predict potential faults in the controller or other components of the multi-controller storage system. It can predict potential faults in advance by learning from historical fault data and system operating status, facilitating preventative measures by maintenance personnel. The workload prediction model can be used to predict the workload of the multi-controller storage system over a future period (i.e., a specified duration), including the number, size, and type of read and write operations. This prediction helps the multi-controller storage system adjust resource allocation in advance to meet expected performance demands and avoid service degradation caused by sudden increases in workload. The security prediction model can be used to predict potential threats, such as attacks or abnormal access patterns, by analyzing controller log information and network traffic.
[0053] Optionally, after the first specified mirror pair is built, the cache state machine can execute a second update process to update the online controller information and mirror pair status information in the multi-controller storage system to reflect the latest cluster architecture and data layout. Specifically, in the event of a controller failure, the corresponding controller's online information will be displayed as offline.
[0054] In one example, assuming the multi-controller storage system is a four-controller system with controllers 1, 2, 3, and 4, in the event of a failure in the mirror pair consisting of controllers 1 and 2, a first fault handling process can be initiated. After completing the interruption and first update processes but not the recovery process, if controller 2 recovers, this change can be detected. The stability of the multi-controller storage system cluster can be determined by checking for new controller failures or specified instabilities. If the cluster is stable, one controller can be selected from controllers 3 and 4 to form a new mirror pair with controller 2 (such as the first specified mirror pair), ensuring data redundancy and rapid recovery of business processing capabilities. After the new mirror pair is established, the cache state machine can be controlled to execute a second update process to update internal state information, reflecting the latest online controller list and mirror pair status. Subsequently, the cache state machine enters the recovery process, reactivating business processing to ensure rapid recovery of data read and write operations and guaranteeing business continuity.
[0055] This embodiment, by executing the second fault handling process at an appropriate stage of the first fault handling process, enables timely response to controller recovery events, avoids long-term interruptions in business processing, and improves business continuity. After a controller recovers, the system first checks whether the cluster is stable. Only after confirming system stability will the business recovery process be executed, thus avoiding data processing risks in unstable states. Through the mirror pair construction process, a suitable controller can be quickly identified from a group of healthy candidate controllers, and a new mirror pair can be built with the recovered controller, ensuring data redundancy and rapid recovery of business processing capabilities.
[0056] As an optional implementation, determining the cluster status of a multi-controller storage system includes: determining whether a second mirror pair exists among the multiple first mirror pairs based on the status information of the first mirror pairs, wherein the first mirror pairs among the multiple first mirror pairs include a specified controller, and the second mirror pair includes a mirror pair with a controller failure; if a second mirror pair exists, determining the cluster status of the multi-controller storage system based on the time of failure of the controller in the second mirror pair.
[0057] It should be noted that the first mirror pair can be a mirror pair in a multi-controller storage system, consisting of a designated controller and another controller, used for data redundancy and concurrent processing. The second mirror pair can be one of the multiple first mirror pairs except for the one where the designated controller fails.
[0058] The timing of a controller failure can be used to indicate the point in time when the controller fails, as well as to assess the scope of the failure and the degree of system recovery.
[0059] Optionally, the cluster state of a multi-controller storage system can include a stable state and an unstable state. In the presence of a second mirror pair, the current cluster's stability is determined based on the timing of the controller failure in the second mirror pair. An unstable state indicates that business processing within the current cluster cannot be maintained. Generally, the stall state can be used to identify an unstable cluster state.
[0060] Optionally, a more precise fault detection mechanism, such as fault backtracking technology based on event logs, can be introduced to more accurately determine the moment of controller failure and further improve the accuracy of cluster status judgment.
[0061] Optionally, if no second mirror pair exists among multiple first mirror pairs, the cluster information is updated using the prior mirror pair information of the specified controller to restore the service processing functionality of the mirror pair in which the specified controller previously resided. If a second mirror pair exists among multiple first mirror pairs, the failure times of the controllers in the second mirror pair excluding the specified controller and the failure time of the specified controller are used to determine whether the services previously processed by the specified controller can be restored. If restoration is possible, the mirror pair construction process is executed.
[0062] By accurately determining the second mirror pair and its failure time through this embodiment, the self-state can be assessed more accurately, avoiding data processing in unstable or potentially dangerous states.
[0063] As an optional implementation, determining the cluster state of the multi-controller storage system based on the failure time of the controller in the second mirror pair includes: obtaining the failure time of the controllers in the second mirror pair excluding the specified controller; if the failure time of the controllers in the second mirror pair excluding the specified controller is earlier than the failure time of the specified controller, determining the cluster state of the multi-controller storage system as a stable state; if the failure time of the controllers in the second mirror pair excluding the specified controller is later than the failure time of the specified controller, determining the cluster state of the multi-controller storage system as an unstable state.
[0064] It should be noted that if the failure time of the controller other than the designated controller in the second mirror pair is earlier than the failure time of the designated controller, then the failure time of the designated controller is later than the failure time of the controller other than the designated controller in the second mirror pair. This confirms that the designated controller in the second mirror pair has performed business processing on the controllers other than the designated controller in the second mirror pair. Since there is no unrecorded business, it can be indicated that there is no unsustainable business processing in the multi-controller storage system, and the cluster state of the multi-controller storage system is determined to be stable. At this point, the business processing of the second mirror pair can be restored by creating a new mirror pair for the designated controller.
[0065] If the controller failure in the second mirror pair occurs after the failure of the specified controller, although the specified controller recovers, the cached data between the failure time of the controllers other than the specified controller in the second mirror pair and the failure time of the specified controller is recorded in the controllers other than the specified controller in the second mirror pair. At this time, the cluster state of the multi-controller storage system is determined to be unstable. Therefore, only the online information of the controllers other than the specified controller in the second mirror pair is restored to inform the other controllers in the cluster that the specified controller has recovered.
[0066] This embodiment avoids building mirror pairs in unstable states by accurately judging the cluster status, thereby reducing unnecessary data synchronization and business interruptions and improving the stability and efficiency of the system.
[0067] As an optional implementation, the above method further includes: executing a third update process when the cluster state of the multi-controller storage system is in an unstable state, wherein the third update process is used to instruct the updating of the online information of the controllers in the multi-controller storage system; and executing a recovery process after the third update process is completed.
[0068] It should be noted that when the cluster status of a multi-controller storage system is unstable, only the online information of controllers in the second mirror pair, excluding the designated controller, is restored to inform the other controllers in the cluster so that the designated controller can be restored. Of course, in a multi-controller storage system, a controller can have multiple mirror pairs, each handling different services. If the designated controller's mirror pair has other mirror pairs besides the second mirror pair, the service processing functions of the other mirror pairs of the designated controller (excluding the second mirror pair) can be restored after updating the online information of the controllers (such as the designated controller) in the multi-controller storage system.
[0069] This embodiment ensures the real-time status of the cluster by forcibly updating the controller's online information under unstable conditions, providing accurate information support for subsequent business recovery and improving the system's response speed and efficiency.
[0070] As an optional implementation, determining a first controller from a set of candidate controllers in a multi-controller storage system includes: obtaining mirror pair status information of candidate controllers in a set of candidate controllers, wherein the mirror pair status information is used to indicate the mirror pair to which the candidate controller belongs; and determining a first controller from a set of candidate controllers based on the mirror pair status information of candidate controllers in a set of candidate controllers, so as to form a first designated mirror pair with a designated controller.
[0071] It should be noted that the mirror pair status information can refer to the information of the mirror pair to which the candidate controller currently belongs, including but not limited to the mirror pair number and the status of the other controller in the mirror pair. Optionally, by obtaining the status information of the candidate controller, a candidate controller, namely the first controller, can be selected from a group of candidate controllers to form a first designated mirror pair with the designated controller. The first designated mirror pair can be used to execute services of a prior mirror pair (such as the second mirror pair) that includes the designated controller and has a controller failure.
[0072] Optionally, the mirror pair status information of the candidate controllers in a set of candidate controllers is parsed to select a first controller from the set of candidate controllers to form a first specified mirror pair with the specified controller.
[0073] In this embodiment, by analyzing the mirror pair status information of candidate controllers, the best controller, namely the first controller, is intelligently selected and paired with the designated controller as a mirror pair. This avoids data redundancy or performance bottlenecks caused by improper mirror pair construction, thereby improving the system's data security and processing efficiency.
[0074] As an optional implementation, the mirror pair status information includes the number of mirror pairs and mirror pair attributes. The mirror pair attributes are used to indicate whether the controller in the mirror pair is a data initiator or a data receiver. Determining a first controller from a group of candidate controllers based on the mirror pair status information of the candidate controllers in the group includes: determining the candidate controller with the fewest mirror pairs in the group as the second controller; if there are no more than one second controller, determining the second controller as the first controller; if there are more than one second controller, determining the frequency of the second controller being a data initiator based on the mirror pair attributes and the number of mirror pairs of the second controllers in the group; and determining the second controller with the lowest frequency of being a data initiator from the group of second controllers as the first controller.
[0075] It's important to note that the number of mirror pairs can refer to the number of mirror pairs a single controller belongs to. This can be used to filter controllers that participate in the fewest mirror pairs, i.e., controllers that bear less data redundancy tasks. This can achieve load balancing by reducing data synchronization and mirroring burden. For example, in a multi-controller storage system, there are three candidate controllers (controller 2, 3, and 4). Controller 2 belongs to one mirror pair, controller 3 belongs to two mirror pairs, and controller 4 also belongs to two mirror pairs. Then, controller 2 is selected to form a mirror pair with the specified controller.
[0076] Optionally, if there is only one second controller, this second controller is directly designated as the first controller, that is, the role (data initiator or receiver) that is ultimately added to the mirror pair.
[0077] Optionally, when there are multiple second controllers, the first controller is determined by the proportion of candidate second controllers among the multiple second controllers that act as data initiators. Specifically, a controller with a lower proportion of data initiators means that it participates in data synchronization relatively less often, and is therefore more suitable to become a new data initiator to maintain system load balancing.
[0078] Optionally, an adaptive mirror pair attribute adjustment mechanism can be introduced to dynamically optimize the mirror pair construction strategy and address potential performance bottlenecks in changing business scenarios. Optionally, performance metrics of each controller, such as CPU utilization, I / O latency, and changes in business models, can be monitored in real time. The real-time monitored performance metrics of each controller are input into a pre-deployed attribute recognition model to obtain the predicted attributes of each controller. Based on the predicted attributes of each controller, the mirror pair attributes of each controller are determined. After adjusting the mirror pair attributes of the controllers, a re-mirroring process is initiated to transfer data from the original initiator to the new receiver, ensuring data redundancy.
[0079] In this embodiment, by comprehensively considering the number and attributes of mirror pairs, the optimal controller, i.e., the first controller, can be intelligently selected and paired with the designated controller as a mirror pair. This avoids data redundancy or performance bottlenecks caused by improper mirror pair construction, thereby improving the system's data security and processing efficiency.
[0080] As an optional implementation, after executing the image pair construction process, the above method further includes: using the first controller as the data receiver, using the specified controller as the data sender, and triggering the execution of the cached data re-mirroring process, wherein the cached data re-mirroring process is used to instruct the specified cached data in the specified controller to be transmitted to the first controller.
[0081] It's important to note that in the cached data re-mirroring process, a designated controller can be used to send cached data to the first controller. The designated controller can be the data initiator in the original mirror pair, or a controller whose data storage responsibility needs to be adjusted due to changes in system state. The cached data re-mirroring process refers to the process in a multi-controller storage system of copying cached data from a designated controller (such as the data initiator) to the first controller (the data receiver). This ensures that data redundancy and high system availability are maintained even in the event of controller failure, recovery, or changes in system state.
[0082] Optionally, the specified cached data in the specified controller can be the cached data recorded by the specified controller when it executes a business event between the time of failure of other controllers (excluding the specified controller) in the second mirror pair (the prior mirror pair of the specified controller) and the time of failure of the specified controller.
[0083] Optionally, to improve data transmission efficiency, data compression can be used to compress the specified cached data, reducing bandwidth requirements and resolving potential data transmission latency issues under limited network resources. Furthermore, to ensure data transmission security, encryption technology can be used to compress the compressed specified cached data.
[0084] Through this embodiment, the data re-mirroring process ensures data redundancy and consistency even when the controller state changes, thus avoiding the potential risk of data loss.
[0085] In one example, suppose the initial mirror pair consists of controllers 0 and 1. However, after a stall state, controller 0 recovers first, while controller 1 has not yet recovered. At this point, controller 0 is considered the designated controller and will assume the role of transmitting cached data. If controller 2 is selected as the first controller from a set of candidate controllers, the cache module will trigger a re-mirroring process when controller 0 determines to transmit data to controller 2. In this process, controller 0 will copy the data in its cache to controller 2, thereby establishing a new mirror pair (0,2) to replace the original (0,1) mirror pair.
[0086] As an optional implementation, determining a first controller from a set of candidate controllers in a multi-controller storage system includes: determining the performance of candidate controllers in the set of candidate controllers; and determining the first controller from the set of candidate controllers based on the performance of candidate controllers in the set of candidate controllers.
[0087] It should be noted that performance metrics may include, but are not limited to, CPU utilization, I / O read / write speed, memory access speed, and network bandwidth. These metrics can be used to evaluate the controller's real-time processing capabilities and its ability to support services. Optionally, the performance of each candidate controller can be analyzed by reading the performance data of a set of candidate controllers, thereby determining a first controller from the set of candidate controllers to form a first designated mirror pair with the specified controller.
[0088] Optionally, a performance prediction model can be used to predict the performance trend of the controller in advance, thereby adjusting the mirror pair construction strategy.
[0089] In this embodiment, by analyzing the performance of candidate controllers, the controller with the best performance is selected as part of the mirror pair, ensuring the efficiency and reliability of mirror pair construction and improving the system's processing efficiency and stability.
[0090] As an optional implementation, determining the performance of a candidate controller in a group of candidate controllers includes: sending performance test commands to each candidate controller in the group of candidate controllers, and recording the response time of the candidate controllers in the group of candidate controllers when executing the performance test commands, wherein the response time of the performance test commands is used to indicate the performance of the candidate controllers in the group of candidate controllers, and the performance test commands include at least one of the following test commands: read performance test command, write performance test command, memory access performance test command, and network test command.
[0091] It should be noted that performance test commands refer to commands issued to candidate controllers that can test and evaluate the performance metrics of the candidate controllers. Performance test commands include at least one of the following: read performance test commands, write performance test commands, memory access performance test commands, and network test commands. Read performance test commands are used to test the speed at which the candidate controller reads data, measuring its ability to make read I / O requests, and are often used to evaluate data access latency. Write performance test commands are used to test the speed at which the candidate controller writes data to storage media, measuring its ability to make write I / O requests. Memory access performance test commands are used to test the efficiency of the controller in accessing and managing memory resources, especially the execution speed of cache operations. Network test commands are used to evaluate the controller's performance in network communication, including data transfer speed, network latency, and bandwidth utilization. Response time is the time required for the controller to return results after executing a performance test command, reflecting its speed and efficiency in handling this type of task.
[0092] This embodiment uses performance test commands to accurately detect the performance metrics of candidate controllers, effectively avoiding errors in mirror pair construction caused by unclear performance metrics, and improving the system's processing efficiency and stability.
[0093] As an optional implementation, determining a first controller from a set of candidate controllers based on the performance of the candidate controllers in the set includes: when there are multiple performance test commands, weighting and summing the multiple response times corresponding to the multiple performance test commands according to the command type of the performance test commands to obtain the total response time of the candidate controllers in the set of candidate controllers; and determining the controller with the smallest total response time of the candidate controllers in the set of candidate controllers as the first controller.
[0094] It should be noted that, considering the varying degrees to which different performance test commands contribute to the overall performance of the controller, a weight value can be assigned to each type of performance test command. Then, the total response time is obtained by summing all response times. The total response time can be used to reflect the overall performance of the controller.
[0095] Optionally, a weight set can be pre-deployed, comprising multiple sets of weight values. These weight values correspond to the proportions of various task types, including read / write tasks and network tasks. Optionally, based on the proportions of task types running in each controller, a set of weight values corresponding to each type is selected from the weight set. This set of weight values includes weight values corresponding to different types of performance test commands. Specifically, the weight set can be generated based on historical running data and a deployed weight prediction model, or it can be set based on empirical values; this application does not impose any limitations on this.
[0096] In this embodiment, by weighted summation of performance test commands, the overall performance of the controller is obtained, which can effectively avoid the deviation in mirror pair construction that may be caused by a single performance indicator, and improve the data security and processing efficiency of the system.
[0097] As an optional implementation, the process of selecting a first controller from a set of candidate controllers may include: First, identifying the number of mirror pairs and the attributes of the candidate controllers in the set of candidate controllers. If there are multiple controllers selected based on the number of mirror pairs and the attributes of the candidate controllers in the set of candidate controllers, further filtering can be performed based on the performance of the remaining controllers. If there are multiple controllers after filtering, a cyclic mirroring method can be adopted, that is, each time a new mirror pair is built, it is selected from the controller sequence in a cyclical order. In this way, even if the performance is the same, the sequential selection can avoid some controllers from being overloaded, which can better cope with the business needs of high concurrency and large data volume, and improve the overall stability and response speed.
[0098] As an optional implementation, if the first fault handling process has not been completed and a designated controller is identified, a second fault handling process of the cache state machine is executed based on the execution phase of the first fault handling process. This includes: if the designated controller is a controller among multiple controllers that does not belong to the designated mirror pair and has triggered a fault, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, a fourth update process is executed, wherein the fourth update process is used to indicate updating the online information of the controllers in the multi-controller storage system; and if the fourth update process has been completed, a recovery process is executed.
[0099] It should be noted that, if the specified controller is one of multiple controllers that does not belong to the specified mirror pair and has triggered a fault, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, the fourth update process is executed first to indicate that the online information of the controller in the multi-controller storage system is updated; if the fourth update process is completed, the recovery process is executed. If the specified controller is one that does not belong to the specified mirror pair and has triggered a fault, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, the second fault handling process may include both the fourth update process and the recovery process.
[0100] This embodiment ensures the real-time nature and accuracy of the controller's online information by executing corresponding update processes at different stages of fault handling, effectively avoiding business processing errors caused by untimely information updates and improving the system's response speed and efficiency.
[0101] As an optional implementation, this application also provides a fault handling method for a multi-controller storage system, such as... Figure 5 As shown, in a multi-controller storage system, if the controllers of the same mirror pair fail simultaneously, the first fault handling process is triggered.
[0102] During the execution of the first fault handling process, it is determined whether to execute the first update process. If the controller fails again without executing the first update process, the controller will execute the first update process and the recovery process. In this scenario, the corresponding fault handling process includes the interrupt process, the first update process, and the recovery process (such as quiesce-ack(STALL)-resume).
[0103] During the execution of the first fault handling process and after the first update process is completed, it is determined whether the recovery process has been completed. If the recovery process is not completed and the controller fails again, the fourth update process and the recovery process are executed. In this scenario, the corresponding fault handling process includes the interrupt process, the first update process, the fourth update process and the recovery process (such as quiesce-ack(1, STALL)-ack(2, STALL)-resume).
[0104] If a controller failure occurs again during the execution of the first fault handling process, the execution of the first update process, and the execution of the recovery process, then the controller will execute the interrupt process, the fifth update process, and the recovery process. In this scenario, the corresponding fault handling process includes the interrupt process, the first update process, the recovery process, the interrupt process, the fifth update process, and the recovery process (e.g., quiesce(1) - ack(1, STALL) - resume(1) - quiesce(2) - ack(2, STALL) - resume(2)).
[0105] As an optional implementation, this application also provides a fault handling method for a multi-controller storage system, such as... Figure 6 As shown, in a multi-controller storage system, if the controllers of the same mirror pair fail simultaneously, and a controller recovers during the execution of the first fault handling procedure, the control will execute the following second fault handling procedure according to the execution stage of the first fault handling procedure. Specifically, as follows:
[0106] Determine if the cluster is in a specified state (e.g., stall state). If the cluster is in the specified state and the first update process has not been executed, execute the first update process and the recovery process. That is, in this scenario, the corresponding fault handling process includes the interrupt process, the first update process, and the recovery process (e.g., quiesce-ack(STALL)-resume).
[0107] When the cluster is in a specified state, executing the first update process, and not executing the recovery process, the fifth update process and the recovery process are executed. In this scenario, the corresponding fault handling processes include the interrupt process, the first update process, the fifth update process, and the recovery process (e.g., quiesce-ack(1, STALL)-ack(2, STALL)-resume). The fifth update process is used to instruct the updating of the online information of the controllers in the multi-controller storage system. Specifically, it updates the online information of the controller to be recovered (e.g., the specified controller) and restores its corresponding mirror pair information.
[0108] When the cluster is in a specified state, executing the first update process, and executing the recovery process, the interrupt process, the sixth update process, and the recovery process are executed. That is, in this scenario, the corresponding fault handling process includes the interrupt process, the first update process, the recovery process, the interrupt process, the sixth update process, and the recovery process (e.g., quiesce(1) - ack(1, STALL) - resume(1) - quiesce(2) - ack(2, STALL) - resume(2)). Among them, the sixth update process is used to instruct the updating of the online information of the controllers in the multi-controller storage system, specifically updating the online information of the controller to be restored (e.g., the specified controller) and restoring its corresponding image pair information.
[0109] If the cluster is not in the specified state and the first update process has not been executed, the seventh update process and recovery process are executed. The seventh update process includes the image pair building process and the second update process. In this scenario, the corresponding fault handling process includes the interruption process, the seventh update process, and the recovery process (e.g., quiesce-ack(CONTRACT (indicating controller failure exit))-resume). Here, ack(CONTRACT) is used to instruct the execution of the image pair building process to synchronize cached data using the rebuilt image pair.
[0110] If the cluster is not in the specified state, the first update process has not been executed, and the recovery process has not been executed, the image pair building process, the second update process, and the recovery process are executed. That is, in this scenario, the corresponding fault handling process includes the interruption process, the first update process, the image pair building process, the second update process, and the recovery process (e.g., quiesce—ack(1, STALL)—ack(2, RECOVER)—resume). Here, ack(RECOVER (instructs the controller to recover) is used to instruct the execution of the image pair building process to synchronize cached data using the rebuilt image pair.
[0111] If the cluster is not in the specified state, the first update process is not executed, and the recovery process is executed, the interruption process, the eighth update process, and the recovery process are executed. The eighth update process includes the image pair building process and the second update process. In this scenario, the corresponding fault handling process includes the interruption process, the first update process, the recovery process, the image pair building process, the second update process, and the recovery process (such as quiesce(1) - ack(1, STALL) - resume(1) - quiesce(2) - ack(2, RECOVER) - resume(2)).
[0112] This optional example demonstrates how introducing precise controls for different stages and scenarios in the fault handling process, such as the first, fourth, and fifth update processes, ensures the storage system can react quickly to sudden failures, avoiding unnecessary business interruptions and enhancing overall system stability and response speed. The sixth and seventh update processes emphasize the importance of updating controller online information and restoring the mirror image. After the controller recovers, the mirror image is used to re-mirror the build process and data, achieving redundant data storage and ensuring data consistency and security.
[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0114] The embodiments of this application also provide a fault handling device for a multi-controller storage system. The multi-controller storage system includes multiple controllers and a cache state machine. The multiple controllers form at least one mirror pair, and one mirror pair in the at least one mirror pair includes two controllers from the multiple controllers. Figure 7 This is a structural block diagram of a fault handling device for a multi-controller storage system according to an embodiment of this application, such as... Figure 7 As shown, the device includes:
[0115] The first execution module 702 is used to control the cache state machine to execute a first fault handling process when a specified mirror pair is detected in at least one mirror pair, wherein the specified mirror pair is a mirror pair in which both controllers are faulty;
[0116] The second execution module 704 is used to execute the second fault handling process of the cache state machine based on the execution stage of the first fault handling process when the first fault handling process has not been completed and the existence of a designated controller is identified, so as to control the multi-controller storage system to continue to perform business processing. The designated controller is either the controller that triggered the recovery among multiple controllers, or the controller that does not belong to the designated mirror pair and triggered the fault among multiple controllers.
[0117] Using the above device, when a specified mirror pair is identified in at least one mirror pair, the cache state machine is controlled to execute a first fault handling process. The specified mirror pair is a pair containing two controllers that have both failed. If the first fault handling process is not completed and a specified controller is identified, a second fault handling process of the cache state machine is executed based on the execution phase of the first fault handling process to control the multi-controller storage system to continue business processing. The specified controller is either the controller that triggered recovery among multiple controllers, or the controller that does not belong to the specified mirror pair and has triggered a failure among multiple controllers. The cache state machine can immediately respond to the fault handling process in case of recovery of a surviving controller or failure of a new controller during the execution of the first fault handling process. This enables rapid recovery of business processing when faced with complex overlapping faults, reducing business interruption time and data processing latency. It solves the problem of long-term business interruption in multi-controller storage systems under fault scenarios in related technologies, significantly enhancing business continuity and system stability.
[0118] Optionally, the execution order of the processes in the first fault handling process is the interruption process, the first update process, and the recovery process, wherein the interruption process is used to indicate the interruption of the business processing of the multi-controller storage system, and the recovery process is used to indicate the recovery of the business processing of the multi-controller storage system; the second execution module 704 includes: a first determining unit, used to determine the cluster state of the multi-controller storage system when it is identified that the specified controller is the controller that triggered recovery among multiple controllers, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed; and a first execution unit, used to execute the image pair construction process when the cluster state of the multi-controller storage system is in a stable state, wherein the image pair construction process is used to indicate the interruption of the business processing of the multi-controller storage system. A set of candidate controllers in the system determines a first controller to form a first designated mirror pair with a designated controller. The first designated mirror pair is used to execute the services of a prior mirror pair that includes the designated controller and has a controller failure. The candidate controllers in the set of candidate controllers are controllers in the multi-controller storage system that are not in a faulty state. A second execution unit is used to execute a second update process after the mirror pair construction process is completed. The second update process is used to instruct the cluster status information of the multi-controller storage system to be updated. The cluster status information of the multi-controller storage system is used to indicate the online information of multiple controllers and the mirror pair status information of multiple controllers. A third execution unit is used to execute a recovery process after the second update process is completed.
[0119] Optionally, the first determining unit is further configured to: determine whether a second mirror pair exists among the multiple first mirror pairs based on the status information of the first mirror pairs among the multiple first mirror pairs, wherein the first mirror pairs among the multiple first mirror pairs include a specified controller, and the second mirror pair includes a mirror pair with a controller failure; if a second mirror pair exists, determine the cluster status of the multi-controller storage system based on the time of failure of the controller in the second mirror pair.
[0120] Optionally, the first determining unit is further configured to: obtain the failure time of the controllers in the second mirror pair other than the specified controller; determine the cluster state of the multi-controller storage system as stable if the failure time of the controllers in the second mirror pair other than the specified controller is earlier than the failure time of the specified controller; and determine the cluster state of the multi-controller storage system as unstable if the failure time of the controllers in the second mirror pair other than the specified controller is later than the failure time of the specified controller.
[0121] Optionally, the second execution module 704 further includes: a fourth execution unit, used to execute a third update process when the cluster state of the multi-controller storage system is in an unstable state, wherein the third update process is used to instruct the updating of the online information of the controllers in the multi-controller storage system; and a fifth execution unit, used to execute a recovery process after the third update process has been completed.
[0122] Optionally, the first execution unit is further configured to: obtain mirror pair status information of a candidate controller in a set of candidate controllers, wherein the mirror pair status information is used to indicate the mirror pair to which the candidate controller belongs; and determine a first controller from the set of candidate controllers based on the mirror pair status information of the candidate controllers in a set of candidate controllers, so as to form a first specified mirror pair with the specified controller.
[0123] Optionally, the mirror pair status information includes the number of mirror pairs and mirror pair attributes, whereby the mirror pair attributes indicate whether the controller in the mirror pair is a data initiator or a data receiver. The first execution unit is further configured to: determine the candidate controller with the fewest number of mirror pairs among a group of candidate controllers as the second controller; if there are not more than one second controller, determine the second controller as the first controller; if there are more than one second controller, determine the frequency at which the second controller is a data initiator among the multiple second controllers based on the mirror pair attributes and the number of mirror pairs among the multiple second controllers; and determine the second controller with the lowest frequency at which it is a data initiator from among the multiple second controllers as the first controller.
[0124] Optionally, the second execution module 704 further includes: a sixth execution unit, used to, after executing the image pair construction process, use the first controller as the data receiver, use the specified controller as the data initiator, and trigger the execution of the cached data re-mirroring process, wherein the cached data re-mirroring process is used to instruct the specified cached data in the specified controller to be transmitted to the first controller.
[0125] Optionally, the first execution unit is further configured to: determine the performance of a candidate controller in a set of candidate controllers; and determine a first controller from the set of candidate controllers based on the performance of the candidate controllers in the set of candidate controllers.
[0126] Optionally, the first execution unit is further configured to: send performance test commands to the candidate controllers in a group of candidate controllers respectively, and record the response time of the candidate controllers in the group of candidate controllers when executing the performance test commands, wherein the response time of the performance test commands is used to indicate the performance status of the candidate controllers in the group of candidate controllers, and the performance test commands include at least one of the following test commands: read performance test command, write performance test command, memory access performance test command, and network test command.
[0127] Optionally, the first execution unit is further configured to: when there are multiple performance test commands, perform a weighted summation of multiple response times corresponding to multiple performance test commands according to the command type of the performance test commands to obtain the total response time of the candidate controllers in a set of candidate controllers; and determine the controller with the smallest total response time of the candidate controllers in the set of candidate controllers as the first controller.
[0128] Optionally, the second execution module 704 further includes: a seventh execution unit, used to execute a fourth update process when the specified controller is a controller that does not belong to the specified mirror pair among multiple controllers and triggers a fault, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, wherein the fourth update process is used to indicate updating the online information of the controller in the multi-controller storage system; and an eighth execution unit, used to execute the recovery process when the fourth update process has been completed.
[0129] For a description of the features in the embodiment corresponding to the fault handling device of the multi-controller storage system, please refer to the relevant description of the embodiment corresponding to the fault handling method of the multi-controller storage system, which will not be repeated here.
[0130] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the fault handling method for a multi-controller storage system.
[0131] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the fault handling method for a multi-controller storage system when it is run.
[0132] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0133] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the fault handling method for a multi-controller storage system.
[0134] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the fault handling method for a multi-controller storage system.
[0135] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0136] The foregoing has provided a detailed description of a fault handling method and electronic device for a multi-controller storage system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A fault handling method for a multi-controller storage system, characterized in that, The multi-controller storage system includes multiple controllers and a cache state machine, the multiple controllers forming at least one mirror pair, and one mirror pair in the at least one mirror pair containing two controllers from the multiple controllers; the method includes: If a specified mirror pair is identified in the at least one mirror pair, the cache state machine is controlled to execute a first fault handling procedure, wherein the specified mirror pair is a mirror pair in which both controllers are faulty; If the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution phase of the first fault handling process to control the multi-controller storage system to continue business processing. The designated controller is either the controller that triggered the recovery among the multiple controllers, or the controller that does not belong to the designated mirror pair and triggered the fault among the multiple controllers.
2. The method according to claim 1, characterized in that, The execution order of the processes in the first fault handling process is the interruption process, the first update process, and the recovery process. The interruption process is used to indicate the interruption of the business processing of the multi-controller storage system, and the recovery process is used to indicate the restoration of the business processing of the multi-controller storage system. When the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution stage of the first fault handling process, including: If the specified controller is identified as the controller that triggered the recovery among the plurality of controllers, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, the cluster status of the multi-controller storage system is determined. When the cluster state of the multi-controller storage system is stable, a mirror pair construction process is executed. The mirror pair construction process is used to instruct the determination of a first controller from a group of candidate controllers in the multi-controller storage system to form a first designated mirror pair with the designated controller. The first designated mirror pair is used to execute the services of a prior mirror pair that includes the designated controller and has a controller failure. The candidate controllers in the group of candidate controllers are controllers in the multi-controller storage system that are not in a failure state. After the mirror pair construction process is completed, a second update process is executed, wherein the second update process is used to instruct the cluster status information of the multi-controller storage system to be updated, and the cluster status information of the multi-controller storage system is used to indicate the online information of the multiple controllers and the mirror pair status information of the multiple controllers. If the second update process is completed, the recovery process is then executed.
3. The method according to claim 2, characterized in that, Determining the cluster status of the multi-controller storage system includes: Based on the status information of the first mirror pairs among the multiple first mirror pairs, it is determined whether there is a second mirror pair among the multiple first mirror pairs, wherein the first mirror pairs among the multiple first mirror pairs include the designated controller, and the second mirror pair has a mirror pair with a controller failure; If the second mirror pair exists, obtain the fault time of the controller in the second mirror pair other than the specified controller; If the failure time of the controller in the second mirror pair other than the failure time of the specified controller is earlier than the failure time of the specified controller, the cluster state of the multi-controller storage system is determined to be a stable state. If the failure time of the controller in the second mirror pair, excluding the specified controller, is later than the failure time of the specified controller, the cluster state of the multi-controller storage system is determined to be unstable.
4. The method according to claim 2, characterized in that, The method further includes: When the cluster state of the multi-controller storage system is unstable, a third update process is executed, wherein the third update process is used to instruct the updating of the online information of the controllers in the multi-controller storage system; After the third update process is completed, the recovery process is executed.
5. The method according to claim 2, characterized in that, The step of determining the first controller from a set of candidate controllers in the multi-controller storage system includes: Obtain the mirror pair status information of the candidate controllers in the group of candidate controllers, wherein the mirror pair status information is used to indicate the mirror pair to which the candidate controller belongs; Based on the mirror pair status information of the candidate controllers in the set of candidate controllers, the first controller is determined from the set of candidate controllers to form the first designated mirror pair with the designated controller.
6. The method according to claim 5, characterized in that, The mirror pair status information includes the number of mirror pairs and mirror pair attributes. The mirror pair attributes are used to indicate whether the controller in the mirror pair is a data initiator or a data receiver. The step of determining the first controller from the set of candidate controllers based on the mirror pair state information of the candidate controllers in the set of candidate controllers includes: The candidate controller with the fewest mirror pairs among the mirror pairs in the group of candidate controllers is determined as the second controller; If the number of second controllers is not multiple, the second controller is identified as the first controller; When there are multiple second controllers, the frequency at which the second controller is the data initiator is determined based on the mirror pair attribute of the second controller among the multiple second controllers and the number of mirror pairs of the second controller among the multiple second controllers; From among the multiple second controllers, the second controller with the lowest frequency of being the data initiator is determined as the first controller.
7. The method according to claim 6, characterized in that, After the image pair construction process is executed, the method further includes: The first controller is used as the data receiver, the designated controller is used as the data initiator, and the cache data re-mirroring process is triggered, wherein the cache data re-mirroring process is used to instruct the designated cache data in the designated controller to be transmitted to the first controller.
8. The method according to claim 2, characterized in that, The step of determining the first controller from a set of candidate controllers in the multi-controller storage system includes: A performance test command is sent to each of the candidate controllers in the group of candidate controllers, and the response time of the candidate controllers in the group of candidate controllers executing the performance test command is recorded. The response time of the performance test command is used to indicate the performance of the candidate controllers in the group of candidate controllers. The performance test command includes at least one of the following test commands: read performance test command, write performance test command, memory access performance test command, and network test command. When there are multiple performance test commands, the response times corresponding to the multiple performance test commands are weighted and summed according to the command type of the performance test commands to obtain the total response time of the candidate controllers in the group of candidate controllers; The controller with the smallest total response time among the candidate controllers in the group of candidate controllers is determined as the first controller.
9. The method according to claim 2, characterized in that, When the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution stage of the first fault handling process, including: If the designated controller is a controller that does not belong to the designated mirror pair among the plurality of controllers and has triggered a fault, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, then the fourth update process is executed, wherein the fourth update process is used to indicate updating the online information of the controllers in the multi-controller storage system. If the fourth update process is completed, the recovery process is then executed.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the fault handling method for the multi-controller storage system as described in any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Multi-control storage cluster operation fault quick self-recovery method and device and storage medium
CN112583909A
Double-control equipment fault processing method and device
CN112911185A
Multi-control storage system, data processing method and device and medium
CN114756176A
Cache data processing method and system under four-control storage device fault
CN115237683A
Data caching method and device, equipment and storage medium
CN115563028A
Cited By
Data storage method and electronic equipment
CN121501572A
Data storage method and electronic device
CN121501572B