Fault handling method of multi-controller storage system and electronic device
By identifying mirror pair failures in a multi-controller storage system and executing the first and second fault handling processes of the cache state machine, the problem of prolonged business interruption caused by simultaneous mirror pair failures is solved, enabling rapid recovery of business processing and system stability.
Patent Information
- Application Number
- CN202511463779.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-10-14
AI Technical Summary
In a multi-controller storage system, when two controllers in the same mirror pair fail simultaneously, business processing is suspended, resulting in prolonged business interruption and affecting system response speed and performance.
In the event of a mirror pair failure, the system controls the cache state machine to execute a first failure handling procedure, and immediately executes a second failure handling procedure before the first failure handling is completed, in response to the recovery of the live controller or the failure of the new controller, to ensure that the system continues to process business.
When faced with a complex set of overlapping faults, it can quickly restore business processing, reduce business interruption time and data processing latency, and improve business continuity and system stability.
Smart Images

Figure CN120929313B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data storage, and particularly relates to a fault processing method of a multi-controller storage system and an electronic device. BACKGROUND
[0002] In the related art, in a multi-controller storage system, controllers are in the form of mirror pairs to improve data redundancy and system availability. The two controllers in a mirror pair back up each other's cache data. When one of the controllers fails, the other controller can take over its work and maintain the continuity of data access. However, when the two controllers in the same mirror pair fail at the same time, business processing is suspended until the failed controller is recovered. However, under the multi-controller storage system, even if other controllers are still healthy, business is stopped, which leads to a long time of business interruption and affects the response speed and performance of the system.
[0003] Therefore, the related art has the problem of long-time business interruption of the multi-controller storage system under a fault scenario. SUMMARY
[0004] The present application provides a fault processing method of a multi-controller storage system and an electronic device to at least solve the problem of long-time business interruption of the multi-controller storage system under a fault scenario in the related art.
[0005] The present application provides a fault processing method of a multi-controller storage system, the multi-controller storage system comprising a plurality of controllers and a cache state machine, the plurality of controllers forming at least one mirror pair, one mirror pair in the at least one mirror pair comprising two controllers in the plurality of controllers; the method comprising:
[0006] In a case where it is identified that there is a specified mirror pair in the at least one mirror pair, controlling the cache state machine to execute a first fault processing procedure, wherein the specified mirror pair is a mirror pair in which both controllers comprised are faulty;
[0007] In a case where it is identified that there is a specified controller, based on an execution stage of the first fault processing procedure, executing a second fault processing procedure of the cache state machine to control the multi-controller storage system to continue business processing, wherein the specified controller is a controller in the plurality of controllers that triggers recovery, or a controller in the plurality of controllers that triggers failure and does not belong to the specified mirror pair.
[0008] The application further provides a fault processing device of a multi-controller storage system, the multi-controller storage system comprising a plurality of controllers and a cache state machine, the plurality of controllers forming at least one mirror pair, one mirror pair in the at least one mirror pair comprising two controllers in the plurality of controllers; the device comprises:
[0009] a first execution module configured to, in a case where it is identified that a specified mirror pair exists in the at least one mirror pair, control the cache state machine to execute a first fault processing procedure, wherein the specified mirror pair is a mirror pair in which both controllers are faulty.
[0010] a second execution module configured to, in a case where the first fault processing procedure is not executed completely and a specified controller exists, execute a second fault processing procedure of the cache state machine based on an execution stage of the first fault processing procedure, so as to control the multi-controller storage system to continue business processing, wherein the specified controller is a controller in the plurality of controllers triggering recovery or a controller in the plurality of controllers triggering fault and not belonging to the specified mirror pair.
[0011] The application further provides an electronic device, comprising a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of the fault processing method of the multi-controller storage system.
[0012] The application further provides a computer readable storage medium, the computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the steps of the fault processing method of the multi-controller storage system.
[0013] The application further provides a computer program product, comprising a computer program, the computer program being executed by a processor to implement the steps of the fault processing method of the multi-controller storage system.
[0014] Through the application, in a case where it is identified that there is a specified mirror pair in the at least one mirror pair, a first fault handling process is performed by the cache state machine, where the specified mirror pair is a mirror pair in which both controllers are faulty; in a case where the first fault handling process is not executed completely and it is identified that there is a specified controller, a second fault handling process of the cache state machine is executed based on an execution stage of the first fault handling process, to control the multi-control storage system to continue business processing, where the specified controller is a controller in the multiple controllers that triggers recovery, or is a controller in the multiple controllers that does not belong to the specified mirror pair and triggers fault. Through the cache state machine, the fault handling process can be immediately responded to in a case where the recovery of a surviving controller or the fault of a new controller occurs during execution of the first fault handling process, so that the business processing can be quickly recovered when facing complex fault superposition, the business interruption time and data processing delay are reduced, and the problem of long business interruption of the multi-control storage system in a fault scenario in the related art is solved, and the business continuity and system stability are significantly enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0016] Figure 1 is a hardware structure block diagram of a fault handling method of a multi-control storage system according to an embodiment of the present application.
[0017] Figure 2 is a flowchart of a fault handling method of a multi-control storage system according to an embodiment of the present application.
[0018] Figure 3 is a schematic diagram of a fault handling process according to an embodiment of the present application.
[0019] Figure 4 is a schematic diagram of a multi-control storage system according to an embodiment of the present application.
[0020] Figure 5 is a schematic diagram of a fault handling method of a multi-control storage system according to an embodiment of the present application.
[0021] Figure 6 is a schematic diagram of another fault handling method of a multi-control storage system according to an embodiment of the present application.
[0022] Figure 7 is a structure block diagram of a fault handling device of a multi-control storage system according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be described clearly and completely in the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.
[0024] It should be noted that, in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0025] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0026] In combination with the specific application environment architecture or specific hardware architecture on which the fault processing method of the multi-control storage system depends, the specific application environment architecture or specific hardware architecture is described here.
[0027] The method embodiments provided in the embodiments of the present application can be executed in a server device or similar computing device. Taking the case of running on a server device, Figure 1 is a hardware structure block diagram of the fault processing method of the multi-control storage system in the embodiments of the present application. As Figure 1 shown, the server device can include one or more (only one is shown in Figure 1 The processor 102 (the processor 102 can include but is not limited to a processing device such as a central processing unit CPU, a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the above-mentioned server device can further include a transmission device 106 for communication function and an input and output device 108. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned server device. For example, the server device can further include more or less components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0028] The memory 104 can be used to store computer programs, such as software programs of application software and modules, for example, a computer program corresponding to the fault processing method of the multi-controller storage system in the embodiments of the present application. The processor 102 performs various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to a server device through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0029] The transmission device 106 is used to receive or send data via a network. The specific examples of the above network can include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF) module which is used to communicate with the Internet in a wireless manner.
[0030] The embodiments of the present application provide a fault processing method of a multi-controller storage system. The method is described in detail in combination with the execution flow of the fault processing method of the multi-controller storage system.
[0031] In the embodiments, a fault processing method of a multi-controller storage system is provided, Figure 2 is a flowchart of the fault processing method of the multi-controller storage system according to the embodiments of the present application, as Figure 2 shown, the method includes the following steps:
[0032] In step S202, in a case where it is identified that there is a specified mirror pair in the at least one mirror pair, a first fault processing flow is executed by controlling the cache state machine, wherein the specified mirror pair is a mirror pair in which both controllers contained therein are faulty.
[0033] In step S204, in a case where the first fault processing flow is not executed completely and it is identified that there is a specified controller, a second fault processing flow of the cache state machine is executed based on an execution stage of the first fault processing flow, so as to control the multi-controller storage system to continue business processing, wherein the specified controller is a controller in the plurality of controllers which triggers recovery, or is a controller in the plurality of controllers which does not belong to the specified mirror pair and triggers a fault.
[0034] Through the above steps, since the cache state machine can respond to the recovery of the surviving controller or the fault handling process under the condition of failure of the new controller during the execution of the first fault handling process, the business processing can be quickly recovered when facing complex fault superposition, the business interruption time and data processing delay are reduced, and the problem of long business interruption of the multi-control storage system in the related art under the fault scenario is solved, and the business continuity and system stability are significantly enhanced.
[0035] It should be noted that the multi-control storage system can include multiple controllers, a cache state machine and a cache module. Among them, the multiple controllers can be the core components of the multi-control storage system, and can cooperate together to improve the performance and availability of the multi-control storage system. The multiple controllers can constitute a cluster of the multi-control storage system. Each of the multiple controllers can be responsible for a part of the (input / output) operation and data management task. The cache state machine can be a component for managing the state and process of the cache module, and can change the behavior of the cache module according to system events and controller states. The cache module can be a component in the multi-control storage system for improving data access speed and efficiency. The multiple controllers in the multi-control storage system can form at least one mirror pair. One mirror pair can include two controllers of the multiple controllers.
[0036] In the embodiment provided in step S202, the specified mirror pair can be a mirror pair in which both controllers fail, and the first fault handling process can be a preset fault handling process in the multi-control storage system. In the case where there is a specified mirror pair in at least one mirror pair, the control cache state machine is triggered to execute the first fault handling process.
[0037] Optionally, in the case where there is a specified mirror pair in at least one mirror pair, a corresponding fault handling event is triggered to be generated. The cache state machine receives and responds to the fault handling event, and triggers the execution of the first fault handling process, wherein the first fault handling process at least includes a suspension process to control the suspension of the business processing of the fault controller.
[0038] Optionally, each controller sends a heartbeat message to other controllers in the system in each designated heartbeat period, and returns an acknowledgement after receiving the heartbeat message. The acknowledgement proves that the sending controller is active in the last designated heartbeat period, and also indicates that the receiving controller itself has the ability to respond to external communication. In the designated heartbeat period, if the controller fails to receive the heartbeat acknowledgement of its mirror or other controllers in the cluster, the controller is marked as a potential fault state. In order to reduce false positives, a fault tolerance threshold can be set, that is, after failing to receive the acknowledgement for several consecutive heartbeat periods, the controller is formally marked as a fault. In a multi-control storage system, the presence of a designated mirror in at least one mirror pair can be determined through the heartbeat message and the fault tolerance threshold. In the case where the designated mirror is determined to exist in at least one mirror pair, the first fault handling process of the cache state machine is triggered to implement the silent cache module, so as to stop the business processing related to the designated mirror pair and prevent data inconsistency.
[0039] In the embodiment provided in step S204, in the process of controlling the cache state machine to execute the first fault handling process, in the case where the presence of the designated controller is detected, the cache state machine does not simply wait for the completion of the first process, but immediately starts the second fault handling process based on the execution stage of the first fault handling process, to ensure that the system can continue the business processing.
[0040] Among them, the designated controller can be the controller triggering the recovery event or the controller triggering the fault (not the controller belonging to the same mirror pair as the two controllers with faults).
[0041] Optionally, when the failed controller is repaired and back online, it can respond to the heartbeat detection of other controllers to determine the recovery of the failed controller, or when the failed controller is repaired and back online, it can actively send a heartbeat signal to inform other components in the system (such as the cluster manager, the cache state machine, etc.) that its state has changed from fault to active. Or when the failed controller is repaired and back online, it can report state information to the cache state machine, and the cache state machine determines the recovery of the failed controller based on the state information.
[0042] In one example, assuming that both controller A and controller B in the designated mirror pair fail, the control cache state machine executes the first fault handling process, and in the execution of the first fault handling process, the failure of controller C is detected. At this time, controller C becomes a "designated controller" because it does not belong to the "designated mirror pair" composed of A and B. The cache state machine executes the second fault handling process according to the current processing stage, instead of waiting for the completion of the first fault handling process, which can ensure that the system can quickly recover the business processing when facing complex fault superposition scenarios, and avoid unnecessary business interruption and data delay.
[0043] By embodiments of the present application, in a case where it is identified that there is a specified mirror pair in the at least one mirror pair, a first fault handling procedure of the cache state machine is controlled to be executed, where the specified mirror pair is a mirror pair in which both controllers included in the mirror pair are faulty; in a case where the first fault handling procedure is not executed completely and it is identified that there is a specified controller, a second fault handling procedure of the cache state machine is executed based on an execution stage of the first fault handling procedure, so as to control the multi-control storage system to continue service processing, where the specified controller is a controller in the plurality of controllers that triggers recovery, or is a controller in the plurality of controllers that triggers fault and does not belong to the specified mirror pair. By the cache state machine, the fault handling procedure can be immediately responded in a case where the recovery of the surviving controller or the fault of the new controller occurs during execution of the first fault handling procedure, so that the service processing can be quickly recovered when facing complex fault superposition, the service interruption time and the data processing delay are reduced, and the problem of long service interruption of the multi-control storage system in the related art under the fault scenario is solved, and the business continuity and the stability of the system are significantly enhanced.
[0044] As an optional implementation, the execution order of the procedures in the first fault handling procedure is an interruption procedure, a first update procedure, and a recovery procedure, where the interruption procedure is used to instruct to interrupt the service processing of the multi-control storage system, and the recovery procedure is used to instruct to recover the service processing of the multi-control storage system. In step S204, in a case where it is identified that the specified controller is a controller in the plurality of controllers that triggers recovery, and the execution stage of the first fault handling procedure indicates that the first update procedure is executed completely but the recovery procedure is not executed, the cluster state of the multi-control storage system is determined; in a case where the cluster state of the multi-control storage system is in a stable state, a mirror pair construction procedure is executed, where the mirror pair construction procedure is used to instruct to determine a first controller from a group of candidate controllers of the multi-control storage system to constitute a first specified mirror pair with the specified controller, where the first specified mirror pair is used to execute service of a previous mirror pair that includes the specified controller and in which there is controller fault, and a candidate controller in the group of candidate controllers is a controller in the multi-control storage system that is not in a fault state; in a case where the mirror pair construction procedure is executed completely, a second update procedure is executed, where the second update procedure is used to instruct to update cluster state information of the multi-control storage system, and the cluster state information of the multi-control storage system is used to instruct online information of the plurality of controllers and mirror pair state information of the plurality of controllers; and in a case where the second update procedure is executed completely, the recovery procedure is executed.
[0045] It should be noted that the first fault processing procedure can include an interruption procedure, a first updating procedure and a recovery procedure, wherein the interruption procedure can be used to instruct to interrupt the service processing of the multi-control storage system, and the recovery procedure can be used to instruct to recover the service processing of the multi-control storage system. Specifically, the first fault processing procedure can refer to Figure 3 After the cache state machine receives the event (such as a fault event) notification, the cache state machine first interrupts the service processing of the multi-control storage system by initiating an interruption (such as queisce) request to ensure the safety of the current operation. Subsequently, after the state machine receives the interruption request, the service mode switching is performed to send an updating (such as an ack) request to the service processing end to realize the service mode switching. Finally, after the state machine receives the updating request, a recovery (such as resume) request is initiated to instruct the service processing end to recover the service processing of the multi-control storage system, and after the service processing end recovers the service processing of the multi-control storage system, the service recovery completion is fed back to the state machine.
[0046] Optionally, in a case where it is determined that there is a specified mirror pair (a mirror pair in which both controllers fail) in the at least one mirror pair, a fault notification event of the cluster is triggered, the cache state machine receives the fault notification event of the cluster, confirms that there is a mirror pair failure, and performs an interruption procedure to suspend the service processing of the cache module. After the interruption procedure, the cache state machine performs a first updating procedure (an acknowledgement stage) to update the cluster information to realize the synchronization of the cluster information.
[0047] Optionally, in a case where it is indicated in the execution stage of the first fault processing procedure that the first updating procedure is completed but the recovery procedure is not performed, if it is determined that there is a specified controller (such as a controller that was previously failed and is recovered), the cluster state of the multi-control storage system is evaluated, for example, whether there is no new failure or unstable factor in the multi-control storage system.
[0048] Optionally, in a case where the specified controller is identified as a controller that triggers recovery among the plurality of controllers, and it is indicated in the execution stage of the first fault processing procedure that the first updating procedure is completed but the recovery procedure is not performed, the second fault processing procedure can include a mirror pair construction procedure, a second updating procedure and a recovery procedure, that is, in a case where the specified controller is identified as a controller that triggers recovery among the plurality of controllers, and it is indicated in the execution stage of the first fault processing procedure that the first updating procedure is completed but the recovery procedure is not performed, the entire fault processing procedure corresponds to the order of: the interruption procedure, the first updating procedure, the mirror pair construction procedure, the second updating procedure and the recovery procedure.
[0049] In a case where it is determined that the cluster state of the multi-control storage system is in a stable state, a mirror pair construction process is performed to select a controller from a set of candidate controllers that are not in a failure state to construct a new mirror pair with the specified controller, i.e., a first specified mirror pair. The first specified mirror pair can be used to perform services of a previous mirror pair that includes the specified controller and in which a controller failure exists. Optionally, in the process of selecting a controller from a plurality of candidate controllers that are not in a failure state, a determination can be made based on load balancing, performance evaluation, and mirror pair grouping rules to ensure data redundancy and service processing capability.
[0050] Optionally, the online information of the plurality of controllers can be used to indicate whether a controller in the multi-control storage system is online. In a case where a controller failure exists, the recorded online information is in an offline state. In a case where a controller failure does not exist, the recorded online information is in an online state. The mirror pair state information of the plurality of controllers can be used to indicate the mirror pair to which each controller in the plurality of controllers belongs. One controller can belong to at least one mirror pair, i.e., one controller can belong to multiple mirror pairs. Taking a four-control storage system as an example, as shown in FIG. 1, there are controller 1, controller 2, controller 3, and controller 4. Controller 1 can form a mirror pair 1 with controller 2, controller 2 can form a mirror pair 2 with controller 3, controller 3 can form a mirror pair 3 with controller 4, and controller 4 can form a mirror pair 4 with controller 1. Different services can be performed in different mirror pairs to which one controller belongs. Figure 4
[0051] Optionally, the set of candidate controllers can be controllers in the multi-control storage system that are not in a failure state and whose performance conditions satisfy a specified performance condition. Optionally, a first pre-trained model can be deployed in each controller in the multi-control storage system. The performance condition of a controller can be determined based on a reflection time length between predicted data output by the first pre-trained model after the training data in a training set is input to the first pre-trained model. The size of the reflection time length is negatively correlated with the performance of the controller. The longer the reflection time length, the poorer the performance of the controller. The shorter the reflection time length, the better the performance of the controller, which can be quickly reflected. At this time, the specified performance condition can be a specified reflection time length. Controllers in the multi-control storage system that are not in a failure state and whose reflection time lengths are less than the specified reflection time length can be selected as candidate controllers. This can ensure that the selected controllers can meet the needs of constructing a new mirror pair in terms of performance and avoid service processing delays or instability caused by low-performance controllers.
[0052] Optionally, the first pre-trained model can or can not be a performance identification model. In the multi-controller storage system, the first pre-trained model can be at least one of the following functional models: a fault prediction model, a workload prediction model, and a security prediction model. The fault prediction model can predict potential faults of controllers or other components in the multi-controller storage system. By learning from historical fault data and system operating states, the model can predict potential faults in advance, allowing maintenance personnel to take preventive measures. The workload prediction model can be used to predict the workload of the multi-controller storage system in a future period (i.e., a specified time period), including the number, size, and type of read and write operations. This prediction helps the multi-controller storage system to adjust resource allocation in advance to meet the expected performance requirements and avoid service degradation due to sudden workload spikes. The security prediction model can be used to predict potential threats such as attacks or abnormal access patterns by analyzing controller log information and network traffic.
[0053] Optionally, after the first specified mirror pair is built, the cache state machine can execute a second update process to update the online controller information and mirror pair state information in the multi-controller storage system to reflect the latest cluster architecture and data layout. In the case where a controller is in a fault state, the online information of the corresponding controller is displayed as an offline state.
[0054] In one example, assume that the multi-controller storage system is a four-controller storage system with controller 1, controller 2, controller 3, and controller 4. In the case where the mirror pair formed by controller 1 and controller 2 fails, the first fault handling process can be started. After the interruption process and the first update process are executed, but before the recovery process is executed, controller 2 is restored. Upon detecting this change, the cache state machine can determine whether the cluster state of the multi-controller storage system is stable by determining whether there is a new controller fault or a specified instability factor in the current cluster. If the cluster state is determined to be stable, a new mirror pair (e.g., the first specified mirror pair) can be formed by selecting one of controller 3 and controller 4 to work with controller 2 to ensure rapid recovery of data redundancy and business processing capacity. After the new mirror pair is built, the cache state machine can execute the second update process to update the internal state information to reflect the latest online controller list and mirror pair state. Subsequently, the cache state machine enters the recovery process to reactivate business processing, ensuring that data read and write operations can be quickly restored and business processing continuity is ensured.
[0055] Through the embodiment, by executing the second fault handling process at a proper stage of the first fault handling process, the controller recovery event can be responded in time, long-term interruption of service processing is avoided, and service continuity is improved. After the controller recovery, it is firstly judged whether the cluster state is in a stable state, and only after confirming that the system is stable, the service recovery process is executed, so that the risk of data processing in an unstable state is avoided. Through the mirror pair construction process, a suitable controller can be quickly determined from a group of healthy candidate controllers, a new mirror pair is constructed with the recovered controller, and rapid recovery of data redundancy and service processing capacity is ensured.
[0056] As an optional implementation, determining the cluster state of the multi-control storage system comprises: determining, according to state information of a first mirror pair in a plurality of first mirror pairs, whether a second mirror pair exists in the plurality of first mirror pairs, wherein the first mirror pair in the plurality of first mirror pairs comprises a specified controller, and the second mirror pair is a mirror pair in which a controller fails; and in the case where the second mirror pair exists, determining the cluster state of the multi-control storage system according to a failure time of the controller in the second mirror pair.
[0057] It should be noted that the first mirror pair can be a mirror pair for data redundancy and concurrent processing composed of the specified controller and another controller in the multi-control storage system. The second mirror pair can be a mirror pair in which a controller of the plurality of first mirror pairs fails.
[0058] The failure time of the controller can be used to indicate the time point at which the controller fails, and can be used to indicate the influence range of the failure and the degree of system recovery.
[0059] Optionally, the cluster state of the multi-control storage system can comprise a stable state and an unstable state. In the case where the second mirror pair exists, it is determined whether the current cluster is in a stable state according to the failure time of the controller in the second mirror pair. The unstable state can indicate that there is a state in which service processing cannot be maintained in the current cluster. Generally, the stall state can be used to identify the unstable state of the cluster.
[0060] Optionally, a more accurate failure detection mechanism can be introduced, such as a failure backtracking technology based on event logs, to more accurately determine the failure time of the controller and further improve the accuracy of the cluster state determination.
[0061] Optionally, in the case where there is no second mirror pair in the plurality of first mirror pairs, the cluster information is updated using the prior mirror pair information of the specified controller to restore the service processing function of the mirror pair in which the specified controller was previously located. In the case where there is a second mirror pair in the plurality of first mirror pairs, whether the service processed by the specified controller in advance can be restored is determined through the failure time of the controller other than the specified controller in the second mirror pair and the failure time of the specified controller, and in the case where it can be restored, a mirror pair construction process is performed.
[0062] Through the embodiment, by accurately determining the second mirror pair and the failure time thereof, the state of the self can be more accurately evaluated, and data processing in an unstable or potentially dangerous state is avoided.
[0063] As an optional implementation, determining the cluster state of the multi-control storage system according to the failure time of the controller in the second mirror pair includes: obtaining the failure time of the controller other than the specified controller in the second mirror pair; in the case where the failure time of the controller other than the specified controller in the second mirror pair is earlier than the failure time of the specified controller, determining that the cluster state of the multi-control storage system is a stable state; and in the case where the failure time of the controller other than the specified controller in the second mirror pair is later than the failure time of the specified controller, determining that the cluster state of the multi-control storage system is an unstable state.
[0064] It should be noted that in the case where the failure time of the controller other than the specified controller in the second mirror pair is earlier than the failure time of the specified controller, at this time, the failure time point of the specified controller is later than the failure time point of the controller other than the specified controller in the second mirror pair, it is determined that the specified controller in the second mirror pair has processed the service of the controller other than the specified controller in the second mirror pair, and at this time, there is no case where the service is not recorded, so it can be indicated that there is no case where the service processing cannot be maintained in the multi-control storage system, and the cluster state of the multi-control storage system is determined to be a stable state. At this time, a new mirror pair can be constructed for the specified controller to restore the service processing of the second mirror pair.
[0065] In the case where the controller failure in the second mirror pair occurs after the failure of the specified controller, at this time, since the specified controller is restored, but the cache data recorded between the failure time of the controller other than the specified controller in the second mirror pair and the failure time of the specified controller is recorded in the controller other than the specified controller in the second mirror pair, at this time, the cluster state of the multi-control storage system is determined to be an unstable state, and therefore only the online information of the controller other than the specified controller in the second mirror pair is restored to inform other controllers in the cluster that the specified controller is restored.
[0066] Through the embodiment, by accurately judging the cluster state, the image pair is avoided to be constructed in the unstable state, thereby unnecessary data synchronization and business interruption are reduced, and the stability and efficiency of the system are improved.
[0067] As an optional implementation, the method further includes: in a case where the cluster state of the multi-control storage system is in the unstable state, performing a third update process, where the third update process is used to instruct to update the online information of the controller in the multi-control storage system; and in a case where the third update process is performed completely, performing a recovery process.
[0068] It should be noted that, in a case where the cluster state of the multi-control storage system is in the unstable state, only the online information of the controller in the second image pair except the specified controller is recovered, to inform other controllers in the cluster that the specified controller is recovered. Of course, in the multi-control storage system, there can be multiple image pairs for one controller, and each image pair can process different services. In a case where there are other image pairs except the second image pair for the specified controller, the service processing function of the specified controller for the other image pairs except the second image pair can be recovered after the online information of the controller (such as the specified controller) in the multi-control storage system is updated.
[0069] Through the embodiment, by forcibly updating the online information of the controller in the unstable state, the real-time performance of the cluster state is ensured, accurate information support is provided for subsequent business recovery, and the response speed and efficiency of the system are improved.
[0070] As an optional implementation, the first controller is determined from a group of candidate controllers in the multi-control storage system, including: obtaining image pair state information of a candidate controller in the group of candidate controllers, where the image pair state information is used to indicate an image pair to which the candidate controller belongs; and determining the first controller from the group of candidate controllers according to the image pair state information of the candidate controller in the group of candidate controllers, to form a first specified image pair with the specified controller.
[0071] It should be noted that the image pair state information can refer to information of the current image pair to which the candidate controller belongs, including but not limited to an image pair number, a state of another controller in the image pair, and the like. Alternatively, by obtaining the state information of the candidate controller, one candidate controller, i.e., the first controller, is selected from the group of candidate controllers to form the first specified image pair with the specified controller, and the first specified image pair can be used to perform the service of the previous image pair (such as the second image pair) including the specified controller and having a controller failure.
[0072] Optionally, the mirror pair state information of the candidate controller in the group of candidate controllers is parsed to select the first controller from the group of candidate controllers to form a first specified mirror pair with the specified controller.
[0073] Through the embodiment, by analyzing the mirror pair state information of the candidate controller, the best controller, i.e., the first controller, is intelligently selected to form a mirror pair with the specified controller, avoiding the problem of insufficient data redundancy or performance bottleneck caused by improper mirror pair construction, and improving the data security and processing efficiency of the system.
[0074] As an optional implementation, the mirror pair state information includes a mirror pair number and a mirror pair attribute, and the mirror pair attribute is used to indicate whether the controller in the mirror pair is a data initiator or a data receiver; and the first controller is determined from the group of candidate controllers according to the mirror pair state information of the candidate controller in the group of candidate controllers, including: determining the candidate controller with the least mirror pair number of the mirror pair to which the candidate controller belongs in the group of candidate controllers as the second controller; in the case where the number of the second controllers is not multiple, determining the second controller as the first controller; in the case where the number of the second controllers is multiple, determining the frequency of the second controller being a data initiator according to the mirror pair attribute of the second controller in the multiple second controllers and the mirror pair number of the second controller in the multiple second controllers; and determining the second controller with the minimum frequency of being a data initiator from the multiple second controllers as the first controller.
[0075] It should be noted that the mirror pair number can refer to the number of mirror pairs to which a controller belongs. It can be used to screen out the controller participating in the least mirror pair, i.e., the controller bearing less data redundancy task, which can achieve the effect of load balancing by reducing data synchronization and mirror burden. For example, in a multi-control storage system, there are currently three controllers as candidates (controllers 2, 3, and 4), among which controller 2 belongs to one mirror pair, controller 3 belongs to two mirror pairs, and controller 4 also belongs to two mirror pairs. Then, controller 2 is selected to form a mirror pair with the specified controller.
[0076] Optionally, in the case where the number of the second controllers is only one, the second controller is directly determined as the first controller, i.e., as the role of finally joining the mirror pair (data initiator or receiver).
[0077] Optionally, in the case where the number of the second controllers is multiple, the first controller is determined through the proportion of the candidate controller in the multiple second controllers being a data initiator. Specifically, the controller with a lower data initiator proportion means that the number of times of participating in data synchronization is relatively small, and thus it is more suitable to become a new data initiator to maintain the load balancing of the system.
[0078] Optionally, the performance bottleneck problem that may occur in the changing business scenario can also be solved by introducing an adaptive mirror pair attribute adjustment mechanism to dynamically optimize the construction strategy of the mirror pair. Optionally, the performance index parameters of each controller, such as CPU utilization, I / O delay, etc., and the change of the business mode are detected in real time. The performance index parameters of each controller are input into the pre-deployed attribute recognition model to obtain the predicted attributes of each controller. Based on the predicted attributes of each controller, the mirror pair attributes of each controller are determined. After adjusting the mirror pair attributes of the controller, the re-mirroring process is started to transmit the data from the original initiator to the new receiver, ensuring the data redundancy.
[0079] Through the embodiment, by comprehensively considering the number of mirror pairs and the mirror pair attributes, the best controller, i.e., the first controller, can be intelligently selected to be a mirror pair with the specified controller, avoiding the data redundancy deficiency or performance bottleneck problem caused by improper mirror pair construction, and improving the data security and processing efficiency of the system.
[0080] As an optional implementation, after the mirror pair construction process is performed, the above method further includes: taking the first controller as a data receiver, taking the specified controller as a data sender, and triggering the execution of a cache data re-mirroring process, wherein the cache data re-mirroring process is used to indicate the transmission of the specified cache data in the specified controller to the first controller.
[0081] It should be noted that in the cache data re-mirroring process, the specified controller can be used to send cache data to the first controller. The specified controller can be the data initiator in the original mirror pair, or a controller whose data storage responsibility needs to be adjusted due to system state changes. The cache data re-mirroring process refers to the process of copying the cache data of the specified controller (such as the data initiator) to the first controller (the data receiver) in the multi-control storage system, which can ensure that the data redundancy and system high availability can be maintained even in the case of controller failure, recovery or system state change.
[0082] Optionally, the specified cache data in the specified controller can be the cache data recorded by the specified controller when performing a business event between the failure time of the other controller in the second mirror pair (the previous mirror pair of the specified controller) and the failure time of the specified controller.
[0083] Optionally, in order to improve the efficiency of data transmission, data compression can be used to compress the specified cache data, reducing the bandwidth requirement of data transmission and solving the data transmission delay problem that may occur under limited network resources. Further, in order to ensure the security during data transmission, encryption technology can be used to compress the specified cache data after data compression.
[0084] Through this embodiment, through the data re-mirroring process, data redundancy and consistency can be guaranteed even when the controller state changes, avoiding potential data loss risks.
[0085] In one example, it is assumed that the initial mirror pair is composed of controllers 0 and 1, but after the stall state, controller 0 recovers first, while controller 1 has not recovered. At this time, controller 0 is regarded as the designated controller and will take on the role of transmitting cache data. In the case of selecting controller 2 as the first controller from a group of candidate controllers, when controller 0 determines to transmit data to controller 2, the cache module will trigger the re-mirroring process. In this process, controller 0 will copy a copy of the data in its cache to controller 2, thereby establishing a new mirror pair relationship (0, 2) to replace the original (0, 1) mirror pair.
[0086] As an optional implementation, determining the first controller from a group of candidate controllers in the multi-controller storage system includes: determining the performance of the candidate controllers in the group of candidate controllers; and determining the first controller from the group of candidate controllers according to the performance of the candidate controllers in the group of candidate controllers.
[0087] It should be noted that the performance can include but is not limited to CPU utilization, I / O read / write rate, memory access speed, and network bandwidth, etc. Through these indicators, the real-time processing capacity and service carrying capacity of the controller can be evaluated. Alternatively, the performance of each candidate controller can be analyzed by reading the performance of the candidate controllers in the group of candidate controllers, so as to determine the first controller from the group of candidate controllers to form the first designated mirror pair with the designated controller.
[0088] Alternatively, a performance prediction model can be used to predict the performance trend of the controller in advance, so as to adjust the construction strategy of the mirror pair.
[0089] Through this embodiment, by analyzing the performance of the candidate controllers, the controller with the best performance is selected as part of the mirror pair, ensuring the efficiency and reliability of the mirror pair construction, and improving the processing efficiency and stability of the system.
[0090] As an optional implementation, the performance of the candidate controller in the group of candidate controllers is determined, including: sending a performance test command to each candidate controller in the group of candidate controllers respectively, and recording the response time of the candidate controller in the group of candidate controllers executing the performance test command, wherein the response time of the performance test command is used to indicate the performance of the candidate controller in the group of candidate controllers, and the performance test command includes at least one of the following test commands: read performance test command, write performance test command, memory access performance test command, network test command.
[0091] It should be noted that the performance test command can be a command sent to the candidate controller, which can test and evaluate the performance indicators of the candidate controller, wherein the performance test command includes at least one of the following test commands: read performance test command, write performance test command, memory access performance test command, network test command. Among them, the read performance test command can be used to test the speed of the candidate controller reading data, to measure its read I / O request capability, and is commonly used to evaluate data access delay. The write performance test command can be used to test the speed of the candidate controller writing data to the storage medium, to measure its write I / O request capability. The memory access performance test command can be used to test the efficiency of the controller accessing and managing memory resources, especially the execution speed of cache operations. The network test command can be used to evaluate the performance of the controller in network communication, including data transmission speed, network delay and bandwidth utilization. The response time can be the time required for the controller to return the result after executing the performance test command, reflecting its speed and efficiency in processing such tasks.
[0092] Through the performance test command, the performance indicators of the candidate controller are accurately detected, the mirror pair construction error caused by unknown performance indicators is effectively avoided, and the processing efficiency and stability of the system are improved.
[0093] As an optional implementation, the first controller is determined from the group of candidate controllers according to the performance of the candidate controller in the group of candidate controllers, including: in the case that the number of performance test commands is multiple, the multiple response times corresponding to the multiple performance test commands are weighted and summed according to the command type of the performance test command, to obtain the total response time of the candidate controller in the group of candidate controllers; the controller with the minimum total response time of the candidate controller in the group of candidate controllers is determined as the first controller.
[0094] It should be noted that by considering that the command types of different performance test commands have different performance contribution degrees in the entire controller, a weight value can be assigned to different types of performance test commands, and then all response times are summed to obtain the total response time. The total response time can be used to reflect the time length of the comprehensive performance of the controller.
[0095] Optionally, a weight set can be pre-deployed, the weight set including a plurality of groups of weight values, the plurality of groups of weight values being weight values corresponding to proportions between a plurality of running task types, the running task types including read-write tasks and network tasks. Optionally, based on proportions between task types running in each controller, a group of weight values corresponding thereto is selected from the weight set, the group of weight values including weight values corresponding to different types of performance test commands. Specifically, the weight set can be generated based on historical running data and a deployed weight prediction model, or can be set based on an empirical value, which is not limited in the present application.
[0096] Through the embodiment, the comprehensive performance of the controller is obtained by weighted summation of the performance test commands, which can effectively avoid the bias of the mirror pair construction caused by a single performance index, and improve the data security and processing efficiency of the system.
[0097] As an optional implementation, the process of selecting the first controller from the group of candidate controllers can include: first, identifying the number and attributes of mirror pairs of the candidate controllers in the group of candidate controllers. In the case where the number of controllers selected by the number and attributes of mirror pairs of the candidate controllers in the group of candidate controllers is multiple, the performance of the remaining controllers can be further selected. In the case where the number of selected controllers is multiple, a cyclic mirror mode can be used, that is, each time a new mirror pair is constructed, the controllers are selected in a cyclic order from the controller sequence. In this way, even if the performance is the same, the cyclic selection can avoid overloading some controllers, better cope with high concurrency and large data business demands, and improve the overall stability and response speed.
[0098] As an optional implementation, in the case where the first fault handling process is not executed, and it is identified that the specified controller exists, the second fault handling process of the cache state machine is executed based on the execution stage of the first fault handling process, including: in the case where the specified controller is a controller in the plurality of controllers that does not belong to the specified mirror pair and triggers a fault, and the execution stage of the first fault handling process indicates that the first update process is executed but the recovery process is not executed, a fourth update process is executed, wherein the fourth update process is used to indicate to update the online information of the controller in the multi-controller storage system; in the case where the fourth update process is executed, the recovery process is executed.
[0099] It should be noted that, in the case that the specified controller is a controller that does not belong to the specified mirror pair and triggers a fault, and the execution stage of the first fault processing procedure indicates that the first update procedure is executed but the recovery procedure is not executed, the fourth update procedure is executed first to indicate updating the online information of the controller in the multi-controller storage system; in the case that the fourth update procedure is executed, the recovery procedure is executed. In the case that the specified controller is a controller that does not belong to the specified mirror pair and triggers a fault, and the execution stage of the first fault processing procedure indicates that the first update procedure is executed but the recovery procedure is not executed, the second fault processing procedure can include the fourth update procedure and the recovery procedure.
[0100] Through the embodiment, by executing the corresponding update procedure in different stages of fault processing, the real-time and accuracy of the controller online information are ensured, the business processing error caused by untimely information updating is effectively avoided, and the response speed and efficiency of the system are improved.
[0101] As an optional implementation, the application further provides a fault processing method of a multi-controller storage system, as shown in Figure 5 In the multi-controller storage system, in the case that the same mirror pair controllers fail at the same time, the first fault processing procedure is triggered to be executed.
[0102] In the process of executing the first fault processing procedure, it is determined whether to execute the first update procedure; in the case that the first update procedure is not executed, the controller fails again, and then the first update procedure and the recovery procedure are controlled to be executed, that is, in this scenario, the corresponding executed fault processing procedure includes the quiesce-ack (STALL)-resume.
[0103] In the process of executing the first fault processing procedure and in the case that the first update procedure is executed, it is determined whether the recovery procedure is executed; in the case that the recovery procedure is not executed and the controller fails again, the fourth update procedure and the recovery procedure are controlled to be executed, that is, in this scenario, the corresponding executed fault processing procedure includes the quiesce-ack (1, STALL)-ack (2, STALL)-resume.
[0104] In case of a controller failure again during the execution of the first failure handling procedure, the execution of the first update procedure, and the execution of the recovery procedure, the interrupt procedure, the fifth update procedure, and the recovery procedure are executed, i.e. in this scenario, the corresponding executed failure handling procedure comprises the interrupt procedure, the first update procedure, the recovery procedure, the interrupt procedure, the fifth update procedure, and the recovery procedure (e.g. quiesce(1)—ack(1,STALL)—resume(1)—quiesce(2)—ack(2,STALL)—resume(2)).
[0105] As an optional implementation, the application further provides a failure handling method of a multi-controller storage system, as shown in the following table. Figure 6 As shown in the table, in the multi-controller storage system, in case of a controller recovery when the same mirror pair of controllers fails simultaneously and the first failure handling procedure is executed, according to the execution stage of the first failure handling procedure, the following second failure handling procedure is executed. Details are shown in the following table.
[0106] It is determined whether the cluster is in a specified state (e.g. the stall state). In case that the cluster is in the specified state and the first update procedure is not executed, the first update procedure and the recovery procedure are executed. I.e. in this scenario, the corresponding executed failure handling procedure comprises the interrupt procedure, the first update procedure, and the recovery procedure (e.g. quiesce—ack(STALL)—resume).
[0107] In case that the cluster is in the specified state, the first update procedure is executed, and the recovery procedure is not executed, the fifth update procedure and the recovery procedure are executed, i.e. in this scenario, the corresponding executed failure handling procedure comprises the interrupt procedure, the first update procedure, the fifth update procedure, and the recovery procedure (e.g. quiesce—ack(1,STALL)—ack(2,STALL)—resume). The fifth update procedure is used to indicate to update the online information of the controllers in the multi-controller storage system, specifically to update the online information of the recovered controller (e.g. the specified controller) and recover the corresponding mirror pair information.
[0108] In the case that the cluster is in the specified state, the first update procedure is executed, and the recovery procedure is executed, the interrupt procedure, the sixth update procedure, and the recovery procedure are executed. That is, in this scenario, the corresponding executed fault handling procedure includes the interrupt procedure, the first update procedure, the recovery procedure, the interrupt procedure, the sixth update procedure, and the recovery procedure (such as quiesce (1) - ack (1, STALL) - resume (1) - quiesce (2) - ack (2, STALL) - resume (2)). The sixth update procedure is used to indicate updating the online information of the controller in the multi-control storage system, specifically, updating the online information of the recovered controller (such as the specified controller) and recovering the corresponding mirror pair information thereof.
[0109] In the case that the cluster is not in the specified state, and the first update procedure is not executed, the seventh update procedure and the recovery procedure are executed. The seventh update procedure includes the mirror pair construction procedure and the second update procedure. That is, in this scenario, the corresponding executed fault handling procedure includes the interrupt procedure, the seventh update procedure, and the recovery procedure (such as quiesce - ack (CONTRACT (used to indicate controller fault exit)) - resume). The ack (CONTRACT) is used to indicate executing the mirror pair construction procedure to use the reconstructed mirror pair to perform cache data synchronization.
[0110] In the case that the cluster is not in the specified state, the first update procedure is not executed, and the recovery procedure is not executed, the mirror pair construction procedure, the second update procedure, and the recovery procedure are executed. That is, in this scenario, the corresponding executed fault handling procedure includes the interrupt procedure, the first update procedure, the mirror pair construction procedure, the second update procedure, and the recovery procedure (such as quiesce - ack (1, STALL) - ack (2, RECOVER) - resume). The ack (RECOVER (indicating controller recovery) is used to indicate executing the mirror pair construction procedure to use the reconstructed mirror pair to perform cache data synchronization.
[0111] In the case that the cluster is not in the specified state, the first update procedure is not executed, and the recovery procedure is executed, the interrupt procedure, the eighth update procedure, and the recovery procedure are executed. The eighth update procedure includes the mirror pair construction procedure and the second update procedure. That is, in this scenario, the corresponding executed fault handling procedure includes the interrupt procedure, the first update procedure, the recovery procedure, the mirror pair construction procedure, the second update procedure, and the recovery procedure (such as quiesce (1) - ack (1, STALL) - resume (1) - quiesce (2) - ack (2, RECOVER) - resume (2)).
[0112] Through the optional example, by introducing precise control for different stages and scenarios in the fault handling process, such as the first update process, the fourth update process, the fifth update process, etc., it is ensured that the storage system can quickly respond when facing sudden failures, avoiding unnecessary business interruption, and enhancing the overall stability and response speed of the system. The sixth update process and the seventh update process emphasize the importance of online information of the update controller and the recovery mirror pair information, which realizes the redundant storage of data through the mirror pair construction process and data re-mirroring after the controller is recovered, and guarantees the consistency and security of data.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better implementation.
[0114] The embodiments of the present application also provide a fault handling device of a multi-controller storage system, the multi-controller storage system comprising a plurality of controllers and a cache state machine, the plurality of controllers comprising at least one mirror pair, one mirror pair in the at least one mirror pair comprising two controllers in the plurality of controllers; Figure 7 is a structural block diagram of a fault handling device of a multi-controller storage system according to an embodiment of the present application, as shown in the figure, the device comprises: Figure 7
[0115] The first execution module 702 is configured to, in a case where it is identified that there is a specified mirror pair in the at least one mirror pair, control the cache state machine to execute a first fault handling process, wherein the specified mirror pair is a mirror pair in which both controllers are faulty.
[0116] The second execution module 704 is configured to, in a case where the first fault handling process is not executed completely and it is identified that there is a specified controller, execute a second fault handling process of the cache state machine based on an execution stage of the first fault handling process, so as to control the multi-controller storage system to continue business processing, wherein the specified controller is a controller in the plurality of controllers which triggers recovery, or is a controller in the plurality of controllers which triggers failure and does not belong to the specified mirror pair.
[0117] In a case where it is identified that the specified mirror pair exists in the at least one mirror pair, the cache state machine is controlled to execute a first fault handling process, where the specified mirror pair is a mirror pair in which both controllers are faulty; in a case where the first fault handling process is not executed completely and it is identified that the specified controller exists, a second fault handling process of the cache state machine is executed based on an execution stage of the first fault handling process, so as to control the multi-control storage system to continue service processing, where the specified controller is a controller triggering recovery in the plurality of controllers, or is a controller triggering fault and not belonging to the specified mirror pair in the plurality of controllers. Through the cache state machine, the fault handling process can be immediately responded in a case where the recovery of the surviving controller or the fault of the new controller occurs during execution of the first fault handling process, so that the service processing can be quickly recovered in the case of complex fault superposition, the service interruption time and the data processing delay are reduced, and the problem of long service interruption of the multi-control storage system in the related art in the fault scenario is solved, and the business continuity and the stability of the system are significantly enhanced.
[0118] Optionally, the execution sequence of the processes in the first fault handling process is an interruption process, a first update process, and a recovery process, where the interruption process is used to instruct to interrupt the service processing of the multi-control storage system, and the recovery process is used to instruct to recover the service processing of the multi-control storage system; the second execution module 704 includes: a first determination unit, configured to determine a cluster state of the multi-control storage system in a case where it is identified that the specified controller is a controller triggering recovery in the plurality of controllers, and the execution stage of the first fault handling process indicates that the first update process is executed completely but the recovery process is not executed; a first execution unit, configured to execute a mirror pair construction process in a case where the cluster state of the multi-control storage system is in a stable state, where the mirror pair construction process is used to instruct to determine a first controller from a group of candidate controllers of the multi-control storage system, so as to form a first specified mirror pair with the specified controller, where the first specified mirror pair is used to execute service of a previous mirror pair including the specified controller and in which a controller is faulty, and the candidate controllers in the group of candidate controllers are controllers not in a fault state in the multi-control storage system; a second execution unit, configured to execute a second update process in a case where the mirror pair construction process is executed completely, where the second update process is used to instruct to update cluster state information of the multi-control storage system, and the cluster state information of the multi-control storage system is used to instruct online information of the plurality of controllers and mirror pair state information of the plurality of controllers; and a third execution unit, configured to execute the recovery process in a case where the second update process is executed completely.
[0119] Optionally, the first determining unit is further configured to: determine, according to the state information of the first mirror pair in the plurality of first mirror pairs, whether a second mirror pair exists in the plurality of first mirror pairs, wherein the first mirror pair in the plurality of first mirror pairs comprises the specified controller, and the second mirror pair is a mirror pair in which a controller fails; and determine, in a case where the second mirror pair exists, a cluster state of the multi-controller storage system according to a failure time of the controller in the second mirror pair.
[0120] Optionally, the first determining unit is further configured to: obtain the failure time of the controller other than the specified controller in the second mirror pair; determine, in a case where the failure time of the controller other than the specified controller in the second mirror pair is earlier than the failure time of the specified controller, the cluster state of the multi-controller storage system to be a stable state; and determine, in a case where the failure time of the controller other than the specified controller in the second mirror pair is later than the failure time of the specified controller, the cluster state of the multi-controller storage system to be an unstable state.
[0121] Optionally, the second execution module 704 further comprises: a fourth execution unit configured to execute a third updating procedure in a case where the cluster state of the multi-controller storage system is in the unstable state, wherein the third updating procedure is used to instruct to update online information of the controller in the multi-controller storage system; and a fifth execution unit configured to execute a recovery procedure in a case where the execution of the third updating procedure is completed.
[0122] Optionally, the first execution unit is further configured to: obtain mirror pair state information of a candidate controller in a group of candidate controllers, wherein the mirror pair state information is used to indicate a mirror pair to which the candidate controller belongs; and determine, according to the mirror pair state information of the candidate controller in the group of candidate controllers, the first controller from the group of candidate controllers to constitute the first specified mirror pair with the specified controller.
[0123] Optionally, the mirror pair state information comprises a mirror pair quantity and a mirror pair attribute, and the mirror pair attribute is used to indicate whether a controller in the mirror pair is a data initiator or a data receiver; and the first execution unit is further configured to: determine, as a second controller, a candidate controller in the group of candidate controllers that belongs to a mirror pair with the least quantity; determine, as the first controller, the second controller in a case where the quantity of the second controller is not multiple; and determine, as the first controller, a second controller in a case where the quantity of the second controller is multiple, according to a frequency at which the second controller in the plurality of second controllers is a data initiator and the mirror pair quantity of the second controller in the plurality of second controllers.
[0124] Optionally, the second execution module 704 further includes a sixth execution unit, configured to, after executing the mirror pair construction process, take the first controller as a data receiver, take the specified controller as a data initiator, and trigger execution of a cache data remirror process, where the cache data remirror process is used to instruct transmission of specified cache data in the specified controller to the first controller.
[0125] Optionally, the first execution unit is further configured to: determine performance conditions of the candidate controllers in the group of candidate controllers; and determine the first controller from the group of candidate controllers according to the performance conditions of the candidate controllers in the group of candidate controllers.
[0126] Optionally, the first execution unit is further configured to: send a performance test command to each of the candidate controllers in the group of candidate controllers, and record response time lengths of the candidate controllers in the group of candidate controllers in executing the performance test command, where the response time length of the performance test command is used to indicate the performance condition of the candidate controller, and the performance test command includes at least one of the following test commands: a read performance test command, a write performance test command, a memory access performance test command, and a network test command.
[0127] Optionally, the first execution unit is further configured to, in a case where the number of performance test commands is multiple, perform weighted summation on multiple response time lengths corresponding to the multiple performance test commands according to command types of the performance test commands, to obtain a total response time length of the candidate controllers in the group of candidate controllers; and determine, as the first controller, a controller with the minimum total response time length among the candidate controllers in the group of candidate controllers.
[0128] Optionally, the second execution module 704 further includes a seventh execution unit, configured to, in a case where the specified controller is a controller that does not belong to the specified mirror pair and triggers a fault among the multiple controllers, and a case where the execution stage of the first fault handling process indicates that the first update process is executed but the recovery process is not executed, execute a fourth update process, where the fourth update process is used to instruct updating of online information of the controllers in the multi-controller storage system; and an eighth execution unit, configured to, in a case where the fourth update process is executed, execute the recovery process.
[0129] The features of the embodiments of the fault handling device of the multi-controller storage system can be referred to the related descriptions of the embodiments of the fault handling method of the multi-controller storage system, which will not be repeated here.
[0130] Embodiments of the present application also provide an electronic device including a memory and a processor, the memory storing a computer program, and the processor being configured to run the computer program to perform the steps in any of the above-mentioned embodiments of the fault handling method of the multi-controller storage system.
[0131] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is arranged to execute the steps in any of the above-mentioned fault processing method embodiments of the multi-control storage system.
[0132] In an example embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media capable of storing computer programs.
[0133] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned fault processing method embodiments of the multi-control storage system.
[0134] The embodiment of the present application further provides another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned fault processing method embodiments of the multi-control storage system.
[0135] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0136] The above describes in detail the fault processing method of the multi-control storage system and the electronic device provided by the present application. The principles and implementation manners of the present application are described by applying specific examples in the present document, and the above description of the examples is only used to help understand the method and core idea of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A fault handling method for a multi-controller storage system, characterized in that, The multi-controller storage system includes multiple controllers and a cache state machine, the multiple controllers forming at least one mirror pair, and one mirror pair in the at least one mirror pair containing two controllers from the multiple controllers; the method includes: If a specified mirror pair is identified in the at least one mirror pair, the cache state machine is controlled to execute a first fault handling procedure, wherein the specified mirror pair is a mirror pair in which both controllers are faulty; If the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution stage of the first fault handling process to control the multi-controller storage system to continue business processing. The designated controller is either the controller that triggered the recovery among the multiple controllers, or the controller that does not belong to the designated mirror pair and triggered the fault among the multiple controllers. The execution order of the processes in the first fault handling process is the interruption process, the first update process, and the recovery process. The interruption process is used to indicate the interruption of the business processing of the multi-controller storage system, and the recovery process is used to indicate the restoration of the business processing of the multi-controller storage system. When the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution stage of the first fault handling process, including: If the specified controller is identified as the controller that triggered the recovery among the plurality of controllers, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, the cluster status of the multi-controller storage system is determined. When the cluster state of the multi-controller storage system is stable, a mirror pair construction process is executed. The mirror pair construction process is used to instruct the determination of a first controller from a group of candidate controllers in the multi-controller storage system to form a first designated mirror pair with the designated controller. The first designated mirror pair is used to execute the services of a prior mirror pair that includes the designated controller and has a controller failure. The candidate controllers in the group of candidate controllers are controllers in the multi-controller storage system that are not in a failure state. After the mirror pair construction process is completed, a second update process is executed, wherein the second update process is used to instruct the cluster status information of the multi-controller storage system to be updated, and the cluster status information of the multi-controller storage system is used to indicate the online information of the multiple controllers and the mirror pair status information of the multiple controllers. If the second update process is completed, the recovery process is executed. The step of determining the cluster status of the multi-controller storage system includes: Based on the status information of the first mirror pairs among the multiple first mirror pairs, it is determined whether there is a second mirror pair among the multiple first mirror pairs, wherein the first mirror pairs among the multiple first mirror pairs include the designated controller, and the second mirror pair has a mirror pair with a controller failure; If the second mirror pair exists, obtain the fault time of the controller in the second mirror pair other than the specified controller; If the failure time of the controller in the second mirror pair other than the failure time of the specified controller is earlier than the failure time of the specified controller, the cluster state of the multi-controller storage system is determined to be a stable state. If the failure time of the controller in the second mirror pair, excluding the specified controller, is later than the failure time of the specified controller, the cluster state of the multi-controller storage system is determined to be unstable.
2. The method according to claim 1, characterized in that, The method further includes: When the cluster state of the multi-controller storage system is unstable, a third update process is executed, wherein the third update process is used to instruct the updating of the online information of the controllers in the multi-controller storage system; After the third update process is completed, the recovery process is executed.
3. The method according to claim 1, characterized in that, The step of determining the first controller from a set of candidate controllers in the multi-controller storage system includes: Obtain the mirror pair status information of the candidate controllers in the group of candidate controllers, wherein the mirror pair status information is used to indicate the mirror pair to which the candidate controller belongs; Based on the mirror pair status information of the candidate controllers in the set of candidate controllers, the first controller is determined from the set of candidate controllers to form the first designated mirror pair with the designated controller.
4. The method according to claim 3, characterized in that, The mirror pair status information includes the number of mirror pairs and mirror pair attributes. The mirror pair attributes are used to indicate whether the controller in the mirror pair is a data initiator or a data receiver. The step of determining the first controller from the set of candidate controllers based on the mirror pair state information of the candidate controllers in the set of candidate controllers includes: The candidate controller with the fewest mirror pairs among the mirror pairs in the group of candidate controllers is determined as the second controller; If the number of second controllers is not multiple, the second controller is identified as the first controller; When there are multiple second controllers, the frequency at which the second controller is the data initiator is determined based on the mirror pair attribute of the second controller among the multiple second controllers and the number of mirror pairs of the second controller among the multiple second controllers; From among the multiple second controllers, the second controller with the lowest frequency of being the data initiator is determined as the first controller.
5. The method according to claim 4, characterized in that, After the image pair construction process is executed, the method further includes: The first controller is used as the data receiver, the designated controller is used as the data initiator, and the cache data re-mirroring process is triggered, wherein the cache data re-mirroring process is used to instruct the designated cache data in the designated controller to be transmitted to the first controller.
6. The method according to claim 1, characterized in that, The step of determining the first controller from a set of candidate controllers in the multi-controller storage system includes: A performance test command is sent to each of the candidate controllers in the group of candidate controllers, and the response time of the candidate controllers in the group of candidate controllers executing the performance test command is recorded. The response time of the performance test command is used to indicate the performance of the candidate controllers in the group of candidate controllers. The performance test command includes at least one of the following test commands: read performance test command, write performance test command, memory access performance test command, and network test command. When there are multiple performance test commands, the response times corresponding to the multiple performance test commands are weighted and summed according to the command type of the performance test commands to obtain the total response time of the candidate controllers in the group of candidate controllers; The controller with the smallest total response time among the candidate controllers in the group of candidate controllers is determined as the first controller.
7. The method according to claim 1, characterized in that, When the first fault handling process has not been completed and a designated controller is identified, the second fault handling process of the cache state machine is executed based on the execution stage of the first fault handling process, including: If the designated controller is a controller that does not belong to the designated mirror pair among the plurality of controllers and has triggered a fault, and the execution phase of the first fault handling process indicates that the first update process has been completed but the recovery process has not been executed, then the fourth update process is executed, wherein the fourth update process is used to indicate updating the online information of the controllers in the multi-controller storage system. If the fourth update process is completed, the recovery process is then executed.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the fault handling method for the multi-controller storage system as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Storage node upgrading method, system and device, storage medium and electronic equipment
CN120523496A
Multi-control storage array, storage system, data processing method and storage medium
WO2025156686A1