Method and apparatus for clearing cached data, electronic device, and storage medium

By replacing the faulty controller with a transition controller in the storage system for mirroring reorganization and data cleaning, the cached data loss caused by concurrent failure during failure recovery is solved, and a storage system with high reliability and rapid recovery is achieved.

CN119718760BActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510217168.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In storage systems, the faulty controller is susceptible to interference from other controller failure events when recovering, resulting in cached data loss and business process interruption, and it is difficult for the existing technology to effectively deal with concurrent failure situations.

Method used

By replacing the faulty controller with the transition controller, mirroring reorganization is performed to form a temporary mirror pair, and after the failure is restored, the data cleaning process is performed. At the same time, monitor whether the controller in the storage system has a concurrent failure, and adjust the data cleaning process based on the execution status and concurrent failure information of the data cleaning process.

Benefits of technology

Ensure that during the failure recovery process, the redundancy of cached data and business continuity are not affected, and can effectively deal with concurrent failures, reduce the impact of failures on the system, and improve the fault tolerance and recovery speed of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718760B_ABST
    Figure CN119718760B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for cache data cleaning, an electronic device, and a storage medium, relating to the technical field of computer data storage. The present disclosure uses a transition controller to replace a faulty controller for mirror recombination and form a temporary mirror pair; it can ensure that the basic data redundancy and read / write functions of the storage system are not severely affected; when the faulty controller recovers, it can perceive complex fault situations in real time; once a concurrent fault occurs, the data cleaning process is adjusted according to the execution status information of the data cleaning process, the teaming information of the temporary mirror pair, and the concurrent fault information of the concurrently faulty controller; it ensures that the data of the transition controller can be effectively cleaned in the case of a concurrent fault, improving the fault tolerance of the storage system; when cleaning the cache data of the transition controller, through an effective fault handling mechanism and precise adjustment of the data cleaning process, the overall performance and efficiency of the storage system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of computer data storage, and in particular to a method and device for cleaning cache data, an electronic device, and a storage medium. Background Art

[0002] In the field of data storage technology, with the explosive growth of data volume and the increasing requirements of business for real-time and reliability of data processing, the performance and stability of storage systems are crucial. In order to meet these requirements, centralized storage systems generally adopt a multi-controller architecture and use back-end interconnection technology to achieve data sharing and synchronization. Cache modules also often use technologies such as circular mirroring to improve data storage and read and write efficiency.

[0003] During the operation of such storage systems, controller failure and recovery are inevitable. When a controller fails, the system needs to readjust the architecture to maintain data redundancy and business continuity, which involves the re-teaming of mirror pairs and data re-mirroring operations. However, in order to speed up the recovery of the controller, some existing cache data synchronization processing solutions use data re-mirroring technology. Although the controller joining time is reduced to a certain extent, in a multi-controller storage system, due to the extremely complex failure scenario, it is very easy to be disturbed by other controller failure events during the data synchronization process, which in turn causes cache data loss and interrupts the business process. At the same time, the traditional simple control process is difficult to adapt to the complex requirements of the multi-controller storage cache data re-mirroring process, which significantly increases the complexity of business processing and further increases the risk of system operation. It can be seen that in the current development of storage technology, how to effectively deal with concurrent failures that occur when the storage system controller is restored, ensure that the cache data re-mirroring process can be correctly controlled when the multi-controller storage device faces cluster controller events during the cache data cleanup process, maintain data consistency between different controllers, and ensure data security, has become a key technical problem that needs to be solved in this field. Summary of the invention

[0004] The present disclosure provides a method and device for cleaning cache data, an electronic device and a storage medium, which are mainly intended to solve the problem that when a faulty controller in a storage system is restored, other controllers fail, resulting in cache data errors.

[0005] According to a first aspect of the present disclosure, a method for cleaning cache data is provided, comprising:

[0006] A transition controller is used to replace a faulty controller in the storage system, and the initial mirror pair where the faulty controller is located is mirrored and reorganized to obtain a first reorganized mirror pair; the repaired faulty controller is added to the first reorganized mirror pair to obtain a temporary mirror pair;

[0007] After the faulty controller that repairs the fault completes data mirroring, remove the transition controller from the temporary mirror pair and perform a data cleaning process on the transition controller;

[0008] Monitor whether concurrent faults occur in the controllers in the storage system and monitor the execution status of the data cleaning process;

[0009] When concurrent faults occur in the controllers in the storage system, adjust the execution process of the data cleaning process according to the execution status information of the data cleaning process and the concurrent fault information of the concurrent fault controllers with concurrent faults; according to the adjusted data cleaning process, perform data cleaning on the transition controller.

[0010] In some embodiments, use a transition controller to replace the faulty controller in the storage system and perform mirror recombination on the initial mirror pair where the faulty controller is located to obtain a first recombined mirror pair, including:

[0011] According to the grouping information of the initial mirror pair where the faulty controller is located, screen a transition controller from the controllers that have not failed to replace the faulty controller;

[0012] Perform mirror recombination on the transition controller and the normal controller in the initial mirror pair where the faulty controller is located to obtain a first recombined mirror pair.

[0013] In some embodiments, according to the grouping information of the initial mirror pair where the faulty controller is located, screening a transition controller from the controllers that have not failed to replace the faulty controller includes:

[0014] Obtain the set of device identifiers of the controllers in the storage system except the faulty controller, where the device identifiers in the set of device identifiers are arranged in order;

[0015] Based on the grouping information, determine the normal device identifier corresponding to the normal controller in the initial mirror pair where the faulty controller is located;

[0016] According to the set of device identifiers and the normal device identifier, find a transition controller to replace the faulty controller.

[0017] In some embodiments, according to the set of device identifiers and the normal device identifier, finding a transition controller to replace the faulty controller includes:

[0018] Find the target device identifier in the set of device identifiers that is the next one after the normal device identifier;

[0019] Use the controller corresponding to the target device identifier as the transition controller.

[0020] In some embodiments, performing mirror recombination on the transition controller and the normal controller in the initial mirror pair where the faulty controller is located to obtain a first recombined mirror pair includes:

[0021] Sort the device identifiers of the transition controller and the normal controller according to the cyclic mirror principle to determine the arrangement order of the transition controller and the normal controller, so as to generate the first reorganized mirror pair.

[0022] In some embodiments, before using the transition controller to replace the faulty controller in the storage system, the method for clearing cached data further includes:

[0023] Monitor the operating status of the controllers in the storage system and identify the faulty controller.

[0024] In some embodiments, the method for clearing cached data further includes:

[0025] Identify the device identifier of the faulty controller, collect and sort the device identifiers of the controllers in the storage system except the faulty controller to generate a set of device identifiers.

[0026] In some embodiments, after adding the faulty controller repaired from the fault to the first reorganized mirror pair to obtain a temporary mirror pair, the method for clearing cached data further includes:

[0027] Control the faulty controller repaired from the fault and the transition controller to perform data mirroring processing.

[0028] In some embodiments, the method for clearing cached data further includes:

[0029] If there is no concurrent fault in the controllers in the storage system, perform data clearing on the transition controller according to the data clearing process.

[0030] In some embodiments, according to the execution status information of the data clearing process and the concurrent fault information of the concurrent faulty controllers with concurrent faults, adjust the execution process of the data clearing process, including:

[0031] Based on the execution status information, determine whether the transition controller starts to execute the data clearing process;

[0032] When the data clearing process has not started, perform mirror reorganization on the mirror pair where the concurrent faulty controller is located to obtain a second reorganized mirror pair;

[0033] Based on the teaming information of the second reorganized mirror pair, determine whether the transition controller belongs to the second reorganized mirror pair to determine whether to continue executing the data clearing process.

[0034] In some embodiments, when the data clearing process has started, the method for clearing cached data further includes:

[0035] Based on the concurrent fault information, determine whether the transition controller is a concurrent faulty controller;

[0036] If the transition controller is a concurrent failure controller, when the transition controller rejoins the mirror pair, the data cleaning process is re-executed for the transition controller;

[0037] If the transition controller is not a concurrent failure controller, after the data cleaning process is executed for the transition controller, data mirroring is performed on the transition controller, and mirror recombination is performed on the mirror pair where the concurrent failure controller is located to obtain a third recombined mirror pair.

[0038] In some embodiments, based on the concurrent failure information, determining whether the transition controller is a concurrent failure controller includes:

[0039] Based on the concurrent failure information, obtaining the device identifier of the concurrent failure controller and the teaming information of the mirror pair where it is located;

[0040] Comparing the device identifier of the concurrent failure controller with the device identifier of the transition controller, and comparing the teaming information of the mirror pair where the concurrent failure controller is located with the teaming information of the temporary mirror pair where the transition controller is located to determine whether the transition controller is a concurrent failure controller.

[0041] In some embodiments, performing mirror recombination on the mirror pair where the concurrent failure controller is located to obtain a second recombined mirror pair includes:

[0042] According to the concurrent failure information, searching for the teaming information of the first failed mirror pair where the concurrent failure controller is located;

[0043] Based on the teaming information of the first failed mirror pair, determining whether the first failed mirror pair is the temporary mirror pair from which the transition controller is removed,

[0044] If the first failed mirror pair is the temporary mirror pair from which the transition controller is removed, pulling the transition controller back into the temporary mirror pair and performing mirror recombination with the normal controller in the temporary mirror pair to obtain a second recombined mirror pair;

[0045] If the first failed mirror pair is not the temporary mirror pair from which the transition controller is removed, performing mirror recombination on the normal controller in the first failed mirror pair and the controller at the next adjacent position to obtain a second recombined mirror pair.

[0046] In some embodiments, based on the teaming information of the second recombined mirror pair, determining whether the transition controller belongs to the second recombined mirror pair to determine whether to continue executing the data cleaning process includes:

[0047] If the transition controller belongs to the second recombined mirror pair, there is no need to continue executing the data cleaning process for the transition controller in the second recombined mirror pair;

[0048] If the transition controller does not belong to the second reorganized mirror pair, continue to execute the data cleaning process for the transition controller that does not belong to the second reorganized mirror pair.

[0049] In some embodiments, mirror reorganization is performed on the mirror pair where the concurrent failure controller is located to obtain a third reorganized mirror pair, including:

[0050] According to the concurrent failure information, find the teaming information of the second failed mirror pair where the concurrent failure controller is located;

[0051] Based on the teaming information of the second failed mirror pair, determine whether the second failed mirror pair is a temporary mirror pair for removing the transition controller,

[0052] If the second failed mirror pair is a temporary mirror pair for removing the transition controller, pull the transition controller back into the temporary mirror pair and perform mirror reorganization with the normal controller in the temporary mirror pair to obtain a third reorganized mirror pair;

[0053] If the second failed mirror pair is not a temporary mirror pair for removing the transition controller, perform mirror reorganization on the normal controller in the second failed mirror pair and the controller adjacent to its next position to obtain a third reorganized mirror pair.

[0054] In some embodiments, the method for cache data cleaning further includes:

[0055] Obtain the status information of each mirror pair in the storage system, and mark the device identifier of the mirror pair in which there is a transition controller that needs to execute the data cleaning process.

[0056] According to the second aspect of the present disclosure, there is provided a cache data cleaning device, including:

[0057] A reorganization unit, configured to replace a failed controller in the storage system with a transition controller, and perform mirror reorganization on the initial mirror pair where the failed controller is located to obtain a first reorganized mirror pair; add the failed controller that has repaired the failure to the first reorganized mirror pair to obtain a temporary mirror pair;

[0058] A removal unit, configured to remove the transition controller from the temporary mirror pair after the failed controller that has repaired the failure completes data mirroring, and execute a data cleaning process on the transition controller;

[0059] A first monitoring unit, configured to monitor whether a concurrent failure occurs in the controller in the storage system, and monitor the execution status of the data cleaning process;

[0060] An adjustment unit, configured to adjust the execution process of the data cleaning process according to the execution status information of the data cleaning process and the concurrent failure information of the concurrent failure controllers that have concurrent failures in the controller in the storage system; and perform data cleaning on the transition controller according to the adjusted data cleaning process.

[0061] In some embodiments, the reorganization unit includes:

[0062] A screening module, configured to screen a transition controller that replaces the failed controller from the controllers that have not failed according to the pairing information of the initial mirror pair where the failed controller is located;

[0063] A reorganization module, configured to perform mirror reorganization on the transition controller and the normal controller in the initial mirror pair where the failed controller is located to obtain a first reorganized mirror pair.

[0064] In some embodiments, the screening module is further configured to:

[0065] Obtain a set of device identifiers of the controllers in the storage system except for the failed controller, where the device identifiers in the set of device identifiers are arranged in sequence;

[0066] Based on the pairing information, determine the normal device identifier corresponding to the normal controller in the initial mirror pair where the failed controller is located;

[0067] According to the set of device identifiers and the normal device identifier, search for a transition controller that replaces the failed controller.

[0068] In some embodiments, searching for a transition controller that replaces the failed controller according to the set of device identifiers and the normal device identifier includes:

[0069] Search for a target device identifier that is the next one after the normal device identifier in the set of device identifiers;

[0070] Use the controller corresponding to the target device identifier as the transition controller.

[0071] In some embodiments, the reorganization module is further configured to:

[0072] According to the circular mirror principle, sort the device identifiers of the transition controller and the normal controller to determine the arrangement order of the transition controller and the normal controller, so as to generate a first reorganized mirror pair.

[0073] In some embodiments, the device for caching data cleaning further includes:

[0074] A second monitoring unit, configured to monitor the running status of the controllers in the storage system and identify the failed controllers before the reorganization unit replaces the failed controllers in the storage system with the transition controllers.

[0075] In some embodiments, the apparatus for caching data cleaning further includes:

[0076] A collection unit, configured to identify the device identifier of the failed controller, collect the device identifiers of the controllers in the storage system except the failed controller, sort them, and generate a set of device identifiers.

[0077] In some embodiments, the apparatus for caching data cleaning further includes:

[0078] A control unit, configured to control the failed controller that has repaired the fault to perform data mirroring processing with the transition controller after the reorganization unit adds the failed controller that has repaired the fault to the first reorganized mirror pair to obtain a temporary mirror pair.

[0079] In some embodiments, the apparatus for caching data cleaning further includes:

[0080] A cleaning unit, configured to perform data cleaning on the transition controller according to the data cleaning process when there is no concurrent failure of the controllers in the storage system.

[0081] In some embodiments, the adjustment unit includes:

[0082] A first judgment module, configured to judge whether the transition controller starts to execute the data cleaning process based on the execution status information;

[0083] A determination module, configured to, when the data cleaning process has not started, perform mirror reorganization on the mirror pair where the concurrent failure controller is located to obtain a second reorganized mirror pair, and judge whether the transition controller belongs to the second reorganized mirror pair based on the teaming information of the second reorganized mirror pair to determine whether to continue executing the data cleaning process.

[0084] In some embodiments, the adjustment unit further includes:

[0085] A second judgment module, configured to judge whether the transition controller is a concurrent failure controller based on the concurrent failure information when the data cleaning process has started;

[0086] An execution module, configured to, if the transition controller is a concurrent failure controller, re-execute the data cleaning process on the transition controller when the transition controller rejoins the mirror pair;

[0087] A processing module, configured to, if the transition controller is not a concurrent failure controller, perform data mirroring processing on the transition controller after executing the data cleaning process on the transition controller, and perform mirror reorganization on the mirror pair where the concurrent failure controller is located to obtain a third reorganized mirror pair.

[0088] In some embodiments, the second judgment module is further configured to:

[0089] Obtain the device identifier of the concurrent fault controller and the teaming information of the mirror pair where it is located based on the concurrent fault information;

[0090] Compare the device identifier of the concurrent fault controller with the device identifier of the transition controller, and compare the teaming information of the mirror pair where the concurrent fault controller is located with the teaming information of the temporary mirror pair where the transition controller is located to determine whether the transition controller is the concurrent fault controller.

[0091] In some embodiments, the determination module is further configured to:

[0092] Search for the teaming information of the first faulty mirror pair where the concurrent fault controller is located according to the concurrent fault information;

[0093] Determine whether the first faulty mirror pair is the temporary mirror pair for removing the transition controller based on the teaming information of the first faulty mirror pair;

[0094] If the first faulty mirror pair is the temporary mirror pair for removing the transition controller, pull the transition controller back into the temporary mirror pair and perform mirror recombination with the normal controller in the temporary mirror pair to obtain the second recombined mirror pair;

[0095] If the first faulty mirror pair is not the temporary mirror pair for removing the transition controller, perform mirror recombination between the normal controller in the first faulty mirror pair and the controller at the next adjacent position to obtain the second recombined mirror pair.

[0096] In some embodiments, the determination module is further configured to:

[0097] If the transition controller belongs to the second recombined mirror pair, there is no need to continue to execute the data cleaning process for the transition controller in the second recombined mirror pair;

[0098] If the transition controller does not belong to the second recombined mirror pair, continue to execute the data cleaning process for the transition controller that does not belong to the second recombined mirror pair.

[0099] In some embodiments, the processing module is further configured to:

[0100] Search for the teaming information of the second faulty mirror pair where the concurrent fault controller is located according to the concurrent fault information;

[0101] Determine whether the second faulty mirror pair is the temporary mirror pair for removing the transition controller based on the teaming information of the second faulty mirror pair;

[0102] If the second faulty mirror pair is the temporary mirror pair for removing the transition controller, pull the transition controller back into the temporary mirror pair and perform mirror recombination with the normal controller in the temporary mirror pair to obtain the third recombined mirror pair;

[0103] If the second failed mirror pair is not the temporary mirror pair for removing the transition controller, mirror recombination is performed between the normal controller in the second failed mirror pair and the controller at the next adjacent position to obtain a third recombined mirror pair.

[0104] In some embodiments, the apparatus for cache data cleaning further includes:

[0105] A marking unit, configured to obtain status information of each mirror pair in the storage system, and mark the device identifier of the transition controller that needs to execute the data cleaning process in each mirror pair.

[0106] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0107] At least one processor; and

[0108] A memory communicatively connected to the at least one processor; wherein,

[0109] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method for cache data cleaning described in the foregoing first aspect.

[0110] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method for cache data cleaning described in the foregoing first aspect.

[0111] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the method for cache data cleaning described in the foregoing first aspect is implemented.

[0112] The present disclosure provides a method and apparatus for cache data cleaning, an electronic device, and a storage medium, relating to the technical field of computer data storage. In the embodiments of the present disclosure, a transition controller is used to replace a faulty controller for mirror recombination, and after the faulty controller is repaired, it is added to the first recombined mirror pair to form a temporary mirror pair; this can ensure that the basic data redundancy and read / write functions of the storage system are not severely affected; when the faulty controller recovers, by continuously monitoring whether concurrent faults occur in the controllers in the storage system and the execution status of the data cleaning process, the system can perceive complex fault situations in real time; once concurrent faults occur, the data cleaning process is adjusted according to the execution status information of the data cleaning process and the concurrent fault information of the concurrently faulty controllers; so as to flexibly cope with complex fault scenarios, reduce the impact of faults on the system, ensure that the data of the transition controller can be effectively cleaned in the case of concurrent faults, ensure that the system can quickly resume normal operation after faults, and improve the fault tolerance of the storage system; when cleaning the cache data of the transition controller, through an effective fault handling mechanism and precise adjustment of the data cleaning process, the storage system can recover to the normal working state faster when facing faults; whether it is a single controller fault or concurrent faults, the system can handle them in an orderly manner, shorten the fault downtime, and improve the overall performance and efficiency of the storage system.

[0113] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0115] Figure 1 is a flowchart of a method for cache data cleaning provided by an embodiment of the present disclosure;

[0116] Figure 2 is a flowchart of a method for mirror pair recombination provided by an embodiment of the present disclosure;

[0117] Figure 3 is a flowchart of another method for cache data cleaning provided by an embodiment of the present disclosure;

[0118] Figure 4 is a flowchart of another method for cache data cleaning provided by an embodiment of the present disclosure;

[0119] Figure 5 is a schematic structural diagram of an apparatus for cache data cleaning provided by an embodiment of the present disclosure;

[0120] Figure 6Schematic structural diagram of another cache data cleaning device provided by an embodiment of the present disclosure. Detailed implementation manners

[0121] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0122] The following describes a cache data cleaning method, device, electronic device, and storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.

[0123] Figure 1 Schematic flowchart of a cache data cleaning method provided by an embodiment of the present disclosure.

[0124] As Figure 1 shown, the method includes the following steps:

[0125] Step 101: Replace the faulty controller in the storage system with a transition controller, and perform mirror recombination on the initial mirror pair where the faulty controller is located to obtain a first recombined mirror pair; add the faulty controller repaired from the fault to the first recombined mirror pair to obtain a temporary mirror pair.

[0126] In an embodiment of the present disclosure, the storage system includes multiple controllers and multiple mirror pairs. In the storage system, in order to achieve functions such as efficient data storage, management, and redundant backup, multiple controllers are usually equipped. These controllers are combined separately through specific policies and mechanisms. Specifically, according to the architecture design of the storage system and the requirements of data processing, different controllers are paired and combined in pairs or according to certain logical rules, thus obtaining multiple mirror pairs. The controllers in each mirror pair cooperate with each other and jointly undertake important tasks such as data storage, reading, synchronization, and fault recovery to ensure the stability, reliability, and data consistency of the storage system. A faulty controller is a controller in the storage system that cannot normally perform its functions such as data storage, processing, and transmission due to hardware damage, software failure, power problems, or other abnormal conditions. For example, a controller that cannot work properly due to a burned-out chip on the circuit board belongs to a faulty controller. A transitional controller is a controller that temporarily replaces a faulty controller and participates in the mirror pair to maintain the normal operation of the system and the continuity of data processing when a certain controller in the storage system fails. It undertakes the corresponding work tasks during the repair period of the faulty controller until the faulty controller resumes normal operation. A temporary mirror pair is a mirror pair composed of a transitional controller and other normal controllers during the repair process of the faulty controller, which is used to maintain the data storage and processing functions during the fault. This mirror pair is temporary, and once the faulty controller is repaired, its composition structure will change. An initial mirror pair is a mirror pair pre-composed in the storage system according to specific rules, which is used for data redundant backup and storage management to ensure data security and availability. Each mirror pair usually consists of two controllers that work together to ensure data consistency.

[0127] In the storage system architecture, multiple controllers that work together are integrated inside. When a certain one of them fails and affects the normal operation of the system, in order to minimize the interference of the fault on data storage and business continuity, the system will quickly activate the fault response mechanism. At this time, a transitional controller will be utilized to accurately replace the faulty controller in the storage system.

[0128] After the transitional controller successfully accesses the system, immediately, a mirror reorganization operation will be performed on the initial mirror pair where the faulty controller originally was. This process involves re-planning and adjusting the data storage structure and data flow direction. Through complex algorithms and logical processing, the data association and backup relationships between each controller are re-combed and constructed, thus successfully obtaining the first reorganized mirror pair. In this new mirror pair, the data storage and reading methods are optimized and adjusted to adapt to the system state after the fault. It should be noted that the first reorganized mirror pair is a classification name for the mirror pair, and does not limit its characteristics such as quantity.

[0129] Subsequently, when the failed controller is repaired and restored to its normal working state, the system will, according to the established process, add the repaired failed controller to the first reorganized mirror pair. This operation is not a simple access, but rather requires re-adapting and adjusting the structure and data synchronization mechanism of the entire mirror pair to ensure that the newly added controller can cooperate seamlessly with other controllers. After this series of operations, a temporary mirror pair is finally obtained. In this temporary mirror pair, stable data interaction and backup relationships are re-established among the controllers, laying a solid foundation for the subsequent stable operation of the storage system.

[0130] Suppose there is a storage system A, which is equipped with four controllers numbered 0, 1, 2, and 3 internally. Based on the cyclic mirror principle, these four controllers are combined in pairs to form four mirror pairs in "State 1": (0,1), (1,2), (2,3), (3,0). During the normal operation of the system, each mirror pair works together to ensure the stability of data storage and processing. These four mirror pairs are the initial mirror pairs when the system is working normally at first.

[0131] When controller 0 in mirror pair 1 (0,1) suddenly fails, to maintain the normal operation of the system, a controller needs to be introduced to temporarily replace the work of controller 0. The transition controller is selected according to the established cyclic mirror principle, that is, based on the device identifier of the normal working controller in the mirror pair where the failed controller is located, the controller with the device identifier one position later is selected as the transition controller.

[0132] In the case of controller 0 failing, the storage system can identify that the controllers in the normal working state are 1, 2, and 3. At this time, mirror pair 1 and mirror pair 4 are affected by the failure of controller 0 and need to use the transition controller to replace the failed controller for mirror reorganization operations. After the reorganization is completed, the combination of the mirror pairs changes to "State 2": (1,2), (1,2), (2,3), (3,1). Among them, mirror pair 1 and mirror pair 4 belong to the first reorganized mirror pairs after reorganization, while mirror pair 2 and mirror pair 3 remain in the state of the initial mirror pairs. After the failed controller 0 repairs the fault, it is added to mirror pair 1 (the first reorganized mirror pair). At this time, the temporary mirror pair consists of (0,1,2), and all the mirror pairs of the storage system are in "State 3": (0,1,2), (1,2), (2,3), (3,1,0). For the convenience of understanding, in this embodiment, an explanation is given for one mirror pair. After controller 0 (the failed controller) is restored, controller 2 (the transition controller) needs to exit and suspend the data caching task in mirror pair 1. After controller 2 exits, the data cached by controller 2 when it was replacing controller 0 needs to be cleared. It should be noted that in this embodiment, for the convenience of understanding, the controllers of the storage system are exemplified by four controllers, but this does not constitute a limitation on the embodiments of the present disclosure.

[0133] It should be noted that the concept of "transition controller" has certain limitations. Taking controller 2 as an example, its roles and functions in mirror pair 1 and mirror pair 2 are significantly different. In mirror pair 1, controller 2 temporarily takes over the work of the faulty controller 0; while in mirror pair 2, controller 2 participates in the work according to the original system settings, without the situation of temporary replacement. Therefore, only in the first reorganized mirror pair and the temporary mirror pair, the corresponding controller is called the transition controller.

[0134] Step 102, after the faulty controller that has repaired the fault completes data mirroring, remove the transition controller from the temporary mirror pair and execute a data cleaning process on the transition controller.

[0135] In the embodiments of the present disclosure, data mirroring is a process of copying data from one storage location to another or multiple storage locations in a storage system to ensure data security, availability, and redundancy. Through data mirroring, when a storage location fails, the system can still obtain data from other mirror locations to ensure business continuity. For example, an important user data is stored in two different controllers at the same time, and the data in these two controllers are mirrored to each other. The data cleaning process is a series of operations to clear the temporary data stored in the transition controller that is no longer needed after the system returns to normal. These temporary data may include intermediate data generated during the transition, temporary copies stored to maintain the operation of the temporary mirror pair, etc.

[0136] During the actual operation of the storage system, when a certain controller unfortunately fails, the system will quickly activate the fault response mechanism, introduce a transitional controller to replace the work of the faulty controller, and form the first reorganized mirror pair with other normal controllers to ensure that the system can continue to store and process data stably. During this period, technicians will repair the faulty controller, troubleshoot the cause of the fault and take corresponding repair measures, such as replacing damaged hardware components, fixing software vulnerabilities, etc. When the faulty controller is repaired, the repaired faulty controller is added to the first reorganized mirror pair to form a temporary mirror pair. In order to enable this controller to participate in the system work normally again, a data mirroring operation is required. The system will copy the latest data stored in other controllers to the repaired faulty controller to ensure that it has a complete and accurate data copy to ensure data consistency. When the data mirroring operation is successfully completed and after the system verifies that the repaired faulty controller can work normally and cooperate with other controllers to process data, at this time the system will remove the previously temporarily added transitional controller from the temporary mirror pair. This operation is to restore the system to the normal mirror pair structure and ensure the stability of the system architecture. Immediately afterwards, the system will execute a data cleaning process on the transitional controller. During the cleaning process, the system will identify and delete all temporary data stored by the transitional controller during the period of replacing the faulty controller. These data are no longer valuable after the system returns to normal. Timely cleaning can release the storage resources of the controller and improve the operation efficiency of the system.

[0137] Suppose there is a storage system composed of controllers 0, 1, 2, and 3. Under normal circumstances, they form mirror pairs in pairs (0,1) (1,2) (2,3) (3,0). When controller 0 fails, the system selects controller 2 as the transitional controller according to the set rules and forms a temporary mirror pair (1,2) with controller 1. After repairing controller 0, the system mirrors the data in controller 2 to controller 0 to ensure that 0 has the latest and complete data. During the recovery process of the faulty controller, the mirror pair of the storage system needs to be switched from "state 3" to "state 1", so the transitional controller in the temporary mirror pair needs to be removed. When it is confirmed that controller 0 can work normally, the system removes controller 2 from the temporary mirror pair (0,1,2) to restore the mirror pair to the original (0,1). Finally, the system executes a data cleaning process on controller C to delete the temporary data stored by 2 during the period of replacing 0, such as some temporary cache data generated during data transmission and processing.

[0138] In this way, when the fault controller recovers, using the form of three replicas (i.e., controllers 0, 1, and 2 exist simultaneously in the temporary mirror pair) can ensure that after the faulty controller is repaired and rejoined to the system, the data of the entire storage system remains consistent and complete. By first performing data mirroring, then removing the transitional controller and cleaning the data, the loss or inconsistency of data is avoided, ensuring the accuracy of business data and providing reliable data services for users.

[0139] Timely cleaning of the temporary data in the transitional controller releases the storage resources and processing resources of the controller. This helps to improve the overall performance of the system, reduce problems such as slow system operation caused by redundant data occupying resources, and enables the system to process new business requests more efficiently. Removing the transitional controller from the temporary mirror pair restores the system to the normal mirror pair structure, enhancing the stability of the system architecture. It avoids potential risks that may be brought about by the long-term existence of the transitional controller, such as compatibility issues with other controllers or system errors caused by role confusion, ensuring the stable operation of the storage system.

[0140] Step 103: Monitor whether concurrent faults occur in the controllers in the storage system and monitor the execution status of the data cleaning process.

[0141] In the embodiments of the present disclosure, a concurrent fault refers to a situation where multiple controllers simultaneously or successively experience faults within a short period during the operation of the storage system. This type of fault is more complex than a single controller fault and may have a greater impact on the normal operation of the storage system. The embodiments of the present disclosure mainly address the superposition of faults in other controllers during the recovery process of the faulty controller. A concurrent fault refers to the superposition of a faulty controller and other controllers in the storage system.

[0142] In the daily operation of the storage system, it is crucial to maintain the stable and reliable operation of the system. To this end, the system continuously monitors the status of each controller. Specifically, the monitoring mechanism collects the working parameters and status information of each controller in real time, such as the CPU usage rate of the controller, the memory occupancy, the data transfer rate, and the error log. By deeply analyzing this information, it is determined whether a controller has failed. Especially when monitoring concurrent faults, the system focuses on the status change trends of multiple controllers. Once it is found that multiple controllers simultaneously exhibit abnormalities, such as a sudden and sharp increase in the CPU usage rate or a sudden interruption in data transfer, it is determined that a concurrent fault has occurred, and an alarm is quickly issued.

[0143] During the execution of the data cleaning process of the fault controller, the system conducts comprehensive monitoring and does not relax the monitoring of concurrent faults. This is because concurrent faults are very likely to have a negative impact on the data cleaning process, reducing the cleaning efficiency and even leading to cleaning failures. Before the data cleaning process starts, the monitoring mechanism not only checks whether the configuration of the cleaning task is accurate, covering the cleaning target range, data screening conditions, etc., but also evaluates whether there are potential risks of concurrent faults in the current system state. If there are potential risks, early warnings will be issued and corresponding countermeasures will be formulated.

[0144] After the data cleaning process starts, on the one hand, the system will track the execution progress of the process in real time, and detailed records of key information such as the amount of data that has been cleaned, the remaining amount of data, and the estimated completion time will be made; on the other hand, the system will continuously monitor the status of each controller. Once a concurrent fault is detected, the data cleaning process will be paused, and the concurrent fault will be processed first to avoid the interference of the concurrent fault on the data cleaning work. After the fault is processed, the cleaning process will be reasonably adjusted and the cleaning work will be resumed according to the system status and data cleaning progress at that time. If errors occur during the cleaning process, such as data reading failures, writing errors, etc., the monitoring system will quickly capture these exceptions and record the error information in detail to facilitate subsequent troubleshooting and repair by technicians, ensuring that the data cleaning process can be successfully completed and maintaining the stable operation of the storage system.

[0145] Step 104, when a concurrent fault occurs in the controller in the storage system, adjust the execution process of the data cleaning process according to the execution status information of the data cleaning process and the concurrent fault information of the concurrent fault controller where the concurrent fault occurs; according to the adjusted data cleaning process, clean the data of the transition controller.

[0146] In the embodiments of the present disclosure, the execution status information is information describing the current state of the data cleaning process, including whether it has been started, the execution progress (the proportion of the amount of data that has been cleaned to the total amount of data, etc.), whether errors or abnormal situations have been encountered (such as data reading failures, writing errors, etc.), whether it has been paused, etc. The concurrent fault information is information related to the controller where the concurrent fault occurs, such as which controllers have failed, the specific types of faults (hardware faults, software faults, etc.), the time when the faults occurred, the severity of the faults, etc.

[0147] During the recovery of the fault controller, it is necessary to clean the cached data in the transition controller. The system will monitor the working status of each controller in real time. Once it detects that a fault occurs in another controller in the storage system, that is, a concurrent fault occurs, it will trigger the corresponding response mechanism. At this time, the system will quickly collect two important aspects of information: on the one hand, it is the execution status information of the current data cleaning process, including whether the data cleaning process has been started. If it has been started, it is necessary to understand which specific step it has reached, whether there have been any errors or exceptions and how they were handled, etc.; on the other hand, it is to collect the detailed fault information of the controller where the concurrent fault occurs, such as the number of the fault controller, the type of the fault (such as whether it is a hardware chip damage or a software program crash), the exact time when the fault occurred, etc.

[0148] Based on the information collected, the system will adjust the execution process of the current data cleaning process. The adjustment methods are diverse. For example, if a part of the data cleaning process has been executed, but the fault of a key concurrent fault controller affects the subsequent data cleaning operations, then the system will pause the current data cleaning process, give priority to handling operations such as fault repair or data migration, etc. After the fault is resolved to a certain extent, it will re-plan the execution steps of the data cleaning process, determine where to continue cleaning and what method to use for cleaning. Or, if the data cleaning process has not been started, and the occurrence of the concurrent fault has caused a large change in the architecture of the storage system, the system will re-evaluate the goals, scope and methods of data cleaning, and formulate a new data cleaning execution process.

[0149] After completing the adjustment of the data cleaning process, the system will perform the data cleaning operation on the transition controller according to the new process. During the cleaning process, the system will still continuously monitor the progress and status of the cleaning to ensure that the data cleaning work can be completed smoothly and will not interfere with other normal operating parts of the storage system.

[0150] When complex concurrent faults occur in the storage system, it can adjust the data cleaning process in a timely manner according to the actual situation, enabling the system to better cope with various emergencies, reducing the impact of faults on the data cleaning work, and improving the overall adaptability and fault tolerance of the system. By reasonably adjusting the data cleaning process, it avoids incomplete or incorrect data cleaning caused by concurrent faults, ensures the correct cleaning of the data in the transition controller, thereby guaranteeing the integrity and consistency of the data in the storage system, and reducing the risk of data loss or inconsistency.

[0151] The present disclosure provides a method for cache data cleaning. In the embodiments of the present disclosure, a transition controller is used to replace a faulty controller for mirror recombination. After the faulty controller is repaired, it is added to the first recombined mirror pair to form a temporary mirror pair, which can ensure that the basic data redundancy and read / write functions of the storage system are not severely affected. When the faulty controller recovers, by continuously monitoring whether concurrent faults occur in the controllers in the storage system and the execution status of the data cleaning process, the system can perceive complex fault situations in real time. Once a concurrent fault occurs, the data cleaning process is adjusted according to the execution status information of the data cleaning process, the pairing information of the temporary mirror pair, and the concurrent fault information of the concurrently faulty controller, so as to flexibly cope with complex fault scenarios, reduce the impact of faults on the system, ensure that the data of the transition controller can be effectively cleaned in the case of concurrent faults, ensure that the system can quickly resume normal operation after a fault, and improve the fault tolerance of the storage system. When cleaning the cache data of the transition controller, through an effective fault handling mechanism and precise adjustment of the data cleaning process, the storage system can recover to the normal working state faster when facing faults. Whether it is a single controller fault or a concurrent fault, the system can handle it in an orderly manner, shorten the fault downtime, and improve the overall performance and efficiency of the storage system.

[0152] Further, in a possible implementation manner of the present disclosure, the storage system is a four-controller storage system. A four-controller storage system is a storage system with four controllers. In the subsequent embodiments, the four-controller storage system is used as an example for illustration. However, it should be clear that this illustration method is not intended to limit the number of controllers in the storage system.

[0153] To clearly illustrate the embodiments of the present disclosure, the embodiments of the present disclosure provide a flowchart of a method for mirror pair recombination.

[0154] As Figure 2 shown, the method includes the following steps:

[0155] Step 201, monitor the operating status of the controllers in the storage system and identify the faulty controller.

[0156] The operating status refers to various status information of the controller during operation, including but not limited to the CPU usage rate of the controller, memory occupancy, data transfer rate, temperature, power status, error log, etc. These status information can reflect the workload, health status, and whether the controller is operating normally of the controller. When the controller shows abnormal conditions and cannot perform its relevant duties such as data storage and processing normally, it becomes a faulty controller. The reasons for the fault may be various, such as hardware damage (such as chip failure, circuit board short circuit, etc.), software errors (such as program crash, driver failure, etc.), network connection problems, etc.

[0157] During the daily operation of the storage system, the running status of each controller is continuously monitored in real time through a preset monitoring mechanism. This monitoring mechanism can obtain the status information of the controller in various ways. For example, the system log of the controller is read regularly through a software program to obtain error information, warning information, etc. that occur during the operation of the controller; sensors are used to monitor the hardware status of the controller in real time, such as a temperature sensor monitors the working temperature of the controller, and a power sensor monitors the stability of the power supply; the performance metrics of the controller can also be obtained through network connection, such as CPU usage, memory occupancy, data transfer rate, etc.

[0158] After obtaining this status information, the monitoring mechanism will analyze and process it. Through preset rules and algorithms, the current status information is compared with the standard status during normal operation. For example, if the CPU usage of the controller exceeds 90% for a long time and is accompanied by a significant decrease in the data transfer rate, it is determined that there is a performance problem with the controller; if frequent hardware error records, such as hard disk read / write errors, are found in the system log, it means that there is a fault in the storage device connected to the controller or the hardware of the controller itself.

[0159] Once the monitoring mechanism determines through analysis that the running status of a certain controller is abnormal and meets the preset fault determination criteria, the controller will be identified as a faulty controller, and an alarm will be sent in a timely manner to notify the system administrator or automatically trigger the corresponding fault handling mechanism.

[0160] Step 202: Identify the device identifier of the faulty controller, collect the device identifiers of the controllers in the storage system except the faulty controller, and sort them to generate a set of device identifiers.

[0161] The device identifier is a unique identity identifier for each controller in the storage system, used to distinguish different controllers, facilitate system management and scheduling, and can be in the form of a digital number, letter code, etc. The set of device identifiers is a set formed by combining multiple device identifiers in a certain order, which is convenient for unified management and operation of these controllers.

[0162] The storage system continuously monitors the running status of each controller through a built-in monitoring mechanism. Once an abnormality is detected in a certain controller, such as frequent error reporting, data transfer interruption, excessive hardware temperature, etc., the system will determine that the controller is a faulty controller and identify the device identifier of the faulty controller. This can be achieved through operations such as querying the configuration file of the controller and interacting with the management interface of the controller to obtain the accurate device identifier.

[0163] After determining the faulty controller, the system starts to traverse all the controllers in the storage system and filters out the other normally operating controllers except the faulty one. For each normal controller, its device identifier is collected by reading its device information, querying the system management database, etc. The device identifiers of the collected normal controllers are sorted according to certain rules. The sorting rules need to be determined according to the type of the device identifier. For example, if the device identifier is a numerical number, it can be sorted in ascending or descending order; if it is an alphabetic code, it can be sorted in alphabetical order. After sorting, these ordered device identifiers are combined into a device identifier set and stored in a specific data structure of the system for convenient subsequent operation calls.

[0164] Suppose a storage system has four controllers with device identifiers C01, C02, C03, and C04 respectively. The system detects that controller C03 has a fault, specifically manifested as being unable to respond to data read and write requests and there are a large number of hardware error prompts in the system log. The system first identifies the device identifier of the faulty controller as C03. Then, it starts to collect the device identifiers of the other controllers except C03, namely C01, C02, and C04. Then, these identifiers are sorted in alphabetical order, and the sorted result is C01, C02, C04. Finally, they are combined into a device identifier set {C01, C02, C04} and stored in the system management database for use when dealing with faults subsequently.

[0165] Step 203: According to the pairing information of the initial mirror pair where the faulty controller is located, filter out the transitional controller that replaces the faulty controller from the controllers that have not failed.

[0166] Step 2031: Obtain the set of device identifiers of the controllers in the storage system except the faulty controller.

[0167] Step 2032: Based on the pairing information, determine the normal device identifier corresponding to the normal controller in the initial mirror pair where the faulty controller is located.

[0168] Step 2033: According to the set of device identifiers and the normal device identifier, find the transitional controller that replaces the faulty controller.

[0169] Furthermore, find the target device identifier that is the next one after the normal device identifier in the set of device identifiers, and use the controller corresponding to the target device identifier as the transitional controller.

[0170] Specifically in step 203, the teaming information is used to describe the relevant information on how each controller forms a mirror pair, including the corresponding relationship of the controllers in each mirror pair, and is the key basis for determining the composition of the mirror pair where the faulty controller is located. A normal controller is a controller in the storage system that can normally execute tasks such as data processing, storage, and transmission without any faults. The normal device identifier is the unique device identifier corresponding to the normal controller, which is used to accurately identify and locate the controller in the system.

[0171] Among them, in step 2031, obtain the set of device identifiers of the controllers in the storage system except the faulty controller. The method can be but is not limited to the following, specifically including:

[0172] There is a set of management mechanisms inside the storage system for recording and managing the device identifiers of all controllers. When a certain controller is detected to have a fault, the system first traverses all the controllers in the storage system through this management mechanism. The system reads the configuration information of each controller one by one, extracts their respective device identifiers from it, and aggregates these identifiers to form a set containing the device identifiers of all other controllers except the faulty controller. This set provides the basic data for subsequent screening of the transition controller.

[0173] Among them, in step 2032, based on the teaming information, determine the normal device identifier corresponding to the normal controller in the initial mirror pair where the faulty controller is located, specifically including:

[0174] According to the pre-recorded teaming information in the system, the system can quickly locate the initial mirror pair where the faulty controller is located. In this initial mirror pair, the system identifies another normally operating controller except the faulty controller and obtains the device identifier corresponding to this normal controller, that is, the normal device identifier. This normal device identifier is the key reference for determining the transition controller.

[0175] Among them, in step 2033, according to the set of device identifiers and the normal device identifier, find the transition controller to replace the faulty controller, specifically including:

[0176] The system compares and analyzes the set of device identifiers obtained in step 2031 with the normal device identifiers determined in step 2032. According to the cyclic mirroring rule, the system searches for the target device identifier in the set of device identifiers that is the next one after the normal device identifier (here, the "next one" is determined according to the encoding rule of the device identifier. For example, if the device identifier is numbered in numerical order, then the next one is the numerical value plus 1; if it is encoded in alphabetical order, then it is the next letter). Once the target device identifier is found, the system can determine the controller corresponding to the target device identifier and select this controller as the transitional controller to replace the faulty controller. Subsequently, the transitional controller will re-team with the normal controller in the initial mirror pair where the faulty controller is located to form a temporary mirror pair, and continue to maintain the data processing and storage functions of the storage system.

[0177] Suppose a storage system has four controllers 0, 1, 2, 3, with device identifiers C00, C01, C02, C03 respectively, and they form initial mirror pairs according to the cyclic mirroring principle, such as (0,1) (1,2) (2,3) (3,0). When controller 3 fails, step 2031: The system obtains the set of controller device identifiers {C00, C01, C02} except controller 3. Step 2032: According to the teaming information, determine the normal controllers 2 and 0 in the initial mirror pairs (2,3) and (3,0) where controller 3 is located, and their corresponding normal device identifiers are C02 and C00. Step 2033: According to the alphabetical order of the device identifiers, search for the next one after C02 in the set of device identifiers, that is, C00; according to the cyclic mirroring principle, when the controller is ranked at the end, the next one is the first one. So for the initial mirror pair (2,3), determine the controller 0 corresponding to C00 as the transitional controller; for the initial mirror pair (3,0), determine the controller 1 corresponding to C01 as the transitional controller. At this time, use controller 0 and 1 to replace the work of controller 3 in the initial mirror pairs respectively, that is, form two first reorganized mirror pairs (2,0) (0,1). At this time, the mirror pairs in the storage system include two initial mirror pairs and two first reorganized mirror pairs, that is, (0,1) (1,2) (2,0) (0,1).

[0178] Step 204, perform mirror reorganization on the transitional controller and the normal controller in the initial mirror pair where the faulty controller is located to obtain the first reorganized mirror pair.

[0179] Furthermore, according to the cyclic mirroring principle, sort the device identifiers of the transitional controller and the normal controller to determine the arrangement order of the transitional controller and the normal controller, so as to generate the first reorganized mirror pair.

[0180] Specifically in step 204, after the storage system detects a controller failure and determines the transition controller, the two controllers that need to perform mirror recombination are identified, namely the normal controller in the initial mirror pair where the transition controller and the failed controller are located. The storage system respectively obtains the device identifiers of the transition controller and the normal controller. These device identifiers are the unique identity identifiers of each controller in the system, which can be in the form of digital numbers, letter codes, etc., and are used to accurately identify and distinguish different controllers. According to the established cyclic mirror principle of the storage system, the device identifiers of the transition controller and the normal controller are sorted. If the cyclic mirror principle is to combine them in alphabetical or numerical order, the system will determine the order of the two device identifiers according to this rule. For example, if the device identifier is a letter code and the cyclic mirror principle is in alphabetical order, then the alphabetical order of the two identifiers is compared to determine their arrangement order in the new mirror pair.

[0181] According to the sorted result, the transition controller and the normal controller are combined together in the determined arrangement order to form the first recombined mirror pair. This new mirror pair will replace the original initial mirror pair containing the failed controller and undertake the tasks of data storage and processing, ensuring that the storage system can continue to operate normally during the failure.

[0182] Based on the example in step 203, for the initial mirror pair (2, 3), the controller 0 corresponding to C00 is determined as the transition controller; for the initial mirror pair (3, 0), the controller 1 corresponding to C01 is determined as the transition controller. At this time, controllers 0 and 1 are respectively used to replace the work of controller 3 in the initial mirror pair, that is, two first recombined mirror pairs (2, 0) and (0, 1) are formed. At this time, the mirror pairs in the storage system include two initial mirror pairs and two first recombined mirror pairs, namely (0, 1), (1, 2), (2, 0), (0, 1).

[0183] To clearly illustrate the embodiments of the present disclosure, the embodiments of the present disclosure provide a flowchart of another method for clearing cached data.

[0184] As Figure 3 shown, this method includes the following steps:

[0185] Step 301, add the failed controller that has repaired the fault to the first recombined mirror pair to obtain a temporary mirror pair.

[0186] Step 302, control the failed controller that has repaired the fault to perform data mirroring with the transition controller.

[0187] Step 303, after the failed controller that has repaired the fault completes data mirroring, remove the transition controller from the temporary mirror pair and execute the data cleaning process on the transition controller.

[0188] Specifically, in steps 301 to 303, in the storage system, when the failed controller is repaired or automatically restored by the system, the storage system detects that the failed controller has returned to the normal working state. The system adds the repaired failed controller to the first recombined mirror pair generated previously. For example, in storage system A, which is internally equipped with four controllers numbered 0, 1, 2, and 3. These four controllers form four mirror pairs in "State 1": (0,1), (1,2), (2,3), and (3,0), which are the initial mirror pairs when the system was initially working properly.

[0189] When controller 0 in mirror pair 1 (0,1) suddenly fails, to maintain the normal operation of the system, a controller needs to be introduced to temporarily replace the work of controller 0. The transition controller is selected according to the established cyclic mirror principle, that is, based on the device identifier of the normally working controller in the mirror pair where the failed controller is located, the controller with the device identifier one position later is selected as the transition controller. When controller 0 fails, the storage system can identify that the controllers in the normal working state are 1, 2, and 3. At this time, mirror pair 1 and mirror pair 4 are affected by the failure of controller 0 and need to use the transition controller to replace the failed controller to perform mirror recombination operations. After recombination, the combination of mirror pairs changes to "State 2": (1,2), (1,2), (2,3), (3,1). Among them, mirror pair 1 and mirror pair 4 belong to the first recombined mirror pairs after recombination, while mirror pair 2 and mirror pair 3 remain in the state of the initial mirror pairs. After the failed controller 0 repairs the fault, it is added to mirror pair 1 (the first recombined mirror pair). At this time, the temporary mirror pair becomes (0,1,2), and all the mirror pairs of the storage system are in "State 3": (0,1,2), (1,2), (2,3), (3,1,0).

[0190] After the temporary mirror pair is formed, the system starts the data mirror processing flow. The system controls the failed controller that has repaired the fault to perform data mirroring with the transition controller. The system first determines the data range that needs to be synchronized, which includes the data newly written or modified by other controllers during the failure period of the failed controller.

[0191] When the failed controller that has repaired the fault completes data mirroring, that is, when the data in the transitional controller is copied to the repaired failed controller completely and correctly, the system confirms that the data synchronization is completed. The system removes the transitional controller from the temporary mirror pair. This operation includes updating the configuration information of the storage system, deleting the transitional controller from the relevant mirror pair configuration, so that it no longer participates in the data storage and processing of this mirror pair. That is, switch the above "State 3" to "State 1". Then, a data cleaning process is executed on the transitional controller. The system first scans the storage area of the transitional controller to identify the temporary data generated during its operation as a substitute for the failed controller. These data are stored in specific directories or data structures, and the system deletes, archives or performs other processing methods on these temporary data according to the pre-set cleaning rules.

[0192] Step 304, monitor whether concurrent faults occur in the controllers in the storage system, and monitor the execution status of the data cleaning process.

[0193] Specifically in Step 304, the working parameters and status information of each controller are collected in real time, such as the CPU usage rate of the controller, memory occupancy, data transfer rate, error log, etc. By analyzing and judging these information, it is determined whether a controller has failed. In particular, for the monitoring of concurrent faults, the system will pay attention to the state change trends of multiple controllers.

[0194] At the same time, the system will also comprehensively monitor the data cleaning process of the failed controller. Before the data cleaning process starts, the monitoring mechanism will check whether the configuration of the cleaning task is correct, including the target scope of cleaning, data screening conditions, etc. When the data cleaning process starts, it will track the execution progress of the process in real time, record information such as the amount of data that has been cleaned, the remaining amount of data, and the estimated completion time. If errors occur during the cleaning process, such as data reading failure, writing error, etc., the monitoring system will promptly capture these abnormal situations and record detailed error information for technicians to troubleshoot and repair.

[0195] Step 305, when concurrent faults occur in the controllers in the storage system, based on the execution status information, judge whether the transitional controller starts to execute the data cleaning process.

[0196] Specifically in step 305, once a concurrent failure is detected, the system triggers an operation to obtain the execution status information of the data cleaning process. The data cleaning process management module in the system is responsible for recording and maintaining the execution status information of the data cleaning process. The data cleaning process management module queries relevant status record files or database tables to obtain detailed information about the data cleaning process of the transition controller. This information includes whether the data cleaning process has been started, if it has been started, what the current execution progress is, whether any errors or exceptions have occurred during the execution process, and whether the process is in a paused state, etc.

[0197] Based on the obtained execution status information, the system determines whether the transition controller has started to execute the data cleaning process. If the execution status information shows that the start flag of the data cleaning process is "started", it indicates that the transition controller has started to execute the data cleaning process; if the start flag is "not started", it means that the transition controller has not started to execute the data cleaning process. For the case where the data cleaning process has started to execute, the system further analyzes the execution progress, error information, etc., to comprehensively understand the current status of the data cleaning process and provide a basis for subsequent decisions. To determine whether the concurrent failure controller will affect the data cleaning process.

[0198] When it is determined that the data cleaning process has not started to execute, step 306 is performed.

[0199] Step 306, perform mirror recombination on the mirror pair where the concurrent failure controller is located to obtain a second recombined mirror pair.

[0200] Furthermore, performing mirror recombination on the mirror pair where the concurrent failure controller is located to obtain a second recombined mirror pair includes: according to the concurrent failure information, searching for the teaming information of the first failed mirror pair where the concurrent failure controller is located; based on the teaming information of the first failed mirror pair, determining whether the first failed mirror pair is a temporary mirror pair for removing the transition controller. If the first failed mirror pair is a temporary mirror pair for removing the transition controller, then pull the transition controller back into the temporary mirror pair and perform mirror recombination with the normal controller in the temporary mirror pair to obtain a second recombined mirror pair; if the first failed mirror pair is not a temporary mirror pair for removing the transition controller, then perform mirror recombination on the normal controller in the first failed mirror pair and the controller adjacent to its next position to obtain a second recombined mirror pair.

[0201] Specifically in step 306, the fault monitoring module running in real time in the storage system continuously monitors the working status of each controller. Once multiple controllers are detected to fail simultaneously, it is determined as a concurrent fault, and the fault monitoring module will collect detailed concurrent fault information. It will record the unique device identifier of each concurrent fault controller to accurately identify the fault source; record the timestamp of the fault occurrence for subsequent fault analysis and system recovery strategy formulation; and will also record in detail the fault types, such as chip overheating and damage at the hardware level, circuit board short circuit, program crash and driver error at the software level, etc., and the impacts of these faults on the current data transmission, storage and processing functions of the system, such as data transmission interruption, data loss risk prompt, etc. The first fault mirror pair is the mirror pair containing the concurrent fault controllers. These mirror pairs, due to the failure of one or more controllers, cannot perform the duties of data storage and backup normally and need to be recombined and adjusted.

[0202] After the system obtains the concurrent fault information, according to the mirror pair configuration database stored internally, it looks up the teaming information of the first fault mirror pair containing the concurrent fault controllers. This configuration database details the composition of each mirror pair. By querying the identifier of the fault controller, the system can quickly locate the first fault mirror pair it belongs to and obtain the relevant information of other controllers in this mirror pair, including their device identifiers, cooperation relationships with the fault controller, etc.

[0203] The system compares the teaming information of the first fault mirror pair found with the relevant records of the temporary mirror pair. The information of the temporary mirror pair is also stored in a specific database table of the system, including the composition history of the temporary mirror pair, the usage record of the transition controller, and the association information with the original fault controller, etc. By comparing key information such as the composition structure of the first fault mirror pair, whether it involves a transition controller, and the removal record of the transition controller, it is determined whether the first fault mirror pair is a temporary mirror pair that removes the transition controller. If the first fault mirror pair is a temporary mirror pair that removes the transition controller, the system first pulls the transition controller back into the temporary mirror pair. This involves updating the system configuration information and re - establishing the communication connection and cooperation relationship between the transition controller and other normal controllers in the temporary mirror pair. Then, the system performs mirror recombination on the transition controller and the normal controllers in the temporary mirror pair. According to the established mirror recombination algorithm and rules, it determines their arrangement order and data interaction method in the new second recombined mirror pair. For example, according to specific priority rules or load - balancing principles, it redistributes the data storage and processing tasks to ensure that the new mirror pair can operate efficiently and stably.

[0204] If the first failed mirror pair is not the temporary mirror pair for removing the transition controller: The system finds the normally operating controller from the first failed mirror pair. According to the mirror reorganization strategy preset in the storage system, it finds the controller after the normal controller (the "controller after" is determined according to the controller numbering rule or the order defined by the system). It performs mirror reorganization between the normal controller and the controller after it. During the reorganization process, the system will reconfigure the data synchronization mechanism, communication protocol, data storage strategy, etc. between the two controllers to generate the second reorganized mirror pair. At the same time, the system will conduct comprehensive testing and verification on the newly generated second reorganized mirror pair to ensure that it can complete the data storage and backup tasks normally and guarantee the stable operation of the storage system after concurrent failures.

[0205] Exemplarily, to recover the failed controller 0, it is necessary to remove the transition controller 2 (in mirror pair 1) and 1 (in mirror pair 4). Subsequently, a data cleaning process is carried out, and all mirror pairs of the storage system are in "status 3": (0, 1, 2) (1, 2) (2, 3) (3, 1, 0). If controller 1 has a concurrent failure at this time (concurrent failure relative to the failed controller), the controllers of the storage system perform mirror reorganization, and the mirror pair grouping is updated to "status 4": (0, 2) (2, 3) (2, 3) (3, 0). At this time, the mirror pairs for mirror reorganization are mirror pair 1 (the second reorganized mirror pair) and mirror pair 2 (the second reorganized mirror pair); mirror pair 3 and mirror pair 4 are restored to the state of the initial mirror pairs. In "status 4", the controller 2 (transition controller) of mirror pair 1 is pulled back into mirror pair 1, so there is no need to execute the data cleaning process; while the removed controller 1 (transition controller) in mirror pair 4 needs to continue to execute the data cleaning process.

[0206] Step 307: Based on the grouping information of the second reorganized mirror pair, determine whether the transition controller belongs to the second reorganized mirror pair to determine whether to continue executing the data cleaning process.

[0207] If the transition controller belongs to the second reorganized mirror pair, there is no need to continue executing the data cleaning process for the transition controller in the second reorganized mirror pair; if the transition controller does not belong to the second reorganized mirror pair, continue to execute the data cleaning process for the transition controller that does not belong to the second reorganized mirror pair.

[0208] Specifically in step 307, the second reorganized mirror pair is the new mirror pair obtained after reorganizing the mirror pair where the concurrent failure controller is located when the storage system has a concurrent failure. Its purpose is to reconstruct the mirror structure of the storage system in complex failure situations and maintain the stability and reliability of data storage and processing.

[0209] After the storage system completes the reorganization of the mirror pair where the concurrent failure controller is located and generates the second reorganized mirror pair, the system extracts the teaming information of this mirror pair from its configuration management module. The configuration management module is responsible for recording and maintaining the detailed configuration data of all mirror pairs in the storage system. The system obtains information such as the device identifiers of each controller in the second reorganized mirror pair, their positional relationships in the mirror pair, and the data synchronization and cooperation rules between them by querying specific data tables or configuration files. This information is stored in a structured form for the system to quickly read and analyze.

[0210] The system compares the device identifier of the transition controller with the controller identifiers in the teaming information of the second reorganized mirror pair one by one. This comparison process can be implemented through a simple string matching algorithm. If a matching identifier is found, it is determined that the transition controller belongs to the second reorganized mirror pair; if no match is found, it is determined that the transition controller does not belong to the second reorganized mirror pair.

[0211] Since the transition controller is still participating in normal work in the second reorganized mirror pair, performing the data cleaning process on it at this time will cause data loss or abnormal system operation. Therefore, the system stops the data cleaning process arrangement for this transition controller and records relevant information in the system log, indicating that the data cleaning is not performed temporarily because the transition controller is still in the working mirror pair. At the same time, the system continuously monitors the working status of the transition controller in the second reorganized mirror pair to ensure that it performs its duties normally until there are new system changes or instructions requiring further processing of it.

[0212] When the system confirms that the transition controller has completed its temporary task and is no longer participating in the current core mirror pair work, in order to optimize the storage system resources and improve the system performance, the system will continue to execute the data cleaning process for this transition controller. The system will start the data cleaning task according to the pre-set data cleaning rules and steps. First, the system scans the storage area of the transition controller to identify temporary data, invalid data, etc. that need to be cleaned. Then, according to the type and storage location of the data, corresponding cleaning methods are adopted, such as directly deleting files, formatting specific storage areas, etc. During the cleaning process, the system monitors the cleaning progress in real time and records any errors or abnormal situations that occur during the cleaning process for subsequent troubleshooting and handling to ensure the smooth completion of the data cleaning process. The data cleaning process referred to in the embodiments of the present disclosure is specifically the invalid data cleaning (discard) of the cache module

[0213] Continuing with the example in step 306, the mirror pairs in the storage system are in "State 4": (0, 2) (2, 3) (2, 3) (3, 0). At this time, the mirror pairs for mirror reorganization are mirror pair 1 (the second reorganized mirror pair) and mirror pair 2 (the second reorganized mirror pair); mirror pair 3 and mirror pair 4 are restored to the state of the initial mirror pairs. In "State 4", the controller 2 (transition controller) of mirror pair 1 is pulled back to mirror pair 1, so there is no need to execute the data cleaning process; while the cleared controller 1 (transition controller) in mirror pair 4 needs to continue to execute the data cleaning process.

[0214] To clearly illustrate the embodiments of the present disclosure, the embodiments of the present disclosure provide a flowchart of another method for caching data cleaning.

[0215] As Figure 4 shown, the method includes the following steps:

[0216] Step 401, adding the failed controller that has repaired the fault to the first reorganized mirror pair to obtain a temporary mirror pair.

[0217] Step 402, controlling the failed controller that has repaired the fault and the transition controller to perform data mirroring.

[0218] Step 403, after the failed controller that has repaired the fault completes data mirroring, removing the transition controller from the temporary mirror pair and performing a data cleaning process on the transition controller.

[0219] Step 404, monitoring whether a concurrent fault occurs in the controllers in the storage system and monitoring the execution status of the data cleaning process.

[0220] Step 405, when a concurrent fault occurs in the controllers in the storage system, based on the execution status information, determining whether the transition controller has started to execute the data cleaning process.

[0221] For the description of step 405, please refer to the description of the above embodiments, and this embodiment will not be elaborated one by one.

[0222] When it is determined that the data cleaning process has started to be executed, step 406 is executed.

[0223] Step 406, based on the concurrent fault information, determining whether the transition controller is a concurrent fault controller.

[0224] Specifically in step 406, based on the concurrent fault information, obtaining the device identifier of the concurrent fault controller and the teaming information of the mirror pair where it is located; comparing the device identifier of the concurrent fault controller with the device identifier of the transition controller, and comparing the teaming information of the mirror pair where the concurrent fault controller is located with the teaming information of the temporary mirror pair where the transition controller is located to determine whether the transition controller is a concurrent fault controller.

[0225] The data cleaning process monitoring module built into the storage system will track the execution of the data cleaning process in real time. Once it detects that the data cleaning process has been started, the monitoring module will trigger the next judgment operation. At the same time, the fault monitoring module of the storage system continuously monitors the operating status of each controller. When it detects that there are concurrent faults in the controllers, the fault monitoring module will record the detailed information of all the concurrently faulty controllers and generate concurrent fault information.

[0226] The system compares the device identifier of the transitional controller with the device identifiers of the concurrently faulty controllers one by one. If the device identifier of the transitional controller is the same as that of a certain concurrently faulty controller and the pairing information of the mirror pair is also the same, then it can be determined that the transitional controller is a concurrently faulty controller. If, after comparison, the device identifier of the transitional controller does not match any of the device identifiers of the concurrently faulty controllers, it means that the transitional controller is not a concurrently faulty controller. If the transitional controller is a concurrently faulty controller, step 407 is executed; if the transitional controller is not a concurrently faulty controller, step 408 is executed.

[0227] Step 407: When the transitional controller rejoins the mirror pair, re - execute the data cleaning process for the transitional controller.

[0228] Specifically in step 407, during the daily operation of the storage system, a controller may fail, thus introducing a transitional controller to maintain the normal operation of the system. When the faulty controller is repaired and ready to rejoin the mirror pair, that is the scenario where the transitional controller rejoins the mirror pair. Before this, the system runs in various states. For example, when the controller recovery is in the second stage and the cache module has initiated the data cleaning process, if the controller executing discard fails, at this time the cluster will update the mirror pair pairing information according to the current online controllers. This is because during the execution of the discard process, after the controller executing discard starts to execute discard, the cache module has removed the controller from the corresponding mirror pair. So when the controller rejoins the cluster again, it will also initiate the discard process for the corresponding uncleaned cache data again during the controller initialization stage.

[0229] Step 408: After executing the data cleaning process for the transitional controller, perform data mirroring on the transitional controller and perform mirror recombination on the mirror pair where the concurrently faulty controller is located to obtain a third recombined mirror pair.

[0230] Further, performing data mirroring on the transition controller and mirror recombination on the mirror pair where the concurrent failure controller is located to obtain a third recombined mirror pair includes: according to the concurrent failure information, finding the teaming information of the second failure mirror pair where the concurrent failure controller is located; based on the teaming information of the second failure mirror pair, determining whether the second failure mirror pair is a temporary mirror pair for removing the transition controller. If the second failure mirror pair is a temporary mirror pair for removing the transition controller, pulling the transition controller back into the temporary mirror pair and performing mirror recombination with the normal controller in the temporary mirror pair to obtain a third recombined mirror pair; if the second failure mirror pair is not a temporary mirror pair for removing the transition controller, performing mirror recombination on the normal controller in the second failure mirror pair and the controller at the next adjacent position to obtain a third recombined mirror pair.

[0231] Specifically in step 408, during the operation of the storage system, when the system detects concurrent failures of multiple controllers, a series of complex and critical countermeasures will be taken, and step 408 is an important part of them.

[0232] After the data cleaning process is completed for the transition controller, to ensure the integrity and consistency of the system data, it is necessary to perform data mirroring on the transition controller and mirror recombination on the mirror pair where the concurrent failure controller is located, so as to obtain a third recombined mirror pair. This operation is crucial for restoring the stable operation of the storage system and ensuring data security. Specifically, when performing this step, it is necessary to work based on the concurrent failure information first. The system will accurately find the teaming information of the second failure mirror pair where the concurrent failure controller is located according to the concurrently collected and recorded failure information, including key data such as which controllers have failed, the type and time of the failure, etc. These teaming information details the combination relationship of each controller in the second failure mirror pair and is an important basis for subsequent operations. Then, based on the obtained teaming information of the second failure mirror pair, the system needs to determine whether this mirror pair is a temporary mirror pair for removing the transition controller. If the second failure mirror pair is a temporary mirror pair for removing the transition controller, the system will pull the transition controller back into the temporary mirror pair. This involves re - establishing the communication connection and data interaction relationship between the transition controller and the normal controller in the temporary mirror pair, and then performing mirror recombination on them according to the established mirror recombination rules to obtain a third recombined mirror pair. If the judgment result is that the second failure mirror pair is not a temporary mirror pair for removing the transition controller, the system will perform mirror recombination on the normal controller in the second failure mirror pair and the controller at the next position (here, "the next position" is determined according to the preset controller sorting rule of the storage system). During the recombination process, the system will re - configure key parameters such as data synchronization policies and communication protocols, and finally obtain a third recombined mirror pair. Through such a rigorous process, the storage system can effectively restore the normal order of data storage and management and ensure the stable operation of the system after concurrent failures occur.

[0233] Exemplarily, the normal state of the mirror pair grouping in the storage system is "State 1": (0,1) (1,2) (2,3) (3,0); after controller 0 fails, the mirror pair grouping in the storage system temporarily becomes "State 2": (1,2) (1,2) (2,3) (3,1); the failed controller 0 joins the first reorganized mirror pair to obtain temporary mirror pairs (0,1,2) and (3,1,0); at this time, the mirror pair grouping in the storage system becomes "State 3": (0,1,2) (1,2) (2,3) (3,1,0). After removing the transition controller 2 from mirror pair 1, a discard process needs to be performed on the transition controller 2. Similarly, after removing the transition controller 1 from mirror pair 4, a discard process needs to be performed on the transition controller 1. Assume that the transition controller 1 fails and exits at this time. For the transition controller 2, it is not a concurrent failure controller that has a concurrent failure. However, since controller 1 fails and exits mirror pair 1, and mirror pair 1 is also the temporary mirror pair from which the transition controller is removed.

[0234] When controller 1 fails and exits, the cluster will immediately update the mirror pair grouping situation. At the same time, the system will record the discard information as (2,32,32,1). This set of information indicates that in the first mirror pair, the controller that performs the discard operation is controller 2; in the fourth mirror pair, the controller that performs the discard operation is controller 1; in the remaining two mirror pairs, there is no controller that performs the discard operation, and here 32 is the default setting representing an invalid value.

[0235] Specifically looking at the first mirror pair, before controller 1 exits, the read and write operations of the data are carried out according to the grouping method of State 1. Therefore, even if controller 2 is re-incorporated into the mirror pair grouping, the data it stores is already old data. In this case, it is necessary to wait for the discard process to be completed before mirroring the cached data from controller 0 to controller 2 again to ensure data consistency and accuracy. When the entire discard process is completed, the system will clear all the recorded discard status information and restore it to the default invalid value state (32,32,32,32) to prepare for subsequent system operations and data management.

[0236] It should be noted that in the embodiments of the present disclosure, the mirror pair grouping information of the storage system includes the sorting relationship between mirror pairs.

[0237] Furthermore, in a possible implementation manner of the embodiments of the present disclosure, the method for clearing cached data further includes: if no concurrent failure occurs in the controllers in the storage system, the transition controllers are cleared of data according to the data clearing process.

[0238] Specifically, during the continuous operation of the storage system, the system continuously monitors the operating status of each controller in real time. If it is determined through monitoring that there is no concurrent failure of the controllers in the storage system, it means that the overall operation of the system is relatively stable. At this time, the data cleaning process of the transition controller can be carried out in an orderly manner according to the pre-set data cleaning process. Before starting the data cleaning process, the system checks whether the relevant configurations are accurate, including the cleaning scope, data screening conditions, etc. After starting, the system tracks the cleaning progress in real time, records the amount of data that has been cleaned and the remaining data, ensures the smooth completion of the cleaning process, thereby releasing the storage resources of the transition controller and improving the overall performance of the storage system.

[0239] Further, in a possible implementation manner of the embodiment of the present disclosure, the method for cleaning cached data further includes: obtaining the status information of each mirror pair in the storage system, and marking the device identifiers of the transition controllers that need to execute the data cleaning process in each mirror pair.

[0240] Specifically, the system periodically scans each mirror pair in the storage system to collect its status information, including data synchronization status, data read / write rate, communication status between controllers, etc. Taking the failure of controller 1 at a certain moment as an example, the cluster will update the mirror pair grouping accordingly, such as updating to "status 4" (0,2) (2,3) (2,3) (3,0), and record the information related to the operation, such as the discard information (2,32,32,1), to clarify the controllers that perform the discard operation in each mirror pair. During this process, the system determines whether there are transition controllers that need to execute the initial cleaning process based on the mirror pair status and related operation records. For example, in the first mirror pair, before the failure of controller 1, data reading and writing were carried out according to the original grouping method, resulting in the data of controller 2 being old data although it was pulled back into the mirror pair again. It needs to wait for the discard process to complete and then re-mirror the cached data. At this time, controller 2 becomes a transition controller that needs to execute the initial cleaning process. Once such a transition controller is determined, the system quickly marks its device identifier for subsequent targeted initial cleaning process. After the discard process is completed, the relevant records will be cleared to invalid values to ensure the timely update of the system status and accurate management of data.

[0241] It should be noted that there may be multiple steps in the embodiments of the present disclosure. For the convenience of description, these steps are numbered, but these numbers are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.

[0242] Corresponding to the above method for clearing cached data, the present invention also provides an apparatus for clearing cached data. Since the apparatus embodiments of the present invention correspond to the above method embodiments, details not disclosed in the apparatus embodiments can be referred to the above method embodiments and will not be elaborated herein.

[0243] Figure 5 FIG. is a schematic structural diagram of an apparatus for clearing cached data provided by an embodiment of the present disclosure, as Figure 5 shown, including:

[0244] A recombination unit 51, configured to replace a faulty controller in a storage system with a transition controller, and perform mirror recombination on an initial mirror pair where the faulty controller is located to obtain a first recombined mirror pair; add the faulty controller that has repaired the fault to the first recombined mirror pair to obtain a temporary mirror pair;

[0245] A removal unit 52, configured to remove the transition controller from the temporary mirror pair after the faulty controller that has repaired the fault completes data mirroring, and perform a data cleaning process on the transition controller;

[0246] A first monitoring unit 53, configured to monitor whether a concurrent fault occurs in the controllers in the storage system, and monitor the execution status of the data cleaning process;

[0247] An adjustment unit 54, configured to, when a concurrent fault occurs in the controllers in the storage system, adjust the execution process of the data cleaning process according to the execution status information of the data cleaning process and the concurrent fault information of the concurrent faulty controllers that have occurred; perform data cleaning on the transition controller according to the adjusted data cleaning process.

[0248] The present disclosure provides an apparatus for cache data cleaning. In the embodiments of the present disclosure, a transition controller is used to replace a faulty controller for mirror recombination. After the faulty controller is repaired, it is added to the first recombined mirror pair to form a temporary mirror pair, which can ensure that the basic data redundancy and read / write functions of the storage system are not severely affected. When the faulty controller recovers, by continuously monitoring whether concurrent faults occur in the controllers in the storage system and the execution status of the data cleaning process, the system can perceive complex fault situations in real time. Once a concurrent fault occurs, the data cleaning process is adjusted according to the execution status information of the data cleaning process, the teaming information of the temporary mirror pair, and the concurrent fault information of the concurrently faulty controller, so as to flexibly cope with complex fault scenarios, reduce the impact of faults on the system, ensure that the data of the transition controller can be effectively cleaned in the case of concurrent faults, ensure that the system can quickly resume normal operation after a fault, and improve the fault tolerance of the storage system. When cleaning the cache data of the transition controller, through an effective fault handling mechanism and precise adjustment of the data cleaning process, the storage system can recover to the normal working state faster when facing faults. Whether it is a single controller fault or a concurrent fault, the system can handle it in an orderly manner, shorten the fault downtime, and improve the overall performance and efficiency of the storage system.

[0249] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the recombination unit 51 includes:

[0250] A screening module 511, configured to screen a transition controller for replacing the faulty controller from the controllers that have not failed according to the teaming information of the initial mirror pair where the faulty controller is located;

[0251] A recombination module 512, configured to perform mirror recombination on the transition controller and the normal controller in the initial mirror pair where the faulty controller is located to obtain a first recombined mirror pair.

[0252] Further, in a possible implementation manner of this embodiment, the screening module 511 is further configured to:

[0253] Obtain a set of device identifiers of the controllers in the storage system except the faulty controller, where the device identifiers in the set of device identifiers are arranged in sequence;

[0254] Based on the teaming information, determine the normal device identifier corresponding to the normal controller in the initial mirror pair where the faulty controller is located;

[0255] According to the set of device identifiers and the normal device identifier, search for a transition controller to replace the faulty controller.

[0256] Further, in a possible implementation manner of this embodiment, searching for a transition controller to replace the faulty controller according to the set of device identifiers and the normal device identifier includes:

[0257] Find the target device identifier that is the one after the normal device identifier in the set of device identifiers.

[0258] Use the controller corresponding to the target device identifier as the transitional controller.

[0259] Further, in a possible implementation manner of this embodiment, the recombination module 512 is further configured to:

[0260] Sort the device identifier of the transitional controller and the device identifier of the normal controller according to the principle of circular mirroring to determine the arrangement order of the transitional controller and the normal controller, so as to generate the first recombination mirror pair.

[0261] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the device for cache data cleaning further includes:

[0262] The second monitoring unit 55 is configured to monitor the running state of the controllers in the storage system and identify the faulty controller before the recombination unit 51 replaces the faulty controller in the storage system with the transitional controller.

[0263] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the device for cache data cleaning further includes:

[0264] The collection unit 56 is configured to identify the device identifier of the faulty controller, collect and sort the device identifiers of the controllers in the storage system except the faulty controller, and generate a set of device identifiers.

[0265] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the device for cache data cleaning further includes:

[0266] The control unit 57 is configured to control the faulty controller that has repaired the fault and the transitional controller to perform data mirroring processing after the recombination unit 51 adds the faulty controller that has repaired the fault to the first recombination mirror pair to obtain a temporary mirror pair.

[0267] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the device for cache data cleaning further includes:

[0268] The cleaning unit 58 is configured to perform data cleaning on the transitional controller according to the data cleaning process when there is no concurrent fault in the controllers in the storage system.

[0269] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the adjustment unit 54 includes:

[0270] The first judgment module 541 is configured to judge whether the transition controller starts to execute the data cleaning process based on the execution status information;

[0271] The determination module 542 is configured to, when the data cleaning process has not started, perform mirror recombination on the mirror pair where the concurrent fault controller is located to obtain a second recombined mirror pair, and judge whether the transition controller belongs to the second recombined mirror pair based on the teaming information of the second recombined mirror pair, so as to determine whether to continue executing the data cleaning process.

[0272] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the adjustment unit 54 further includes:

[0273] The second judgment module 543 is configured to, when the data cleaning process has started, judge whether the transition controller is a concurrent fault controller based on the concurrent fault information;

[0274] The execution module 544 is configured to, if the transition controller is a concurrent fault controller, re-execute the data cleaning process on the transition controller when the transition controller rejoins the mirror pair;

[0275] The processing module 545 is configured to, if the transition controller is not a concurrent fault controller, perform data mirroring on the transition controller after executing the data cleaning process on the transition controller, and perform mirror recombination on the mirror pair where the concurrent fault controller is located to obtain a third recombined mirror pair.

[0276] Further, in a possible implementation manner of this embodiment, the second judgment module 543 is further configured to:

[0277] Obtain the device identifier of the concurrent fault controller and the teaming information of the mirror pair where it is located based on the concurrent fault information;

[0278] Compare the device identifier of the concurrent fault controller with the device identifier of the transition controller, and compare the teaming information of the mirror pair where the concurrent fault controller is located with the teaming information of the temporary mirror pair where the transition controller is located, so as to determine whether the transition controller is a concurrent fault controller.

[0279] Further, in a possible implementation manner of this embodiment, the determination module 542 is further configured to:

[0280] Find the teaming information of the first fault mirror pair where the concurrent fault controller is located according to the concurrent fault information;

[0281] Based on the teaming information of the first fault mirror pair, determine whether the first fault mirror pair is the temporary mirror pair for removing the transition controller;

[0282] If the first faulty mirror pair is a temporary mirror pair for removing the transition controller, then pull the transition controller back into the temporary mirror pair and perform mirror recombination with the normal controller in the temporary mirror pair to obtain a second recombined mirror pair;

[0283] If the first faulty mirror pair is not a temporary mirror pair for removing the transition controller, then perform mirror recombination on the normal controller in the first faulty mirror pair with the controller at the next adjacent position to obtain a second recombined mirror pair.

[0284] Further, in a possible implementation manner of this embodiment, the determining module 542 is further configured to:

[0285] If the transition controller belongs to the second recombined mirror pair, there is no need to continue executing the data cleaning process for the transition controller in the second recombined mirror pair;

[0286] If the transition controller does not belong to the second recombined mirror pair, continue to execute the data cleaning process for the transition controller that does not belong to the second recombined mirror pair.

[0287] Further, in a possible implementation manner of this embodiment, the processing module 545 is further configured to:

[0288] According to the concurrent fault information, find the teaming information of the second faulty mirror pair where the concurrent fault controller is located;

[0289] Based on the teaming information of the second faulty mirror pair, determine whether the second faulty mirror pair is a temporary mirror pair for removing the transition controller;

[0290] If the second faulty mirror pair is a temporary mirror pair for removing the transition controller, then pull the transition controller back into the temporary mirror pair and perform mirror recombination with the normal controller in the temporary mirror pair to obtain a third recombined mirror pair;

[0291] If the second faulty mirror pair is not a temporary mirror pair for removing the transition controller, then perform mirror recombination on the normal controller in the second faulty mirror pair with the controller at the next adjacent position to obtain a third recombined mirror pair.

[0292] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the device for caching data cleaning further includes:

[0293] A marking unit 59, configured to obtain the status information of each mirror pair in the storage system and mark the device identifiers of the transition controllers that need to execute the data cleaning process in each mirror pair.

[0294] It should be noted that the foregoing explanations of the method embodiments also apply to the device in this embodiment, with the same principle, and are not limited in this embodiment.

[0295] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0296] An embodiment of the present disclosure also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments for clearing cached data.

[0297] An embodiment of the present disclosure also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above method embodiments for clearing cached data when running.

[0298] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROMs for short), random access memories (RAMs for short), external hard drives, magnetic disks, or optical discs.

[0299] An embodiment of the present disclosure also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above method embodiments for clearing cached data.

[0300] An embodiment of the present disclosure also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above method embodiments for clearing cached data.

[0301] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0302] The above has introduced in detail a method for clearing cached data. Specific examples are used in this article to elaborate on the principle and implementation manner of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present disclosure, several improvements and modifications can be made to the present disclosure, and these improvements and modifications also fall within the protection scope of the claims of the present disclosure.

Claims

1. A method for cleaning cache data, characterized in that: The method is applied to a storage system, and the method comprises: Using a transition controller to replace the faulty controller in the storage system, and reorganizing the mirror image of the initial mirror pair where the faulty controller is located to obtain a first reorganized mirror pair; adding the repaired faulty controller to the first reorganized mirror pair to obtain a temporary mirror pair; After the faulty controller that has been repaired completes data mirroring, the transition controller is removed from the temporary mirror pair, and a data cleanup process is performed on the transition controller; Monitoring whether a concurrent failure occurs in a controller in the storage system, and monitoring the execution status of the data cleaning process; When a concurrent failure occurs to a controller in the storage system, judging whether the transition controller starts to execute the data cleaning process according to the execution status information of the data cleaning process; When the data cleaning process is not started, determining whether to execute the data cleaning process based on concurrent fault information of the concurrent fault controller where the concurrent fault occurs; When the data cleaning process has begun to be executed, it is determined whether the transition controller is the concurrent fault controller based on the concurrent fault information; if the transition controller is the concurrent fault controller, then when the transition controller rejoins the mirror pair, the data cleaning process is re-executed on the transition controller; if the transition controller is not the concurrent fault controller, the data cleaning process continues to be executed on the transition controller; and data is cleaned on the transition controller according to the adjusted data cleaning process.

2. The method for cleaning cache data according to claim 1, characterized in that: The method of replacing the faulty controller in the storage system with a transition controller and reorganizing the initial mirror pair where the faulty controller is located to obtain a first reorganized mirror pair includes: Selecting the transition controller to replace the failed controller from controllers that have not failed according to the teaming information of the initial mirror pair where the failed controller is located; The transition controller is mirrored and reorganized with a normal controller in the initial mirror pair where the faulty controller is located to obtain the first reorganized mirror pair.

3. The method for cleaning cache data according to claim 2, characterized in that: The step of selecting the transition controller to replace the faulty controller from controllers that have not failed according to the teaming information of the initial mirror pair in which the faulty controller is located comprises: Acquire a set of device identifications of controllers in the storage system except the faulty controller, wherein the device identifications in the set of device identifications are arranged in order; Based on the team information, determine a normal device identifier corresponding to a normal controller in the initial mirror pair where the faulty controller is located; The transition controller that replaces the faulty controller is searched for according to the device identification set and the normal device identification.

4. The method for cleaning cache data according to claim 3, characterized in that: The searching, according to the device identification set and the normal device identification, for the transition controller that replaces the faulty controller comprises: The target device identifier that is one digit after the normal device identifier in the device identifier set is searched, and the controller corresponding to the target device identifier is used as the transition controller.

5. The method for cleaning cache data according to claim 2, characterized in that: The step of reorganizing the transition controller with a normal controller in the initial mirror pair where the faulty controller is located to obtain the first reorganized mirror pair includes: According to the circular mirroring principle, the device identification of the transition controller and the device identification of the normal controller are sorted, and the arrangement order of the transition controller and the normal controller is determined to generate the first recombined mirror pair.

6. The method for cleaning cache data according to claim 1, characterized in that: Before replacing the failed controller in the storage system with the transition controller, the cache data clearing method further includes: The operating status of the controller in the storage system is monitored to identify the faulty controller.

7. The method for cleaning cache data according to claim 3, characterized in that: The method for cleaning cache data also includes: The device identification of the faulty controller is identified, and the device identifications of controllers in the storage system other than the faulty controller are collected and sorted to generate the device identification set.

8. The method for cleaning cache data according to claim 1, characterized in that: After adding the faulty controller whose fault is repaired to the first reorganized mirror pair to obtain a temporary mirror pair, the cache data cleaning method further includes: The fault controller that controls the repair fault performs data mirroring with the transition controller.

9. The method for cleaning cache data according to claim 1, characterized in that: The method for cleaning cache data also includes: If no concurrent failure occurs to the controller in the storage system, data cleaning is performed on the transition controller according to the data cleaning process.

10. The method for cleaning cache data according to claim 1, characterized in that: The determining whether to execute the data cleaning process based on the concurrent fault information of the concurrent fault controller where the concurrent fault occurs includes: Reorganize the mirror image of the mirror pair where the concurrent fault controller is located to obtain a second reorganized mirror image pair; Based on the teaming information of the second recombined mirror pair, it is determined whether the transition controller belongs to the second recombined mirror pair to determine whether to continue to execute the data cleaning process.

11. The method for cleaning cache data according to claim 1, characterized in that: After determining that the transition controller is not the concurrent fault controller and continuing to execute the data cleaning process on the transition controller, the cache data cleaning method further includes: Data mirroring is performed on the transition controller, and mirror reorganization is performed on the mirror pair where the concurrent fault controller is located to obtain a third reorganized mirror pair.

12. The method for cleaning cache data according to claim 1, characterized in that: The determining whether the transition controller is the concurrent fault controller based on the concurrent fault information includes: Based on the concurrent fault information, obtaining the device identification of the concurrent fault controller and the team information of the mirror pair to which it belongs; The device identification of the concurrent fault controller is compared with the device identification of the transition controller, and the team information of the mirror pair where the concurrent fault controller is located is compared with the team information of the temporary mirror pair where the transition controller is located to determine whether the transition controller is the concurrent fault controller.

13. The method for cleaning cache data according to claim 10, characterized in that: The mirror reorganization of the mirror pair where the concurrent fault controller is located to obtain a second reorganized mirror pair includes: According to the concurrent fault information, searching for teaming information of the first fault mirror pair where the concurrent fault controller is located; Based on the teaming information of the first faulty mirror pair, determining whether the first faulty mirror pair is a temporary mirror pair for removing the transition controller, If the first faulty mirror pair is a temporary mirror pair from which the transition controller is removed, the transition controller is pulled back into the temporary mirror pair and mirrored with a normal controller in the temporary mirror pair to obtain the second reconstituted mirror pair; If the first faulty mirror pair is not a temporary mirror pair after the transition controller is removed, a normal controller in the first faulty mirror pair is mirrored and reassembled with a subsequent controller to obtain the second reassembled mirror pair.

14. The method for cleaning cache data according to claim 13, characterized in that: The step of judging whether the transition controller belongs to the second reorganized mirror pair based on the teaming information of the second reorganized mirror pair to determine whether to continue to execute the data cleaning process includes: If the transition controller belongs to the second reorganized mirror pair, there is no need to continue to execute the data cleaning process for the transition controller in the second reorganized mirror pair; If the transition controller does not belong to the second recombinant mirror pair, the data cleaning process continues to be performed on the transition controller that does not belong to the second recombinant mirror pair.

15. The method for cleaning cache data according to claim 11, characterized in that: The mirror reorganization of the mirror pair where the concurrent fault controller is located to obtain a third reorganized mirror pair includes: According to the concurrent fault information, searching for teaming information of the second fault mirror pair where the concurrent fault controller is located; Based on the teaming information of the second faulty mirror pair, determining whether the second faulty mirror pair is a temporary mirror pair for removing the transition controller, If the second faulty mirror pair is a temporary mirror pair from which the transition controller is removed, the transition controller is pulled back into the temporary mirror pair and mirrored with a normal controller in the temporary mirror pair to obtain the third reconstituted mirror pair; If the second faulty mirror pair is not a temporary mirror pair after the transition controller is removed, a normal controller in the second faulty mirror pair is mirrored and reassembled with a subsequent controller to obtain the third reassembled mirror pair.

16. The method for cleaning cache data according to claim 11, characterized in that: The method for cleaning cache data also includes: The status information of each mirror pair in the storage system is obtained, and the device identifier of the transition controller that needs to execute the data cleaning process is marked in each mirror pair.

17. A device for cleaning cache data, characterized in that: The device is applied to a storage system, and comprises: A reorganization unit, configured to replace a faulty controller in the storage system with a transition controller, so as to reorganize the mirror image of the initial mirror pair where the faulty controller is located, and obtain a first reorganized mirror pair; and to add the repaired faulty controller to the first reorganized mirror pair, and obtain a temporary mirror pair; A removal unit, configured to remove the transition controller from the temporary mirror pair after the repaired faulty controller completes data mirroring, and perform a data cleanup process on the transition controller; A first monitoring unit, configured to monitor whether a concurrent failure occurs in a controller in the storage system, and to monitor an execution status of the data cleaning process; an adjustment unit, configured to determine whether the transition controller starts to execute the data cleaning process according to the execution status information of the data cleaning process when a concurrent failure occurs in the controller in the storage system; When the data cleaning process is not started, determining whether to execute the data cleaning process based on concurrent fault information of the concurrent fault controller where the concurrent fault occurs; When the data cleaning process has begun to be executed, it is determined whether the transition controller is the concurrent fault controller based on the concurrent fault information; if the transition controller is the concurrent fault controller, then when the transition controller rejoins the mirror pair, the data cleaning process is re-executed on the transition controller; if the transition controller is not the concurrent fault controller, the data cleaning process continues to be executed on the transition controller; and data is cleaned on the transition controller according to the adjusted data cleaning process.

18. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the cache data cleaning method described in any one of claims 1-16.

19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the cache data cleaning method according to any one of claims 1-16.

20. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for cleaning cache data according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Data caching method and device, equipment and storage medium

    CN115563028A