A fault handling method and electronic device

By using a fault prediction model in a dual-active data center to identify and respond to faults in advance, the problems of long fault response time and inflexible response strategies in existing technologies are solved, enabling rapid and flexible fault recovery.

CN121349750BActive Publication Date: 2026-03-31INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511892379.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-31
Estimated Expiration
2045-12-15

AI Technical Summary

Technical Problem

Existing active-active technologies cannot proactively predict faults, resulting in long fault response times and inflexible response strategies, making it impossible to select the best response strategy based on the fault type.

Method used

By acquiring real-time performance data from a dual-active data center, a pre-trained fault prediction model is used to predict fault types, and corresponding response strategies are executed before faults occur, including early response to node faults, link faults, and storage media faults.

Benefits of technology

Significantly reduces fault recovery time, improves the flexibility and efficiency of fault response, reduces response time to the millisecond level, and enhances the reliability and business continuity of active-active systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349750B_ABST
    Figure CN121349750B_ABST
Patent Text Reader

Abstract

The application discloses a fault processing method and electronic equipment, and relates to the technical field of computers, and comprises the following steps: identifying in advance a fault that a target data center may encounter according to real-time performance data of the target data center, and executing a corresponding response strategy in advance according to the fault type, so that the fault recovery time can be significantly reduced, the technical problems that the fault response time is long and the fault response strategy is not flexible are solved, and the technical effects of reducing the response time and flexibly responding are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a fault handling method and an electronic device. Background Technology

[0002] Active-active architecture is a highly available disaster recovery architecture that ensures rapid takeover of services in the event of a single site failure by synchronizing data and business load in real time between the primary and backup data centers, significantly improving business continuity. However, existing active-active technologies are essentially passive response mechanisms, with high availability relying on the detection, assessment, and switchover process after a failure occurs. This approach cannot predict failures and only begins responding after a failure occurs, resulting in prolonged business interruption and consequently, long failure response times. Secondly, the response strategy is simplistic, typically limited to primary / backup switching, and cannot select the optimal execution strategy based on the failure type, leading to inflexible failure response strategies. Therefore, current active-active architectures lack timeliness and intelligence in failure handling, urgently requiring a failure handling method that can proactively predict and accurately respond to failures. Summary of the Invention

[0003] This application provides a fault handling method and an electronic device to at least solve the problems of long fault response time and inflexible fault response strategies in related technologies.

[0004] This application provides a fault handling method, including:

[0005] Acquire real-time performance data of the target data center, where the target data is the primary or backup data center in a dual-active data center system. The real-time performance data includes: node data of the target data center, link data of the transmission link between the primary and backup data centers, and storage media data of the target data center.

[0006] Real-time performance data is input into a pre-trained fault prediction model. The fault prediction model compares the real-time performance data and performance change data with multiple set thresholds corresponding to the set fault type to predict the fault type and obtain the prediction results output by the fault prediction model. The performance change data is obtained by performing time series analysis on the real-time performance data.

[0007] If the prediction results indicate that a target failure will occur in the target data center within a set future time, then the failure response strategy corresponding to the target failure will be executed before the target failure occurs. The target failure includes at least one of the following: a node failure determined based on node data, a link failure determined based on link data, and a storage medium failure determined based on storage medium data.

[0008] This application also provides a fault handling apparatus, including:

[0009] The acquisition unit is used to acquire real-time performance data of the target data center. The target data is the primary data center or the backup data center in a dual-active data center system. The real-time performance data includes: node data of the target data center, link data of the transmission link between the primary data center and the backup data center, and storage media data of the target data center.

[0010] The fault prediction unit is used to input real-time performance data into a pre-trained fault prediction model, compare the real-time performance data and performance change data with multiple set thresholds corresponding to the set fault type to predict the fault type, and obtain the prediction results output by the fault prediction model.

[0011] The strategy execution unit is used to execute the fault response strategy corresponding to the target fault before the target fault occurs, if the prediction result indicates that the target data center will experience a target fault within a set future time. The target fault includes at least one of the following: a node fault determined based on node data, a link fault determined based on link data, and a storage medium fault determined based on storage medium data.

[0012] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described fault handling methods when executing the computer program.

[0013] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described fault handling methods.

[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described fault handling methods.

[0015] This application enables the early identification of potential faults in the target data center based on real-time performance data, and allows for the implementation of corresponding response strategies according to the fault type. This significantly reduces fault recovery time, solves the technical problems of long fault response time and inflexible fault response strategies, and achieves the technical effects of reducing response time and providing flexible response. Attached Figure Description

[0016] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1A flowchart illustrating a fault handling method provided in an embodiment of this application;

[0018] Figure 2 A flowchart illustrating an early response to node failures provided in an embodiment of this application;

[0019] Figure 3 This is a schematic diagram of a storage medium failure early response process provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0021] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] The specific application environment architecture or specific hardware architecture on which the execution of the fault handling method depends is described here.

[0024] Before describing the embodiments, key terms involved will be explained first, wherein,

[0025] LUN (Logical Unit Number) is a number used in a storage system to identify the smallest logical storage unit that can be accessed by a host. Essentially, it is a "logical disk" that abstracts physical storage resources (such as hard drives and RAID groups) for the server (host) to identify and use.

[0026] IOPS (Input / Output Operations Per Second) is a core performance indicator that measures the ability of storage devices (such as hard drives, SSDs, and storage arrays) to handle read and write requests. It directly reflects the "response speed" and "concurrency processing efficiency" of the storage system.

[0027] PCIe (Peripheral Component Interconnect Express) is a packet-switched, point-to-point, serial high-speed differential signal interconnect technology.

[0028] RAID (Redundant Array of Independent Disks) is a technology that combines multiple independent physical disks in a specific way (RAID level) to form a single logical disk.

[0029] RTO (Recovery Time Objective) refers to the maximum acceptable time from the interruption of a system or business process to its necessary restoration after a disaster or disruption event.

[0030] RPO (Recovery Point Objective) refers to the maximum amount of data loss that a business system can tolerate in the event of a disaster or disruption, and is usually measured in time.

[0031] FC link (Fibre Channel) refers to an end-to-end physical connection and logical communication path established between two Fibre Channel network devices through optical fiber or copper cable media.

[0032] SCSI (Internet Small Computer System Interface) refers to an end-to-end logical communication path over a TCP / IP network that transmits block-level data by encapsulating SCSI commands.

[0033] SCSI (Small Computer System Interface) is a set of standard protocols used for physical connections and data transfer between computers and their peripheral devices (mainly storage devices).

[0034] Active-active is a computer disaster recovery solution. Its implementation mode allows two data centers, a primary and a backup, to simultaneously handle user services. In this mode, the primary and backup data centers serve as backups for each other and perform real-time data backup. It is mainly a technical architecture used to improve data reliability and business continuity. Its core is to ensure that if one data center fails, the other data center can seamlessly take over the business by synchronizing data in real time between two or more data centers.

[0035] Currently, the high availability architecture of a dual-active system means that both data centers in the system can handle business operations, and when either data center processes business data, the data is simultaneously updated to the other data center. Current dual-active systems can respond to most faults and handle them proactively. For example, in the event of a single node failure, the dual-active system will not interrupt business operations; it can resume processing after a simple response to the node failure. Another example is a single-site failure: if the primary data center fails, the dual-active system will immediately switch to the backup data center, which will take over all business operations, and the dual-active system will switch to single-write mode, i.e., synchronization will be paused.

[0036] While active-active storage architectures offer high availability and strong disaster recovery capabilities, current active-active architectures respond to failures reactively, meaning they only act after a failure has occurred. When an active-active system encounters a failure, it must first determine the type of failure before responding, lacking corresponding failure prediction methods. There is significant room for improvement in RTO (Recovery Time Objective). Furthermore, current active-active systems respond passively to failures, lacking the ability to select the best response based on the failure type. Moreover, most current active-active systems still rely on a single, fixed approach when dealing with failures: switching services to a backup data center, resulting in inflexible response strategies.

[0037] This application provides a fault handling method that, based on real-time performance data such as the current controller behavior and IO traffic behavior of the target data center, predicts the most likely fault scenarios and proactively initiates active-active response based on the most probable fault scenarios. This reduces RTO while improving the reliability of active-active response, solving the "one-size-fits-all" approach of prioritizing master-slave switching when encountering faults in existing active-active systems, and improving response flexibility. The following embodiments, combined with the execution flow of the fault handling method, provide a detailed description of the method.

[0038] Figure 1 This is a flowchart illustrating a fault handling method provided in an embodiment of this application, applied to a fault handling device, specifically including as follows: Figure 1 The following steps are shown:

[0039] S101. Obtain real-time performance data of the target data center.

[0040] The target data is the primary or backup data center in a dual-active data center system. The real-time performance data includes: node data of the target data center, link data of the transmission link between the primary and backup data centers, and storage media data of the target data center.

[0041] In an understandable application scenario, two sites form a dual-active cluster / system. Each site has a data center, which can be either a primary or backup data center. For example, the first site corresponds to the primary data center, and the second site corresponds to the backup data center. Each site consists of multiple nodes, and the nodes of the two sites are generally in a 2+2 or 4+4 mode. Each data center has a corresponding storage controller, which usually refers to the CPU module and its firmware / software inside the node, responsible for performing specific I / O processing, RAID calculation, cache management, and other tasks.

[0042] Understandably, the target data center is either the primary or backup data center in a dual-active system, meaning it uses a dual-active architecture, or it could be the primary or backup data center in another cluster. A pre-built fault prediction module can be started in the primary data center. This module predicts fault scenarios and notifies the dual-active system to respond or prepare in advance based on these scenarios. The fault prediction module obtains real-time performance data (i.e., configuration stream data) from the target data center. This real-time performance data includes node data, link data between the primary and backup data centers, and storage media data. The configuration stream data includes storage cluster status, disk information, and link configuration. The subsequent fault prediction module can then establish storage cluster health status information based on this configuration stream data. Simultaneously, a basic fault prediction model is established. The default log setting is for non-core business operations; this can be adjusted later to include core and non-core business operations based on the business scenario. Afterward, the module begins monitoring the cluster status. When the fault prediction module predicts a specific type of fault, it notifies the dual-active system to respond in advance, ensuring that RPO and business operations are not affected.

[0043] S102. Input the real-time performance data into the pre-trained fault prediction model. The fault prediction model compares the real-time performance data and performance change data with multiple set thresholds corresponding to the set fault type to predict the fault type, and obtains the prediction results output by the fault prediction model.

[0044] The target data center corresponds to at least one site, and each site consists of multiple nodes. The performance change data is obtained by performing time-series analysis on real-time performance data.

[0045] Understandably, based on the above S101, real-time performance data is input into the fault prediction model. The fault prediction model extracts performance change data, which includes at least the time-series changes in performance indicators, based on the real-time performance data. It compares the real-time performance data and performance change data with the magnitudes of multiple set thresholds corresponding to different fault types, evaluates the current abnormal situation of the dual-active system, predicts the possible fault scenarios in the target data center, and obtains the prediction results output by the fault prediction model. The prediction results include the target faults that will occur in the target data center within a set future time and the fault location. The target faults are at least one of the following basic faults: node faults, link faults, and storage media faults. The basic fault characteristics of each basic fault are described in the following embodiments.

[0046] Optionally, real-time performance data is input into a pre-trained fault prediction model. The fault prediction model compares the real-time performance data and performance change data with multiple preset thresholds corresponding to a set fault type to predict the fault type, and obtains the prediction results output by the fault prediction model, including:

[0047] Based on node data, if the fault prediction model determines that the load of the target controller on the target node corresponding to the target data center is greater than the first set threshold and there is no high input / output trigger, the memory utilization of the target controller is greater than the second set threshold, the response latency carried by the target controller exceeds the third set threshold and there is no high input / output response latency of the corresponding other controllers, the target controller log shows a bus error, and / or the controller corresponding to another data center receives an alarm message indicating that the heartbeat interval of the controller corresponding to the target data center has been extended, then the prediction result of the node failure in the target data center and the target node being a faulty node within the set future time is output.

[0048] Understandably, node failures are categorized into single-node failures and multi-node failures. A single-node failure refers to a system composed of multiple nodes where only one node fails, while all other nodes function normally. A multi-node failure refers to a system composed of multiple nodes where two or more nodes fail simultaneously or sequentially within a certain time window. Single-node failures are a common risk that system redundancy design aims to eliminate. Therefore, the following embodiments use single-node failure as an example for detailed explanation, such as a single-node failure of the storage controller in the primary or backup data center.

[0049] Understandably, a single node failure will cause the current business IO in the active-active system to drop by a few seconds. This drop-off time is used for the silent bitmap synchronization of the active-active system and the recovery of IO forwarded to the failed node. For example, IO in the synchronization process needs to reselect the target node. The basic characteristics of a single-node failure include: the target controller on the target node in the target data center has a load greater than a first set threshold and no high I / O triggers, for example, the target controller's CPU load is greater than 95% for 5 consecutive minutes without high IO triggers; the target controller's memory utilization is greater than a second set threshold, for example, the controller's memory utilization is greater than 95% for 3 consecutive minutes; the response latency carried by the target controller exceeds a third set threshold and the corresponding other controllers do not have high I / O response latency, for example, the LUN IO response latency carried by the controller exceeds 500ms and the other controller nodes in the same storage pool do not have high LUN IO response latency; a bus error appears in the target controller's log, for example, a PCIe bus error appears in the controller log; and / or, other data center corresponding controllers receive an alarm message indicating an extended heartbeat interval from the target data center corresponding controller, for example, the "primary data center controller heartbeat interval" received by the backup data center controller is extended from 1s to >5s, indicating a single-node failure in the primary data center controller. Other possible basic characteristics of single-node failures can be set by the user according to their needs and are not limited here.

[0050] Understandably, if different fundamental fault characteristics are simultaneously met or at least partially met within a certain period, it can be considered that a single node failure has occurred in the target data center controller. The met characteristics can be determined by the user according to their needs. The target data center is designated as the faulty data center, the target node as the faulty node, and the target controller as the faulty controller. Subsequently, the prediction results for the target data center experiencing node failure and the target node becoming the faulty node within a set future timeframe are output.

[0051] Understandably, when the fault prediction model predicts that a node will fail next, it proactively notifies the active-active system of the data center and specific node that will soon fail. This allows the active-active system to perform bitmap synchronization in advance. This bitmap is referred to as the first bitmap. The system then simulates the situation after a node failure, making the bitmap of the mirror node valid before the predicted failure, and ceasing to forward I / O to the failed node. The failed node also stops handling I / O traffic. The first bitmap refers to the bitmap used for data mirroring between nodes, representing the data synchronization strategy between them.

[0052] Optionally, real-time performance data is input into a pre-trained fault prediction model. The fault prediction model compares the real-time performance data and performance change data with multiple preset thresholds corresponding to a set fault type to predict the fault type, and obtains the prediction results output by the fault prediction model, including:

[0053] Based on link data, if the fault prediction model determines that the target link utilization rate is greater than the fourth set threshold, the bandwidth utilization fluctuation of the target link increases by the fifth set threshold under abnormal business growth, the data synchronization delay between the primary data center and the backup data center is the sixth set threshold, the data synchronization queue length of the primary data center is greater than the seventh set threshold and continues to grow, the input burst of the primary data center does not match the bandwidth growth rate of the target link, and / or the frame loss rate or retransmission rate of the target link is greater than the eighth set threshold, then the prediction result of the target data center experiencing a link failure within a set future time and the target link being a failed link will be output. Here, the target link refers to the link between the primary data center and the backup data center used for data synchronization.

[0054] Understandably, link failure refers to a failure triggered by link congestion between the primary and backup data centers in a dual-active system. The bandwidth of the links used for data synchronization between the primary and backup data centers (e.g., FC links, iSCSI links) is saturated, causing synchronization latency to exceed the RPO threshold, or even preventing the dual-active system from synchronizing I / O to the backup data center. In this failure scenario, the dual-active RPO is no longer guaranteed. Specifically, the basic failure characteristics of link failures caused by link congestion between the primary and backup data centers include: target link utilization exceeding a fourth set threshold (e.g., actual link utilization exceeding 98% for 5 minutes); bandwidth utilization fluctuation of the target link increasing by a fifth set threshold under abnormal business growth (e.g., bandwidth utilization fluctuation increasing by at least 30% within 1 minute under abnormal business growth); data synchronization latency between the primary and backup data centers exceeding a sixth set threshold (e.g., data synchronization latency > 100ms, while normal dual-active system synchronization latency should be < 10ms); and the primary data center data synchronization queue length exceeding a certain threshold. The seventh threshold is set and continues to increase. For example, the length of the data synchronization queue in the main data center is >500 IOs and continues to increase; the input burst volume in the main data center does not match the bandwidth growth rate of the target link. For example, the write IO burst volume in the main data center does not match the bandwidth growth rate of the link. In this case, it indicates that the bandwidth of the main data center is mostly used for business processing and the data synchronization bandwidth is difficult to support. For example, if the write business data increases by 100% and the link bandwidth increases by 200%, it indicates that there is invalid synchronization traffic; and / or, the frame loss rate or retransmission rate of the target link is greater than the eighth threshold. For example, the "frame loss rate" of the FC link is >0.1%, or the "TCP retransmission rate" of the iSCSI link is >0.5%.

[0055] Understandably, if different basic fault characteristics are simultaneously met or at least partially met within a certain period, it is considered that a link failure has occurred in the target data center controller. The specific characteristics that need to be met can be determined by the user. The target data center is designated as the faulty data center, and the target link is designated as the faulty link. Subsequently, the prediction results of the target data center experiencing a link failure within a set future time and the target link being a faulty link are output.

[0056] Understandably, the basic fault characteristics of the above-mentioned link failure can accurately reflect whether the link congestion is caused by a sudden business interruption or by invalid synchronization traffic. The fault prediction module will notify the active-active system based on the prediction results so that the active-active system can respond in advance, without having to wait until the link congestion causes the synchronization IO to time out (i.e., the link failure occurs) before responding. The timeout IO causes the active-active system to switch to single write. Since the fault type is distinguished, the active-active system will determine whether it needs to switch to single write based on the fault scenario.

[0057] Optionally, real-time performance data is input into a pre-trained fault prediction model. The fault prediction model compares the real-time performance data and performance change data with multiple preset thresholds corresponding to a set fault type to predict the fault type, and obtains the prediction results output by the fault prediction model, including:

[0058] Based on media data, if the fault prediction model determines that the target storage medium has an additional set number of mapped sectors, an uncorrectable sector appears in the target storage medium, the load of other storage media increases to a second set value due to the input / output transfer of the target storage medium, the target storage medium has a pre-reconstruction flag or an imminent reconstruction flag, the synchronization latency carried by the target storage medium is greater than a ninth set threshold, and / or the target link write latency increases to a third set value due to the blockage of the target storage medium, then the prediction results of the target data center experiencing storage medium failure and the target storage medium being a failure storage medium will be output. The target storage medium is used to store data in the primary data center and the backup data center.

[0059] Understandably, underlying storage media failures are generally hard drive failures. Normally, such failures will be categorized into triggering storage pool reconstruction or the LUN to which the hard drive belongs becoming OFFLINE, depending on the severity of the failure. However, storage pool reconstruction will consume the synchronization bandwidth of the dual-active system, while OFFLINE will directly cause the dual-active system to switch to single-write. Specifically, the basic characteristics of underlying storage media failures include: an additional set number of mapped sectors in the target storage media, for example, at least 10 new remapped sectors within 24 hours; uncorrectable sectors appearing in the target storage media; the load on other storage media increasing to a second set value due to the input / output shift of the target storage media, for example, the load on other hard IOs in the same RAID group increasing by 15%+ due to the IO shift of the failed disk; the target storage media displaying a pre-reconstruction flag or an imminent reconstruction flag; the synchronization latency carried by the target storage media exceeding a ninth set threshold, for example, the synchronization latency of a dual-active LUN carried by the failed disk >300ms, while the normal synchronization latency should be less than 10ms; and / or, the target link write latency increasing to a third set value due to the blockage of the target storage media, for example, the synchronization link write latency in a dual-active system increasing by 10% due to the IO blockage of the failed hard disk.

[0060] Understandably, if the aforementioned basic fault characteristics are simultaneously met or at least partially met within a certain period, it is considered that a storage media failure (e.g., hard drive failure) has occurred in the target data center. The specific characteristics that need to be met can be determined by the user. The target data center is designated as the faulty data center, and the target storage medium is designated as the faulty storage medium. Subsequently, the system outputs the predicted results of the target data center experiencing a storage media failure within a set future time and the target storage medium being designated as the faulty storage medium, and informs the active-active system to respond in advance.

[0061] Optionally, after inputting real-time performance data into a pre-trained fault prediction model and obtaining the prediction results output by the fault prediction model, the method further includes:

[0062] The prediction accuracy for different fault types is calculated based on the prediction results and historical prediction results. If the prediction accuracy is lower than the tenth set threshold, the set threshold used to predict different fault types in the fault prediction model is updated.

[0063] Understandably, the fault prediction model in the fault prediction module will automatically update based on each fault situation to ensure prediction accuracy. It calculates the prediction accuracy for different fault types based on current and historical prediction results. For example, if the prediction accuracy for a certain fault type drops by 3%, the default value (the set thresholds related to the basic fault characteristics mentioned above) is updated based on the current fault data. For instance, if the prediction accuracy for controller node faults drops by 3% within 24 consecutive hours, the basic fault characteristics corresponding to the node fault will be updated. For example, if all node faults encountered in the current data center show that the controller CPU load is greater than 90% for 5 consecutive minutes instead of 95%, then the basic fault characteristic of the fault prediction model will be updated to 90%, thereby improving the prediction accuracy for this fault scenario.

[0064] Understandably, the fault prediction model distinguishes fault types, classifying faults into local faults (e.g., hard drive failure) and global faults (e.g., link congestion, controller node failure). Different fault scenarios require different pre-processing methods. For example, local faults such as hard drive failures become localized, so that the way a dual-active system deals with faults is no longer simply by switching to a backup data center.

[0065] S103. If the prediction result indicates that the target data center will experience a target failure within a set future time, then execute the failure response strategy corresponding to the target failure before the target failure occurs.

[0066] The target fault includes at least one of the following: a node fault determined based on node data, a link fault determined based on link data, and a storage medium fault determined based on storage medium data.

[0067] Understandably, based on the above S102, when the prediction result indicates that the target data center will experience a target failure within a set future time, the active-active system is notified to respond in advance according to the prediction result. Furthermore, the failure prediction module will interface with the arbitration module to inform the arbitration procedure in advance when a site-level failure is predicted, enabling the arbitration procedure to more quickly arbitrate the data center in the active-active system to take over the services. In an active-active architecture within the same storage cluster, the arbitration can also be informed whether a new cluster configuration node needs to be selected in advance based on the site failure. Specifically, for node failures, the active-active system responds in advance by executing corresponding strategies, such as reassembling the bitmap in advance and switching services in advance to take over the failed node, ensuring a significant reduction in RTO during node failures. For link congestion, the active-active system responds in advance by executing corresponding strategies, switching to redundant links and suspending non-core services to free up synchronization bandwidth. For storage media failures, the active-active system responds in advance by executing corresponding strategies, such as performing data migration in advance, notifying RAID for verification in advance, and reducing the IO ratio of the failed disk in advance, avoiding the impact of a single disk failure on the active-active system. This will be explained in detail in the following embodiments.

[0068] Optionally, before the target failure occurs, execute the corresponding failure response strategy, including:

[0069] If the prediction indicates that a node failure will occur in the target data center within a set future time, then before the node failure occurs, the bitmap used to characterize data synchronization between nodes is reorganized to redetermine the master node's takeover of the target node's services and to suspend the allocation of services to the target node; or, before the node failure occurs, the bitmap used to characterize data synchronization between nodes is reorganized to redetermine the master node's takeover of the target node's services, and after the node failure occurs, the allocation of services to the target node is suspended.

[0070] Understandably, if a node failure in the target data center is predicted, the bitmap used for data mirroring between nodes is reorganized in advance before the failure occurs. A new node is identified to take over the services of the target node (which is now the node about to fail), and this new node replaces the target node as the primary node in one segment of the bitmap and the backup node in another. Simultaneously, the allocation of new services to the target node is suspended; that is, at least some services on the target node are migrated to the new node in advance, and no further service allocation is performed. After this advance preparation, when the target node actually fails, there is no need to reorganize the bitmap again. Furthermore, because the new node takes over the IO services of the target node (now the failed node) in advance, the target node actually has no IO services during the failure. Therefore, the active-active system can respond and recover quickly. This significantly reduces the work required by the active-active system after a failure, significantly reducing RTO time and service interruption time. Moreover, because the new node takes over the IO services of the failed node in advance, the failed node has no actual services or only a small amount of actual services during the failure, which also significantly improves data reliability.

[0071] Understandably, before the target failure occurs, the active-active system executes at least some of the fault response strategies in advance to minimize the time to operation (RTO). For example, bitmap reassembly is completed before the node failure occurs, and after the node failure, the new node takes over all services of the failed node and no longer assigns services to the failed node.

[0072] For example, the four nodes in the main data center are responsible for the bitmap from 0 to 1 / 4 (i.e., the first segment), from 1 / 4 to 2 / 4 (i.e., the second segment), from 2 / 4 to 3 / 4 (i.e., the third segment), and from 3 / 4 to the end of the bitmap (i.e., the fourth segment). The nodes are then mirrored in pairs, with each pair acting as a primary and backup node. In the first segment, node 1 is the primary node managing that bitmap, and node 2 is the backup node for the mirror. Nodes 1 and 2 are responsible for the first segment, and so on. Nodes 2 and 3 are responsible for the second segment, and nodes 3 and 4 are responsible for the third segment. Nodes 4 and 1 handle the last segment, with node 1 acting as the backup node. When a failure of node 1 is predicted, the mirror pairs are reassembled, with nodes 2 and 3 responsible for the first and second segments, nodes 3 and 4 for the third segment, and nodes 4 and 2 for the last segment. When a failure of node 1 is predicted, the active-active system will preemptively reassemble the bitmap, making node 2 the master node of the first bitmap segment and taking over the services of node 1. In other words, node 2 is changed from a backup node of the first bitmap segment to the master node. Node 3 mirrors the data of the first bitmap segment in advance and becomes the backup node for that segment. Simultaneously, node 2 mirrors the data of the last bitmap segment in advance and becomes the backup node for that segment. When node 1 actually fails, there is no need to reassemble the bitmap again. Furthermore, because node 2 has taken over the IO services of the first bitmap segment in advance, node 1 actually has no IO services during the failure, allowing the active-active system to respond and recover quickly. In this system, node 1 is the failed node, and node 2 is the new node; the method for determining the new node after the failed node is not limited.

[0073] Understandably, as shown in Table 1, the original node mirror pairs corresponding to the target data center are 1-2, 2-3, 3-4, and 4-1. 1-2 refers to node 1 and node 2 as mirror pairs in the first segment of the bitmap, with node 1 as the primary node for processing business operations and node 2 as the backup node. Other mirror pairs are not elaborated upon. The LUN bitmap corresponds to the actual physical address of the entire LUN's storage, used to confirm the specific write or read location of business I / O. If a failure of node 1 is predicted, the bitmap is reassembled in advance to form new node mirror pairs: 2-3, 2-3, 3-4, and 4-2. In the first segment of the bitmap, node 2 processes business operations, and node 3 serves as a backup.

[0074] Table 1:

[0075]

[0076] Optionally, before the target failure occurs, execute the corresponding failure response strategy, including:

[0077] If the prediction indicates that a link failure will occur in the target data center within a set future time, and the link failure is caused by link congestion due to a sudden surge in business activity, then before the link failure occurs, an alarm message will be issued indicating that the current business in the target data center exceeds the current load capacity. If the current business is a core business and the pressure will not decrease within the first set time, the backup link of the target link will be activated. If link congestion still exists even after activating the backup link, data synchronization between the primary data center and the backup data center will be suspended. If the load carried by the target link decreases to the first set value after the second set time, then incremental synchronization between the primary data center and the backup data center will be activated, and the determined difference data will be synchronized to the backup data center. At the same time, data synchronization between the primary data center and the backup data center will also be activated.

[0078] Understandably, if a link failure between the primary and backup data centers is predicted, before the failure occurs, the system analyzes the basic failure characteristics to determine whether the link congestion is caused by a sudden surge in business activity or by invalid synchronization traffic. If the link failure is due to a sudden surge in business activity, i.e., the target failure is caused by increased write activity at the front end of the business (e.g., FC link "frame loss rate" > 0.1%), an alert is issued before the failure occurs to indicate that the current business activity has exceeded the capacity of the active-active system, informing users to reduce business pressure. If the current business activity is critical and the pressure will not decrease in the short term, the active-active system will activate redundant links to attempt to alleviate the pressure. If congestion persists even after the redundant links are activated, the active-active system will switch to single-write mode, where the primary data center takes over the business independently and temporarily stops synchronizing data to the backup data center to reduce the impact on front-end business. By default, the system checks the front-end business IO load every 10 minutes. When the business IO load decreases, it initiates incremental synchronization between the primary data center and the backup data center, synchronizing all the differences between the primary and backup data centers to the backup data center. At the same time, the dual-active system switches to dual-write, which means that data is written to both the primary and backup data centers, while single-write means that data is written only to the primary data center.

[0079] Optionally, before the target failure occurs, execute the corresponding failure response strategy, including:

[0080] If the prediction indicates that a link failure will occur in the target data center within a set future time, and the link failure is caused by link congestion due to invalid traffic synchronization, then before the link failure occurs, update the bitmap between the primary data center and the backup data center; analyze the synchronization logs between the primary data center and the backup data center to identify invalid data that has been synchronized; initiate the synchronization of incremental data (excluding invalid data) between the primary data center and the backup data center; and suspend the data synchronization of non-core services between the primary data center and the backup data center.

[0081] Understandably, when a current failure (target failure) is predicted to be caused by invalid synchronization traffic (e.g., iSCSI link "TCP retransmission rate" > 0.5%), the active-active system will update the second bitmap between the primary and backup data centers. The second bitmap refers to the bitmap between the data centers. Subsequently, the synchronization logs are analyzed to identify identical data between the primary and backup data centers. Synchronizing identical data constitutes invalid synchronization. Invalid traffic (e.g., duplicate synchronization of active-active metadata, temporary file synchronization) is located, discarded, and incremental synchronization is re-initiated. Simultaneously, non-core services (e.g., log synchronization) are paused to release bandwidth, waiting for the link to clear up before resuming these paused services, ensuring link congestion is alleviated.

[0082] Understandably, the active-active system can respond in advance to predicted link failures caused by various reasons, rather than waiting for link congestion to cause synchronous I / O timeouts. Instead, it directly switches to single-write mode due to timeout I / O, resolving the issue of inflexible response strategies. Furthermore, by differentiating fault types, the active-active system can determine whether to switch to single-write mode based on the fault scenario.

[0083] Optionally, before the target failure occurs, execute the corresponding failure response strategy, including:

[0084] If the prediction indicates that a storage media failure will occur in the target data center within a set future time, then before the storage media failure occurs, online data migration is initiated to migrate data from the primary data center and / or backup data center to the backup storage media of the target storage media; the number of input / output operations per second of the target storage media is adjusted to the fourth set value; and the verification information of the independent disk redundancy array where the target storage media is located is calculated based on the data of the currently normal storage media, so as to directly reuse the verification information after the storage media failure occurs.

[0085] Understandably, when a hard drive failure is predicted as the target fault, the system proactively identifies the faulty hard drive and its associated LUN before the failure occurs. An online data migration is initiated in the background, migrating the data to a spare hard drive in the same RAID group as the target hard drive. The migration rate can be controlled within 50MB / s to ensure no impact on services. Simultaneously, the RAID group is notified to begin pre-calculation of verification information. Based on the current healthy hard drive data, the verification information of the RAID group containing the faulty drive is calculated in advance. After the failure, this information is directly reused to reduce RAID reconstruction time and further reduce the bandwidth consumption of the active-active system. The IOPS ratio of the faulty drive is proactively adjusted to reduce its IO load. The remaining load of the faulty drive is transferred to other hard drives in the same RAID group. This proactive fault response ensures that the active-active system is unaware of such faults, and active-active services are unaffected.

[0086] Understandably, integrating hardware and link layers into the active-active disaster recovery solution provides solutions for hard drive failures and link congestion, unlike related solutions where active-active is taken over by the application layer or network layer.

[0087] Understandably, the fault prediction module is an addition based on a dual-active architecture, but the node fault prediction and hard disk fault prediction of the fault prediction module can be applied to other scenarios.

[0088] The fault handling method disclosed herein enables active-active systems to predict faults in advance. Compared to traditional active-active systems, early fault identification and proactive response significantly reduce RTO. Traditional active-active systems typically respond to faults "after the fact," meaning they only respond when a fault occurs. The fault prediction module, however, predicts faults 10 to 60 seconds in advance and responds accordingly. By the time a fault occurs, the main services have already switched over, reducing RTO from seconds (10-30 seconds) to milliseconds (500ms). Furthermore, because fault types are identified, some fault scenarios can be addressed proactively without directly impacting front-end IO services. For example, in the case of node failure, proactive response mirrors data to a backup node, allowing the backup node to respond to services ahead of time. When the failed node truly goes offline, there are no ongoing services, allowing it to quickly complete the offline process and restore the cluster to a stable state more quickly. This significantly reduces RTO and improves service reliability. For example, in the case of link congestion, a traditional active-active system would passively wait for synchronous I / O timeouts before switching to single-write. With a fault prediction module, the cause of the link congestion can be identified, and a proactive response can be initiated. Furthermore, pre-defined service priorities can free up some synchronous bandwidth. For active-active systems, this is no longer a simple matter of waiting for synchronous I / O timeouts due to link congestion before switching to single-write. Another example is hard drive failure. Instead of waiting for the hard drive to go offline before switching to single-write, active-active systems can proactively migrate data, avoiding volume offline scenarios. Since this processing is done in the background, the active-active system is unaware of such failures and does not need to respond, ensuring the availability and reliability of the active-active system. The response is also more flexible and efficient.

[0089] Based on the above embodiments, Figure 2 A flowchart illustrating an early response to node faults provided in this application embodiment specifically includes, as follows: Figure 2 The following steps are shown:

[0090] 1) The fault prediction module predicts that node 1 is about to fail; 2) The active-active system is notified to take precautions in advance; 3) The mirror node performs bitmap synchronization in advance; 4) The first segment bitmap is reassembled from 1-2 to 2-3, and the fourth segment bitmap is reassembled from 4-1 to 4-2; 5) After the bitmap is reassembled, node 2 takes over the business of node 1, and node 1 no longer undertakes business or allocates business to node 1; 6) Node 1 actually fails; 7) The active-active system responds quickly, without needing to reassemble the bitmap again or interrupting business, significantly reducing response time.

[0091] Understandably, the implementation details of 1) to 7) above are as described in the above embodiments and will not be repeated here.

[0092] Based on the above embodiments, Figure 3 This application provides a schematic flowchart of an early response to storage medium failures, specifically including the following: Figure 3 The following steps are shown:

[0093] 1) Detect the faulty disk; 2) Pre-migrate the data on the faulty disk; 3) Pre-calculate RAID verification; 4) Transfer IO load; 5) Push maintenance and operation early warning; 6) The replacement of the faulty disk is ready.

[0094] Understandably, the implementation details of 1) to 6) above are as described in the above embodiments and will not be repeated here.

[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0096] Embodiments of this application also provide a fault handling apparatus, wherein,

[0097] The acquisition unit is used to acquire real-time performance data of the target data center. The target data is the primary data center or the backup data center in a dual-active data center system. The real-time performance data includes: node data of the target data center, link data of the transmission link between the primary data center and the backup data center, and storage media data of the target data center.

[0098] The fault prediction unit is used to input real-time performance data into a pre-trained fault prediction model, compare the real-time performance data and performance change data with multiple set thresholds corresponding to the set fault type to predict the fault type, and obtain the prediction results output by the fault prediction model. The performance change data is obtained by performing time-series analysis on the real-time performance data.

[0099] The strategy execution unit is used to execute the fault response strategy corresponding to the target fault before the target fault occurs, if the prediction result indicates that the target data center will experience a target fault within a set future time. The target fault includes at least one of the following: a node fault determined based on node data, a link fault determined based on link data, and a storage medium fault determined based on storage medium data.

[0100] The target data center corresponds to at least one site, and each site consists of multiple nodes.

[0101] Optionally, the fault prediction unit is used for:

[0102] Based on node data, if the fault prediction model determines that the load of the target controller on the target node corresponding to the target data center is greater than the first set threshold and there is no high input / output trigger, the memory utilization of the target controller is greater than the second set threshold, the response latency carried by the target controller exceeds the third set threshold and there is no high input / output response latency of the corresponding other controllers, the target controller log shows a bus error, and / or the controller corresponding to another data center receives an alarm message indicating that the heartbeat interval of the controller corresponding to the target data center has been extended, then the prediction result of the node failure in the target data center and the target node being a faulty node within the set future time is output.

[0103] Optionally, the policy execution unit is used for:

[0104] If the prediction indicates that a node failure will occur in the target data center within a set future time, then before the node failure occurs, the bitmap used to characterize data synchronization between nodes is reorganized to redetermine the master node's takeover of the target node's services and to suspend the allocation of services to the target node; or, before the node failure occurs, the bitmap used to characterize data synchronization between nodes is reorganized to redetermine the master node's takeover of the target node's services, and after the node failure occurs, the allocation of services to the target node is suspended.

[0105] Optionally, the fault prediction unit is used for:

[0106] Based on link data, if the fault prediction model determines that the target link utilization rate is greater than the fourth set threshold, the bandwidth utilization fluctuation of the target link increases by the fifth set threshold under abnormal business growth, the data synchronization delay between the primary data center and the backup data center is the sixth set threshold, the data synchronization queue length of the primary data center is greater than the seventh set threshold and continues to grow, the input burst of the primary data center does not match the bandwidth growth rate of the target link, and / or the frame loss rate or retransmission rate of the target link is greater than the eighth set threshold, then the prediction result of the target data center experiencing a link failure within a set future time and the target link being a failed link will be output. Here, the target link refers to the link between the primary data center and the backup data center used for data synchronization.

[0107] Optionally, the policy execution unit is used for:

[0108] If the prediction results indicate that a link failure will occur in the target data center within a set future time, and the link failure is caused by link congestion due to a sudden surge in business, then an alarm message will be issued before the link failure occurs, indicating that the current business of the target data center exceeds the current pressure capacity.

[0109] If the current business is the core business and the pressure will not decrease within the first set time, activate the backup link of the target link;

[0110] If link congestion still exists when activating the backup link, suspend data synchronization between the primary data center and the backup data center;

[0111] If the load on the target link drops to the first set value after the second set time, incremental synchronization between the primary data center and the backup data center will be initiated, and the determined difference data will be synchronized to the backup data center. At the same time, data synchronization between the primary data center and the backup data center will be initiated.

[0112] Optionally, the policy execution unit is used for:

[0113] If the prediction indicates that a link failure will occur in the target data center within a set future time, and the link failure is due to link congestion caused by invalid traffic synchronization, then update the bitmap between the primary data center and the backup data center before the link failure occurs.

[0114] Analyze the synchronization logs between the primary and backup data centers to identify invalid data that has already been synchronized;

[0115] Initiate the synchronization of incremental data, excluding invalid data, between the primary and backup data centers;

[0116] Suspend data synchronization between the primary and backup data centers for non-core business operations.

[0117] Optionally, the fault prediction unit is used for:

[0118] Based on media data, if the fault prediction model determines that the target storage medium has an additional set number of mapped sectors, an uncorrectable sector appears in the target storage medium, the load of other storage media increases to a second set value due to the input / output transfer of the target storage medium, the target storage medium has a pre-reconstruction flag or an imminent reconstruction flag, the synchronization latency carried by the target storage medium is greater than a ninth set threshold, and / or the target link write latency increases to a third set value due to the blockage of the target storage medium, then the prediction results of the target data center experiencing storage medium failure and the target storage medium being a failure storage medium will be output. The target storage medium is used to store data in the primary data center and the backup data center.

[0119] Optionally, the policy execution unit is used for:

[0120] If the forecast indicates that the target data center will experience a storage media failure within a set future time, then before the storage media failure occurs, initiate an online data migration to migrate the data in the primary data center and / or backup data center to the backup storage media of the target storage media.

[0121] Adjust the target storage medium's input / output operations per second to the fourth preset value;

[0122] The verification information of the independent disk redundancy array where the target storage medium is located is calculated based on the data of the current normal storage medium, so that the verification information can be directly reused after the storage medium fails.

[0123] Optionally, the fault prediction device is also used for:

[0124] Calculate the prediction accuracy for different fault types based on the prediction results and historical prediction results;

[0125] If the prediction accuracy is lower than the tenth set threshold, then update the set thresholds used to predict different fault types in the fault prediction model.

[0126] For a description of the features in the embodiment corresponding to the fault handling device, please refer to the relevant description of the embodiment corresponding to the fault handling method, which will not be repeated here.

[0127] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described fault handling method embodiments.

[0128] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described fault handling method embodiments when it is run.

[0129] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0130] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described fault handling method embodiments.

[0131] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described fault handling method embodiments.

[0132] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0133] The foregoing has provided a detailed description of a fault handling method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A failure handling method characterized by, The method comprises: obtaining real-time performance data of a target data center, wherein the target data center is a primary data center or a backup data center in a dual-active data center system, and the real-time performance data comprises node data of the target data center, link data of a transmission link between the primary data center and the backup data center, and storage medium data of the target data center; inputting the real-time performance data into a pre-trained fault prediction model, extracting performance change data reflecting time sequence changes of performance indicators from the real-time performance data by the fault prediction model, comparing the real-time performance data and the performance change data with a plurality of set thresholds corresponding to set fault types to predict the fault types, and obtaining a prediction result output by the fault prediction model; if the prediction result indicates that the target data center will occur a target fault within a set future time, executing a fault handling strategy corresponding to the target fault before the target fault occurs, wherein the target fault comprises at least one of a node fault determined based on the node data, a link fault determined based on the link data, and a storage medium fault determined based on the storage medium data, and the fault handling strategy comprises at least one of bitmap reorganization, link switching, and data pre-migration.

2. The method of claim 1, wherein, The target data center corresponds to at least one site, and each site is composed of a plurality of nodes. The method of inputting the real-time performance data into a pre-trained fault prediction model, comparing the real-time performance data and the performance change data with a plurality of set thresholds corresponding to set fault types to predict the fault types by the fault prediction model, and obtaining a prediction result output by the fault prediction model comprises: based on the node data, if it is determined by the fault prediction model that the load of a target controller on a target node corresponding to the target data center is greater than a first set threshold and there is no high input / output trigger, the memory usage of the target controller is greater than a second set threshold, the response delay of the target controller exceeds a third set threshold and there is no high input / output response delay of the corresponding remaining controllers, the target controller has a bus error in the log, and / or the controller corresponding to the other data center receives an alarm message that the heartbeat interval of the controller corresponding to the target data center is prolonged, the prediction result that the target data center will occur a node fault within a set future time and the target node is a fault node is output.

3. The method of claim 2, wherein, The method of executing a fault handling strategy corresponding to the target fault before the target fault occurs comprises: if the prediction result indicates that the target data center will occur a node fault within the set future time, reorganizing a bitmap used to represent data synchronization between nodes to re-determine that a primary node takes over the business of the target node before the node fault occurs, and suspending the allocation of business to the target node; or reorganizing a bitmap used to represent data synchronization between nodes to re-determine that a primary node takes over the business of the target node before the node fault occurs, and suspending the allocation of business to the target node after the node fault occurs.

4. The method of claim 1, wherein, The real-time performance data is input into a pre-trained fault prediction model, the real-time performance data and the performance change data are compared with a plurality of set threshold values corresponding to set fault types through the fault prediction model to predict a fault type, and a prediction result output by the fault prediction model is obtained, and the prediction result output by the fault prediction model comprises: Based on the link data, if it is determined through the fault prediction model that the target link utilization rate is greater than a fourth set threshold value, the bandwidth utilization rate fluctuation amplitude of the target link increases by a fifth set threshold value in the case of abnormal business growth, the data synchronization delay between the primary data center and the backup data center is greater than a sixth set threshold value, the primary data center data synchronization queue length is greater than a seventh set threshold value and continuously increases, the primary data center input burst volume and the bandwidth growth amplitude of the target link do not match, and / or the frame loss rate or retransmission rate of the target link is greater than an eighth set threshold value, the prediction result that the target data center will occur link fault in a set future time and the target link is a fault link is output, wherein the target link refers to a link between the primary data center and the backup data center for data synchronization.

5. The method of claim 4, wherein, The target fault corresponding fault handling strategy is executed before the target fault occurs, and the target fault corresponding fault handling strategy comprises: If the prediction result indicates that the target data center will occur link fault in the set future time, and the link fault is caused by link congestion due to business burst, alarm information that the current business of the target data center exceeds the current bearable pressure range is sent before the link fault occurs; In the case that the current business is core business and the pressure will not decrease within a first set time, a backup link of the target link is started; In the case that there is still link congestion after starting the backup link, data synchronization between the primary data center and the backup data center is suspended; If the load carried by the target link decreases to a first set value after a second set time, incremental synchronization between the primary data center and the backup data center is started, and the determined difference data is synchronized to the backup data center, and data synchronization between the primary data center and the backup data center is started.

6. The method of claim 4, wherein, The target fault corresponding fault handling strategy is executed before the target fault occurs, and the target fault corresponding fault handling strategy comprises: If the prediction result indicates that the target data center will occur link fault in the set future time, and the link fault is caused by link congestion due to invalid traffic synchronization, the bitmap between the primary data center and the backup data center is updated before the link fault occurs; The synchronization log between the primary data center and the backup data center is analyzed to determine the invalid data that has been synchronized; Incremental data other than the invalid data is initiated to be synchronized between the primary data center and the backup data center; Data synchronization of non-core business between the primary data center and the backup data center is suspended.

7. The method of claim 1, wherein, The real-time performance data is input into a pre-trained fault prediction model, the real-time performance data and the performance change data are compared with a plurality of set threshold values corresponding to set fault types through the fault prediction model to predict the fault types, and a prediction result output by the fault prediction model is obtained, and the prediction result output by the fault prediction model comprises: Based on the storage medium data, if it is determined through the fault prediction model that a new set number of mapping sectors in a target storage medium, uncorrectable sectors in the target storage medium, a load of other storage media increased to a second set value due to input and output transfer of the target storage medium, a pre-reconstruction flag or a reconstruction imminent flag in the target storage medium, a synchronization time delay of the target storage medium greater than a ninth set threshold value, and / or a target link write delay increased to a third set value due to blocking of the target storage medium, a prediction result that the target data center will have a storage medium fault and the target storage medium is a fault storage medium within a set future time is output, wherein the target storage medium is used to store data of the primary data center and the backup data center.

8. The method of claim 7, wherein, The fault handling strategy corresponding to the target fault is executed before the target fault occurs, and the fault handling strategy corresponding to the target fault comprises: If the prediction result indicates that the target data center will have a storage medium fault within the set future time, online data migration is started before the storage medium fault occurs, and data in the primary data center and / or the backup data center is migrated to a backup storage medium of the target storage medium; The input and output operation number per second of the target storage medium is adjusted to a fourth set value; Based on data of a current normal storage medium, check information of a redundant array of independent disks in which the target storage medium is located is calculated, so that the check information is directly reused after the storage medium fault occurs.

9. The method of claim 1, wherein, After the prediction result output by the fault prediction model is obtained, the method further comprises: According to the prediction result and historical prediction results, a prediction accuracy rate of different fault types is calculated; If the prediction accuracy rate is lower than a tenth set threshold value, set threshold values for predicting different fault types in the fault prediction model are updated.

10. An electronic device, comprising: The method comprises: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the fault handling method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Fault prediction method, device and equipment for off-site active-active architecture and storage medium

    CN117648223A