Cluster disaster recovery method and device based on hyper-converged cluster

By using intelligent fault detection models in hyperconverged clusters, real-time detection and migration of virtual machines to target nodes, the problems of low fault detection accuracy and improper resource scheduling are solved, and disaster recovery capabilities and business recovery efficiency are improved.

CN120474897AInactive Publication Date: 2025-08-12JINAN INSPUR DATA TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510971870.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The accuracy of fault detection in the hyperconverged cluster is low and the resource scheduling mechanism is poor, resulting in insufficient disaster recovery capabilities.

Method used

The intelligent fault detection model is used to detect hyperconverged cluster nodes in real time, determine the target node based on the storage load information, data replica location and node location information, and migrate the virtual machines of the potential fault node to the target node.

Benefits of technology

Accurate prediction and automatic response to potential failures are achieved, resource scheduling is optimized, and business recovery is ensured quickly, improving the disaster recovery capacity and resource utilization efficiency of hyper-converged clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474897A_ABST
    Figure CN120474897A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster disaster recovery method and device based on a hyper-converged cluster, and relates to the field of computers, and the method comprises the steps: carrying out the fault detection of a plurality of nodes in the hyper-converged cluster through an intelligent fault detection model, and enabling the intelligent fault detection model to be deployed on a management node of the hyper-converged cluster, the plurality of nodes comprise the management node; when it is detected that a potential fault node exists in the multiple nodes, a target node is determined in the multiple nodes according to node information of the multiple nodes, and the node information comprises storage load information, data copy position information of a data copy corresponding to the node and node position information; the target node is different from the potential fault node; and migrating a target virtual machine corresponding to the potential fault node to the target node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage, and in particular to a cluster disaster recovery method and device based on a hyper-converged cluster. Background Art

[0002] With the rapid development of information technology, enterprises are increasingly dependent on data centers, making data security and business continuity crucial. Hyperconverged architectures are widely used in enterprise data center construction because they integrate computing, storage, networking, and server virtualization resources and technologies into a single unit device, enabling seamless modular horizontal expansion and forming a unified resource pool.

[0003] In a hyperconverged architecture, traditional disaster recovery solutions based on hyperconverged stretched clusters present several challenges. First, in terms of fault detection and recovery, existing technologies mostly employ rule-based fault detection methods, relying on preset thresholds and simple logical judgments to identify faults. This approach cannot cope with complex and volatile business scenarios and system operating conditions, making it difficult to predict potential failures in advance. For example, when business loads suddenly fluctuate abnormally, or when system resources (such as CPU and memory) exhibit complex trends over a period of time, rule-based detection methods often fail to accurately and promptly determine whether a failure risk exists. This results in reactive response after a failure has occurred, severely impacting business continuity. Furthermore, traditional solutions often require manual intervention during the fault recovery process, performing a series of complex operations such as manually switching services to a backup site and manually configuring relevant parameters. This is not only time-consuming but also prone to human error, further extending business interruption. Furthermore, in terms of resource scheduling mechanisms, there is a lack of close coordination between upper-layer virtual machine resource scheduling and underlying distributed storage. When migrating virtual machines, they often only consider the load on computing resources (such as CPU and memory), while ignoring key factors such as storage load (such as IOPS and bandwidth) and the location of data replicas. This can lead to severe performance degradation after the VM is migrated to the target node due to storage bottlenecks or excessive data transmission latency. For example, when migrating a VM to a node with an already heavy storage load, frequent I / O waits may occur, resulting in slow application response or even failure to operate properly. Furthermore, improper resource scheduling can increase pressure on network bandwidth, further impacting the stability and reliability of the entire system.

[0004] Currently, no effective solution has been proposed to address the problems of low fault detection accuracy and poor resource scheduling mechanism in hyper-converged clusters, which lead to poor disaster recovery capabilities of hyper-converged clusters. Summary of the Invention

[0005] The present application provides a cluster disaster recovery method and device based on a hyper-converged cluster, so as to at least solve the problems in the related art of low fault detection accuracy and poor resource scheduling mechanism in the hyper-converged cluster, which lead to poor disaster recovery capability of the hyper-converged cluster.

[0006] The present application provides a cluster disaster recovery method based on a hyper-converged cluster, comprising: performing fault detection on multiple nodes in the hyper-converged cluster through an intelligent fault detection model, wherein the intelligent fault detection model is deployed on a management node of the hyper-converged cluster, and the multiple nodes include the management node; in the case where a potential fault node is detected among the multiple nodes, determining a target node among the multiple nodes based on node information of the multiple nodes, wherein the node information includes: storage load information, data copy location information of a data copy corresponding to the node, node location information, and the target node is different from the potential fault node; and migrating a target virtual machine corresponding to the potential fault node to the target node.

[0007] The present application also provides a cluster disaster recovery device based on a hyper-converged cluster, comprising: a detection module, used to perform fault detection on multiple nodes in the hyper-converged cluster through an intelligent fault detection model, wherein the intelligent fault detection model is deployed on the management node of the hyper-converged cluster, and the multiple nodes include the management node; a determination module, used to determine a target node among the multiple nodes based on node information of the multiple nodes when a potential fault node is detected among the multiple nodes, wherein the node information includes: storage load information, data copy location information of the data copy corresponding to the node, node location information, and the target node is different from the potential fault node; a migration module, used to migrate the target virtual machine corresponding to the potential fault node to the target node.

[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned cluster disaster recovery methods based on a hyper-converged cluster when executing the computer program.

[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned cluster disaster recovery methods based on a hyper-converged cluster are implemented.

[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned cluster disaster recovery methods based on a hyper-converged cluster when executed by a processor.

[0011] Through the present application, fault detection is performed on multiple nodes in a hyper-converged cluster in real time through an intelligent fault detection model, wherein the intelligent fault detection model is deployed on the management node of the hyper-converged cluster, and these multiple nodes include the management node; if a potential fault node is detected among these nodes, a target node is determined among these nodes based on the node information of these nodes, wherein the node information includes: storage load information, data copy location information of the data copy corresponding to the node, node location information, and the target node is not the same node as the potential fault node; then the target virtual machine corresponding to the potential fault node is migrated to the target node; the above scheme provides a smarter and more efficient disaster recovery method, which can not only accurately predict and automatically respond to potential faults, but also optimize resource scheduling, ensuring that when a fault occurs, the business can be quickly restored, minimizing terminal time, while maintaining good operating performance, and improving the overall disaster recovery capability and resource utilization efficiency of the hyper-converged cluster; thereby solving the problem in related technologies that the fault detection accuracy in the hyper-converged cluster is low and the resource scheduling mechanism is poor, resulting in poor disaster recovery capability of the hyper-converged cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 This is a hardware structure block diagram of a cloud server according to an embodiment of the present application, which is a cluster disaster recovery method based on a hyper-converged cluster;

[0014] Figure 2 This is a flow chart of a cluster disaster recovery method based on a hyper-converged cluster according to an embodiment of the present application;

[0015] Figure 3 This is a flow chart of a disaster recovery method based on a hyper-converged stretched cluster according to an embodiment of the present application;

[0016] Figure 4 This is a structural block diagram of a cluster disaster recovery device based on a hyper-converged cluster according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0019] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the cluster disaster recovery method based on the hyper-converged cluster depends, the specific application environment architecture or specific hardware architecture is described here.

[0021] The method embodiments provided in the embodiments of the present application can be executed in a cloud server or similar computing device. Taking running on a cloud server as an example, Figure 1 This is a hardware structure diagram of a cloud server based on a cluster disaster recovery method of a hyper-converged cluster according to an embodiment of the present application. Figure 1 As shown, the cloud server may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The cloud server may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the cloud server. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0022] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the startup method of the operating system in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to a cloud server via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0023] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a cloud server's communication provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the internet wirelessly.

[0024] In this embodiment, a cluster disaster recovery method based on a hyper-converged cluster is provided, including but not limited to being applied to cloud servers. Figure 2 Flowchart of a cluster disaster recovery method based on a hyper-converged cluster according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps S202-S206:

[0025] Step S202: performing fault detection on multiple nodes in the hyper-converged cluster using an intelligent fault detection model, wherein the intelligent fault detection model is deployed on a management node of the hyper-converged cluster, and the multiple nodes include the management node;

[0026] Step S204: When a potential faulty node is detected among the multiple nodes, a target node is determined among the multiple nodes based on node information of the multiple nodes, wherein the node information includes: storage load information, data replica location information of a data replica corresponding to the node, and node location information, and the target node is different from the potential faulty node;

[0027] Step S206: Migrate the target virtual machine corresponding to the potential fault node to the target node.

[0028] Through the above steps, fault detection is performed on multiple nodes in a hyper-converged cluster in real time through an intelligent fault detection model, wherein the intelligent fault detection model is deployed on the management node of the hyper-converged cluster, and these multiple nodes include the management node; if a potential fault node is detected among these nodes, a target node is determined among these nodes based on the node information of these nodes, wherein the node information includes: storage load information, data copy location information of the data copy corresponding to the node, node location information, and the target node is not the same node as the potential fault node; then the target virtual machine corresponding to the potential fault node is migrated to the target node; the above scheme provides a smarter and more efficient disaster recovery method, which can not only accurately predict and automatically respond to potential faults, but also optimize resource scheduling, ensuring that when a fault occurs, the business can be quickly restored, minimizing terminal time, while maintaining good operating performance, and improving the overall disaster recovery capability and resource utilization efficiency of the hyper-converged cluster; thereby solving the problem in related technologies that the fault detection accuracy in the hyper-converged cluster is low and the resource scheduling mechanism is poor, resulting in poor disaster recovery capability of the hyper-converged cluster.

[0029] In an exemplary embodiment, determining a target node from among the multiple nodes based on the node information of the multiple nodes includes: obtaining the node information of the multiple nodes from a storage management center; screening out a plurality of candidate nodes from the multiple nodes based on the multiple storage load information, wherein the storage load of the multiple candidate nodes is lower than a first preset threshold; and determining the target node from among the multiple candidate nodes based on the first node position information corresponding to the multiple candidate nodes.

[0030] Within the framework of the exemplary embodiment, the process of determining the target node is streamlined and focused on the intelligent disaster recovery mechanism of the hyperconverged cluster. The process begins by centrally obtaining detailed node information for each node from the storage management center, including storage load information and node geographic location. Subsequently, an algorithm is implemented to select candidate nodes from the large pool of nodes based on this storage load information. These nodes all maintain storage load levels below a first preset threshold, ensuring that the target node has sufficient storage resources to accommodate the migrated virtual machines.

[0031] Next, based on the location information of the first candidate node, the system further precisely locates the node closest to the data replica as the target node. This step cleverly combines storage performance considerations with network topology advantages to significantly reduce data movement latency and ensure efficient operation of the virtual machine after migration.

[0032] Storage load information, specifically the node's IOPS (input and output operations per second) and bandwidth usage, is a key indicator of the real-time performance of storage devices or systems. In a hyperconverged architecture, storage load balancing is crucial, directly impacting the operational efficiency and user experience of virtual machines. The first preset threshold serves as a safety boundary for storage load, determining whether a node can accommodate additional virtual machines without performance degradation.

[0033] The comprehensive use of node information reflects the subtlety of hyper-converged cluster management. Through intelligent decision-making, the optimal path planning for virtual machine migration is achieved, which not only considers the immediate status of storage resources, but also the location information of data copies, aiming to build a highly reliable and high-performance disaster recovery solution.

[0034] Optionally, the target node is determined from the multiple candidate nodes based on the first node position information of the multiple candidate nodes, including: obtaining the first data copy position information of the target data copy used by the target virtual machine; calculating the distance values between the multiple candidate nodes and the target data copy based on the multiple first node position information and the first data copy position information; and determining the target node from the multiple candidate nodes based on the distance values, wherein the distance value corresponding to the target node is the smallest among the multiple distance values.

[0035] The target node selection process deeply incorporates location information considerations to optimize data access paths and ensure excellent performance after virtual machine migration. When the process starts, it first locates the data copy that the target virtual machine relies on, namely the target data copy, and then accurately obtains the location information of its first data copy. Next, the geographic distance calculation is performed on the candidate node group. Specifically, based on the first node location information of each candidate node, the system accurately calculates the geographic distance from each candidate node to the target data copy. This calculation process is the key to the decision-making mechanism, aiming to identify the node closest to the data copy in the network topology.

[0036] Ultimately, the system selects the node with the shortest geographical distance to the target data replica from the candidate list—that is, the node with the smallest distance value—as the target node. This strategy effectively overcomes the data transmission delays that can occur with traditional migration methods, ensuring that after migration, the virtual machine can access the required data replica directly or via a high-speed network path, significantly improving operational efficiency and user experience.

[0037] First, data replica location information, which includes the exact storage location of the data replica within the hyperconverged cluster, including the node ID and specific storage location, is a crucial parameter in resource scheduling and disaster recovery. Maintaining this detailed location information enables the system to quickly locate data replicas, enabling efficient data access and migration. Geographic distance calculation is based on the hyperconverged architecture's network topology. An algorithm analyzes node connections and transmission paths, quantifying the network distance from each node to the data replica, providing a key basis for decision-making.

[0038] When a virtual machine needs to be migrated, the migration decision module obtains storage load information and data replica location information for each node from the storage management center. First, nodes with storage loads below a preset threshold are selected as candidate nodes. Then, for each candidate node, the distance between it and the data replica currently used by the virtual machine is calculated (this can be calculated based on the network topology and node location information). The node closest to the data replica and with the lowest storage load is selected as the target node. If multiple nodes meet the requirements, the optimal target node is further selected based on other factors, such as the node's remaining computing resources and network connection quality.

[0039] Optionally, the target node is determined from the multiple candidate nodes based on the distance value, including: when the number of first target nodes is one, the first target node is determined as the target node, wherein the distance value corresponding to the first target node is the smallest among the multiple distance values; when the number of first target nodes is multiple, a second target node is determined from the multiple first target nodes based on the remaining computing resources and network connection quality of the multiple first target nodes, and the second target node is determined as the target node, wherein the node information includes the remaining computing resources and the network connection quality.

[0040] When refining the target node selection logic, the embodiment demonstrates a two-tiered decision path. First, if the only candidate node exhibits the minimum distance value, it is directly selected as the target node, ensuring the shortest data access path. This is a direct application of the principle of geographic proximity, aiming to reduce data transmission latency and optimize virtual machine performance.

[0041] Secondly, when multiple candidate nodes meet the minimum distance requirement, the system turns to a deeper analysis of resources and network quality. At this point, based on the detailed remaining computing resource status and network connection quality indicators in the node information, the system performs a secondary screening, selecting a second target node with both ample computing power and high-quality network connectivity, and ultimately finalizing it as the target node. This decision-making process not only consolidates the advantages of geographical proximity, but also further integrates computing and network considerations, ensuring that after migration, the virtual machine enjoys the shortest data access path while benefiting from high-performance computing resources and a stable network environment, comprehensively guaranteeing business continuity and user experience.

[0042] Remaining computing resources refer to the available CPU capacity, memory size, and other computing power indicators on a node, reflecting the node's ability to handle additional workloads. Network connection quality, encompassing key dimensions such as bandwidth, latency, and packet loss rate, serves as the basis for evaluating data transmission efficiency and stability. Through a comprehensive assessment of these two dimensions, the system can make more informed node selections, improving the overall effectiveness of disaster recovery migration.

[0043] Optionally, before obtaining the node information of the multiple nodes from the storage management center, the method also includes: monitoring the IOPS values and bandwidth usage of the multiple nodes in real time through a storage load monitoring tool, and updating the storage load information of the multiple nodes in real time in the storage management center based on the IOPS values and the bandwidth usage, wherein the multiple nodes are all deployed with the storage load monitoring tool; and, updating the data copy location information of the data copies corresponding to the multiple nodes in real time in the storage management center.

[0044] In the pre-processing phase of acquiring node information, the implementation incorporates a dual mechanism of real-time monitoring and information updates, ensuring that system decisions are based on the most accurate and up-to-date operational status. Specifically, all nodes are equipped with storage load monitoring tools that continuously track the node's IOPS (input and output operations per second, a key metric for storage device access speed) and bandwidth usage, providing immediate feedback to the storage management center. Subsequently, the storage management center continuously updates the storage load information based on the real-time data received, ensuring its freshness and validity, and providing a solid data foundation for subsequent node screening.

[0045] The Storage Management Center also maintains a dynamic database of data replica location information, updating data replica location information in real time. This database tracks the precise location of each data replica, including the node and network path it resides on. This real-time update of information is crucial for migrating virtual machines to optimal locations. It ensures the system accurately understands the distribution of all data replicas when making migration decisions, enabling optimal selection based on geographic location and storage availability.

[0046] Through the above embodiments, the system can not only respond to changes in storage resource status in real time, but also dynamically grasp data distribution, providing refined resource management and data positioning capabilities for the implementation of intelligent disaster recovery strategies.

[0047] It's important to note that each node in a hyperconverged cluster is equipped with a storage load monitoring tool to monitor its IOPS and bandwidth usage in real time. This monitoring data is aggregated and sent to the storage management center, which maintains a storage load information table and updates the storage load status of each node in real time. Furthermore, within the distributed storage system, a data replica location information table is maintained, recording the node and site information for each data replica. This table is updated promptly when changes occur to data replicas.

[0048] Optionally, after migrating the target virtual machine corresponding to the potential fault node to the target node, the method further includes: when it is determined that the migration of the target virtual machine is completed, inspecting the running status of the target virtual machine on the target node; when it is determined that the running status fails the inspection, re-migrating the target virtual machine to other nodes, wherein the multiple nodes include the other nodes, and the other nodes do not include the potential fault node and the target node.

[0049] After completing the migration of the target virtual machine on the potential fault node, the embodiment further emphasizes the importance of post-migration verification and adjustment. Once the system confirms that the process of migrating the virtual machine to the target node is successfully completed, it will immediately start the operation verification program to closely monitor the operation status of the target virtual machine on the new node to ensure a smooth transition and that the performance meets the expected standards. If the test results show that the operation status of the virtual machine does not meet the standards, that is, the operation fails to pass the test, the system will quickly take remedial measures and reschedule the virtual machine to other nodes. This series of operations automatically avoids the identified potential fault nodes and previously tried target nodes, aiming to find a more suitable host node until the target virtual machine runs stably in the new environment.

[0050] Operational verification is essentially a comprehensive assessment of the performance and compatibility of the VM after migration, covering key metrics such as CPU utilization, memory usage, and network connection stability to ensure efficient operation of the VM on the target node. The decision to re-migrate demonstrates the system's flexibility and redundancy in disaster recovery. Even if the initial selection fails to meet expectations, the strategy can be quickly adjusted to ensure unimpeded business continuity, avoid resource waste, and optimize overall system performance. This mechanism significantly enhances the robustness of the disaster recovery solution and its reliability in practical applications.

[0051] Based on the above steps, after migrating the target virtual machine corresponding to the potential fault node to the target node, the method further includes: adding a fault mark to the potential fault node to prompt the target object to repair the potential fault node; isolating the potential fault node from the hyper-converged cluster to eliminate the interference of the potential fault node on the hyper-converged cluster.

[0052] After the target virtual machine is migrated to a secure node, the implementation focuses on identifying and isolating potentially faulty nodes to facilitate system maintenance and ensure the stability and efficiency of the overall cluster operation. The system then applies a fault mark to the potentially faulty node. This mark not only intuitively reminds the maintenance team to pay attention to the node's health, but also urges prompt scheduling of repairs to ensure that hardware issues are promptly resolved. At the same time, potentially faulty nodes are strategically isolated from the hyper-converged cluster to prevent potential instability from affecting other cluster members, ensuring continuous business operations and data security.

[0053] Fault marking is an internal system alert mechanism, typically in the form of a label or status indicator. It accurately identifies nodes in the system that may experience problems, providing clues for rapid location during subsequent troubleshooting and hardware maintenance. Isolation refers to the system temporarily removing a node from the network or cluster when it detects a potential failure or a confirmed failure, halting its external services to prevent the faulty node from negatively impacting cluster performance, data consistency, or security. This mechanism embodies the proactive defense approach of disaster recovery strategies. Through timely isolation, it ensures cluster robustness and business continuity, making it an essential component of building a stable and reliable hyperconverged architecture.

[0054] Optionally, fault detection is performed on multiple nodes in a hyper-converged cluster through an intelligent fault detection model, including: collecting node indicator data of the multiple nodes, wherein the node indicator data is used to indicate the health status of the node; performing data preprocessing on the multiple node indicator data to obtain multiple preprocessed node indicator data; and performing fault detection on the multiple preprocessed node indicator data through the intelligent fault detection model to determine whether the potential fault node exists.

[0055] This example succinctly outlines the core process of an intelligent fault detection model for identifying potentially faulty nodes in a hyperconverged cluster. After the model is launched, the system quickly conducts a comprehensive scan of each node in the cluster, collecting a variety of node metrics, including CPU utilization, memory usage, network bandwidth, and I / O operation frequency. These data sets form a barometer of node health. This collected raw data then undergoes a preprocessing step to remove non-critical factors, normalize the numerical range, and convert it into a format suitable for model analysis. This preprocessed node metric data is then converted.

[0056] Subsequently, intelligent fault detection models played a key role, applying machine learning algorithms to deeply analyze this pre-processed data, identifying underlying problem patterns and accurately pinpointing nodes in the cluster that might experience failures. This series of actions not only demonstrated the proactive role of technology in fault prevention but also provided solid data support for subsequent disaster recovery decisions.

[0057] Node metric data is a combination of various indicators regularly collected by the system that reflect the operational status of nodes. These indicators include, but are not limited to, hardware resource consumption, network communication status, and storage access efficiency. They serve as direct evidence for assessing node health. Data preprocessing, a necessary step before data analysis, involves data cleaning, formatting, outlier handling, and missing value filling. This process aims to improve data quality and consistency, ensure the accuracy and efficiency of model training, and more effectively identify potential sources of failure in the cluster.

[0058] Based on the above steps, after performing fault detection on the multiple preprocessed node indicator data through the intelligent fault detection model to determine whether the potential fault node exists, the method also includes: when it is determined that the potential fault node exists, sending a fault warning signal to the management node to instruct the management node to migrate the target virtual machine, wherein the fault warning signal carries the fault category information of the potential fault node.

[0059] Once the intelligent fault detection model identifies a potential faulty node, subsequent processes are rapidly initiated, transmitting a fault warning signal to the management node. This signal carries fault classification information—the nature and type of the fault—providing clear evidence for fault identification. Upon receiving the signal, the management node immediately interprets the fault classification information to guide specific disaster recovery operations, such as virtual machine migration, ensuring unimpeded services and data security.

[0060] Fault warning signals are alerts sent to management nodes when the system detects signs of a failure that could impact stability. These signals include not only the identification of the faulty node but also the type of fault, such as hardware failure, software anomaly, or network outage. This detailed information helps management nodes quickly locate problems and implement precise policies. As the control center of the hyperconverged cluster, the management node is responsible for receiving and responding to warning signals from the intelligent fault detection model, performing operations such as virtual machine migration, resource scheduling, fault isolation, and subsequent fault repair, ensuring the cluster can respond quickly to potential threats and maintaining business continuity and data security. This mechanism demonstrates the intelligent and automated nature of disaster recovery strategies, effectively improving the system's self-recovery capabilities and operational efficiency.

[0061] Optionally, data preprocessing is performed on the multiple node indicator data to obtain multiple preprocessed node indicator data, including: data cleaning is performed on the multiple node indicator data to obtain multiple cleaned node indicator data; data format conversion is performed on the multiple cleaned node indicator data to obtain multiple converted node indicator data, wherein the multiple converted node indicator data are all in the target format; data normalization is performed on the multiple converted node indicator data to obtain the multiple preprocessed node indicator data.

[0062] When processing node indicator data, the system first performs a deep cleansing of the various collected indicators, eliminating noise and outliers to ensure data purity, resulting in cleaned node indicator data. The system then converts this cleaned data into a pre-set target format to ensure consistency and readability for subsequent analysis. Finally, data normalization is performed to adjust the data range to a consistent scale, preventing biased model analysis caused by varying values. This completes the preprocessing process and generates preprocessed node indicator data, providing accurate and standardized input for the intelligent fault detection model.

[0063] Data cleaning is the initial step in data preprocessing, aiming to remove redundant, erroneous, or irrelevant data from collected data to improve its quality and reliability. Data format conversion ensures that all indicator data meets the input requirements of the intelligent model, addressing the issue of diverse data sources and formats, and enabling effective data identification and processing within the model. Data normalization, by scaling data values to a specified range, such as 0 to 1, eliminates the impact of numerical differences between indicators on model training, ensuring the fairness and accuracy of model analysis results. This series of preprocessing steps collectively builds a solid foundation for data preparation and ensures the efficient operation of intelligent models.

[0064] Optionally, before performing fault detection on the multiple preprocessed node indicator data through the intelligent fault detection model, the method also includes: obtaining historical fault data of the multiple nodes; performing data preprocessing on the multiple historical fault data to obtain multiple preprocessed historical fault data; and training the intelligent fault detection model through the multiple preprocessed historical fault data.

[0065] Before using the intelligent fault detection model to analyze pre-processed node indicator data to identify signs of failure, the system must first undergo a robust model training process. This preparatory stage involves collecting historical fault records for each node. This historical fault data covers detailed information such as fault type, duration, scope of impact, and handling results. Subsequently, the acquired historical fault data is pre-processed, including data cleaning, format unification, and normalization, to produce a pre-processed historical fault dataset. Using this dataset as training material, the intelligent model learns and grasps the patterns and regularities of fault occurrence, ultimately developing an intelligent fault detection model with accurate fault prediction capabilities, laying a solid algorithmic foundation for real-time fault monitoring.

[0066] Historical fault data is a collection of fault instances accumulated by the system. It records the details of past failure events, including the time, cause, impact, and handling process. It is key information for training models to identify future failure modes. Preprocessed historical fault data is cleaned to remove clutter, converted to a unified data format, and then normalized to make the dataset more organized and easier for intelligent models to learn and understand, thereby improving training efficiency and the model's predictive accuracy. Intelligent fault detection models, trained on historical fault data, learn to identify fault characteristics. They can quickly analyze node indicator data in real-time monitoring and accurately determine failure risks. They are a critical technical tool for maintaining stable operations in modern data centers. This training process ensures the reliability and effectiveness of the model in practical applications, providing strong support for fault prevention and timely response.

[0067] This application utilizes AI technology to build an intelligent fault detection model. The AI model is trained by collecting a large amount of historical fault data and service feature data. This data includes, but is not limited to, resource load data such as CPU usage, memory utilization, service load, and network bandwidth utilization, as well as site health data such as network outages and storage accessibility. During the training process, deep learning algorithms, such as neural networks, are used to enable the model to learn complex relationships and patterns between data. The trained AI model can analyze system operating data in real time and predict potential failure points. When signs of a possible failure are detected, the disaster recovery process is automatically triggered. For example, if the model predicts that a node's CPU usage will consistently exceed a warning threshold for a period of time, accompanied by insufficient memory resources and abnormal network bandwidth utilization, and signs of unstable network connectivity at the node's site, the model will determine that the node is at high risk of failure and immediately initiate the disaster recovery process. During the disaster recovery process, the system automatically executes actions based on pre-set policies. For example, affected virtual machines can be quickly migrated to other healthy sites or nodes, while the faulty node is isolated and diagnosed for subsequent repair.

[0068] The embodiment of the present application provides an optional disaster recovery method based on a hyper-converged extended cluster (equivalent to the above-mentioned hyper-converged cluster), such as Figure 3 As shown in the figure, the specific implementation process includes:

[0069] Step 1: Data collection agent;

[0070] Deploy data collection agents on each node of the hyper-converged cluster.

[0071] Step 2: Collect resource load and site health data in real time;

[0072] Step 3: Data is transmitted to the processing center;

[0073] Step 4: Perform data cleaning / denoising / normalization preprocessing at the data processing center to ensure data accuracy and consistency, providing a high-quality data foundation for subsequent model training and analysis;

[0074] Step 5. Select a deep learning framework to build an AI model, such as TensorFlow or PyTorch, and build a neural network model. The model structure can include an input layer, multiple hidden layers, and an output layer.

[0075] Step 6: The input layer receives preprocessed data;

[0076] Step 7: The hidden layer learns the data feature patterns through complex neuron connections and activation functions;

[0077] Step 8: The output layer outputs the fault prediction results of the potential fault points;

[0078] Step 9: During the model training process, continuously adjust the model parameters, such as weights and biases, to improve the model's prediction accuracy. Use cross-validation and other techniques to evaluate and optimize the model to ensure that the model has good generalization ability.

[0079] Step 10: Deploy the trained AI model to the management node;

[0080] Step 11: Receive node operation data in real time;

[0081] Step 12: The model analyzes the data in real time to determine whether a potential fault point (equivalent to the potential fault node) is detected; if so, execute step 13; otherwise, return to step 11;

[0082] Step 13: Send an early warning signal to the management system (equivalent to the above-mentioned management node);

[0083] Step 14: The management system triggers the disaster recovery process;

[0084] Step 15: Select a backup site or node according to the preset strategy;

[0085] Step 16: Automatically migrate virtual machines and isolate faulty nodes using automated scripts and tools.

[0086] Obviously, the embodiments described above are only part of the embodiments of the present application, rather than all the embodiments. In order to better understand the above method, the above process is described below in conjunction with the embodiments, but it is not intended to limit the technical solutions of the embodiments of the present application. Specifically:

[0087] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0088] The embodiment of the present application also provides a cluster disaster recovery device based on a hyper-converged cluster, Figure 4 A cluster disaster recovery device based on a hyper-converged cluster according to an embodiment of the present application is provided. Figure 4 As shown, the device includes:

[0089] a detection module 42, configured to perform fault detection on a plurality of nodes in a hyper-converged cluster using an intelligent fault detection model, wherein the intelligent fault detection model is deployed on a management node of the hyper-converged cluster, and the plurality of nodes includes the management node;

[0090] a determination module 44 configured to, when detecting that a potential faulty node exists among the plurality of nodes, determine a target node among the plurality of nodes based on node information of the plurality of nodes, wherein the node information includes: storage load information, data replica location information of a data replica corresponding to the node, and node location information, wherein the target node is different from the potential faulty node;

[0091] The migration module 46 is configured to migrate the target virtual machine corresponding to the potential fault node to the target node.

[0092] Through the above device, fault detection is performed on multiple nodes in a hyper-converged cluster in real time through an intelligent fault detection model, wherein the intelligent fault detection model is deployed on the management node of the hyper-converged cluster, and these multiple nodes include the management node; if a potential fault node is detected among these nodes, a target node is determined among these nodes based on the node information of these nodes, wherein the node information includes: storage load information, data copy location information of the data copy corresponding to the node, node location information, and the target node is not the same node as the potential fault node; then the target virtual machine corresponding to the potential fault node is migrated to the target node; the above scheme provides a smarter and more efficient disaster recovery method, which can not only accurately predict and automatically respond to potential faults, but also optimize resource scheduling, ensuring that when a fault occurs, the business can be quickly restored, minimizing terminal time, while maintaining good operating performance, and improving the overall disaster recovery capability and resource utilization efficiency of the hyper-converged cluster; thereby solving the problem in related technologies that the fault detection accuracy in the hyper-converged cluster is low and the resource scheduling mechanism is poor, resulting in poor disaster recovery capability of the hyper-converged cluster.

[0093] Optionally, the determination module 44 is further used to obtain node information of the multiple nodes from the storage management center; filter out multiple candidate nodes from the multiple nodes based on the multiple storage load information, wherein the storage load of the multiple candidate nodes is lower than a first preset threshold; and determine the target node from the multiple candidate nodes based on the first node position information corresponding to the multiple candidate nodes.

[0094] Optionally, the above-mentioned determination module 44 is also used to obtain the first data copy location information of the target data copy used by the target virtual machine; calculate the distance values between the multiple candidate nodes and the target data copy based on the multiple first node location information and the first data copy location information; determine the target node from the multiple candidate nodes based on the distance value, wherein the distance value corresponding to the target node is the smallest among the multiple distance values.

[0095] Optionally, the above-mentioned determination module 44 is used to, when the number of first target nodes is one, determine the first target node as the target node, wherein the distance value corresponding to the first target node is the smallest among the multiple distance values; when the number of first target nodes is multiple, determine the second target node from the multiple first target nodes based on the remaining computing resources and network connection quality of the multiple first target nodes, and determine the second target node as the target node, wherein the node information includes the remaining computing resources and the network connection quality.

[0096] Optionally, the above-mentioned determination module 44 is also used to monitor the IOPS values and bandwidth usage of the multiple nodes in real time through a storage load monitoring tool, and update the storage load information of the multiple nodes in real time in the storage management center according to the IOPS values and the bandwidth usage, wherein the multiple nodes are all deployed with the storage load monitoring tool; and, update the data copy location information of the data copies corresponding to the multiple nodes in real time in the storage management center.

[0097] Optionally, the above-mentioned migration module 46 is also used to test the operation status of the target virtual machine on the target node when it is determined that the migration of the target virtual machine is completed; and when it is determined that the operation status fails the test, re-migrate the target virtual machine to other nodes, wherein the multiple nodes include the other nodes, and the other nodes do not include the potential fault node and the target node.

[0098] Optionally, the migration module 46 is further configured to add a fault mark to the potential fault node to prompt the target object to repair the potential fault node; and isolate the potential fault node from the hyper-converged cluster to eliminate interference of the potential fault node on the hyper-converged cluster.

[0099] Optionally, the above-mentioned detection module 42 is also used to collect node indicator data of the multiple nodes, wherein the node indicator data is used to indicate the health status of the node; perform data preprocessing on the multiple node indicator data to obtain multiple preprocessed node indicator data; and perform fault detection on the multiple preprocessed node indicator data through the intelligent fault detection model to determine whether there is a potential fault node.

[0100] Optionally, the above-mentioned detection module 42 is also used to send a fault warning signal to the management node when it is determined that the potential fault node exists, so as to instruct the management node to migrate the target virtual machine, wherein the fault warning signal carries the fault category information of the potential fault node.

[0101] Optionally, the above-mentioned detection module 42 is also used to perform data cleaning on the multiple node indicator data to obtain multiple cleaned node indicator data; perform data format conversion on the multiple cleaned node indicator data to obtain multiple converted node indicator data, wherein the multiple converted node indicator data are all in the target format; perform data normalization on the multiple converted node indicator data to obtain the multiple preprocessed node indicator data.

[0102] Optionally, the detection module 42 is further configured to obtain historical fault data of the plurality of nodes; perform data preprocessing on the plurality of historical fault data to obtain a plurality of preprocessed historical fault data; and train the intelligent fault detection model using the plurality of preprocessed historical fault data.

[0103] For descriptions of features in the embodiments corresponding to the cluster disaster recovery device based on a hyper-converged cluster, please refer to the relevant descriptions of the embodiments corresponding to the cluster disaster recovery method based on a hyper-converged cluster, which will not be repeated here.

[0104] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the cluster disaster recovery method based on a hyper-converged cluster.

[0105] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the cluster disaster recovery method based on a hyper-converged cluster when running.

[0106] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0107] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned cluster disaster recovery method embodiments based on a hyper-converged cluster are implemented.

[0108] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned cluster disaster recovery method embodiments based on a hyper-converged cluster.

[0109] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0110] The above is a detailed introduction to a cluster disaster recovery method and device based on a hyper-converged cluster provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A cluster disaster recovery method based on a hyper-converged cluster, characterized in that: include: Performing fault detection on multiple nodes in a hyper-converged cluster using an intelligent fault detection model, wherein the intelligent fault detection model is deployed on a management node of the hyper-converged cluster, and the multiple nodes include the management node; In a case where a potential fault node is detected among the multiple nodes, determining a target node among the multiple nodes according to node information of the multiple nodes, wherein the node information includes: storage load information, data copy location information of a data copy corresponding to the node, and node location information, and the target node is different from the potential fault node; Migrate the target virtual machine corresponding to the potential fault node to the target node.

2. The cluster disaster recovery method based on a hyper-converged cluster according to claim 1, characterized in that: Determining a target node from the multiple nodes according to the node information of the multiple nodes includes: Obtaining node information of the plurality of nodes from a storage management center; Filtering a plurality of candidate nodes from the plurality of nodes according to the plurality of storage load information, wherein the storage loads of the plurality of candidate nodes are lower than a first preset threshold; The target node is determined from the multiple candidate nodes according to the first node position information corresponding to the multiple candidate nodes.

3. The cluster disaster recovery method based on a hyper-converged cluster according to claim 2, characterized in that: Determining the target node from the multiple candidate nodes according to the first node position information corresponding to the multiple candidate nodes includes: Obtaining first data copy location information of a target data copy used by the target virtual machine; Calculate the distance values between the plurality of candidate nodes and the target data copy according to the plurality of first node position information and the first data copy position information; The target node is determined from the multiple candidate nodes according to the distance value, wherein the distance value corresponding to the target node is the smallest among the multiple distance values.

4. The cluster disaster recovery method based on a hyper-converged cluster according to claim 3, characterized in that: Determining the target node from the multiple candidate nodes according to the distance value includes: In a case where the number of the first target node is one, determining the first target node as the target node, wherein the distance value corresponding to the first target node is the smallest among the multiple distance values; In the case where there are multiple first target nodes, a second target node is determined from the multiple first target nodes based on the remaining computing resources and network connection quality of the multiple first target nodes, and the second target node is determined as the target node, wherein the node information includes the remaining computing resources and the network connection quality.

5. The cluster disaster recovery method based on a hyper-converged cluster according to claim 2, characterized in that: Before obtaining the node information of the plurality of nodes from the storage management center, the method further includes: monitoring the IOPS values and bandwidth usage of the plurality of nodes in real time through a storage load monitoring tool, and updating storage load information of the plurality of nodes in real time in the storage management center according to the IOPS values and the bandwidth usage, wherein the storage load monitoring tool is deployed on each of the plurality of nodes; and The data copy location information of the data copies corresponding to the multiple nodes is updated in real time in the storage management center.

6. The cluster disaster recovery method based on a hyper-converged cluster according to claim 1, characterized in that: After migrating the target virtual machine corresponding to the potential fault node to the target node, the method further includes: When it is determined that the migration of the target virtual machine is complete, checking the running status of the target virtual machine on the target node; If it is determined that the operation status fails the inspection, the target virtual machine is re-migrated to another node, wherein the multiple nodes include the other nodes, and the other nodes do not include the potential fault node and the target node.

7. The cluster disaster recovery method based on a hyper-converged cluster according to claim 1, characterized in that: After migrating the target virtual machine corresponding to the potential fault node to the target node, the method further includes: Adding a fault mark to the potential fault node to prompt the target object to repair the potential fault node; The potential fault node is isolated from the hyper-converged cluster to eliminate interference of the potential fault node on the hyper-converged cluster.

8. The cluster disaster recovery method based on a hyper-converged cluster according to claim 1, characterized in that: Fault detection for multiple nodes in a hyper-converged cluster is performed using an intelligent fault detection model, including: Collecting node indicator data of the plurality of nodes, wherein the node indicator data is used to indicate the health status of the nodes; Performing data preprocessing on the plurality of node indicator data to obtain a plurality of preprocessed node indicator data; Fault detection is performed on the plurality of pre-processed node indicator data using the intelligent fault detection model to determine whether there is a potential fault node.

9. The cluster disaster recovery method based on a hyper-converged cluster according to claim 8, characterized in that: After performing fault detection on the plurality of pre-processed node indicator data using the intelligent fault detection model to determine whether there is a potential fault node, the method further includes: When it is determined that the potential fault node exists, a fault warning signal is sent to the management node to instruct the management node to migrate the target virtual machine, wherein the fault warning signal carries fault category information of the potential fault node.

10. The cluster disaster recovery method based on a hyper-converged cluster according to claim 8, characterized in that: Performing data preprocessing on the plurality of node indicator data to obtain a plurality of preprocessed node indicator data, including: performing data cleaning on the plurality of node indicator data to obtain a plurality of cleaned node indicator data; Performing data format conversion on the plurality of cleaned node indicator data to obtain a plurality of converted node indicator data, wherein the plurality of converted node indicator data are all in a target format; Data normalization processing is performed on the plurality of converted node indicator data to obtain the plurality of preprocessed node indicator data.

11. The cluster disaster recovery method based on a hyper-converged cluster according to claim 8, characterized in that: Before performing fault detection on the plurality of pre-processed node indicator data using the intelligent fault detection model, the method further includes: Obtaining historical fault data of the multiple nodes; performing data preprocessing on the plurality of historical fault data to obtain a plurality of preprocessed historical fault data; The intelligent fault detection model is trained using the plurality of pre-processed historical fault data.

12. A cluster disaster recovery device based on a hyper-converged cluster, characterized in that: include: a detection module, configured to perform fault detection on a plurality of nodes in a hyper-converged cluster using an intelligent fault detection model, wherein the intelligent fault detection model is deployed on a management node of the hyper-converged cluster, and the plurality of nodes includes the management node; a determination module, configured to, when detecting that a potential fault node exists among the multiple nodes, determine a target node among the multiple nodes based on node information of the multiple nodes, wherein the node information includes: storage load information, data copy location information of a data copy corresponding to the node, and node location information, wherein the target node is different from the potential fault node; A migration module is used to migrate the target virtual machine corresponding to the potential fault node to the target node.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

14. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 11 when executing the computer program.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Virtual machine migration system, method and device

    CN107544839A

  • High-availability implementation method based on hyper-converged double nodes

    CN110912991A

  • Distributed system fault judgment and recovery method, cloud operating system applying method and computing platform

    CN118331779A

  • Automatic maintenance method and device of container and computer program product

    CN119806880A