Single machine room fault recovery laas cluster processing system

By designing a single-unit failure recovery laas cluster processing system, the single point of failure risk, detection delay and inflexible resource scheduling in the IaaS cluster processing system are solved, rapid failure recovery and business continuity are achieved, and the system's disaster recovery capabilities and resource utilization efficiency are improved.

CN120342837APending Publication Date: 2025-07-18GUANGXI POWER GRID CORP
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510301999.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing IaaS cluster processing system has problems such as high single point failure risk, delayed fault detection response, inflexible resource scheduling and insufficient cross-computer room collaboration capabilities, resulting in service interruption and resource waste.

Method used

A single-computer room fault recovery LaAS cluster processing system is designed, including a fault detection module, a fault assessment module, a resource scheduling module, a service migration module, a traffic scheduling module, a redundant deployment module and a cross-computer room collaboration module. Through real-time monitoring, dynamic adjustment and cross-computer room collaboration, rapid fault recovery and service migration are achieved.

Benefits of technology

It realizes rapid identification of potential failures, dynamic resource scheduling, reduce service interruption time, and improves system disaster recovery capabilities, ensuring business continuity and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342837A_ABST
    Figure CN120342837A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing and server hosting technology, and discloses a single machine room fault recovery laas cluster processing system, comprising a fault detection module which is responsible for monitoring the health state of each node in a cluster in real time, analyzing system logs and performance indexes, and quickly identifying potential faults; the fault evaluation module is used for receiving the output of the fault detection module, evaluating a fault influence range and severity and providing a basis for a fault recovery strategy; the resource scheduling module is used for dynamically adjusting resource allocation according to a fault assessment result and selecting an optimal healthy node for service migration; and the service migration module is responsible for smoothly migrating the service on the fault node to the healthy node, and reducing the service interruption time by adopting a live migration technology. The invention provides a single-computer-room fault recovery laas cluster processing system, and solves the problems that a cluster processing system in the prior art is high in single-point fault risk, delayed in fault detection response, inflexible in resource scheduling and insufficient in cross-computer-room cooperative capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of cloud computing and server hosting, and particularly to a single data center fault recovery IaaS cluster processing system. Background Art

[0002] A single data center fault recovery IaaS cluster processing system is a comprehensive system that can quickly and automatically detect, diagnose, and recover server faults in an IaaS cluster in a single data center environment. It can ensure that when a single data center fault occurs, the cluster service can be quickly restored to guarantee business continuity. Currently, with the popularization of cloud computing technology, more and more enterprises and applications are deployed on the IaaS platform. In the existing technology, although the IaaS cluster processing system already has a certain degree of fault recovery ability, it still faces the following technical problems:

[0003] Single point of failure risk: There are key components or services in the system that are not redundantly deployed. Once these single points fail, the overall service will be interrupted.

[0004] Fault detection and response delay: It relies on manual or inefficient automated means for fault detection, with a long response time, which affects the fault recovery efficiency.

[0005] Inflexible resource scheduling: When a fault occurs, the resource scheduling strategy lacks intelligence and cannot be dynamically adjusted according to the real-time load situation, which may lead to uneven resource allocation or waste.

[0006] Insufficient cross-data center collaboration ability: There is a lack of an effective cross-data center collaboration mechanism. When a single data center fails, it is difficult to quickly migrate services to other data centers.

[0007] To solve the above problems, a single data center fault recovery IaaS cluster processing system is proposed in this application. Summary of the Invention

[0008] Based on the technical problems existing in the background art, the present invention proposes a single data center fault recovery IaaS cluster processing system.

[0009] A single data center fault recovery IaaS cluster processing system proposed by the present invention includes:

[0010] Fault detection module: responsible for real-time monitoring of the health status of each node in the cluster, analyzing system logs and performance metrics, and quickly identifying potential faults.

[0011] Fault assessment module: receives the output of the fault detection module, evaluates the scope and severity of the fault impact, and provides a basis for the fault recovery strategy.

[0012] Resource scheduling module: according to the fault assessment result, dynamically adjusts resource allocation, and selects the best healthy node for service migration.

[0013] Service Migration Module: Responsible for smoothly migrating services on faulty nodes to healthy nodes, and adopting hot migration technology to reduce service interruption time;

[0014] Traffic Scheduling Module: Automatically adjusts network traffic routing according to service migration situations to ensure that user requests can be correctly forwarded to new service nodes;

[0015] Redundant Deployment Module: Responsible for redundant deployment of key components and services to ensure that the system can still provide complete services in case of single-point failures;

[0016] Cross-Data Center Collaboration Module: Establishes cross-data center communication mechanisms to achieve sharing of fault information and collaboration in service migration between data centers, and improves the disaster tolerance of the overall system;

[0017] Operation and Maintenance Management Module: Provides a unified operation and maintenance management interface, supports fault recording, analysis, reporting, and policy optimization, and improves operation and maintenance efficiency;

[0018] The fault detection module is connected to the fault assessment module, the fault assessment module is connected to the resource scheduling module, the resource scheduling module is connected to the service migration module, the service migration module is connected to the traffic scheduling module, the cross-data center collaboration module is respectively connected to the resource scheduling module and the service migration module, the redundant deployment module is connected to all other modules, and the operation and maintenance management module is connected to all other modules.

[0019] Preferably, the fault detection module includes:

[0020] Log Collection Unit: Responsible for collecting system logs from each node in the cluster, including application logs, system logs, and security logs;

[0021] Performance Monitoring Unit: Real-time monitors the performance metrics of each node in the cluster, and these metrics can reflect the health status and performance bottlenecks of the nodes;

[0022] Machine Learning Analysis Unit: Analyzes logs and performance metrics using an SVM classifier to identify abnormal patterns and potential faults;

[0023] Alarm Notification Unit: When potential faults are detected, sends alarm messages to operation and maintenance personnel via email, text message, and system notification for timely response and handling.

[0024] Preferably, the fault assessment module includes:

[0025] Fault Information Receiving Unit: Receives fault information from the fault detection module, including fault type, location, and time;

[0026] Risk assessment unit: Analyze the fault information using Bayesian network to evaluate the impact scope, severity, and possible losses of the fault;

[0027] Strategy recommendation unit: Provide recommendations for subsequent fault recovery strategies based on the risk assessment results, including which faults to handle first and what recovery measures to take.

[0028] Preferably, the resource scheduling module includes:

[0029] Resource monitoring unit: Real-time monitor the resource usage of each node in the cluster, including CPU, memory, storage, and network resources;

[0030] Greedy algorithm unit: Dynamically schedule resources using the greedy algorithm and select the best healthy node for service migration;

[0031] Scheduling execution unit: Execute specific resource scheduling operations according to the scheduling decision of the greedy algorithm unit, such as starting new nodes, shutting down faulty nodes, and migrating services.

[0032] Preferably, the service migration module includes:

[0033] Service status monitoring unit: Monitor the service status running on the faulty node to ensure that the service is in a stable state before migration;

[0034] Hot migration technology unit: Smoothly migrate the services on the faulty node to a healthy node using hot migration technology;

[0035] Migration verification unit: After migration, verify the migrated services to ensure that the services run normally on the new node and meet the performance requirements.

[0036] Preferably, the traffic scheduling module includes:

[0037] Traffic monitoring unit: Real-time monitor the flow direction and distribution of network traffic to ensure that user requests can be correctly forwarded to service nodes;

[0038] Routing update unit: Automatically update the network traffic routing information according to the service migration situation;

[0039] Load balancing unit: Reasonably allocate network traffic among multiple healthy nodes to achieve load balancing and improve the overall performance and reliability of the system.

[0040] Preferably, the redundant deployment module includes:

[0041] Component identification unit: Identify the key components and services in the system, and the failures of these components and services may have a serious impact on the entire system;

[0042] Redundant Deployment Policy Unit: Formulate redundant deployment policies to determine which components and services need to be redundantly deployed, as well as the methods and degrees of redundant deployment;

[0043] Automated Deployment Unit: Automatically deploy redundant components and services according to the redundant deployment policy.

[0044] Preferably, the cross-data center collaboration module includes:

[0045] Cross-data center Communication Unit: Establish a cross-data center communication mechanism to achieve information exchange and collaborative work between different data centers;

[0046] Fault Information Sharing Unit: Share fault information between data centers, including fault types, locations, and impact scopes;

[0047] Service Migration Collaboration Unit: Coordinate resources and service migration operations between different data centers when cross-data center service migration is required.

[0048] Preferably, the operation and maintenance management module includes:

[0049] Operation and Maintenance Interface Unit: Provide a unified operation and maintenance management interface to facilitate operation and maintenance personnel to view system status, manage resources, and configure policies;

[0050] Fault Record Unit: Record fault information in the system, including the time of fault occurrence, handling process, and handling results;

[0051] Analysis and Report Unit: Statistically analyze the fault information and generate a fault analysis report;

[0052] Policy Optimization Unit: Optimize fault detection, assessment, and recovery policies based on the fault analysis report and operation and maintenance experience to improve the overall operation and maintenance efficiency and reliability of the system.

[0053] The above technical solutions of the present invention have the following beneficial technical effects:

[0054] 1. Through the implementation of the redundant deployment module, the backup of key components and services is ensured. When the primary node fails, the backup node can immediately take over the work, avoiding the interruption of the overall service caused by a single point of failure. The redundant mechanism makes the fault recovery process more automated and rapid, reducing the need for manual intervention;

[0055] 2. The fault detection module can automatically analyze system logs and performance metrics to quickly identify potential faults, which is more accurate and efficient than manual or traditional automated means. And the fault assessment module can quickly evaluate the impact scope and severity of the fault, providing a basis for subsequent fault recovery strategies, shortening the time from fault detection to response;

[0056] 3. The resource scheduling module can dynamically adjust resource allocation according to the real-time load situation, ensuring that resources are fully utilized, avoiding uneven resource allocation or waste, and through dynamic resource adjustment, the system can better cope with sudden traffic or load peaks and maintain a high-performance operating state;

[0057] 4. By establishing a cross-data center communication mechanism and a service migration coordination mechanism through the cross-data center coordination module, when a single data center fails, services can be quickly migrated to other data centers to ensure service continuity and availability. And the service migration module adopts hot migration technology, reducing service interruption time and improving the user experience. Brief Description of the Drawings

[0058] Figure 1 It is a module block diagram of a single data center failure recovery laaS cluster processing system proposed by the present invention. Detailed Embodiments

[0059] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the detailed embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0060] As Figure 1 shown, a single data center failure recovery laaS cluster processing system proposed by the present invention includes:

[0061] Fault detection module: responsible for real-time monitoring of the health status of each node in the cluster, analyzing system logs and performance metrics, and quickly identifying potential faults;

[0062] Fault assessment module: receives the output of the fault detection module, evaluates the scope and severity of the fault impact, and provides a basis for the fault recovery strategy;

[0063] Resource scheduling module: according to the fault assessment results, dynamically adjusts resource allocation and selects the best healthy node for service migration;

[0064] Service migration module: responsible for smoothly migrating the services on the faulty node to the healthy node, and adopting hot migration technology to reduce service interruption time;

[0065] Traffic scheduling module: according to the service migration situation, automatically adjusts the network traffic routing to ensure that user requests can be correctly forwarded to the new service node;

[0066] Redundant deployment module: responsible for redundant deployment of key components and services to ensure that the system can still provide complete services in case of a single point of failure;

[0067] Cross-data center collaboration module: Establish a cross-data center communication mechanism to achieve the sharing of fault information between data centers and the collaboration of service migration, improving the disaster tolerance of the overall system;

[0068] Operation and maintenance management module: Provide a unified operation and maintenance management interface, support fault recording, analysis, reporting, and policy optimization, and improve operation and maintenance efficiency;

[0069] The fault detection module is connected to the fault assessment module, the fault assessment module is connected to the resource scheduling module, the resource scheduling module is connected to the service migration module, the service migration module is connected to the traffic scheduling module, the cross-data center collaboration module is respectively connected to the resource scheduling module and the service migration module, the redundant deployment module is connected to all other modules, and the operation and maintenance management module is connected to all other modules.

[0070] It should be noted that: The fault detection module is connected to the fault assessment module: The fault detection module sends the detected fault information, including fault type, time, and location, to the fault assessment module through the API interface or message queue. After receiving this information, the fault assessment module conducts further analysis and evaluation; The fault assessment module is connected to the resource scheduling module: The fault assessment module shares the assessment results, including the fault impact range and severity, with the resource scheduling module through the internal API interface or database. The resource scheduling module formulates resource scheduling policies based on this information; The resource scheduling module is connected to the service migration module: After determining the target node for service migration, the resource scheduling module sends a migration instruction to the service migration module through the API interface, and the service migration module executes the specific migration operation; The service migration module is connected to the traffic scheduling module: During or after the migration process, the service migration module notifies the traffic scheduling module to update the network traffic routing information through the API interface to ensure that user requests can be correctly forwarded to the new service node; The redundant deployment module is connected to all other modules: The redundant deployment module maintains an indirect connection with all other modules through system configuration and automated deployment tools. It is responsible for dynamically adding redundant components and services during the system design phase or runtime to enhance the fault tolerance of the system. This connection is more based on the system architecture and configuration level rather than direct API calls; The cross-data center collaboration module is connected to all modules involved in service migration: The cross-data center collaboration module maintains connections with the resource scheduling module and the service migration module through cross-data center network protocols such as VPN and MPLS, and specific communication interfaces such as RESTful API. When cross-data center service migration is required, the cross-data center collaboration module is responsible for coordinating resource and service migration operations between different data centers; The operation and maintenance management module is connected to all modules: As the control center of the system, the operation and maintenance management module interacts with all other modules through a unified operation and maintenance management interface and internal API interfaces. It collects the operation data, fault information, and performance metrics of each module and provides functions such as fault recording, analysis, reporting, and policy optimization.

[0071] In specific embodiments, the fault detection module includes:

[0072] A log collection unit: responsible for collecting system logs from each node in the cluster, including application logs, system logs, and security logs;

[0073] A performance monitoring unit: real-time monitors the performance metrics of each node in the cluster, and these metrics can reflect the health status and performance bottlenecks of the nodes;

[0074] A machine learning analysis unit: uses an SVM classifier to analyze logs and performance metrics, and identifies abnormal patterns and potential faults;

[0075] An alarm notification unit: when potential faults are detected, sends alarm messages to the operation and maintenance personnel via email, text message, and system notification, so as to respond and process in a timely manner.

[0076] It should be noted that the fault detection algorithm uses an SVM classifier, and its decision function is:

[0077] \(f(x)=\text{sign}\left(\sum_{i = 1}^{n}\alpha_i y_i K(x,x_i)+b\right)\), where \(x\) is the input feature vector, \(y_i\) is the sample label, \(\alpha_i\) is the support vector coefficient, \(K(x,x_i)\) is the kernel function, and \(b\) is the bias term;

[0078] By using the SVM classifier to perform real-time analysis on system logs and performance metrics, abnormal patterns and potential faults can be identified, which is more intelligent and accurate than traditional rule-based or threshold-based methods;

[0079] Implementation process:

[0080] The log collection unit collects system logs;

[0081] The performance monitoring unit monitors performance metrics;

[0082] The machine learning analysis unit uses the SVM algorithm to train the collected data, establishes a classification model, and classifies new data in real-time to determine whether there are abnormalities;

[0083] Calculation process: SVM finds the optimal classification hyperplane by maximizing the classification margin, and divides the data set into normal and abnormal categories;

[0084] Example: When the CPU usage rate of a certain node suddenly soars, the performance monitoring unit captures this change, and the machine learning analysis unit uses the SVM model to analyze this change. If it is judged as abnormal, an alarm will be triggered.

[0085] In a specific embodiment, the fault assessment module includes:

[0086] A fault information receiving unit: Receives fault information from the fault detection module, including fault type, location, and time;

[0087] A risk assessment unit: Analyzes the fault information using a Bayesian network to evaluate the scope of influence, severity, and potential losses of the fault;

[0088] A strategy recommendation unit: Provides recommendations for subsequent fault recovery strategies based on the risk assessment results, including which faults to prioritize and what recovery measures to take.

[0089] It should be noted that: By performing probabilistic inference on the fault information using a Bayesian network, it is possible to evaluate the scope of influence, severity, and potential losses of the fault, providing a quantitative basis for the recovery strategy;

[0090] For example: When a certain server crashes, the risk assessment unit calculates through the Bayesian network that the probability of this fault causing service interruption is 80%, and the expected loss is XX yuan per minute.

[0091] In a specific embodiment, the resource scheduling module includes:

[0092] A resource monitoring unit: Monitors the resource usage of each node in the cluster in real time, including CPU, memory, storage, and network resources;

[0093] A greedy algorithm unit: Dynamically schedules resources using the greedy algorithm to select the best healthy node for service migration;

[0094] A scheduling execution unit: Executes specific resource scheduling operations according to the scheduling decision of the greedy algorithm unit, such as starting a new node, shutting down a faulty node, and migrating services.

[0095] It should be noted that the resource scheduling algorithm uses the greedy algorithm to select the node with the largest remaining capacity for service migration. The formula is:

[0096] \text{Node}_{\text{best}}=\arg\max_j\left(\text{Capacity}_j-\text{Load}_j\right)\] where \(\text{Capacity}_j\) is the total capacity of node \(j\), and \(\text{Load}_j\) is the current load of node \(j\);

[0097] The greedy algorithm can quickly select the current optimal healthy node for service migration in resource scheduling;

[0098] Implementation process:

[0099] The resource monitoring unit monitors the cluster resources in real time;

[0100] The greedy algorithm unit selects the most suitable healthy node according to the resource usage and service migration requirements;

[0101] Calculation process: The greedy algorithm selects the local optimal solution each time and gradually constructs the global solution, such as selecting the node with the most remaining resources for migration;

[0102] Example: When a service needs to be migrated from a faulty node, the greedy algorithm unit calculates the remaining resources of each healthy node and selects the node with the richest resources for migration.

[0103] In a specific embodiment, the service migration module includes:

[0104] The service status monitoring unit: monitors the service status running on the faulty node to ensure that the service is in a stable state before migration;

[0105] The live migration technology unit: uses the live migration technology to smoothly migrate the service on the faulty node to a healthy node;

[0106] The migration verification unit: after the migration is completed, verifies the migrated service to ensure that the service runs normally on the new node and meets the performance requirements.

[0107] It should be noted that: The live migration technology can smoothly migrate the service to a new node without interrupting the service, reducing the service interruption time and user perception;

[0108] Implementation process:

[0109] The service status monitoring unit monitors the service status;

[0110] The live migration technology unit performs the migration operation after ensuring the service stability;

[0111] Calculation process: The live migration technology involves synchronization and optimization in multiple aspects such as memory and network. The specific calculation process depends on the specific implementation method and system architecture;

[0112] Example: When a certain application service needs to be migrated from a faulty node, the live migration technology unit smoothly migrates the service to another healthy node with almost no perception.

[0113] In a specific embodiment, the traffic scheduling module includes:

[0114] The traffic monitoring unit: monitors the flow direction and distribution of network traffic in real time to ensure that user requests can be correctly forwarded to the service node;

[0115] The routing update unit: automatically updates the network traffic routing information according to the service migration situation;

[0116] Load balancing unit: Reasonably allocate network traffic among multiple healthy nodes to achieve load balancing, improve the overall performance and reliability of the system.

[0117] It should be noted that: Automatically adjust the network traffic routing according to the service migration situation to achieve load balancing, improving the overall performance and reliability of the system;

[0118] Implementation process:

[0119] The traffic monitoring unit monitors network traffic;

[0120] The routing update unit updates the routing information according to the service migration situation;

[0121] The load balancing unit distributes traffic among multiple healthy nodes

[0122] For example: When a certain service node has too high a load, the load balancing unit automatically forwards some traffic to other low-load nodes.

[0123] In a specific embodiment, the redundant deployment module includes:

[0124] Component identification unit: Identify key components and services in the system, and the failures of these components and services may have a serious impact on the entire system;

[0125] Redundant deployment strategy unit: Develop a redundant deployment strategy to determine which components and services need to be redundantly deployed, as well as the method and degree of redundant deployment;

[0126] Automated deployment unit: Automatically deploy redundant components and services according to the redundant deployment strategy.

[0127] It should be noted that: Through the automated deployment strategy, the redundant deployment of key components and services is achieved, improving the disaster tolerance ability of the system;

[0128] For example: For database services, the redundant deployment module will automatically deploy one or more backup database instances to ensure that the service can be quickly taken over in case of a primary database failure.

[0129] In a specific embodiment, the cross-data center collaboration module includes:

[0130] Cross-data center communication unit: Establish a cross-data center communication mechanism to achieve information exchange and collaborative work between different data centers;

[0131] Fault information sharing unit: Share fault information among data centers, including fault type, location, and impact scope;

[0132] Service migration coordination unit: When cross-data center service migration is required, it coordinates resources and service migration operations between different data centers.

[0133] In a specific embodiment, the operation and maintenance management module includes:

[0134] Operation and maintenance interface unit: Provides a unified operation and maintenance management interface, facilitating operation and maintenance personnel to view the system status, manage resources, and configure policies;

[0135] Fault recording unit: Records fault information in the system, including the fault occurrence time, handling process, and handling results;

[0136] Analysis and reporting unit: Statistically analyzes the fault information and generates a fault analysis report;

[0137] Policy optimization unit: Optimizes fault detection, evaluation, and recovery policies based on the fault analysis report and operation and maintenance experience, improving the overall operation and maintenance efficiency and reliability of the system.

[0138] The working principle of this system: First, the fault detection module monitors the health status of cluster nodes in real time. By combining system logs and performance metric analysis, potential faults are quickly identified. Subsequently, the fault evaluation module receives the output from the detection module, evaluates the scope and severity of the fault, and provides a decision basis for subsequent fault recovery policies. The resource scheduling module then dynamically adjusts resource allocation based on the evaluation results, preferentially selects healthy nodes for service migration. The service migration module uses hot migration technology to smoothly migrate services on the faulty node to a healthy node, minimizing service interruption time. The traffic scheduling module automatically adjusts network traffic routing according to the actual situation of service migration to ensure that user requests can be accurately forwarded to the new service node. In addition, the redundant deployment module is responsible for the redundant deployment of key components and services to cope with single-point failures and ensure the continuity of system services. The cross-data center coordination module realizes the sharing of fault information between data centers and the coordination of service migration by establishing a cross-data center communication mechanism, further enhancing the overall disaster tolerance ability of the system. Finally, the operation and maintenance management module provides a unified operation and maintenance management interface, supporting the recording, analysis, reporting, and policy optimization of faults, thus significantly improving operation and maintenance efficiency and system stability.

[0139] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principles of the present invention, and do not constitute a limitation to the present invention. Therefore, any modifications, equivalent replacements, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modification examples falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

Claims

1. A single computer room fault recovery laas cluster processing system, characterized in that, It includes: Fault detection module: Responsible for real-time monitoring of the health status of each node in the cluster, analyzing system logs and performance metrics, and quickly identifying potential faults; Fault assessment module: Receives the output of the fault detection module, evaluates the scope and severity of the fault impact, and provides a basis for the fault recovery strategy; Resource scheduling module: Dynamically adjusts resource allocation according to the fault assessment results, and selects the best healthy node for service migration; Service migration module: Responsible for smoothly migrating the services on the faulty node to the healthy node, and adopting hot migration technology to reduce service interruption time; Traffic scheduling module: Automatically adjusts the network traffic routing according to the service migration situation to ensure that user requests can be correctly forwarded to the new service node; Redundant deployment module: Responsible for the redundant deployment of key components and services to ensure that the system can still provide complete services in case of a single-point failure; Cross-data center collaboration module: Establishes a cross-data center communication mechanism to achieve fault information sharing and service migration collaboration between data centers, and improves the overall disaster tolerance of the system; Operation and maintenance management module: Provides a unified operation and maintenance management interface, supports fault recording, analysis, reporting, and strategy optimization, and improves operation and maintenance efficiency; The fault detection module is connected to the fault assessment module, the fault assessment module is connected to the resource scheduling module, the resource scheduling module is connected to the service migration module, the service migration module is connected to the traffic scheduling module, the cross-data center collaboration module is respectively connected to the resource scheduling module and the service migration module, the redundant deployment module is connected to all other modules, and the operation and maintenance management module is connected to all other modules.

2. The single-computer room fault recovery laas cluster processing system according to claim 1, wherein The fault detection module includes: Log collection unit: Responsible for collecting system logs from each node in the cluster, including application logs, system logs, and security logs; Performance monitoring unit: Real-time monitors the performance metrics of each node in the cluster, and these metrics can reflect the health status and performance bottlenecks of the nodes; Machine learning analysis unit: Uses an SVM classifier to analyze the logs and performance metrics to identify abnormal patterns and potential faults; Alarm notification unit: When a potential fault is detected, sends alarm information to the operation and maintenance personnel via email, SMS, and system notifications for timely response and handling.

3. A single-computer room fault recovery laaS cluster processing system according to claim 2, wherein The fault assessment module includes: Fault information receiving unit: Receives fault information from the fault detection module, including fault type, location, and time; Risk assessment unit: Analyzes the fault information using a Bayesian network to evaluate the impact scope, severity, and possible losses of the fault; Strategy recommendation unit: Provides recommendations for subsequent fault recovery strategies according to the risk assessment results, including which faults to prioritize and what recovery measures to take.

4. A single-computer room fault recovery laaS cluster processing system according to claim 3, wherein, The resource scheduling module includes: Resource monitoring unit: Real-time monitors the resource usage of each node in the cluster, including CPU, memory, storage, and network resources; Greedy algorithm unit: Dynamically schedules resources using a greedy algorithm to select the best healthy node for service migration; Scheduling execution unit: Executes specific resource scheduling operations according to the scheduling decision of the greedy algorithm unit, such as starting a new node, shutting down a faulty node, and migrating services.

5. The single-computer room fault recovery laaS cluster processing system according to claim 4, characterized in that, The service migration module includes: Service Status Monitoring Unit: Monitor the service status running on the faulty node to ensure that the service is in a stable state before migration; Hot Migration Technology Unit: Use hot migration technology to smoothly migrate the services on the faulty node to a healthy node; Migration Verification Unit: After the migration is completed, verify the migrated services to ensure that the services run properly on the new node and meet the performance requirements.

6. A single-computer room fault recovery laaS cluster processing system according to claim 5, characterized in that, The traffic scheduling module includes: Traffic Monitoring Unit: Real-time monitor the flow direction and distribution of network traffic to ensure that user requests can be correctly forwarded to the service nodes; Routing Update Unit: Automatically update the network traffic routing information according to the service migration situation; Load Balancing Unit: Reasonably allocate network traffic among multiple healthy nodes to achieve load balancing and improve the overall performance and reliability of the system.

7. A single-computer room fault recovery laas cluster processing system according to claim 6, characterized in that, The redundant deployment module includes: Component Identification Unit: Identify the key components and services in the system, and the failures of these components and services may have a serious impact on the entire system; Redundant Deployment Strategy Unit: Develop a redundant deployment strategy to determine which components and services need to be redundantly deployed, as well as the methods and degrees of redundant deployment; Automated Deployment Unit: Automatically deploy redundant components and services according to the redundant deployment strategy.

8. A single-computer room fault recovery laaS cluster processing system according to claim 7, wherein, The cross-data center collaboration module includes: Cross-data center Communication Unit: Establish a cross-data center communication mechanism to achieve information exchange and collaborative work between different data centers; Fault Information Sharing Unit: Share fault information between data centers, including fault type, location, and impact scope; Service Migration Collaboration Unit: Coordinate the resources and service migration operations between different data centers when cross-data center service migration is required.

9. A single-computer room fault recovery laas cluster processing system according to claim 8, characterized in that, The operation and maintenance management module includes: Operation and Maintenance Interface Unit: Provide a unified operation and maintenance management interface to facilitate operation and maintenance personnel to view the system status, manage resources, and configure policies; Fault Record Unit: Record the fault information in the system, including the fault occurrence time, processing process, and processing result; Analysis Report Unit: Statistically analyze the fault information and generate a fault analysis report; Strategy Optimization Unit: Optimize the fault detection, evaluation, and recovery strategies according to the fault analysis report and operation and maintenance experience to improve the overall operation and maintenance efficiency and reliability of the system.

Citation Information

Cited By

  • Intelligent cloud desktop resource scheduling method and system fused with fault recovery

    CN121029388A

  • Intelligent cloud desktop resource scheduling method and system with fusion of fault recovery

    CN121029388B