Virtualized environment automatic fault recovery method and system based on embedded Hypervisor

Through real-time monitoring and automated analysis, the embedded hypervisor system can identify and recover from faults, solving the downtime problem caused by manual intervention in traditional methods and improving the stability and reliability of the virtualized environment.

CN120743589APending Publication Date: 2025-10-03ZHONGLING ZHIXING (CHENGDU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510558956.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional embedded hypervisor fault recovery methods require manual intervention, which is time-consuming and error-prone, resulting in extended system downtime, affecting business continuity and user experience.

Method used

Collect system performance indicator data in real time through monitoring protocols, identify abnormal data using preset alarm rules, analyze the cause of failures through log analysis and event correlation, and automatically execute recovery strategies, including cross-node migration and snapshot rollback, and record the processing process to reduce manual intervention.

Benefits of technology

It achieves automated fault detection and recovery, reduces system downtime, improves the stability and reliability of the virtualization environment, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743589A_ABST
    Figure CN120743589A_ABST
Patent Text Reader

Abstract

The invention discloses an embedded Hypervisor-based virtualization environment automatic fault recovery method and system, and the method comprises the steps: collecting system performance index data in real time through a monitoring protocol, carrying out the abnormal data recognition of the system performance index data through a preset alarm rule, and obtaining system abnormal index data, analyzing the system abnormal index data to obtain alarm information; analyzing the alarm information based on log analysis, performance data analysis and event association to obtain a fault reason; executing a predefined recovery strategy according to the fault reason; monitoring the recovered system performance index data by using a health examination tool Zabbix, calling and verifying the service state of the virtual machine through an API, if verification fails, carrying out cross-node migration, migrating the faulted virtual machine to a healthy node, and rolling back to a recent stable state based on a snapshot mechanism; according to the invention, faults can be automatically detected, and fault positioning and recovery can be automatically carried out when the faults are detected, so that the downtime of the system is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an automated fault recovery method and system for a virtualized environment based on an embedded Hypervisor. Background Art

[0002] With the widespread adoption of cloud computing and virtualization technologies, embedded hypervisors are playing an increasingly important role in managing multiple operating system instances, such as Android or Linux. While these embedded hypervisors improve resource utilization and flexibility, they also introduce new challenges, particularly in fault management and recovery. Traditional fault recovery methods often require manual intervention, which is time-consuming and error-prone, leading to extended system downtime and impacting business continuity and user experience. Therefore, developing a system that can automatically detect, locate, and recover from faults is crucial to improving the stability and reliability of virtualized environments. Summary of the Invention

[0003] The object of the present invention is to provide a method and system for automatic fault recovery in a virtualized environment based on an embedded Hypervisor to solve the above problems.

[0004] A first aspect of the present invention provides a method for automated fault recovery in a virtualized environment based on an embedded hypervisor, comprising:

[0005] Collecting system performance indicator data in real time through a monitoring protocol, identifying abnormal data on the system performance indicator data using preset alarm rules to obtain system abnormal indicator data, and analyzing the system abnormal indicator data to obtain alarm information;

[0006] Analyze alarm information based on log analysis, performance data analysis, and event correlation to determine the cause of the fault;

[0007] Execute a predefined recovery strategy based on the cause of the failure;

[0008] Use the health check tool Zabbix to monitor system performance indicators after recovery. Verify the virtual machine service status through API calls. If verification fails, perform cross-node migration to migrate the faulty virtual machine to a healthy node and roll back to the most recent stable state based on the snapshot mechanism.

[0009] Use Bbox to record the troubleshooting process.

[0010] Furthermore, the monitoring protocol collects system performance indicator data in real time, including:

[0011] Collect system performance indicator data through SNMP monitoring protocol and preset object identifiers;

[0012] The system performance indicator data includes but is not limited to: CPU utilization, memory occupancy, disk I / O rate, network throughput, and log information.

[0013] Furthermore, the system performance indicator data is identified as abnormal data using a preset alarm rule to obtain system abnormal indicator data, and the system abnormal indicator data is analyzed to obtain alarm information, including:

[0014] The alarm rules include:

[0015] Static threshold rule: compares system indicator data with custom thresholds and triggers an alarm when the system performance indicator data exceeds the custom threshold;

[0016] Dynamic anomaly detection algorithm rules: Analyze system indicator data through machine learning models and trigger alarms when system indicator data deviates from the baseline;

[0017] Alarm information includes: system abnormality indicator data, affected components, and current system status.

[0018] Furthermore, based on log analysis, performance data analysis, and event correlation, the alarm information is analyzed to obtain the cause of the fault, including:

[0019] Use ELK Stack to analyze system metrics logs and identify abnormal software.

[0020] Use performance monitoring tools to analyze system abnormality indicator data, affected components, and the current system status to identify resource bottleneck components;

[0021] The event correlation engine is used to build a fault propagation graph, and the rule engine and dynamic reasoning mechanism are combined to analyze the system abnormality indicator data to identify whether the virtual machine is running abnormally.

[0022] Furthermore, a predefined recovery strategy is executed according to the cause of the failure, including:

[0023] When the software fails, restart the virtual machine;

[0024] When a component encounters a performance bottleneck, Docker is used to isolate the faulty component and reallocate resources.

[0025] When a virtual machine runs abnormally, the virtual machine can be migrated and its status rolled back by using the API of the virtualization management platform.

[0026] A second aspect of the present invention provides an automated fault recovery system for a virtualized environment based on an embedded hypervisor, comprising:

[0027] Fault detection module: used to collect system performance index data in real time through monitoring protocol, identify abnormal data of the system performance index data according to preset alarm rules to obtain system abnormal index data, and analyze the system abnormal index data to obtain alarm information;

[0028] Fault location module: used to analyze alarm information based on log analysis, performance data analysis and event correlation to determine the cause of the fault;

[0029] Fault recovery module: used to execute a predefined recovery strategy according to the fault cause;

[0030] Fault Verification Module: This module uses the health check tool Zabbix to monitor system performance indicators after recovery and verifies the virtual machine service status through API calls. If verification fails, it performs cross-node migration, migrating the faulty virtual machine to a healthy node and rolling back to the most recent stable state based on the snapshot mechanism.

[0031] Management module: used to record the fault handling process through Bbox.

[0032] The present invention has at least the following beneficial effects:

[0033] The present invention provides an automated fault recovery method for a virtualized environment based on an embedded hypervisor. The method monitors the embedded hypervisor and virtual machines in real time through a monitoring protocol and collects system performance indicator data. The method then uses preset alarm rules to identify abnormal data in the system performance indicator data to detect abnormal behavior or performance degradation in the system, and obtains system abnormal indicator data. The abnormal indicator data is analyzed to generate alarm information. Log analysis, performance data analysis, and system status information are then used to locate the specific location and cause of the fault. A predefined recovery strategy is automatically executed based on the specific location and cause of the fault. After the system recovers, the health check tool Zabbix is ​​used to monitor the recovered system performance indicator data. The virtual machine service status is verified through API calls. If the verification fails, a cross-node migration is performed, migrating the faulty virtual machine to a healthy node and rolling it back to the most recent stable state based on a snapshot mechanism. Finally, the entire fault handling process is recorded through Bbox and the results are fed back to the administrator for subsequent analysis and optimization. From real-time system performance indicator data collection, anomaly detection, fault location, policy execution, and recovery verification, all steps in the method are automatically triggered by preset rules, eliminating the need for manual intervention and reducing human intervention delays.

[0034] The present invention also provides an automated fault recovery system in a virtualized environment based on an embedded hypervisor. This system can automatically detect faults and automatically locate and recover them when a fault is detected, thereby reducing system downtime. The system includes a fault detection module, a fault location module, a fault recovery module, a fault verification module, and a management module. Through continuous monitoring, rapid location, and automatic recovery, the stability and reliability of the virtualized environment are improved. The system supports mainstream hypervisors such as KVM, VMware, and Xen, and can be seamlessly integrated into an enterprise's existing virtualization architecture. Each functional module in the system can be independently deployed and expanded, supporting on-demand customization of fault handling strategies, thereby reducing operation and maintenance costs.

[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flow chart of an automated fault recovery method for a virtualized environment based on an embedded hypervisor provided by the present invention;

[0037] Figure 2 This is a schematic diagram of an automated fault recovery system for a virtualized environment based on an embedded Hypervisor provided by the present invention. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0039] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.

[0040] Example 1: Combination Figure 1 To illustrate this embodiment,

[0041] This embodiment provides an automated fault recovery method for a virtualized environment based on an embedded hypervisor, including:

[0042] Collecting system performance indicator data in real time through a monitoring protocol, identifying abnormal data on the system performance indicator data using preset alarm rules to obtain system abnormal indicator data, and analyzing the system abnormal indicator data to obtain alarm information;

[0043] Analyze alarm information based on log analysis, performance data analysis, and event correlation to determine the cause of the fault;

[0044] Execute a predefined recovery strategy based on the cause of the failure;

[0045] Use the health check tool Zabbix to monitor system performance indicators after recovery. Verify the virtual machine service status through API calls. If verification fails, perform cross-node migration to migrate the faulty virtual machine to a healthy node and roll back to the most recent stable state based on the snapshot mechanism.

[0046] Use Bbox to record the troubleshooting process.

[0047] A monitoring protocol monitors the embedded hypervisor and virtual machines in real time and collects system performance metrics. Pre-set alarm rules identify anomalies in this data to detect abnormal behavior or performance degradation in the system. This data is then analyzed to generate alarm information. Log analysis, performance data analysis, and system status information are then used to locate the specific location and cause of the fault. Pre-defined recovery strategies are then automatically executed based on the specific location and cause of the fault. After the system recovers, the health check tool Zabbix monitors the recovered system performance metrics. API calls are used to verify the virtual machine service status. If verification fails, a cross-node migration is performed, migrating the faulty virtual machine to a healthy node and rolling it back to the most recent stable state based on a snapshot mechanism. Finally, Bbox records the entire fault handling process and provides feedback to the administrator for subsequent analysis and optimization. This method automatically triggers all steps, from real-time system performance metric data collection, anomaly detection, fault location, policy execution, and recovery verification, based on pre-set rules. This eliminates the need for manual intervention and reduces delays.

[0048] Furthermore, the monitoring protocol collects system performance indicator data in real time, including:

[0049] Collect system performance indicator data through SNMP monitoring protocol and preset object identifiers;

[0050] The system performance indicator data includes but is not limited to: CPU utilization, memory occupancy, disk I / O rate, network throughput, and log information.

[0051] This method uses the underlying monitoring capabilities of the embedded hypervisor to achieve real-time collection of system performance indicator data and second-level alarms, significantly shortening the fault discovery and processing time window, thereby reducing the risk of business interruption.

[0052] Furthermore, the system performance indicator data is identified as abnormal data using a preset alarm rule to obtain system abnormal indicator data, and the system abnormal indicator data is analyzed to obtain alarm information, including:

[0053] The alarm rules include:

[0054] Static threshold rule: compares system indicator data with custom thresholds and triggers an alarm when the system performance indicator data exceeds the custom threshold;

[0055] Static threshold rules support user-defined thresholds for key metrics (e.g., CPU utilization ≥ 80%) to meet basic monitoring needs. Set reasonable thresholds for each monitored system metric, such as triggering alarms when CPU utilization reaches 100%, memory utilization exceeds 95% for more than 5 minutes, or disk I / O latency exceeds 50ms.

[0056] Dynamic anomaly detection algorithm rules: Analyze system indicator data through machine learning models and trigger alarms when system indicator data deviates from the baseline;

[0057] Time series analysis algorithms such as ARIMA (Autoregressive Integrated Moving Average) are used to predict the trends of system performance indicators and compare them with actual data to identify anomalies.

[0058] Alarm information includes: system abnormality indicator data, affected components, and current system status.

[0059] Furthermore, based on log analysis, performance data analysis, and event correlation, the alarm information is analyzed to obtain the cause of the fault, including:

[0060] Use ELK Stack to analyze system metrics logs and identify abnormal software.

[0061] Use ELK Stack (Elasticsearch, Logstash, Kibana) for log analysis, collect and analyze system logs to find error messages and abnormal behavior, and use log management tools such as Splunk to collect and analyze application logs and identify abnormal software.

[0062] Use performance monitoring tools to analyze system abnormality indicator data, affected components, and the current system status to identify resource bottleneck components;

[0063] Use performance monitoring tools (such as Prometheus) to analyze system abnormal indicator data, affected components, and the current status of the system to locate resource bottlenecks or overloaded components.

[0064] The event correlation engine is used to build a fault propagation graph, and the rule engine and dynamic reasoning mechanism are combined to analyze the system abnormality indicator data to identify whether the virtual machine is running abnormally.

[0065] The event correlation engine is used to correlate system abnormality indicator data and identify complex event chains that may lead to system failures. Event correlation rules are used to correlate and analyze system abnormality indicator data from different sources to identify whether the virtual machine is running abnormally.

[0066] Furthermore, a predefined recovery strategy is executed according to the cause of the failure, including:

[0067] When the software fails, restart the virtual machine;

[0068] When a component encounters a performance bottleneck, the Docker containerization technology is used to isolate the faulty component and reallocate resources. For performance bottleneck components, the system automatically triggers the Docker containerization isolation and resource reallocation strategies (such as dynamically adjusting vCPU / memory quotas), thereby improving resource utilization and system resilience.

[0069] When a virtual machine runs abnormally, the virtual machine can be migrated and its status rolled back by using the API of the virtualization management platform.

[0070] Example 2: Combination Figure 2 To illustrate this embodiment,

[0071] This embodiment is an automated fault recovery system for a virtualized environment based on an embedded hypervisor, including:

[0072] Fault detection module 10: used to collect system performance index data in real time through a monitoring protocol, identify abnormal data on the system performance index data according to preset alarm rules to obtain system abnormal index data, and analyze the system abnormal index data to obtain alarm information;

[0073] Fault location module 20: used to analyze alarm information based on log analysis, performance data analysis and event correlation to obtain the cause of the fault;

[0074] Fault recovery module 30: configured to execute a predefined recovery strategy according to the fault cause;

[0075] Fault Verification Module 40: Uses the health check tool Zabbix to monitor system performance indicators after recovery and verifies the virtual machine service status through API calls. If verification fails, cross-node migration is performed to migrate the faulty virtual machine to a healthy node and roll back to the most recent stable state based on the snapshot mechanism.

[0076] Management module 50: used to record the fault handling process through Bbox.

[0077] The present invention also provides an automated fault recovery system in a virtualized environment based on an embedded hypervisor. This system can automatically detect faults and automatically locate and recover them when a fault is detected, thereby reducing system downtime. The system includes a fault detection module, a fault location module, a fault recovery module, a fault verification module, and a management module. Through continuous monitoring, rapid location, and automatic recovery, the stability and reliability of the virtualized environment are improved. The system supports mainstream hypervisors such as KVM, VMware, and Xen, and can be seamlessly integrated into an enterprise's existing virtualization architecture. Each functional module in the system can be independently deployed and expanded, supporting on-demand customization of fault handling strategies, thereby reducing operation and maintenance costs.

[0078] It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, product, or apparatus. Without further limitation, the phrase "comprises a..." does not preclude the presence of additional identical elements in the process, method, product, or apparatus comprising the elements. Terms such as "first," "second," and the like are used to designate names and do not imply any particular order. The above illustrative description of the present invention and its embodiments is non-limiting. The present invention may be embodied in other specific forms without departing from the spirit or essential features of the present invention. The drawings illustrate only one embodiment of the present invention; the actual structure is not limited thereto. Any reference numerals in the claims should not limit the claims to which they relate. Therefore, if a person skilled in the art is inspired by this and, without departing from the spirit of the present invention, devises structures and embodiments similar to the present invention without inventiveness, they shall fall within the scope of protection of this patent.

Claims

1. A virtualization environment automatic fault recovery method based on embedded hypervisor, characterized in that: include: Collecting system performance indicator data in real time through a monitoring protocol, identifying abnormal data on the system performance indicator data using preset alarm rules to obtain system abnormal indicator data, and analyzing the system abnormal indicator data to obtain alarm information; Analyze alarm information based on log analysis, performance data analysis, and event correlation to determine the cause of the fault; Execute a predefined recovery strategy based on the cause of the failure; Use the health check tool Zabbix to monitor system performance indicators after recovery. Verify the virtual machine service status through API calls. If verification fails, perform cross-node migration to migrate the faulty virtual machine to a healthy node and roll back to the most recent stable state based on the snapshot mechanism. Use Bbox to record the troubleshooting process.

2. The method for automated fault recovery in a virtualized environment based on an embedded Hypervisor according to claim 1, wherein: Collect system performance indicator data in real time through monitoring protocols, including: Collect system performance indicator data through SNMP monitoring protocol and preset object identifiers; The system performance indicator data includes but is not limited to: CPU utilization, memory occupancy, disk I / O rate, network throughput, and log information.

3. The method for automated fault recovery in a virtualized environment based on an embedded hypervisor according to claim 2, wherein: The system performance indicator data is identified as abnormal data using a preset alarm rule to obtain system abnormal indicator data, and the system abnormal indicator data is analyzed to obtain alarm information, including: The alarm rules include: Static threshold rule: compares system indicator data with custom thresholds and triggers an alarm when the system performance indicator data exceeds the custom threshold; Dynamic anomaly detection algorithm rules: Analyze system indicator data through machine learning models and trigger alarms when system indicator data deviates from the baseline; Alarm information includes: system abnormality indicator data, affected components, and current system status.

4. The method for automated fault recovery in a virtualized environment based on an embedded hypervisor according to claim 3, wherein: Analyze alarm information based on log analysis, performance data analysis, and event correlation to determine the cause of the fault, including: Use ELK Stack to analyze system metrics logs and identify abnormal software. Use performance monitoring tools to analyze system abnormality indicator data, affected components, and the current system status to identify resource bottleneck components; The event correlation engine is used to build a fault propagation graph, and the rule engine and dynamic reasoning mechanism are combined to analyze the system abnormality indicator data to identify whether the virtual machine is running abnormally.

5. The method for automated fault recovery in a virtualized environment based on an embedded Hypervisor according to claim 4, wherein: Execute predefined recovery strategies based on the cause of the failure, including: When the software fails, restart the virtual machine; When a component encounters a performance bottleneck, Docker is used to isolate the faulty component and reallocate resources. When a virtual machine runs abnormally, the virtual machine can be migrated and its status rolled back by using the API of the virtualization management platform.

6. An automated fault recovery system for a virtualized environment based on an embedded hypervisor, characterized in that: include: Fault detection module: used to collect system performance index data in real time through monitoring protocol, identify abnormal data of the system performance index data according to preset alarm rules to obtain system abnormal index data, and analyze the system abnormal index data to obtain alarm information; Fault location module: used to analyze alarm information based on log analysis, performance data analysis and event correlation to determine the cause of the fault; Fault recovery module: used to execute a predefined recovery strategy according to the fault cause; Fault Verification Module: This module uses the health check tool Zabbix to monitor system performance indicators after recovery and verifies the virtual machine service status through API calls. If verification fails, it performs cross-node migration, migrating the faulty virtual machine to a healthy node and rolling back to the most recent stable state based on the snapshot mechanism. Management module: used to record the fault handling process through Bbox.