Virtual machine emergency processing method and device during host machine downtime, medium and equipment

By automating the acquisition of virtual machine inventory and drift status assessment, manual operations are replaced, achieving standardized and efficient emergency handling of virtual machines. This solves the problem of insufficient integration of existing tools into the bank's operation and maintenance system and meets the high requirements of the financial industry.

CN121858201APending Publication Date: 2026-04-14CHINA CITIC BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA CITIC BANK CO LTD
Filing Date
2025-11-27
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing virtualization management and automation tools cannot achieve closed-loop management of the entire process from information acquisition and automated detection to batch application recovery in the bank's operation and maintenance system, making it difficult to meet the high requirements of the financial industry for business continuity and stability.

Method used

By automating the acquisition of a list of virtual machines and key information of virtual machines on a crashed host, determining the drift status of each virtual machine, restarting virtual machines that have not drifted, and filtering virtual machines whose applications have not started, standardized recovery operations are performed, replacing manual list compilation and hierarchical notification, thus achieving automated emergency handling.

Benefits of technology

It significantly reduces the time required for information collection, improves the accuracy and efficiency of drift status judgment, ensures the standardization and reliability of emergency operations, improves the efficiency of batch application recovery, reduces the risk of manual operation, and meets the high requirements of the financial industry for business continuity and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858201A_ABST
    Figure CN121858201A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual machine emergency processing method and device during host machine downtime, a medium and equipment, and relates to the technical field of computers.The method comprises the steps that a virtual machine list and virtual machine key information corresponding to a downtime host machine are obtained, and the virtual machine list comprises multiple virtual machines; based on the virtual machine key information, whether each virtual machine drifts or not is judged one by one; aiming at the first virtual machine which does not drift, restarting the first virtual machine after checking the fault of the downtime host machine; and in the drifting second virtual machine and / or the restarted first virtual machine, screening a target virtual machine of which the application is not started, and executing an application recovery operation on the target virtual machine according to an application recovery processing method matched with the target virtual machine. According to the method and the device, automatic emergency processing of the virtual machine can be realized when the host machine crashes, and high requirements of the financial industry on business continuity and stability are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, medium and device for emergency handling of virtual machines when the host machine crashes. Background Technology

[0002] In the fintech field, virtualization technology has been widely adopted in the IT architectures of banks and other financial institutions due to its advantages such as improved resource utilization and reduced operational costs. With the distributed expansion of banking operations, the deployment model of virtual machines across multiple branches and systems has become mainstream. The host machine, as the core carrier of virtual machines, directly affects the continuity of various banking services. However, during host machine operation, unexpected situations such as hardware failures and software anomalies may cause downtime, leading to a large number of virtual machines deployed on it experiencing batch anomalies. Rapid emergency response is required to minimize business interruption losses; therefore, virtual machine emergency response technology has become a crucial component of the bank's virtualization operation and maintenance system.

[0003] Currently, the industry's emergency response to virtual machine crashes primarily relies on operations and maintenance personnel to manually complete the entire process. Specifically, after a host machine crashes, operations and maintenance personnel need to manually compile a list of virtual machines on that host machine, confirm the running status of each virtual machine and its corresponding application, and notify relevant branches and application managers level by level via telephone, instant messaging, or email before coordinating subsequent recovery operations. Although some institutions have introduced virtualization management tools or open-source automation tools, existing virtualization management tools and automation tools lack deep integration with the bank's operations and maintenance system. They cannot cover the entire closed-loop management process from information acquisition and automation detection to batch application recovery, making it difficult to meet the financial industry's high requirements for business continuity and stability. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, medium and device for emergency handling of virtual machines when the host machine crashes, which can realize automated emergency handling of virtual machines when the host machine crashes, meeting the high requirements of the financial industry for business continuity and stability.

[0005] According to a first aspect of this application, an emergency handling method for virtual machines when the host machine crashes is provided, including: Obtain the list of virtual machines and key information of the virtual machines corresponding to the crashed host machine. The list of virtual machines contains multiple virtual machines, and the key information of the virtual machines includes at least the virtual machine name, purpose attribute, operating system type and the name of the organizational unit to which it belongs. Based on the key information of the virtual machines, determine one by one whether each virtual machine has drifted; For the first virtual machine that did not drift, after troubleshooting the faulty host machine, the first virtual machine is restarted; In the second virtual machine that experienced the drift and / or the first virtual machine after the restart, target virtual machines with applications that have not been started are selected, and application recovery operations are performed on the target virtual machines according to the application recovery handling method matched to the target virtual machines.

[0006] According to a second aspect of this application, an emergency handling device for virtual machines when the host machine crashes is provided, comprising: The acquisition module is used to acquire a list of virtual machines and key information of the virtual machines corresponding to the crashed host machine. The list of virtual machines contains multiple virtual machines, and the key information of the virtual machines includes at least the virtual machine name, purpose attribute, operating system type and the name of the organizational unit to which it belongs. The judgment module is used to determine whether each virtual machine has drifted based on the key information of the virtual machine; The restart module is used to restart the first virtual machine that has not experienced drift after troubleshooting the faulty host machine. The execution module is used to filter out target virtual machines with applications not started in the second virtual machine that has drifted and / or the first virtual machine after restarting, and perform application recovery operations on the target virtual machines according to the application recovery handling method matched to the target virtual machines.

[0007] According to a third aspect of this application, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the above-described virtual machine emergency handling method when the host machine crashes.

[0008] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described virtual machine emergency handling method when the host machine crashes.

[0009] By employing the above technical solutions, the virtual machine emergency handling method, apparatus, medium, and equipment provided in this application for host machine crashes automatically acquire a list of virtual machines corresponding to the crashed host machine, along with key virtual machine information including virtual machine name, purpose attributes, and operating system type. This replaces the traditional manual list compilation by maintenance personnel, effectively solving the problem of low efficiency in manual list compilation and significantly shortening the time spent on initial information collection. By determining whether each virtual machine has drifted based on its key information, the tedious process of manually confirming the status of each virtual machine is avoided, reducing omissions and errors that are prone to occur in manual judgment and improving the accuracy and efficiency of drift status judgment. For the first virtual machine that has not drifted, it is restarted after troubleshooting the crashed host machine. This standardized fault diagnosis and restart process replaces traditional manual handling. This automated operation addresses the inconsistency in handling methods across different branches, ensuring the standardization of emergency operations. In the second virtual machine that experiences drift and / or the first virtual machine after a restart, it automatically detects the target virtual machine where the application is not running and performs recovery according to the matching application recovery method. This replaces the inefficient manual method of checking application status one by one and notifying responsible personnel at each level. It improves the efficiency of batch application recovery, avoids delays in cross-branch communication, and ensures the reliability of application recovery through standardized recovery operations. Ultimately, it achieves efficient and standardized management from information acquisition and status assessment to scenario-based handling, effectively compensating for the shortcomings of existing tools and insufficient integration with the bank's operation and maintenance system. This meets the high requirements of the financial industry for business continuity and stability, while reducing the risks of manual operation and improving the controllability of emergency response.

[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an emergency handling method for virtual machines when the host machine crashes, provided in an embodiment of this application, is shown. Figure 2 A flowchart illustrating an emergency handling method for virtual machines when the host machine crashes, according to another embodiment of this application, is shown. Figure 3 This illustration shows a structural schematic diagram of a virtual machine emergency handling device provided in an embodiment of this application when the host machine crashes. Detailed Implementation

[0012] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0013] Currently, the industry's emergency response to virtual machine crashes primarily relies on operations and maintenance personnel to manually complete the entire process. Specifically, after a host machine crashes, operations and maintenance personnel need to manually compile a list of virtual machines on that host machine, confirm the running status of each virtual machine and its corresponding application, and notify relevant branches and application managers level by level via telephone, instant messaging, or email before coordinating subsequent recovery operations. Although some institutions have introduced virtualization management tools or open-source automation tools, existing virtualization management tools and automation tools lack deep integration with the bank's operations and maintenance system. They cannot cover the entire closed-loop management process from information acquisition and automation detection to batch application recovery, making it difficult to meet the financial industry's high requirements for business continuity and stability.

[0014] To address the aforementioned technical problems, embodiments of the present invention provide an emergency handling method for virtual machines when the host machine crashes, such as... Figure 1 As shown, the method includes: Step 110: Obtain the list of virtual machines and key information of the virtual machines corresponding to the crashed host machine. The list of virtual machines contains multiple virtual machines. The key information of the virtual machines includes at least the virtual machine name, purpose attribute, operating system type and the name of the organization unit to which it belongs.

[0015] The virtual machine list is a structured list that records the basic information of multiple virtual machines deployed on a failed host machine; the virtual machine name is the name information used to uniquely identify each virtual machine, ensuring accurate location of the target virtual machine in subsequent query, judgment and other operations; the purpose attribute is an attribute identifier that distinguishes the business purpose of the virtual machine, which can usually be divided into production type (carrying core business applications) and office type (carrying daily office tools); the operating system type refers to the type of operating system installed on the virtual machine (such as Windows, Linux); the organizational unit name is the name information that clearly indicates the management department or organization to which the virtual machine belongs (such as the operation and maintenance department of a branch or the technology department of a region).

[0016] In this embodiment of the disclosure, a pre-configured automated script can be used to connect to the configuration management database in the operation and maintenance system. After detecting a host machine crash event, the script automatically reads the IP address of the crashed host machine and sends a related query command to the database. After the database responds, it returns the original data of all virtual machines on the host machine. The script then filters and organizes the original data to generate a structured list containing multiple virtual machines according to preset rules, and simultaneously extracts the core information of each virtual machine.

[0017] This technology replaces manual list compilation and information extraction with automated scripts, which can significantly shorten the time spent on initial information collection after the host machine crashes and avoid the problems of list omissions and information errors that are prone to occur in manual operations. Furthermore, by standardizing the extraction of key information from virtual machines, it provides a unified data standard for subsequent emergency response in different scenarios, effectively solving the problems of information dispersion and inefficient process connection in the traditional manual mode.

[0018] Step 120: Based on the key information of the virtual machine, determine whether each virtual machine has drifted.

[0019] In this embodiment of the disclosure, the content used to uniquely identify each virtual machine in the key information of the virtual machine can be used as the retrieval basis. The system connects to the virtualization management platform and inputs this unique identifier to obtain the core information of the host machine where each virtual machine is currently running. Simultaneously, the core information of the failed host machine is retrieved. Through automated comparison logic, the core information of the host machine where each virtual machine is currently running is verified one by one with the core information of the failed host machine. This determines whether each virtual machine has migrated from its original failed host machine to another healthy host machine, forming a drift status determination result for each virtual machine. By relying on the key information of the virtual machine to achieve automated one-by-one determination, the tedious operation of manually logging into the virtualization management platform to query and compare the host machine information of each virtual machine can be avoided, significantly improving the efficiency of drift status determination.

[0020] In specific application scenarios, if it is determined that all virtual machines in the virtual machine list are first virtual machines that have not drifted, the operation in step 130 of the embodiment can be performed for each first virtual machine. After restarting the first virtual machine, the subsequent step 140 of the embodiment can be performed. If it is determined that the virtual machine list contains both first virtual machines that have not drifted and second virtual machines that have drifted, the operation in step 130 of the embodiment can be performed for each first virtual machine, and the operation in step 140 of the embodiment can be performed for each second virtual machine. If it is determined that all virtual machines in the virtual machine list are second virtual machines that have drifted, step 130 of the embodiment can be skipped, and the operation in step 140 of the embodiment can be performed for each second virtual machine.

[0021] Step 130: For the first virtual machine that did not drift, after troubleshooting the faulty host machine, restart the first virtual machine.

[0022] In this embodiment of the disclosure, the first virtual machine that has been determined not to have drifted can be locked first. The hardware monitoring interface and system log of the downed host machine can be connected through an automated fault diagnosis tool to detect and locate the fault type of the host machine. After the fault is repaired and the host machine restores its normal operation capability, a batch restart command is sent to all the first virtual machines that have not drifted through the virtualization management platform to ensure that these virtual machines are restarted on the repaired original host machine and restored to the basic operating state, laying the foundation for subsequent application status detection and business recovery.

[0023] By first troubleshooting and repairing the faulty host machine before restarting the non-migrated virtual machines, the logic of blindly restarting virtual machines when the host machine fault is not resolved can avoid startup failures or repeated failures caused by the first virtual machine restart, thus ensuring the success rate of the first virtual machine restart. At the same time, the targeted restart operation for non-migrated virtual machines does not require waiting for the virtual machine migration process, which can reduce unnecessary time consumption and quickly restore these virtual machines to basic operation, buying time for subsequent application layer recovery. This can effectively shorten the overall emergency response cycle, reduce the risk of business interruption caused by the long-term inability of virtual machines to run, and improve the reliability and efficiency of emergency response.

[0024] Step 140: In the second virtual machine that has drifted and / or the first virtual machine after restarting, select the target virtual machine with applications that have not started, and perform application recovery operation on the target virtual machine according to the application recovery handling method matched to the target virtual machine.

[0025] The first virtual machine refers to a virtual machine that did not migrate after the host machine crashed, and was restarted after the original crashed host machine was troubleshooted and repaired; its operating environment remains the repaired original host machine. The second virtual machine refers to a virtual machine that was migrated from the original crashed host machine to another healthy host machine to continue running after the host machine crashed; its operating environment has been separated from the original failed host machine. The target virtual machine refers to the virtual machine with application startup issues selected from the second virtual machine that experienced migration and / or the first virtual machine after restarting; it is the specific target of subsequent application recovery operations. The application recovery handling method refers to a preset repair plan for different application startup scenarios, which is usually stored in association with information such as the virtual machine IP and associated application name, and is used to guide the application recovery operation of the target virtual machine.

[0026] In this embodiment of the disclosure, for the second virtual machine that has drifted and / or the first virtual machine after restarting, an application status detection operation can be automatically performed (such as checking whether the process of the associated application exists and whether the service port is connected). The virtual machine whose application has not started is selected from the two types of virtual machines as the target virtual machine. Then, the script calls the preset application operation and maintenance ledger and matches the corresponding application recovery and disposal method (such as specific service restart commands and configuration repair steps) according to the IP of the target virtual machine, the name of the associated application, etc. Finally, according to the matching result, the script automatically executes the recovery operation or pushes standardized recovery instructions to the responsible personnel to complete the application recovery of the target virtual machine.

[0027] In summary, the virtual machine emergency handling method provided by this invention, when the host machine crashes, automates the acquisition of a list of virtual machines corresponding to the crashed host machine, along with key virtual machine information including virtual machine name, purpose attributes, and operating system type. This replaces the traditional manual list compilation by maintenance personnel, effectively solving the problem of low efficiency in manual list compilation and significantly shortening the time spent on initial information collection. By determining whether each virtual machine has drifted based on its key information, the tedious process of manually confirming the status of each virtual machine is avoided, reducing omissions and errors that are prone to occur in manual judgment and improving the accuracy and efficiency of drift status judgment. For the first virtual machine that has not drifted, it is restarted after troubleshooting the crashed host machine. This standardized fault diagnosis and restart process replaces the arbitrary operations of traditional manual handling. It can resolve inconsistencies in handling methods across different branches, ensuring the standardization of emergency operations. In the second virtual machine that has drifted and / or the first virtual machine after a restart, it automatically detects the target virtual machine where the application has not started and performs recovery according to the matching application recovery handling method. This replaces the inefficient mode of manually checking the application status one by one and notifying the responsible persons at each level. It can not only improve the efficiency of batch application recovery and avoid delays in cross-branch communication, but also ensure the reliability of application recovery through standardized recovery operations. Ultimately, it can achieve efficient and standardized management from information acquisition and status judgment to scenario-based handling, effectively making up for the deficiencies in the integration of existing tools with the bank's operation and maintenance system, meeting the high requirements of the financial industry for business continuity and stability, while reducing the risks of manual operation and improving the controllability of emergency response.

[0028] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, and to fully illustrate the implementation methods of this embodiment, this embodiment also provides another emergency handling method for virtual machines when the host machine crashes, such as... Figure 2 As shown, the method includes: Step 210: Obtain the list of virtual machines and key information of the virtual machines corresponding to the crashed host machine.

[0029] For embodiments of this disclosure, step 210 may include the following steps: Step 210-1: Enter the first IP address of the crashed host into the configuration management database and trigger the virtual machine data query command. The virtual machine data query command is used to indicate the query for virtual machine data associated with the first IP address.

[0030] In this embodiment of the disclosure, a pre-configured emergency response automation script can be executed to input the first IP address corresponding to the confirmed downtime host machine into the configuration management database, and simultaneously trigger a preset virtual machine data query instruction. The query instruction uses the first IP address as the core search condition, and explicitly instructs the configuration management database to filter and return all virtual machine-related data that have an affiliation with the first IP address. This data can cover key information such as virtual machine name, virtual machine IP address, operating system type, and organizational unit to which it belongs.

[0031] Step 210-2: Receive the virtual machine data corresponding to the crashed host machine sent by the configuration management database response.

[0032] In this embodiment of the disclosure, after inputting the first IP address of the crashed host into the configuration management database and triggering a virtual machine data query command, the configuration management database will retrieve the internally stored IT resource association data based on the first IP address in the command, filter out all virtual machine records that have an affiliation with the crashed host, perform field normalization (such as standardizing the field format of virtual machine name, operating system type, etc.) and integrity verification on these records (ensuring that no key information is missing), and then actively push the normalized virtual machine data corresponding to the crashed host to the emergency response system through a preset standardized data transmission interface.

[0033] Step 210-3: Using a Python script, generate a virtual machine list containing at least one virtual machine based on the virtual machine data according to preset filtering rules, and synchronously obtain key information of each virtual machine in the virtual machine list.

[0034] In this embodiment of the disclosure, after receiving the virtual machine data corresponding to the downed host machine returned by the configuration management database, the pre-written Python script automatically loads preset filtering rules (such as retaining production and office virtual machines, excluding test environment virtual machines, and extracting core fields such as virtual machine name, IP address, and purpose attribute). The raw virtual machine data is cleaned (invalid fields are removed, and missing key information is filled in) and filtered through the data parsing module, and a virtual machine list containing at least one virtual machine is generated according to a preset format (such as a structured format of virtual machine name-IP address-purpose attribute). At the same time, during the process of generating the list, the script will synchronously extract key information (including virtual machine name, purpose attribute, operating system type, and organizational unit name, etc.) from the raw data of each virtual machine, and associate and bind this key information with the corresponding virtual machine in the list, and store it in the structured database of the emergency response system.

[0035] Step 220: Based on the key information of the virtual machine, determine whether each virtual machine has drifted.

[0036] For embodiments of this disclosure, step 220 may include the following steps: Step 220-1: Based on the virtual machine name in the virtual machine's key information, input the unique identifier of each virtual machine into the virtualization management platform and query the second IP address of its current host machine.

[0037] In this embodiment of the disclosure, the name of each virtual machine can be extracted from the acquired key information of the virtual machine, and the virtual machine name can be mapped to its corresponding unique identifier in the virtualization management platform. Through a pre-configured automated interactive script, the standard interface of the virtualization management platform is connected, and the unique identifier of each virtual machine is passed to the virtualization management platform one by one, triggering a query request for the current running location of each virtual machine. After receiving the request, the virtualization management platform retrieves the virtual machine-host machine association data stored internally, locates and returns the network address of the host machine where each virtual machine is currently running, i.e., the second IP address.

[0038] Step 220-2: Compare the second IP address obtained from the query with the first IP address of the crashed host machine.

[0039] In this embodiment of the disclosure, after obtaining the second IP address of the host machine where each virtual machine is currently located through the virtualization management platform, the automated script retrieves the first IP address of the crashed host machine from the temporary data module of the emergency response system. According to the one-to-one correspondence, the second IP address and the first IP address of each virtual machine are passed to the preset IP comparison algorithm module. The algorithm module automatically compares the two IP addresses through string matching logic. Specifically, it can compare the network segment and host portion of the two. If the network segment and host portion of the two are the same, it can be determined that the second IP address and the first IP address are consistent. If the network segment and host portion of the two are different, it can be determined that the second IP address and the first IP address are inconsistent.

[0040] Step 220-3: If the second IP address is the same as the first IP address, then it is determined that the virtual machine has not drifted.

[0041] Step 220-4: If the second IP address and the first IP address are inconsistent, it is determined that the virtual machine has completed the migration.

[0042] Step 230: For the first virtual machine that did not drift, after troubleshooting the faulty host machine, restart the first virtual machine.

[0043] In this embodiment of the disclosure, virtual machines that have not experienced drift can be selected from the set of virtual machines that have completed the drift determination as the first virtual machine. By using a preset fault diagnosis tool (such as a hardware monitoring script or a system log analysis program) to connect to the management interface of the crashed host machine, the hardware status (such as CPU, memory, and hard disk operation), system processes, and network configuration of the host machine can be automatically detected to locate the cause of the fault that caused the host machine to crash. After the fault is repaired and the host machine is confirmed to have restored its normal operating capability, a restart command is sent to all the first virtual machines that have not experienced drift through the batch operation interface of the virtualization management platform, so that all the first virtual machines will perform a restart action.

[0044] Step 240: For the second virtual machine that has drifted and / or the first virtual machine after restarting, execute a one-click check script on the corresponding host through the Entegor platform to perform detection operations on the virtual machine restart status and the corresponding application startup status.

[0045] The Entegor platform is a digital operation and maintenance automation execution platform. It has the functions of remotely connecting to the host, calling and executing scripts, and collecting and storing test results. It is used to support the automated detection and emergency response process in the operation and maintenance process. The one-click check script is a preset automated detection script (such as a shell script) that can be executed on the target host to obtain the virtual machine startup time, query the application process status, and check the service port connectivity. It is the core tool for completing the virtual machine and application status detection. The virtual machine restart status detection operation includes: for the second virtual machine that has drifted and / or the first virtual machine after restarting, query the startup time of each second virtual machine, compare the startup time with the host machine downtime. If the startup time is after the downtime, it is determined that it has been restarted; otherwise, it is determined that it has not been restarted. The application startup status detection operation includes: according to the service name of the associated application in the virtual machine's key information (i.e., the service identifier of the application running on the virtual machine), query the process status of the associated application to determine whether the target application process exists; synchronously check the connectivity of the associated service port of the associated application. If the target application process exists and the associated service port is connected, it is determined that the application has started; otherwise, it is determined that it has not started. The target application process refers to the application running process corresponding to the application service name associated with the virtual machine's key information; the associated service port refers to the network port used by the associated application to provide services to the outside world.

[0046] Step 250: Identify the virtual machines whose corresponding virtual machine restart status is "not restarted" or whose corresponding virtual machine restart status is "restarted" and whose application startup status is "not started" as the target virtual machines.

[0047] In this embodiment of the disclosure, for the second virtual machine that has drifted and / or the first virtual machine that has been restarted, virtual machine status data generated by the Python emergency detection script in the database of the emergency response system can be retrieved respectively. Then, the automated filtering script loads the preset judgment logic to perform condition verification on the status data of each virtual machine: if the restart status of a virtual machine is marked as not restarted, it is directly determined that it meets the target virtual machine filtering conditions; if the restart status is marked as restarted, the application startup status data of the virtual machine is further read. If the application startup status is not started, it is also determined that it meets the filtering conditions; the script marks all the second virtual machines that have drifted and / or the first virtual machines that have been restarted that meet any of the above conditions as target virtual machines.

[0048] This technology replaces manual verification of virtual machine status with automated logical screening, which can significantly improve the identification efficiency of target virtual machines, especially in multi-virtual machine scenarios, and can significantly shorten the screening time. The dual-condition coverage of not restarted and restarted but application not started can avoid the omission of scenarios that may occur during manual screening, and ensure that all virtual machines with abnormal operation are included in the scope of handling.

[0049] Step 260: Perform application recovery operations on the target virtual machine according to the application recovery handling method matched to the target virtual machine.

[0050] In one possible implementation of this disclosure, step 260 may include the following steps: based on the organizational unit name in the preset maintenance log and key information of the virtual machine, locate the responsible personnel corresponding to the target virtual machine and create a cross-level emergency connection. The preset maintenance log pre-stores the association relationship and communication method between each organizational unit and the responsible personnel. Send a standardized handling instruction to the responsible personnel through a preset communication interface. The standardized handling instruction is used to instruct the responsible personnel to perform application recovery operations on the target virtual machine based on the cross-level emergency connection. The standardized handling instruction includes the target virtual machine IP, the associated application name, and the application recovery handling method.

[0051] The pre-set maintenance ledger refers to a pre-established and stored maintenance management database. Its core content is the correspondence between each organizational unit (such as departments or branches) and responsible personnel, while also recording the contact information of responsible personnel (such as mobile phone numbers, corporate IM accounts, email addresses, etc.) for quickly locating the relevant responsible persons for fault handling. The name of the affiliated organizational unit refers to the name of the management organization to which the virtual machine belongs, recorded in the key information of the virtual machine. It is used to associate the responsible personnel in the pre-set maintenance ledger and clarify the responsible subject for fault handling. Cross-level emergency connection refers to a real-time communication channel (such as online meetings, dedicated communication groups, etc.) created to coordinate the emergency handling of responsible personnel at different levels and departments. It can break down organizational hierarchy barriers and ensure that responsible personnel can quickly synchronize information and coordinate operations. Standardized handling instructions refer to standardized instructions containing key information required for emergency recovery. They usually cover the target virtual machine IP (locating the object), the associated application name (clarifying the recovery object), and the application recovery handling method (guiding the operation), which is used to unify the recovery operation standards and avoid the arbitrariness of manual operation.

[0052] This technology automates the matching of pre-set maintenance logs with organizational unit names, replacing the inefficient manual process of searching for responsible persons at each level, significantly reducing the time required to locate them. The automatic creation of cross-level emergency connections avoids delays caused by manual coordination in establishing communication groups, ensuring rapid action by responsible personnel. Standardized handling instructions clearly include key information needed for recovery, avoiding errors or ambiguities in manual information transmission and ensuring the accuracy and standardization of recovery operations. The entire process, through automated association, connection, and instruction push, reduces manual intervention and mitigates the risk of recovery delays due to poor communication or incomplete information. Simultaneously, it standardizes emergency response procedures, ensuring consistency across different organizational units and responsible persons, effectively improving application recovery efficiency and shortening business interruption time.

[0053] As another possible implementation, step 260 of the embodiment may also include the following steps: automatically classifying the target virtual machine according to preset classification rules based on the key information of the virtual machine and marking the disposal priority; determining the application recovery script adapted to the target virtual machine based on the classification results and disposal priority; and automatically performing application recovery operations on the target virtual machine through the application recovery script.

[0054] The preset classification rules refer to predefined criteria for classifying target virtual machines, typically based on usage attributes (production / office) and operating system type (Windows / Linux). These rules automate the grouping of target virtual machines, providing a basis for matching application recovery scripts. The handling priority refers to the recovery order level set according to the maximum tolerable downtime of the target virtual machine, which is directly related to the importance of the services hosted by the virtual machine. Higher service importance corresponds to a shorter maximum tolerable downtime and a higher handling priority; conversely, lower service importance corresponds to a longer maximum tolerable downtime and a lower handling priority. High-priority virtual machines will receive script resources and execution permissions first, ensuring rapid recovery of critical services with short tolerable downtime and minimizing losses from service interruptions. Application recovery scripts are automated programs (such as Shell scripts and PowerShell scripts) written for specific operating systems and application types. They contain instructions for restarting application processes, repairing configuration parameters, and verifying service status, and can be directly executed on the target virtual machine to achieve application recovery.

[0055] Accordingly, when automatically classifying target virtual machines and marking their disposal priorities based on key virtual machine information and preset classification rules, the implementation steps may include: grouping at least one target virtual machine based on the usage attribute in the key virtual machine information, classifying each target virtual machine into a production virtual machine group or an office virtual machine group; further subdividing all target virtual machines in the production virtual machine group and the office virtual machine group into Windows system virtual machines and Linux system virtual machines, respectively, according to the operating system type in the key virtual machine information, forming four virtual machine groups; and marking the corresponding disposal priorities for the target virtual machines in the four virtual machine groups according to preset priority setting rules.

[0056] The disposal priority is used to guide the allocation of subsequent testing resources and the order of recovery operations, ensuring that core business virtual machines are disposed of first. For example, production-type Windows system virtual machines and production-type Linux system virtual machines can be set to the highest priority, and office-type Windows system virtual machines and office-type Linux system virtual machines can be set to the second highest priority, according to actual application needs.

[0057] In specific application scenarios, as a preferred approach, Python scripts can automatically connect to various data storage nodes throughout the entire emergency response process, accurately extracting key data from the system. This data may include inventory retrieval time, detection results, recovery operation records, and verification results. The inventory retrieval time originates from the timestamp record when the script connects to the CMDB to obtain the virtual machine inventory. Detection results cover virtual machine drift judgment results, restart status detection conclusions, and application startup status judgment data. Recovery operation records contain specific information on two implementation methods (such as communication records of responsible personnel during handling, standardized instruction sending time, automatically executed application recovery script name, and execution duration). Verification results integrate technical and business verification. Furthermore, the Python script can structure and organize this key data according to a preset standardized template, automatically generating an emergency response report document containing charts. Subsequently, through the preset API interface of the operation and maintenance management system, the report document is uploaded to a designated archiving directory. The system synchronously records the document number, archiving time, and associated fault event ID, while updating the emergency response process status to closed loop. This completes the closed-loop management of the entire process from fault response to data aggregation, document generation, and archiving, providing comprehensive data support for subsequent fault review and audit tracing.

[0058] In summary, the technical solution in this application first obtains virtual machine data by inputting the IP address of the failed host machine into the configuration management database, and then uses Python scripts to generate a list and key information. This automation replaces manual processing, significantly improving the efficiency of initial information collection. By querying the current host machine IP address based on the virtual machine name and comparing it with the IP address of the failed host machine, accurate judgment of the drift status is achieved, reducing human error and omissions. For non-drifted virtual machines, troubleshooting and restarting are performed; for drifted virtual machines and restarted virtual machines, restart and application status are checked, comprehensively covering two core scenarios to ensure no abnormal scenarios are missed. Standardized instructions containing key information are pushed to the responsible person, or priority is marked according to purpose and operating system and matched with recovery scripts for automatic execution, achieving standardization and flexibility in application recovery. The entire process is automated through Python scripts, significantly shortening emergency response time, reducing the risk of manual operation, and ensuring business continuity and stability.

[0059] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides an emergency handling device for virtual machines when the host machine crashes, such as... Figure 3 As shown, the device includes: an acquisition module 31, a judgment module 32, a restart module 33, and an execution module 34.

[0060] The acquisition module 31 can be used to acquire the list of virtual machines and key information of the virtual machines corresponding to the crashed host. The list of virtual machines contains multiple virtual machines, and the key information of the virtual machines includes at least the virtual machine name, purpose attribute, operating system type and the name of the organizational unit to which it belongs. The judgment module 32 can be used to determine whether each virtual machine has drifted based on key information of the virtual machine; Restart module 33 can be used to restart the first virtual machine that has not drifted after troubleshooting the faulty host machine; The execution module 34 can be used to filter out target virtual machines with applications not started in the second virtual machine that has drifted and / or the first virtual machine after restarting, and perform application recovery operations on the target virtual machine according to the application recovery handling method matched to the target virtual machine.

[0061] In some embodiments of this application, the acquisition module 31 can be specifically used to input the first IP address of the crashed host into the configuration management database and trigger a virtual machine data query instruction. The virtual machine data query instruction is used to indicate the query for virtual machine data associated with the first IP address; receive the virtual machine data corresponding to the crashed host sent by the configuration management database; use a Python script to generate a virtual machine list containing at least one virtual machine based on the virtual machine data according to preset filtering rules; and synchronously acquire the key information of each virtual machine in the virtual machine list.

[0062] In some embodiments of this application, the judgment module 32 can be specifically used to input the unique identifier of each virtual machine into the virtualization management platform based on the virtual machine name in the virtual machine key information, and query the second IP address of its current host machine; compare the queried second IP address with the first IP address of the crashed host machine; if the second IP address and the first IP address are consistent, it is determined that the virtual machine has not drifted; if the second IP address and the first IP address are inconsistent, it is determined that the virtual machine has completed the drift.

[0063] In some embodiments of this application, the execution module 34 can be specifically used to execute a one-click check script on the corresponding host through the Entegor platform to perform detection operations on the virtual machine restart status and the corresponding application startup status for the second virtual machine that has drifted and / or the first virtual machine after restarting; the virtual machines whose corresponding virtual machine restart status is not restarted, or whose corresponding virtual machine restart status is restarted and the application startup status is not started, are identified as target virtual machines; wherein, the virtual machine restart status detection operation includes: for the second virtual machine that has drifted and / or the first virtual machine after restarting, querying the startup time of each virtual machine, comparing the startup time with the host machine downtime, if the startup time is after the downtime, it is determined that it has been restarted, otherwise it is determined that it has not been restarted; the application startup status detection operation includes: according to the service name of the virtual machine's corresponding associated application, querying the process status of the associated application to determine whether the target application process exists; synchronously detecting the connectivity of the associated service port of the associated application, if the target application process exists and the associated service port is connected, it is determined that the application has started, otherwise it is determined that it has not started.

[0064] In some embodiments of this application, the execution module 34 can also be used to locate the responsible personnel corresponding to the target virtual machine based on the organizational unit name in the preset operation and maintenance ledger and the key information of the virtual machine, and to create a cross-level emergency connection. The preset operation and maintenance ledger stores the association relationship and communication method between each organizational unit and the responsible personnel. The standardized handling instruction is sent to the responsible personnel through a preset communication interface. The standardized handling instruction is used to instruct the responsible personnel to perform application recovery operation on the target virtual machine based on the cross-level emergency connection. The standardized handling instruction includes the target virtual machine IP, the associated application name, and the application recovery handling method.

[0065] In some embodiments of this application, the execution module 34 can also be used to automatically classify the target virtual machine according to preset classification rules based on the key information of the virtual machine and mark the disposal priority; determine the application recovery script adapted to the target virtual machine based on the classification result and disposal priority; and automatically perform application recovery operation on the target virtual machine through the application recovery script.

[0066] In some embodiments of this application, the execution module 34 can also be used to group at least one target virtual machine based on the purpose attribute in the virtual machine key information, classifying each target virtual machine into a production virtual machine group or an office virtual machine group; according to the operating system type in the virtual machine key information, all target virtual machines in the production virtual machine group and the office virtual machine group are further subdivided into Windows system virtual machines and Linux system virtual machines, forming four types of virtual machine groups; and the target virtual machines in the four types of virtual machine groups are marked with corresponding disposal priorities according to preset priority setting rules.

[0067] It should be noted that other corresponding descriptions of the functional units involved in the virtual machine emergency handling device when the host machine crashes provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.

[0068] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The above describes an emergency handling method for virtual machines when the host machine crashes.

[0069] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0070] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 3 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The above describes an emergency handling method for virtual machines when the host machine crashes.

[0071] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0072] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0073] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0074] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.

[0075] This invention automates the acquisition of a list of virtual machines (VMs) corresponding to a failed host machine, along with key VM information including VM name, purpose, and operating system type. This replaces the traditional manual list-building process by maintenance personnel, effectively solving the problem of low efficiency in manual list building and significantly reducing the time spent on initial information collection. By determining whether each VM has drifted based on its key information, the invention avoids the tedious process of manually checking the status of each VM, reducing omissions and errors that are prone to occur in manual judgment, and improving the accuracy and efficiency of drift status assessment. For the first VM that has not drifted, it is restarted after troubleshooting the failed host machine. This standardized fault diagnosis and restart process replaces the arbitrary operations of traditional manual handling, enabling solutions for different branch operations. To address inconsistencies in procedures, this system ensures standardized emergency operations. In the second virtual machine that experienced drift and / or the first virtual machine after a restart, it automatically detects target virtual machines where applications are not running and performs recovery according to the matching application recovery procedures. This replaces the inefficient manual process of checking application status on each machine and notifying responsible parties at each level. This improves the efficiency of batch application recovery, avoids delays in cross-branch communication, and ensures the reliability of application recovery through standardized operations. Ultimately, it achieves efficient and standardized management from information acquisition and status assessment to scenario-based handling, effectively compensating for the shortcomings of existing tools and insufficient integration with the bank's operations and maintenance system. This meets the high requirements of the financial industry for business continuity and stability, while reducing the risks of manual operations and improving the controllability of emergency response.

[0076] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.

[0077] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A method for emergency handling of virtual machines when the host machine crashes, characterized in that, include: Obtain the list of virtual machines and key information of the virtual machines corresponding to the crashed host machine. The list of virtual machines contains multiple virtual machines, and the key information of the virtual machines includes at least the virtual machine name, purpose attribute, operating system type and the name of the organizational unit to which it belongs. Based on the key information of the virtual machines, determine one by one whether each virtual machine has drifted; For the first virtual machine that did not drift, after troubleshooting the faulty host machine, the first virtual machine is restarted; In the second virtual machine that experienced the drift and / or the first virtual machine after the restart, target virtual machines with applications that have not been started are selected, and application recovery operations are performed on the target virtual machines according to the application recovery handling method matched to the target virtual machines.

2. The method according to claim 1, characterized in that, The process of obtaining the list of virtual machines and key information about the virtual machines corresponding to the downed host includes: Input the first IP address of the crashed host into the configuration management database and trigger a virtual machine data query instruction, which is used to indicate the query for virtual machine data that matches the first IP address; Receive virtual machine data corresponding to the crashed host machine sent in response from the configuration management database; Using a Python script, a virtual machine list containing at least one virtual machine is generated based on the virtual machine data according to preset filtering rules, and key information of each virtual machine in the virtual machine list is obtained synchronously.

3. The method according to claim 1, characterized in that, Based on the aforementioned key information about the virtual machines, determine one by one whether each virtual machine has drifted, including: Based on the virtual machine name in the key information of the virtual machine, input the unique identifier of each virtual machine into the virtualization management platform and query the second IP address of its current host machine; The second IP address obtained from the query is compared with the first IP address of the crashed host machine; If the second IP address is the same as the first IP address, then it is determined that the virtual machine has not drifted. If the second IP address and the first IP address are inconsistent, it is determined that the virtual machine has completed the migration.

4. The method according to claim 1, characterized in that, In the second virtual machine where the drift occurred and / or the first virtual machine after a reboot, filter for target virtual machines where applications have not started, including: For the second virtual machine that has drifted and / or the first virtual machine after a restart, execute a one-click check script on the corresponding host through the Entegor platform to perform detection operations on the virtual machine restart status and the corresponding application startup status; The virtual machines whose corresponding restart status is "not restarted" or whose corresponding restart status is "restarted" and whose application startup status is "not started" are identified as the target virtual machines; The virtual machine restart status detection operation includes: for the second virtual machine that drifted and / or the first virtual machine after restarting, querying the startup time of each virtual machine, comparing the startup time with the host machine downtime, and if the startup time is after the downtime, it is determined that it has been restarted; otherwise, it is determined that it has not been restarted. The application startup status detection operation includes: according to the service name of the application associated with the virtual machine, querying the process status of the associated application to determine whether the target application process exists; synchronously detecting the connectivity of the associated service port of the associated application, if the target application process exists and the associated service port is connected, it is determined that the application has started; otherwise, it is determined that it has not started.

5. The method according to claim 4, characterized in that, The step of performing application recovery operations on the target virtual machine according to the application recovery handling method matched to the target virtual machine includes: Based on the pre-set operation and maintenance log and the name of the organizational unit to which the virtual machine belongs in the key information of the virtual machine, the responsible personnel corresponding to the target virtual machine are located, and a cross-level emergency connection is created. The pre-set operation and maintenance log stores the association relationship and communication method between each organizational unit and the responsible personnel. A standardized handling instruction is sent to the responsible personnel through a preset communication interface. The standardized handling instruction is used to instruct the responsible personnel to perform an application recovery operation on the target virtual machine based on the cross-level emergency connection. The standardized handling instruction includes the target virtual machine IP, the associated application name, and the application recovery handling method.

6. The method according to claim 4, characterized in that, The step of performing application recovery operations on the target virtual machine according to the application recovery handling method matched to the target virtual machine includes: Based on the key information of the virtual machine, the target virtual machine is automatically classified according to a preset classification rule, and its processing priority is marked. Based on the classification results and the aforementioned processing priority, determine the application recovery script adapted to the target virtual machine; The application recovery script automatically performs application recovery operations on the target virtual machine.

7. The method according to claim 6, characterized in that, The automatic classification of the target virtual machine based on the key information of the virtual machine according to a preset classification rule, and the marking of its processing priority, includes: Based on the purpose attribute in the key information of the virtual machine, at least one of the target virtual machines is grouped, and each target virtual machine is classified into a production virtual machine group or an office virtual machine group. Based on the operating system type in the key information of the virtual machine, all target virtual machines in the production virtual machine group and the office virtual machine group are further subdivided into Windows system virtual machines and Linux system virtual machines, forming four types of virtual machine groups. According to the preset priority setting rules, the target virtual machines in the four types of virtual machine groups are marked with the corresponding disposal priorities.

8. An emergency handling device for virtual machines when the host machine crashes, characterized in that, include: The acquisition module is used to acquire a list of virtual machines and key information of the virtual machines corresponding to the crashed host machine. The list of virtual machines contains multiple virtual machines, and the key information of the virtual machines includes at least the virtual machine name, purpose attribute, operating system type and the name of the organizational unit to which it belongs. The judgment module is used to determine whether each virtual machine has drifted based on the key information of the virtual machine; The restart module is used to restart the first virtual machine that has not experienced drift after troubleshooting the faulty host machine. The execution module is used to filter out target virtual machines with applications not started in the second virtual machine that has drifted and / or the first virtual machine after restarting, and to perform application recovery operations on the target virtual machines according to the application recovery handling method matched to the target virtual machines.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.