Equipment exception processing method, electronic equipment, computer storage medium and computer program product

By monitoring the accelerator status while the virtual machine is running, adding a backup accelerator or migrating the virtual machine, the complex recovery problem of the virtual machine working state caused by accelerator exceptions is solved, and the automated recovery and simplified operation of the accelerator are achieved.

CN120508414APending Publication Date: 2025-08-19HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410183522.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the cloud service background, when the accelerator device status of the virtual machine is abnormal, the working status of the virtual machine is affected, and the existing recovery process is complicated and cumbersome.

Method used

Monitor the accelerator device status while the virtual machine of the local server is running, and add the backup accelerator in the local server as an exclusive instance to the virtual machine in the event of an exception, load its instance driver, or migrate the virtual machine to the local server and add the backup accelerator when the accelerator is abnormal.

Benefits of technology

The working state recovery process of the virtual machine is simplified, the automatic recovery of the accelerator is realized, and manual operation steps are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508414A_ABST
    Figure CN120508414A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an equipment exception handling method, electronic equipment, a computer storage medium and a computer program product. The equipment exception handling method comprises the steps of monitoring an equipment state of a current accelerator in a running state of a virtual machine of a local server; when the equipment state indicates that the current accelerator is abnormal, adding a standby accelerator installed in the local server as an exclusive instance to the virtual machine; and loading the instance driving program of the standby accelerator to the virtual machine. According to the embodiment of the invention, the exclusive instance of the accelerator can be automatically recovered, so that the recovery process of the working state of the virtual machine is simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular to a device exception handling method, electronic equipment, computer storage medium, and computer program product. Background Art

[0002] In the current cloud service backend, the cloud server host is equipped with accelerators such as GPUs to allocate the accelerator's device capabilities to virtual machines, thereby improving the data processing efficiency of the virtual machines in the host and providing more cloud service computing power for virtual machine tenants.

[0003] When an accelerator's device status is abnormal or faulty, the virtual machine's operating status is affected. To restore the virtual machine's operating status, the VM tenant must shut down the VM on the cloud service frontend, assign a healthy accelerator to the VM on the cloud service backend, and then restart the VM on the cloud service frontend. However, this virtual machine recovery process is complex. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a device exception handling method, an electronic device, a computer storage medium, and a computer program product to at least partially solve the above-mentioned problems.

[0005] According to a first aspect of an embodiment of the present invention, a method for handling device exceptions is provided, comprising: monitoring the device status of a current accelerator while a virtual machine of a local server is running; when the device status indicates that an exception has occurred in the current accelerator, adding a backup accelerator installed in the local server as an exclusive instance to the virtual machine; and loading an instance driver of the backup accelerator into the virtual machine.

[0006] According to a second aspect of an embodiment of the present invention, a method for handling device exceptions is provided, comprising: when the device status of a current accelerator of a remote server indicates that an exception has occurred in the current accelerator, migrating the virtual machine of the remote server to a local server; adding a backup accelerator installed in the local server as an exclusive instance to the virtual machine; and loading an instance driver of the backup accelerator into the virtual machine.

[0007] According to a third aspect of an embodiment of the present invention, a device exception handling apparatus is provided, comprising: a status monitoring module, configured to monitor the device status of a current accelerator while a virtual machine of a local server is running; an instance creation module, configured to add a backup accelerator installed in the local server as an exclusive instance to the virtual machine when the device status indicates that the current accelerator has an exception; and an instance driver module, configured to load an instance driver program of the backup accelerator into the virtual machine.

[0008] According to a fourth aspect of an embodiment of the present invention, a device exception handling apparatus is provided, comprising: a virtual machine migration module, which migrates the virtual machine of the remote server to a local server when the device status of the current accelerator of the remote server indicates that the current accelerator has an exception; an instance creation module, which adds the backup accelerator installed in the local server as an exclusive instance to the virtual machine; and an instance driver module, which loads the instance driver of the backup accelerator into the virtual machine.

[0009] According to a fifth aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect or the second aspect.

[0010] According to a sixth aspect of an embodiment of the present invention, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect or the second aspect is implemented.

[0011] According to a seventh aspect of an embodiment of the present invention, a computer program product is provided, comprising a computer program / instruction, which implements the method described in the first aspect or the second aspect when executed by a processor.

[0012] In the solution of the embodiment of the present invention, the backup accelerator installed in the local server is added to the virtual machine as an exclusive instance, which enables the virtual machine to discover the exclusive instance of the backup accelerator, and then load the instance driver of the backup accelerator into the virtual machine, which enables the virtual machine to access the backup accelerator and automatically restore the exclusive instance of the accelerator, thereby simplifying the process of restoring the working status of the virtual machine. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0014] Figure 1 The present invention is a schematic block diagram of a server system to which some embodiments of the present invention are applicable.

[0015] Figure 2 The present invention is a flowchart of the steps of a device exception handling method according to some embodiments of the present invention.

[0016] Figure 3 The following is a flowchart of the steps of a device exception handling method according to some other embodiments of the present invention.

[0017] Figure 4 The following is a flowchart of the steps of a device exception handling method according to some other embodiments of the present invention.

[0018] Figure 5 Schematic block diagram of a device exception handling apparatus according to some other embodiments of the present invention.

[0019] Figure 6 Schematic block diagram of a device exception handling apparatus according to some other embodiments of the present invention.

[0020] Figure 7 Schematic diagram of the structure of electronic devices according to other embodiments of the present invention. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention.

[0022] The specific implementation of the embodiment of the present invention is further described below with reference to the accompanying drawings of the embodiment of the present invention.

[0023] In a server system, in addition to processors and memory such as CPUs, computing resources in the server can also be configured with accelerators for high-performance or artificial intelligence computing, for example, in the form of multiple accelerator cards. In some cases, an accelerator such as a single GPU card can be allocated to a virtual machine as an exclusive instance, thereby simplifying the accelerator virtualization process and greatly improving the data processing efficiency of the virtual machine. For example, the accelerator can be a device such as a GPU, NPU, or DPU. In some cases, a server can be configured with 4 or 8 GPU cards, so that the GPU computing resources can be quickly and efficiently allocated to each virtual machine in the server.

[0024] When the accelerator's device status is abnormal or faulty, the operating status of the virtual machine will be affected. Especially when the entire accelerator is assigned to a virtual machine, restoring the virtual machine's operating status requires the VM tenant to shut down the VM on the cloud service frontend, assign a healthy accelerator to the VM on the cloud service backend, and then restart the VM on the cloud service frontend.

[0025] Since the above-mentioned recovery process of the virtual machine is relatively complicated, various embodiments of the present invention provide a series of solutions to simplify the recovery process of the working status of the virtual machine.

[0026] Figure 1 FIG2 shows a server system to which some embodiments of the present invention are applicable. The server system includes a management node 20 and various servers 10 as computing nodes. The server 10 can be one or more physical machines. Figure 1 In the example, each server is implemented as a physical machine. A physical machine, acting as a host, can be configured with one or more virtual machines. Management node 20 can control the server's devices through a control module in each server 10. In some examples, the control module in server 10 can function as part of a local agent for management node 20 in server 10.

[0027] Specifically, the server serving as the host machine can manage the virtual machines by installing and configuring a virtualization management program, which includes a front-end management program and a back-end management program.

[0028] The front-end hypervisor generally refers to the part of the virtual machine software where users interact directly. It provides the user interface and management tools for creating, configuring, and managing virtual machines. The front-end hypervisor offers an intuitive graphical interface that allows users to easily interact with virtual machines and manage various settings and operations of the virtualized environment.

[0029] In virtual machine software, a backend hypervisor refers to the underlying virtualization technology or emulator used to run virtual machines on a host computer. This backend software is responsible for providing core virtualization functionality, including virtual hardware simulation, resource allocation and management, and interaction with the host hardware. This can be virtualization technologies such as KVM or QEMU, or embedded virtual machine monitors such as those found in VirtualBox and VMware. In some examples, the control module can be implemented using a backend hypervisor.

[0030] Figure 2 The present invention is a flowchart of the steps of a device exception handling method according to some embodiments of the present invention. Figure 2 The device exception handling method is applied to the local server. Specifically, the device exception handling method includes:

[0031] S210: Monitoring the device status of the current accelerator in the running state of the virtual machine of the local server.

[0032] It should be understood that the number of accelerators herein may be one or more accelerators, such as a GPU card, a DPU card, or an NPU card. As an example, one accelerator corresponds to one accelerator instance.

[0033] It should also be understood that the driver of the current accelerator can be installed through the virtualization component of the local server's operating system, and the exclusive instance of the current accelerator can be allocated to the virtual machine. For example, when the local server's operating system is a Linux operating system, the virtualization component can be Virtual Function I / O.

[0034] S220: When the device status indicates that the current accelerator is abnormal, add the backup accelerator installed in the local server as an exclusive instance to the virtual machine.

[0035] It should be understood that a control module such as QEMU can simulate hot-plugging and hot-unplugging of the instance of the standby accelerator from the virtual machine. The hot-plugging operation can add the instance of the standby accelerator to the virtual machine, and the hot-unplugging operation can remove the instance of the standby accelerator from the virtual machine.

[0036] Furthermore, after adding the instance of the standby accelerator to the virtual machine, if the instance driver of the standby accelerator is installed in the virtual machine, the virtual machine can access the standby accelerator. After removing the instance of the standby accelerator from the virtual machine, regardless of whether the instance driver of the standby accelerator is installed in the virtual machine, the virtual machine cannot access the standby accelerator.

[0037] S230: Load the instance driver of the standby accelerator into the virtual machine.

[0038] It should be understood that loading the instance driver of the standby accelerator into the virtual machine enables the virtual machine to access the instance of the standby accelerator in the management module through its own operating system, and the instance of the standby accelerator can access the standby accelerator through the operating system of the local server where the virtual machine is located.

[0039] In the solution of the embodiment of the present invention, the backup accelerator installed in the local server is added to the virtual machine as an exclusive instance, which enables the virtual machine to discover the exclusive instance of the backup accelerator, and then load the instance driver of the backup accelerator into the virtual machine, which enables the virtual machine to access the backup accelerator and automatically restore the exclusive instance of the accelerator, thereby simplifying the process of restoring the working status of the virtual machine.

[0040] Combine Figure 1 To describe Figure 2 A device exception handling method according to an embodiment of the present invention. Figure 1The monitoring module, exception handling module, and control module in the system can be implemented through the above-mentioned back-end management program as part of the virtual machine management program. Specifically, the monitoring module is used to monitor the device status of each accelerator, and the exception handling module is used to determine the corresponding exception handling method for the exception type. When the exception handling module determines to execute the switching of the backup accelerator, the control module is used to execute the hot plug operation or hot pull operation of the accelerator instance from the virtual machine, the migration of the virtual machine, and the installation control of the accelerator instance driver. It should be understood that the monitoring module, the control module, and the exception handling module can serve as a local agent of the management node 20 in the server 10.

[0041] In other embodiments, as an example of adding a standby accelerator installed on a local server as an exclusive instance to a virtual machine, the standby accelerator's device driver can be installed in the local server's operating system. Then, through the operating system's virtualization component, an exclusive instance of the standby accelerator is generated, allowing the virtual machine's control module on the local server to access the exclusive instance of the standby accelerator. In this way, the standby accelerator is exposed to the local server's control module for monitoring the virtual machine through the virtual machine component, concisely implementing virtualization of the standby accelerator.

[0042] Specifically, the virtualization component can be Virtual Function I / O, which installs the device driver of the standby accelerator into the operating system (for example, the Linux operating system) through the virtualization component, transfers the operating system's control authority over the standby accelerator to the control module of the virtual machine, and realizes the virtualization instance of the standby accelerator.

[0043] Furthermore, as an example of loading the instance driver of the standby accelerator into the virtual machine, a discovery event of the virtual machine's exclusive instance of the standby accelerator can be generated by the control module, and then, in response to the discovery event, the instance driver of the standby accelerator is loaded into the operating system of the virtual machine. Thus, the discovery event is generated by reusing the management capability of the device instance of the control module, and then the discovery event is used to guide the process of loading the instance driver of the standby accelerator into the operating system of the virtual machine, which is compatible with the configuration of the control module while improving the efficiency of device exception handling. For example, generating a discovery event of the virtual machine's exclusive instance of the standby accelerator through the control module can simulate the hot plug operation of the standby accelerator instance to the virtual machine for the control module. Since the control module can manage the virtual machine's access to the virtualized instance of the standby accelerator, the discovery event can be generated by means of communication between the control module and the management module that manages the slot of the standby accelerator in the operating system. For example, the discovery event can be a discovery event forwarded to the control module in response to the simulation of the management module when the standby accelerator is inserted into the slot.

[0044] Furthermore, as an example of loading an instance driver of a standby accelerator into a virtual machine, the current accelerator can be updated to the standby accelerator, thereby further making it compatible with the software configuration of the current accelerator. For example, the control module may monitor the standby accelerator as the current accelerator. Furthermore, when the device driver of the standby accelerator is installed in the operating system of the local server, the address space of the standby accelerator needs to be mapped to the address space of the operating system of the local server. When the instance driver of the standby accelerator is installed in the operating system of the virtual machine itself, the address space of the operating system of the local server needs to be mapped to the address space of the operating system of the virtual machine. Therefore, the device status monitoring of the accelerator can be continued by replacing the address space of the standby accelerator with the address space of the current accelerator.

[0045] In other embodiments, the device exception handling method further includes: allocating the current accelerator as an exclusive instance to the virtual machine, and when the device status indicates that the current accelerator has experienced an exception, removing the exclusive instance of the current accelerator from the virtual machine, thereby facilitating the unbinding of the virtual machine from the instance of the current accelerator. Furthermore, when the backup accelerator is not installed on the local server, migrating the virtual machine to an off-site server where the backup accelerator is located, so that the backup accelerator is allocated as an exclusive instance to the virtual machine. In other words, when the backup accelerator is installed on an off-site server, removing the exclusive instance of the current accelerator from the virtual machine facilitates the migration of the virtual machine to the off-site server.

[0046] Specifically, as an example of removing the current exclusive instance of the accelerator from the virtual machine, the virtual machine's control module on the local server can generate a removal event for the virtual machine's exclusive instance of the current accelerator. Then, in response to the removal event, the current exclusive instance of the accelerator is removed from the virtual machine. This reuses the control module's device instance management capabilities to generate a removal event for the current exclusive instance of the accelerator.

[0047] It should be understood that after removing the backup accelerator instance from a virtual machine, the virtual machine will no longer be able to access the backup accelerator, regardless of whether the backup accelerator instance driver is installed in the virtual machine. Furthermore, after removing the backup accelerator instance from the virtual machine, the virtual machine is disconnected from the underlying computing resource instance, facilitating the migration of the virtual machine from a local server to a remote server and maintaining compatibility with the technical architecture for virtual machine migration.

[0048] In other embodiments, the device exception handling method further includes: sending a bus reset instruction to the current accelerator via the processor of the local server, wherein the bus reset instruction instructs the current accelerator to perform a reset operation. Thus, sending the bus reset instruction facilitates recovery from a partial fault of the current accelerator.

[0049] Specifically, the accelerator can be restored through a secondary bus reset (SBR), for example, by sending a reset signal on a bus (eg, PCIe) connected to the failed accelerator (eg, the current accelerator), causing the accelerator to perform a reset operation through the reset signal.

[0050] Furthermore, after performing the reset operation, the accelerator's device driver can be restarted, and an access test can be performed on the accelerator through the local server's operating system. If the access test passes, the accelerator instance can be made visible to the virtual machine through the virtual machine's control module, and the accelerator's instance driver can be loaded into the virtual machine's operating system. In one example, when the failed accelerator instance is removed from the virtual machine, the accelerator's instance driver can be uninstalled from the virtual machine's operating system, and after the above-mentioned access test passes, the accelerator's instance driver can be installed into the virtual machine's operating system. In another example, before performing an access test on the accelerator through the local server's operating system, the accelerator's instance driver can be kept installed in the virtual machine's operating system, so that after the access test passes, the accelerator instance in the virtual machine can be restored.

[0051] In addition, when the access test passes, if the exclusive instance of the standby accelerator has been added to the virtual machine, the instance driver of the standby accelerator can be stopped from being loaded into the operating system of the virtual machine. In addition, when the access test passes, if the instance driver of the accelerator is loaded into the operating system of the virtual machine, the instance of the standby accelerator can be hot-plugged from the virtual machine, and the instance driver of the restored accelerator (i.e., the current accelerator) can be reloaded into the operating system of the virtual machine, thereby switching the standby accelerator to the restored current accelerator.

[0052] Alternatively, if the access test fails, the instance driver of the accelerator may be uninstalled from the operating system of the virtual machine, and other recovery and repair methods of the accelerator, such as manual repair, may be initiated.

[0053] In other embodiments, the device exception handling method further includes migrating the virtual machine to a remote server where the backup accelerator is located, so that the backup accelerator on the remote server is allocated as an exclusive instance to the virtual machine. This integrates the migration of the virtual machine server into the automated accelerator recovery process. Even if the backup server installed on the local server becomes unavailable, the backup accelerator on the remote server can be used to automatically recover the accelerator.

[0054] Specifically, when the backup accelerator installed on the local server is unavailable, a query request for a backup accelerator can be sent to the management node. The management node then uses the accelerator installation list of each server to identify the currently available backup accelerators and the remote server where the backup accelerator is located. In other words, after the accelerator's device driver is installed on the server, the server's operating system sends the accelerator's identification and usage status to the management node via the control module. The management node then associates the accelerator's identification with the server's identification to manage and maintain the installation and usage status of each accelerator.

[0055] Furthermore, the management node can obtain the operating parameters of the virtual machine in the local server (for example, through the monitoring module) and forward the operating parameters of the virtual machine to the remote server (for example, the monitoring module of the remote server), so that the operating system of the remote server creates and starts the virtual machine based on the operating parameters of the virtual machine, thereby achieving the migration of the virtual machine from the local server to the remote server, that is, the server where the virtual machine is located has an alternative accelerator installed. It should be understood that local server and remote server are relative concepts, and the local server can also be regarded as a remote server of the remote server.

[0056] Furthermore, the process of installing the candidate accelerator in the local server may be implemented by installing a device driver of the candidate accelerator into the operating system of the local server, and allocating a virtualized instance of the candidate accelerator to a virtual machine (eg, via a control module).

[0057] In other embodiments, the device exception handling method further includes: starting an application program that accesses the standby accelerator in the virtual machine, thereby directly restarting the application program to achieve recovery of the application program without restarting the virtual machine.

[0058] Specifically, the control module may send an application startup instruction to the virtual machine (e.g., in response to the completion of loading the instance driver). For example, the control module parses the application identifier and process parameters from the virtual machine's operating parameters, and in response to the completion of loading the instance driver, starts the application based on the process parameters of the application corresponding to the identifier.

[0059] Figure 3 The following is a flowchart of the steps of a device exception handling method according to some other embodiments of the present invention. Figure 3 The device exception handling methods include:

[0060] S310: When the device status of the current accelerator of the remote server indicates that the current accelerator is abnormal, migrate the virtual machine of the remote server to the local server.

[0061] S320: Add the standby accelerator installed in the local server to the virtual machine as an exclusive instance.

[0062] S330: Load the instance driver of the standby accelerator into the virtual machine.

[0063] It should be understood that the "local server" and "remote server" in this article are relative concepts, that is, the "local server" is the remote server of the "remote server". Figure 2 In the example, you can migrate the virtual machine to a remote server where the backup accelerator is located, so that the backup accelerator on the remote server is allocated to the virtual machine as an exclusive instance. Figure 3 In the example, the virtual machine of the remote server is migrated to the local server, and the standby accelerator installed in the local server is added to the virtual machine as an exclusive instance.

[0064] It should also be understood that in order to achieve the migration of the virtual machine of the remote server to the local server, a control module such as QEMU can obtain the operating parameters of the virtual machine from the operating system of the virtual machine through inter-process communication, and send the operating parameters of the virtual machine to the management node. In other words, the management node can obtain the operating parameters of the virtual machine in the local server, and forward the operating parameters of the virtual machine to the remote server (for example, the monitoring module of the remote server), so that the operating system of the remote server creates and starts the virtual machine based on the operating parameters of the virtual machine, thereby achieving the migration of the virtual machine from the local server to the remote server, that is, the server where the virtual machine is located is installed with an alternative accelerator. It should be understood that local server or remote server is a relative concept, and the local server is also a remote server of the remote server.

[0065] In this embodiment, when the device status of the current accelerator of the remote server indicates that the current accelerator is abnormal, the virtual machine of the remote server is migrated to the local server, and the backup accelerator is added to the virtual machine as an exclusive instance, so that the virtual machine can access the backup accelerator installed in the local server across servers, and then the instance driver of the backup accelerator is loaded into the virtual machine, so that the virtual machine in the remote server can access the backup accelerator in the local server. Therefore, the exclusive instance of the accelerator in the remote server is automatically restored, thereby simplifying the recovery process of the working status of the virtual machine.

[0066] The following will be combined Figure 4 The device exception handling methods according to other embodiments of the present invention are described in detail. Figure 4 The device exception handling methods include:

[0067] Step S410: Monitor the status of the current accelerator, and then the process proceeds to step S420. For example, the status of the current accelerator can be monitored through the operating system of the local server. If the local server cannot normally access the current accelerator, the recovery of the current accelerator is started.

[0068] Step S420: Determine whether the current accelerator is abnormal. If so, the process proceeds to step S430; if not, the process proceeds to step S410.

[0069] Step S430: Determine whether the current accelerator is allocated to the virtual machine as an independent instance. If so, the process proceeds to step S450; if not, the process proceeds to step S440. For example, whether the current accelerator is an independent instance can be determined by a control module such as QEMU. For example, when the current accelerator instance is virtualized to the virtual machine's control module, the control module can obtain description information of the current accelerator instance. The description information can indicate that the current accelerator as a whole is virtualized as an exclusive instance and allocated to the virtual machine.

[0070] Step S440: Execute cold migration of the virtual machine. For example, the monitoring module can obtain the running status of the virtual machine and send the running status of the virtual machine to the management node. The management node provides the running status of the virtual machine to an off-site server with a backup accelerator, and creates and starts the virtual machine in the off-site server based on the running status to achieve migration of the virtual machine. For example, because the virtual machine has a large dependency on the resources it accesses, when the current accelerator is not allocated to the virtual machine as an exclusive instance, for example, when multiple accelerator instances are allocated to the virtual machine, there is a software dependency between the non-faulty accelerator and the virtual machine. Therefore, the tenant of the virtual machine can be notified to shut down the virtual machine, thereby cold migrating the virtual machine to the off-site server, thereby achieving more reliable device recovery through cold migration.

[0071] Step S450: Determine whether a backup accelerator is installed in the local server. If yes, the process proceeds to step S470; if no, the process proceeds to step S460. For example, a query request for a backup accelerator can be sent to the management node. The management node determines whether a currently available backup accelerator exists in the local server's accelerator installation list based on each server's accelerator installation list. That is, after the accelerator's device driver is installed in the server, the server's operating system sends the accelerator's identification and usage status to the management node via the control module. The management node then associates the accelerator's identification with the server's identification to manage and maintain the installation and usage status of each accelerator.

[0072] Step S460: Execute the hot-plug operation of the current accelerator from the virtual machine, hot-migrate the virtual machine to the remote server where the backup accelerator is located, and perform the hot-plug operation of the backup accelerator to the virtual machine, and then the process proceeds to step S480. For example, the instance of the backup accelerator can be removed from the virtual machine to implement the hot-plug operation. It should be understood that after the instance of the backup accelerator is removed from the virtual machine, the virtual machine cannot access the backup accelerator regardless of whether the instance driver of the backup accelerator is installed in the virtual machine. For another example, generating a discovery event of the virtual machine's exclusive instance of the backup accelerator through the control module can simulate the hot-plug operation of the instance of the backup accelerator to the virtual machine for the control module. It should also be understood that the operation of migrating the virtual machine from the local server to the remote server can adopt the description method in the above-mentioned various embodiments, which will not be repeated here.

[0073] Step S470: Hot-plug the current accelerator from the virtual machine and hot-plug the backup accelerator into the virtual machine. The process then proceeds to step S480. For example, the backup accelerator instance can be removed from the virtual machine to implement the hot-plug operation. For another example, generating a discovery event indicating that the virtual machine has exclusive access to the backup accelerator instance can simulate a hot-plug operation of the backup accelerator instance into the virtual machine for the control module. It should be understood that if the backup accelerator exists on the local server, there is no need to migrate the virtual machine.

[0074] Step S480: Reload the instance driver of the standby accelerator. It should be understood that the method for loading the instance driver can adopt the method described in the above embodiments, which will not be repeated here.

[0075] Figure 5 Schematic block diagram of a device exception handling apparatus according to some other embodiments of the present invention. Figure 5 The device exception handling device includes:

[0076] The status monitoring module 510 monitors the device status of the current accelerator in the running state of the virtual machine of the local server.

[0077] The instance creation module 520 adds the backup accelerator installed in the local server as an exclusive instance to the virtual machine when the device status indicates that the current accelerator is abnormal.

[0078] The instance driver module 530 loads the instance driver of the standby accelerator into the virtual machine.

[0079] In the solution of the embodiment of the present invention, the backup accelerator installed in the local server is added to the virtual machine as an exclusive instance, which enables the virtual machine to discover the exclusive instance of the backup accelerator, and then load the instance driver of the backup accelerator into the virtual machine, which enables the virtual machine to access the backup accelerator and automatically restore the exclusive instance of the accelerator, thereby simplifying the process of restoring the working status of the virtual machine.

[0080] In other embodiments, the instance creation module is specifically used to: install the device driver of the backup accelerator in the local server into the operating system of the local server; generate an exclusive instance of the backup accelerator through the virtualization component of the operating system, so that the control module of the virtual machine in the local server accesses the exclusive instance of the backup accelerator.

[0081] In other embodiments, the instance driver module is specifically configured to: generate, through the control module, a discovery event of the virtual machine's exclusive instance of the standby accelerator; and load the instance driver of the standby accelerator into the operating system of the virtual machine in response to the discovery event.

[0082] In some other embodiments, the instance driving module is specifically configured to update the current accelerator to the standby accelerator.

[0083] In other embodiments, the device exception handling apparatus further includes: an instance removal module, configured to: allocate the current accelerator as an exclusive instance to the virtual machine; and remove the exclusive instance of the current accelerator from the virtual machine when the device status indicates that the current accelerator is abnormal.

[0084] In other embodiments, the instance removal module is specifically used to: generate a removal event of the virtual machine's exclusive instance of the current accelerator through the control module of the virtual machine in the local server; and remove the exclusive instance of the current accelerator from the virtual machine in response to the removal event.

[0085] In some other embodiments, the device exception handling apparatus further includes: a device reset module, which sends a bus reset instruction to the current accelerator through the processor of the local server, and the bus reset instruction instructs the current accelerator to perform a reset operation.

[0086] In some other embodiments, the device exception handling apparatus further includes: an application startup module, which starts an application program for accessing the standby accelerator in the virtual machine.

[0087] Figure 6 Schematic block diagram of a device exception handling apparatus according to some other embodiments of the present invention. Figure 6 The device exception handling device includes:

[0088] The virtual machine migration module 610 migrates the virtual machine of the remote server to the local server when the device status of the current accelerator of the remote server indicates that the current accelerator is abnormal.

[0089] The instance creation module 620 adds the standby accelerator installed in the local server to the virtual machine as an exclusive instance.

[0090] The instance driver module 630 loads the instance driver of the standby accelerator into the virtual machine.

[0091] In this embodiment, when the device status of the current accelerator of the remote server indicates that the current accelerator is abnormal, the virtual machine of the remote server is migrated to the local server, and the backup accelerator is added to the virtual machine as an exclusive instance, so that the virtual machine can access the backup accelerator installed in the local server across servers, and then the instance driver of the backup accelerator is loaded into the virtual machine, so that the virtual machine in the remote server can access the backup accelerator in the local server. Therefore, the exclusive instance of the accelerator in the remote server is automatically restored, thereby simplifying the recovery process of the working status of the virtual machine.

[0092] The specific implementation of each module in the device can refer to the corresponding description of the corresponding steps in the above method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-described device and module can refer to the corresponding process description in the above method embodiment, and will not be repeated here.

[0093] Reference Figure 7 , shows a schematic structural diagram of an electronic device according to another embodiment of the present invention. The specific embodiment of the present invention does not limit the specific implementation of the electronic device.

[0094] like Figure 7 As shown, the electronic device may include: a processor (processor) 702 for executing a program 710 , a communication interface (Communications Interface) 704 , a memory (memory) 706 , and a communication bus 708 .

[0095] The processor, the communication interface, and the memory communicate with each other via a communication bus.

[0096] Communication interface, used to communicate with other electronic devices or servers.

[0097] The processor is used to execute the program, and specifically can execute the relevant steps in the above method embodiment.

[0098] Specifically, the program may include program codes including computer operation instructions.

[0099] The processor may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or different types of processors, such as one or more CPUs and one or more ASICs.

[0100] Memory is used to store programs. The memory may include high-speed RAM memory, and may also include non-volatile memory (non-volatile memory), such as at least one disk storage.

[0101] The program may include multiple computer instructions. Specifically, the program may enable the processor to execute operations corresponding to each device exception handling method described in any of the aforementioned multiple method embodiments through the multiple computer instructions.

[0102] The specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, and will not be repeated here.

[0103] An embodiment of the present invention further provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments. The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk.

[0104] An embodiment of the present invention further provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to each device exception handling method in the above-mentioned multiple method embodiments.

[0105] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0106] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0107] The method according to the embodiment of the present invention described above can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.

[0108] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present invention.

[0109] The above implementation methods are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Ordinary technicians in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the scope of patent protection of the embodiments of the present invention should be defined by the claims.

Claims

1. A method for handling device abnormalities, comprising: Monitor the current accelerator device status while the virtual machine of the local server is running; When the device status indicates that the current accelerator is abnormal, adding the backup accelerator installed in the local server as an exclusive instance to the virtual machine; An instance driver of the standby accelerator is loaded into the virtual machine.

2. The method according to claim 1, wherein Adding the standby accelerator installed in the local server to the virtual machine as an exclusive instance includes: Installing a device driver of the standby accelerator in the local server into the operating system of the local server; An exclusive instance of the standby accelerator is generated through a virtualization component of the operating system, so that a control module of the virtual machine in the local server accesses the exclusive instance of the standby accelerator.

3. The method according to claim 2, wherein: Loading the instance driver of the standby accelerator into the virtual machine includes: generating, by the control module, a discovery event of the virtual machine's exclusive instance of the standby accelerator; In response to the discovery event, an instance driver of the standby accelerator is loaded into the operating system of the virtual machine.

4. The method according to claim 1, wherein The method further comprises: Allocating the current accelerator as an exclusive instance to the virtual machine; When the device status indicates that an abnormality occurs in the current accelerator, the exclusive instance of the current accelerator is removed from the virtual machine.

5. The method according to claim 4, wherein Removing the exclusive instance of the current accelerator from the virtual machine includes: generating, by the control module of the virtual machine in the local server, a removal event of the virtual machine's exclusive instance of the current accelerator; In response to the removal event, the exclusive instance of the current accelerator is removed from the virtual machine.

6. The method according to claim 4, wherein: The method further comprises: A bus reset instruction is sent to the current accelerator through the processor of the local server, where the bus reset instruction instructs the current accelerator to perform a reset operation.

7. The method according to claim 1, wherein The method further comprises: An application program for accessing the standby accelerator is started in the virtual machine.

8. A method for handling device abnormalities, comprising: When the device status of the current accelerator of the remote server indicates that the current accelerator is abnormal, migrating the virtual machine of the remote server to the local server; adding the standby accelerator installed in the local server as an exclusive instance to the virtual machine; An instance driver of the standby accelerator is loaded into the virtual machine.

9. An electronic device comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, where the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 8.

10. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

11. A computer program product comprising a computer program / instruction, which implements the method according to any one of claims 1 to 8 when executed by a processor.