Device anomaly processing method, electronic device, computer storage medium, and computer program product

By using the backup accelerator as an exclusive instance in a local or off-site server and loading the instance driver, the complex recovery of the virtual machine's working state caused by accelerator exception in the cloud service background is solved, and the automated recovery of the accelerator and efficient operation of the virtual machine are realized.

WO2025177103A1PCT designated stage Publication Date: 2025-08-28CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/051209
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2025-02-05
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

In the cloud service background, when the accelerator device of the cloud server is abnormal or faulty, the working status of the virtual machine is affected, and the existing recovery process is complicated and cumbersome.

Method used

In a local or off-site server, by monitoring the accelerator status, using a backup accelerator as an exclusive instance to add it to the virtual machine, and loading the instance drivers to achieve automated recovery of the accelerator.

Benefits of technology

Simplifies the working state recovery process of the virtual machine and improves the efficiency and reliability of device exception handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025051209_28082025_PF_FP_ABST
    Figure IB2025051209_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a device anomaly processing method, an electronic device, a computer storage medium, and a computer program product. The device anomaly processing method comprises: monitoring a device state of a current accelerator under a running state of a virtual machine of a local server; when the device state indicates that the current accelerator is anomalous, adding a standby accelerator installed in the local server to the virtual machine as an exclusive instance; and loading an instance driving program of the standby accelerator to the virtual machine. The embodiments of the present disclosure can automatically recover the exclusive instance of the accelerator, thereby simplifying the recovery process of a working state of the virtual machine.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Device Abnormality Handling Method, Electronic Device, Computer Storage Medium, and Computer Program Product This disclosure claims priority to Chinese patent application number 202410183522.3, filed with the China Patent Office on February 19, 2024, entitled "Device Abnormality Handling Method, Electronic Device, Computer Storage Medium, and Computer Program Product," the entire contents of which are incorporated herein by reference. Technical Field: Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a device abnormality handling method, electronic device, computer storage medium, and computer program product. Background: In current cloud service backends, cloud server host machines are equipped with accelerators such as GPUs (Graphics Processing Units). This allocates the accelerator's device capabilities to virtual machines, improving the data processing efficiency of the virtual machines in the host machine and providing more cloud service computing power to virtual machine tenants. When the accelerator's device status is abnormal or faulty, the operating status of the virtual machine is affected. To restore the virtual machine's operating status, the virtual machine tenant must shut down the virtual machine on the cloud service frontend, assign a healthy accelerator to the virtual machine on the cloud service backend, and then restart the virtual machine on the cloud service frontend. However, the above-mentioned virtual machine recovery process is relatively complex. In view of this, embodiments of the present disclosure provide a device exception handling method, an electronic device, a computer storage medium, and a computer program product to at least partially address the aforementioned issues. According to a first aspect of an embodiment of the present disclosure, a device exception handling method is provided, comprising: monitoring the device status of a current accelerator while a virtual machine on a local server is running; when the device status indicates that the current accelerator has experienced an exception, adding a backup accelerator installed on the local server as an exclusive instance to the virtual machine; and loading an instance driver for the backup accelerator into the virtual machine. According to a second aspect of an embodiment of the present disclosure, a device exception handling method is provided, comprising: when the device status of a current accelerator on a remote server indicates that the current accelerator has an exception, migrating a virtual machine on the remote server to a local server; adding a backup accelerator installed on the local server as an exclusive instance to the virtual machine; and loading an instance driver for the backup accelerator into the virtual machine.According to a third aspect of an embodiment of the present disclosure, a device exception handling apparatus is provided, comprising: a status monitoring module for monitoring the device status of a current accelerator while a virtual machine on a local server is running; an instance creation module for adding a backup accelerator installed on the local server as an exclusive instance to the virtual machine when the device status indicates that the current accelerator has experienced an exception; and an instance driver module for loading an instance driver program for the backup accelerator into the virtual machine. According to a fourth aspect of an embodiment of the present disclosure, a device exception handling apparatus is provided, comprising: a virtual machine migration module for migrating the virtual machine on a remote server to a local server when the device status of the current accelerator on the remote server indicates that the current accelerator has experienced an exception; an instance creation module for adding the backup accelerator installed on the local server as an exclusive instance to the virtual machine; and an instance driver module for loading the instance driver program for the backup accelerator into the virtual machine. According to a fifth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is configured to store at least one executable instruction, wherein the executable instruction causes the processor to perform operations corresponding to the method described in the first or second aspect. According to a sixth aspect of an embodiment of the present disclosure, a computer storage medium is provided, storing a computer program, which, when executed by a processor, implements the method described in the first or second aspect. According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program / instructions, which, when executed by a processor, implements the method described in the first or second aspect. In the solution of the embodiment of the present disclosure, a backup accelerator installed on a local server is added to a virtual machine as an exclusive instance, enabling the virtual machine to discover the exclusive instance of the backup accelerator. The instance driver of the backup accelerator is then loaded into the virtual machine, enabling the virtual machine to access the backup accelerator and automatically restoring the exclusive instance of the accelerator, thereby simplifying the process of restoring the virtual machine's operating status. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments described in the embodiments of the present disclosure. Those skilled in the art can also obtain other drawings based on these drawings.Figure 1 is a schematic block diagram of a server system applicable to some embodiments of the present disclosure; Figure 2 is a flowchart of the steps of a device exception handling method according to some embodiments of the present disclosure; Figure 3 is a flowchart of the steps of a device exception handling method according to other embodiments of the present disclosure; Figure 4 is a flowchart of the steps of a device exception handling method according to other embodiments of the present disclosure; Figure 5 is a schematic block diagram of a device exception handling apparatus according to other embodiments of the present disclosure; Figure 6 is a schematic block diagram of a device exception handling apparatus according to other embodiments of the present disclosure; and Figure 7 is a schematic structural diagram of an electronic device according to other embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS To help those skilled in the art better understand the technical solutions in the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a portion of the embodiments of the present disclosure, and are not exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure should fall within the scope of protection of the embodiments of the present disclosure. The specific implementation of the embodiments of the present disclosure will be further described below in conjunction with the accompanying drawings. In server systems, computing resources, in addition to processors such as CPUs (Central Processing Units) and memory, can also include accelerators for high-performance or artificial intelligence computing, for example, in the form of multiple accelerator cards. In some cases, an accelerator, such as a single GPU card, can be allocated to a virtual machine as an exclusive instance, simplifying the accelerator virtualization process while significantly improving the virtual machine's data processing efficiency. For example, the accelerator can be a GPU, NPU (Neural Processing Unit), or DPU (Data Processing Unit). In some cases, a server can be configured with four or eight GPU cards, allowing GPU computing resources to be quickly and efficiently allocated to each virtual machine in the server. If the accelerator's device status is abnormal or faulty, the virtual machine's operating status will be affected. Especially when the entire accelerator is assigned to a virtual machine, if you want to restore the working status of the virtual machine, the virtual machine tenant needs to shut down the virtual machine on the cloud service front end, assign an accelerator with a normal device status to the virtual machine on the cloud service back end, and then start the virtual machine on the cloud service front end.Because the aforementioned virtual machine recovery process is relatively complex, various embodiments of the present disclosure provide a series of solutions to simplify the virtual machine operating status recovery process. Figure 1 illustrates a server system applicable to some embodiments of the present disclosure. The server system includes a management node 20 and various servers 10 serving as compute nodes. Servers 10 can be one or more physical machines. In the example of Figure 1 , each server is implemented as a physical machine. Physical machines, acting as hosts, can be configured with one or more virtual machines. The management node 20 can control server devices through a control module within each server 10. In some examples, the control module within a server 10 can function as a local agent for the management node 20 within the server 10. Specifically, a server serving as a host can manage virtual machines by installing and configuring a virtualization hypervisor. A virtualization hypervisor includes a front-end hypervisor and a back-end hypervisor. The front-end hypervisor generally refers to the portion where users directly interact with the virtual machine software, providing a user interface and management tools for creating, configuring, and managing virtual machines. The front-end hypervisor provides an intuitive graphical interface, allowing users to easily interact with virtual machines and manage various settings and operations within the virtualized environment. In virtual machine software, a backend hypervisor refers to the underlying virtualization technology or emulator used to run virtual machines on a host machine. The backend software is responsible for providing core virtualization functions, including virtual hardware simulation, resource allocation and management, and interaction with host hardware. These can be virtualization technologies such as KVM (Kernel-based Virtual Machine) or QEMU (Quick Emulator), or embedded virtual machine monitors such as those in Virtua I Box and VMware. In some examples, the control module can be implemented using a backend hypervisor. Figure 2 is a flowchart of the steps of a device exception handling method according to some embodiments of the present disclosure. The device exception handling method in Figure 2 is applied to a local server. Specifically, the device exception handling method includes the following steps: S210: Monitoring the device status of the current accelerator while the virtual machine on the local server is running. It should be understood that the number of accelerators herein may refer to one or more accelerators, such as GPU cards, DPU cards, or NPU cards. As an example, one accelerator corresponds to one accelerator instance.It should also be understood that the driver of the current accelerator can be installed through the virtualization component of the operating system of the local server, and the exclusive instance of the current accelerator can be allocated to the virtual machine. For example, when the operating system of the local server is a Linux operating system, the virtualization component can be Virtua I Function I / O. o

[0002] S220: When the device status indicates that the current accelerator has experienced an abnormality, the backup accelerator installed in the local server is added to the virtual machine as an exclusive instance. It should be understood that a control module, such as QEMU, can simulate hot-plugging and hot-unplugging the backup accelerator instance from the virtual machine. Hot-plugging can add the backup accelerator instance to the virtual machine, while hot-unplugging can remove the backup accelerator instance from the virtual machine. Furthermore, after the backup accelerator instance is added to the virtual machine, if the backup accelerator instance driver is installed in the virtual machine, the virtual machine can access the backup accelerator. After the backup accelerator instance is removed from the virtual machine, the virtual machine cannot access the backup accelerator regardless of whether the backup accelerator instance driver is installed in the virtual machine.

[0003] S230: Load the instance driver of the backup accelerator into the virtual machine. It should be understood that loading the instance driver of the backup accelerator into the virtual machine enables the virtual machine to access the instance of the backup accelerator in the management module through its own operating system. The instance of the backup accelerator can then access the backup accelerator via the operating system of the local server where the virtual machine resides. In the solution of the embodiment of the present disclosure, the backup accelerator installed on the local server is added to the virtual machine as an exclusive instance, enabling the virtual machine to discover the exclusive instance of the backup accelerator. The instance driver of the backup accelerator is then loaded into the virtual machine, enabling the virtual machine to access the backup accelerator, automatically restoring the exclusive instance of the accelerator, and thus simplifying the process of restoring the virtual machine's operating status. The device exception handling method of the embodiment of FIG2 is described in conjunction with FIG1 . The monitoring module, exception handling module, and control module in FIG1 can be implemented via the aforementioned backend hypervisor as part of the virtual machine hypervisor. Specifically, the monitoring module is used to monitor the device status of each accelerator, and the exception handling module is used to determine the corresponding exception handling method based on the exception type. When the exception handling module determines to execute a switchover of the standby accelerator, the control module is configured to perform hot-insertion or hot-removal of the accelerator instance from the virtual machine, migration of the virtual machine, and installation control of the accelerator instance driver. It should be understood that the monitoring module, control module, and exception handling module can serve as local agents of the management node 20 in the server 10. In other embodiments, as an example of adding a standby accelerator installed in a local server to a virtual machine as an exclusive instance, the device driver for the standby accelerator in the local server can be installed in the local server's operating system. Then, through the operating system's virtualization component, an exclusive instance of the standby accelerator is generated, allowing the virtual machine's control module in the local server to access the exclusive instance of the standby accelerator. Thus, the standby accelerator is exposed to the local server's control module for monitoring the virtual machine through the virtual machine component, concisely implementing virtualization of the standby accelerator. Specifically, the virtualization component may be Virtual Function I / O. The virtualization component is used to install the device driver of the standby accelerator into an operating system (e.g., a Linux operating system), transfer the operating system's control authority over the standby accelerator to the control module of the virtual machine, and implement a virtualized instance of the standby accelerator. Furthermore, as an example of loading the instance driver of the standby accelerator into the virtual machine, the control module may generate a discovery event indicating that the virtual machine has exclusive access to the standby accelerator instance. In response to the discovery event, the instance driver of the standby accelerator is loaded into the operating system of the virtual machine.Thus, by reusing the control module's device instance management capabilities, a discovery event is generated. This discovery event then guides the loading of the backup accelerator instance driver into the virtual machine's operating system, maintaining compatibility with the control module's configuration while improving device exception handling efficiency. For example, by generating a discovery event for the virtual machine's exclusive access to the backup accelerator instance, the control module can simulate a hot-plug operation of the backup accelerator instance into the virtual machine. Because the control module manages virtual machine access to the virtualized instance of the backup accelerator, discovery events can be generated through communication between the control module and the management module in the operating system that manages the backup accelerator's slot. For example, the discovery event can simulate the discovery event forwarded to the control module by the management module in response to the backup accelerator being inserted into the slot. Furthermore, as an example of loading the backup accelerator instance driver into the virtual machine, the current accelerator can be updated to the backup accelerator, further enhancing compatibility with the current accelerator's software configuration. For example, the control module may monitor the backup accelerator as the current accelerator. Furthermore, when installing the backup accelerator's device driver into the local server's operating system, the backup accelerator's address space needs to be mapped to the local server's operating system's address space. When installing the backup accelerator instance driver into the virtual machine's operating system, the address space of the local server's operating system needs to be mapped to the virtual machine's operating system address space. Therefore, by replacing the backup accelerator's address space with the current accelerator's address space, accelerator device status monitoring can continue. In other embodiments, the device exception handling method further includes: allocating the current accelerator as an exclusive instance to the virtual machine, and when the device status indicates an exception with the current accelerator, removing the exclusive instance of the current accelerator from the virtual machine, thereby facilitating the unbinding of the virtual machine from the current accelerator instance. Furthermore, if the backup accelerator is not installed on the local server, the virtual machine is migrated to a remote server where the backup accelerator is located, and the backup accelerator is allocated to the virtual machine as an exclusive instance. In other words, when the backup accelerator is installed on a remote server, removing the exclusive instance of the current accelerator from the virtual machine facilitates migration of the virtual machine to the remote server. Specifically, as an example of removing the current exclusive instance of the accelerator from the virtual machine, the virtual machine's control module on the local server can generate an event for removing the current exclusive instance of the accelerator from the virtual machine. Then, in response to the removal event, the current exclusive instance of the accelerator is removed from the virtual machine. This reuses the control module's device instance management capabilities to generate the removal event for the current exclusive instance of the accelerator.It should be understood that after the backup accelerator instance is removed from the virtual machine, the virtual machine cannot access the backup accelerator, regardless of whether the backup accelerator instance driver is installed in the virtual machine. Furthermore, after the backup accelerator instance is removed from the virtual machine, the association between the virtual machine and the underlying computing resource instance is disassociated, facilitating the migration of the virtual machine from a local server to a remote server and maintaining compatibility with the technical architecture for virtual machine migration. In other embodiments, the device exception handling method further includes: sending a bus reset instruction to the current accelerator via the local server's processor, where the bus reset instruction instructs the current accelerator to perform a reset operation. Thus, sending the bus reset instruction facilitates partial recovery from a fault in the current accelerator. Specifically, the accelerator can be restored via a secondary bus reset (SBR). For example, a reset signal is sent on a bus (e.g., PCIe) connected to the faulty accelerator (e.g., the current accelerator), causing the accelerator to perform a reset operation via the reset signal. Furthermore, after performing the reset operation, the accelerator's device driver can be restarted, and an access test can be performed on the accelerator via the local server's operating system. If the access test passes, the virtual machine's control module can make the accelerator instance visible to the virtual machine, and the accelerator instance driver can be loaded into the virtual machine's operating system. In one example, when the failed accelerator instance is removed from the virtual machine, the accelerator instance driver can be uninstalled from the virtual machine's operating system. After the access test passes, the accelerator instance driver can be installed into the virtual machine's operating system. In another example, before performing the access test on the accelerator via the local server's operating system, the accelerator instance driver can remain installed in the virtual machine's operating system. After the access test passes, the accelerator instance in the virtual machine can be restored. Furthermore, if a dedicated instance of the backup accelerator has already been added to the virtual machine when the access test passes, loading the backup accelerator instance driver into the virtual machine's operating system can be stopped. Furthermore, if the access test passes and the accelerator instance driver is loaded into the virtual machine's operating system, the standby accelerator instance can be hot-unplugged from the virtual machine and the instance driver of the restored accelerator (i.e., the current accelerator) can be reloaded into the virtual machine's operating system, switching the standby accelerator to the restored current accelerator. Alternatively, if the access test fails, the accelerator instance driver can be uninstalled from the virtual machine's operating system and other accelerator recovery and repair methods, such as manual repair, can be initiated.In other embodiments, the device exception handling method further includes migrating the virtual machine to a remote server where the backup accelerator is located, allocating the remote server's backup accelerator as an exclusive instance to the virtual machine. This integrates virtual machine server migration into the accelerator's automated recovery process. Even when the backup server installed on the local server is unavailable, the accelerator can be automatically recovered using the backup accelerator on the remote server. Specifically, when the backup accelerator installed on the local server is unavailable, a query request for the backup accelerator can be sent to the management node. The management node identifies currently available candidate accelerators from the accelerator installation list of each server and determines the remote server where the candidate accelerator is located. Specifically, after the accelerator's device driver is installed on the server, the server's operating system transmits the accelerator's identifier and usage status to the management node via the control module. The management node associates the accelerator identifier with the server identifier to manage and maintain the installation and usage status of each accelerator. Furthermore, the management node can (for example, via a monitoring module) obtain the operating parameters of the virtual machine on the local server and forward the virtual machine's operating parameters to the remote server (for example, the monitoring module of the remote server). This causes the operating system of the remote server to create and start the virtual machine based on the virtual machine's operating parameters, thereby migrating the virtual machine from the local server to the remote server. This means that the server where the virtual machine resides has an alternative accelerator installed. It should be understood that the terms "local server" and "remote server" are relative terms, and the local server also serves as a remote server of a remote server. Furthermore, the process of installing the alternative accelerator on the local server can be implemented by installing the device driver for the alternative accelerator into the operating system of the local server and allocating a virtualized instance of the alternative accelerator to the virtual machine (for example, via a control module). In other embodiments, the device exception handling method further includes starting an application in the virtual machine that accesses the backup accelerator. Therefore, directly restarting the application achieves application recovery without restarting the virtual machine. Specifically, the control module can send an application startup instruction to the virtual machine (for example, in response to the completion of loading the instance driver). For example, the control module parses the application identifier and process parameters from the virtual machine's operating parameters and, in response to the instance driver's loading completion, starts the application based on the process parameters of the application corresponding to the identifier. FIG3 is a flowchart of the steps of a device exception handling method according to other embodiments of the present disclosure. The device exception handling method of FIG3 includes:

[0004] S310: When the device status of the current accelerator of the remote server indicates that the current accelerator is abnormal, migrate the virtual machine of the remote server to the local server.

[0005] S320: Add the standby accelerator installed in the local server to the virtual machine as an exclusive instance.

[0006] S330: Loading the instance driver of the backup accelerator into the virtual machine. It should be understood that the "local server" and "remote server" herein are relative concepts; that is, the "local server" is the remote server of the "remote server." For example, in the example of FIG2 , the virtual machine can be migrated to the remote server where the backup accelerator is located, so that the backup accelerator of the remote server is allocated to the virtual machine as an exclusive instance. In the example of FIG3 , the virtual machine of the remote server is migrated to the local server, and the backup accelerator installed in the local server is added to the virtual machine as an exclusive instance. It should also be understood that to implement the migration of the virtual machine from the remote server to the local server, a control module such as QEMU can obtain the virtual machine's operating parameters from the virtual machine's operating system through inter-process communication and send the virtual machine's operating parameters to the management node. That is, the management node can obtain the operating parameters of the virtual machine on the local server and forward them to the remote server (e.g., the monitoring module of the remote server). This enables the operating system of the remote server to create and start the virtual machine based on the operating parameters of the virtual machine, thereby migrating the virtual machine from the local server to the remote server. In other words, the server where the virtual machine resides has a backup accelerator installed. It should be understood that the terms "local server" and "remote server" are relative terms, and the local server also serves as a remote server for a remote server. In this embodiment, when the device status of the current accelerator on the remote server indicates an abnormality, the virtual machine on the remote server is migrated to the local server, and the backup accelerator is added to the virtual machine as an exclusive instance, enabling the virtual machine to access the backup accelerator installed on the local server across servers. Furthermore, the instance driver for the backup accelerator is loaded into the virtual machine, enabling the virtual machine on the remote server to access the backup accelerator on the local server. Consequently, the exclusive instance of the accelerator on the remote server is automatically restored, simplifying the process of restoring the operating status of the virtual machine. The device exception handling method according to other embodiments of the present disclosure will be described in detail below with reference to FIG. The device abnormality handling method of FIG4 includes: Step S410: monitoring the current state of the accelerator, and then, the process proceeds to step S420 oFor example, the status of the current accelerator can be monitored by the operating system of the local server. If the local server cannot access the current accelerator normally, the recovery of the current accelerator is started. Step S420: Determine whether the current accelerator is abnormal. If yes, the process proceeds to step S430. If no, the process proceeds to step S410. o Step S430: Determine whether the current accelerator is assigned to the virtual machine as an independent instance. If so, the process proceeds to step S450; if not, the process proceeds to step S440. For example, a control module such as QEMU can determine whether the current accelerator is an independent instance. For example, when the current accelerator instance is virtualized to the virtual machine's control module, the control module can obtain description information of the current accelerator instance. This description information may indicate that the current accelerator as a whole is virtualized as an exclusive instance and assigned to the virtual machine. Step S440: Perform cold migration of the virtual machine. For example, the monitoring module can obtain the operating status of the virtual machine and send the operating status to the management node. The management node provides the operating status to the remote server with the backup accelerator and creates and starts the virtual machine on the remote server based on the operating status to implement virtual machine migration. For example, because a virtual machine has a significant dependency on the resources it accesses, if the current accelerator is not assigned to the virtual machine as an exclusive instance, for example, when multiple accelerator instances are assigned to the virtual machine, there may be software dependencies between the non-faulty accelerator and the virtual machine. Therefore, the virtual machine's tenant can be notified to shut down the virtual machine, thereby cold migrating the virtual machine to a remote server. This allows for more reliable device recovery through cold migration. Step S450: Determine whether a backup accelerator is installed in the local server. If so, the process proceeds to step S470; if not, the process proceeds to step S460. o For example, a query request for a backup accelerator can be sent to the management node. The management node determines whether a currently available backup accelerator exists on the local server from the accelerator installation list of each server. That is, after the accelerator's device driver is installed on the server, the server's operating system sends the accelerator's identifier and usage status to the management node via the control module. The management node associates the accelerator's identifier with the server's identifier to manage and maintain the installation and usage status of each accelerator. Step S460: Hot-unplug the current accelerator from the virtual machine, hot-migrate the virtual machine to the off-site server where the backup accelerator is located, and hot-plug the backup accelerator into the virtual machine. The process then proceeds to step S480. oFor example, the instance of the backup accelerator can be removed from the virtual machine to implement a hot-plug operation. It should be understood that after the instance of the backup accelerator is removed from the virtual machine, the virtual machine cannot access the backup accelerator, regardless of whether the instance driver for the backup accelerator is installed in the virtual machine. For another example, generating a discovery event for the virtual machine's exclusive use of the backup accelerator instance by the control module can simulate a hot-plug operation of the backup accelerator instance into the virtual machine for the control module. It should also be understood that the operation of migrating a virtual machine from a local server to a remote server can employ the methods described in the various embodiments above and will not be further described here. Step S470: Execute a hot-plug operation of the current accelerator from the virtual machine and a hot-plug operation of the backup accelerator into the virtual machine. The process then proceeds to step S480. For example, the instance of the backup accelerator can be removed from the virtual machine to implement a hot-plug operation. For another example, generating a discovery event for the virtual machine's exclusive use of the backup accelerator instance by the control module can simulate a hot-plug operation of the backup accelerator instance into the virtual machine for the control module. It should be understood that if the backup accelerator exists on the local server, there is no need to perform virtual machine migration. Step S480: Reload the instance driver of the backup accelerator. It should be understood that the method for loading the instance driver can be the same as described in the various embodiments above and will not be repeated here. Figure 5 is a schematic block diagram of a device exception handling apparatus according to other embodiments of the present disclosure. The device exception handling apparatus in Figure 5 includes: a status monitoring module 510, which monitors the device status of the current accelerator while the virtual machine on the local server is running. An instance creation module 520, which adds the backup accelerator installed on the local server as an exclusive instance to the virtual machine when the device status indicates an exception has occurred with the current accelerator. An instance driver module 530, which loads the instance driver of the backup accelerator into the virtual machine. In the solution of the embodiments of the present disclosure, adding the backup accelerator installed on the local server as an exclusive instance to the virtual machine enables the virtual machine to discover the exclusive instance of the backup accelerator. The instance driver of the backup accelerator is then loaded into the virtual machine, enabling the virtual machine to access the backup accelerator and automatically restoring the exclusive instance of the accelerator, thereby simplifying the process of restoring the virtual machine's operating status. In other embodiments, the instance creation module is specifically configured to: install a device driver for the standby accelerator in the local server into the operating system of the local server; and generate an exclusive instance of the standby accelerator through a virtualization component of the operating system, so that the control module of the virtual machine in the local server accesses the exclusive instance of the standby accelerator.In other embodiments, the instance driver module is specifically configured to: generate, via the control module, a discovery event indicating that the virtual machine has access to the exclusive instance of the backup accelerator; and, in response to the discovery event, load the instance driver for the backup accelerator into the operating system of the virtual machine. In other embodiments, the instance driver module is specifically configured to: update the current accelerator to the backup accelerator. In other embodiments, the device exception handling apparatus further comprises: an instance removal module configured to: assign the current accelerator as an exclusive instance to the virtual machine; and, when the device status indicates that the current accelerator has an exception, remove the exclusive instance of the current accelerator from the virtual machine. In other embodiments, the instance removal module is specifically configured to: generate, via the control module of the virtual machine in the local server, a removal event indicating that the virtual machine has access to the exclusive instance of the current accelerator; and, in response to the removal event, remove the exclusive instance of the current accelerator from the virtual machine. In other embodiments, the device exception handling apparatus further comprises: a device reset module configured to, via the processor of the local server, send a bus reset instruction to the current accelerator, wherein the bus reset instruction instructs the current accelerator to perform a reset operation. In other embodiments, the device exception handling apparatus further includes: an application startup module for launching an application in the virtual machine that accesses the backup accelerator. Figure 6 is a schematic block diagram of a device exception handling apparatus according to other embodiments of the present disclosure. The device exception handling apparatus in Figure 6 includes: a virtual machine migration module 610 for migrating the virtual machine of the remote server to the local server when the device status of the current accelerator on the remote server indicates that the current accelerator has experienced an exception. An instance creation module 620 for adding the backup accelerator installed in the local server as an exclusive instance to the virtual machine. An instance driver module 630 for loading the instance driver of the backup accelerator into the virtual machine. In this embodiment, when the device status of the current accelerator on the remote server indicates an abnormality, the virtual machine on the remote server is migrated to the local server, and the backup accelerator is added to the virtual machine as an exclusive instance, enabling the virtual machine to access the backup accelerator installed on the local server across servers. Furthermore, the instance driver for the backup accelerator is loaded into the virtual machine, enabling the virtual machine on the remote server to access the backup accelerator on the local server. Consequently, the exclusive instance of the accelerator on the remote server is automatically restored, simplifying the process of restoring the virtual machine's operating status. The specific implementation of each module in the apparatus can be found in the corresponding descriptions of the corresponding steps in the above-mentioned method embodiment, and corresponding beneficial effects are achieved, so detailed description is omitted here.Those skilled in the art will clearly understand that, for ease of description and brevity, the specific operating processes of the devices and modules described above can refer to the corresponding process descriptions in the aforementioned method embodiments and will not be repeated here. Referring to FIG7 , a schematic diagram of the structure of an electronic device according to another embodiment of the present disclosure is shown. The specific embodiments of the present disclosure do not limit the specific implementation of the electronic device. As shown in FIG7 , the electronic device may include: a processor 702 for executing a program 710, a communication interface 704, a memory 706, and a communication bus 708. The processor, communication interface, and memory communicate with each other via the communication bus. The communication interface is used to communicate with other electronic devices or servers. The processor is used to execute the program, specifically, to perform the relevant steps in the aforementioned method embodiments. Specifically, the program may include program code, which includes computer operating instructions. The processor may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present disclosure. The intelligent device includes one or more processors, which may be of the same type, such as one or more CPUs, or different types, such as one or more CPUs and one or more ASICs. A memory is used to store programs. The memory may include high-speed RAM memory or non-volatile memory, such as at least one disk drive. The program may include multiple computer instructions. Specifically, the program may cause the processor to perform operations corresponding to the device exception handling methods described in any of the aforementioned method embodiments. The specific implementation of each step in the program can be found in the corresponding descriptions of the steps and units in the aforementioned method embodiments, and corresponding beneficial effects are achieved, which are not described in detail here. Those skilled in the art will clearly understand that, for ease and brevity of description, the specific operating processes of the devices and modules described above can be found in the corresponding descriptions of the aforementioned method embodiments, which are not described in detail here. The presently disclosed embodiments also provide a computer storage medium storing a computer program, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments.The computer storage medium includes, but is not limited to, a compact disc (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk. The present disclosure also provides a computer program product comprising computer instructions that instruct a computing device to perform operations corresponding to each device exception handling method in the aforementioned multiple method embodiments. Furthermore, it should be noted that the user-related information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, sample data used for model training, data used for analysis, stored data, and displayed data, etc.) involved in the present disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with relevant regulations and standards, and corresponding operation portals are provided for the user to choose to authorize or reject. It should be noted that, depending on implementation needs, the various components / steps described in the present disclosure can be split into more components / steps, or two or more components / steps or partial operations of a component / step can be combined into a new component / step to achieve the objectives of the present disclosure. The above-mentioned method according to the embodiment of the present disclosure may be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium downloaded via a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)). It will be understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods described herein are implemented.Furthermore, when a general-purpose computer accesses the code for implementing the methods described herein, the execution of the code transforms the general-purpose computer into a special-purpose computer for executing the methods described herein. Those skilled in the art will appreciate that the units and method steps described in the various examples in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals skilled in the art may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present disclosure. The above embodiments are intended only to illustrate the embodiments of the present disclosure and are not intended to limit them. Persons skilled in the relevant technical fields may make various changes and modifications without departing from the spirit and scope of the embodiments of the present disclosure. Therefore, all equivalent technical solutions are also within the scope of the embodiments of the present disclosure, and the scope of patent protection for the embodiments of the present disclosure shall be defined by the claims.

Claims

Claims 1. A device exception handling method, comprising: Monitor the current accelerator device status while the virtual machine of the local server is running; When the device status indicates that the current accelerator is abnormal, the standby accelerator installed in the local server is added to the virtual machine as an exclusive instance; and the instance driver of the standby accelerator is loaded into the virtual machine.

2. The method according to claim 1, wherein: Adding the standby accelerator installed in the local server as an exclusive instance to the virtual machine includes: installing a device driver for the standby accelerator in the local server into an operating system of the local server; and generating an exclusive instance of the standby accelerator through a virtualization component of the operating system, so that a control module of the virtual machine in the local server accesses the exclusive instance of the standby accelerator.

3. The method according to claim 2, wherein: Loading the instance driver of the standby accelerator into the virtual machine includes: generating, by the control module, a discovery event of the virtual machine's exclusive instance of the standby accelerator; and loading the instance driver of the standby accelerator into the operating system of the virtual machine in response to the discovery event.

4. The method according to any one of claims 1 to 3, wherein: The method further includes: allocating the current accelerator as an exclusive instance to the virtual machine; and removing the exclusive instance of the current accelerator from the virtual machine when the device status indicates that an abnormality occurs in the current accelerator.

5. The method according to claim 4, wherein: Removing the exclusive instance of the current accelerator from the virtual machine includes: generating, by a control module of the virtual machine in the local server, a removal event of the virtual machine's exclusive instance of the current accelerator; and removing the exclusive instance of the current accelerator from the virtual machine in response to the removal event.

6. The method according to claim 4, wherein: The method further includes: sending a bus reset instruction to the current accelerator by a processor of the local server, wherein the bus reset instruction instructs the current accelerator to perform a reset operation.

7. The method according to any one of claims 1 to 6, wherein: The method further includes: starting an application program for accessing the standby accelerator in the virtual machine.

8. A device exception handling method, comprising: When the device status of the current accelerator of the remote server indicates that the current accelerator is abnormal, migrating the virtual machine of the remote server to the local server; The standby accelerator installed in the local server is added to the virtual machine as an exclusive instance; and an instance driver of the standby accelerator is loaded into the virtual machine.

9. An electronic device, comprising: Processor, memory, communication interface and communication bus, the processor, the The memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, where the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 8.

10. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

11. A computer program product, comprising a computer program / instruction, which implements the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Virtual machine migration method and system

    CN111736943A

  • Virtual machine migration method and device, upgrading method and server

    CN115599494A

  • Network card virtualization method based on DPU cross-card link aggregation

    CN116319303A

  • Method and device for accessing storage node and computer equipment

    CN116560785A

  • Graphic processing unit GPU scheduling method and device and storage medium

    CN117331704A