Failure processing method, device and storage medium
By setting up a pickup component on the bus device, receiving access requests instead of the bus device and generating failure response messages, the host and virtual machine downtime caused by bus device failure is solved, and the virtual machine is protected.
Patent Information
- Application Number
- PCT/IB2024/063167
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-03
AI Technical Summary
In the cloud computing field, when a bus device fails, the host cannot respond to non-delivered access requests, resulting in downtime, which leads to virtual machine exceptions, causing difficulties in operation and maintenance work.
Set up a pickup component on the bus device to receive access requests issued by the host instead of the bus device, and generate a response message representing the failure, and send it to the host to trigger a protection policy to prevent the host from crashing.
It effectively avoids the impact of bus device failure on the virtual machine, prevents the virtual machine from being directly downtime, and reduces the impact of failure on the virtual machine.
Smart Images

Figure IB2024063167_03072025_PF_FP_ABST
Abstract
Description
[0001] A Fault Handling Method, Device, and Storage Medium. This disclosure claims priority to Chinese patent application No. 202311873214.5, filed with the China Patent Office on December 29, 2023, entitled "A Fault Handling Method, Device, and Storage Medium," the entire contents of which are incorporated herein by reference. Technical Field: This disclosure relates to the field of cloud computing, and more particularly to a fault handling method, device, and storage medium. Background: Virtualization technology is widely used in the field of cloud computing. Based on virtualization technology, multiple virtual machines can be created on a host machine and managed using a virtual machine manager (VMM) or other means. Various bus devices (PCI / PCIE devices) are often required on a host machine. Furthermore, in many cases, the host machine needs to initiate non-posted access requests to the bus devices. Bus resources can only be released normally after receiving a response message from the bus device. Currently, if a bus device fails and is unable to respond to non-posted access requests, the host machine will crash due to the non-posted access requests timing out and not being responded to. This will also cause all virtual machines on the host machine to crash due to the host machine's downtime, causing all virtual machines to become abnormal and significantly complicating subsequent operations and maintenance. SUMMARY Various aspects of the present disclosure provide a fault handling method, device, and storage medium to reduce the impact of bus device failures on virtual machines. An embodiment of the present disclosure provides a fault handling method adapted for a proxy component provided on a bus device installed on a host machine. The method comprises: receiving access requests initiated by the host machine on behalf of the bus device when a preset type of failure occurs on the bus device; if a non-posted access request is received, generating a response message for the access request indicating that the bus device has failed; and sending the response message to the host machine to trigger the host machine to perform protection processing on the virtual machines running on it according to a preset protection policy. An embodiment of the present disclosure also provides a fault handling method, which is suitable for a host machine, a virtual machine running on the host machine, a bus device installed on the host machine, and a proxy component provided on the bus device. The method includes: initiating a non-delivered access request to the bus device in accordance with the bus protocol; if a response message returned by the proxy component is received, indicating that a fault has occurred in the bus device, protection processing is performed on the virtual machine in accordance with a preset protection strategy; wherein the response message is issued by the proxy component on behalf of the bus device when a preset type of fault occurs in the bus device.An embodiment of the present disclosure further provides a bus device having a proxy component, which is configured to execute one or more computer instructions to: receive an access request initiated by a host machine on behalf of the bus device in the event of a preset type of failure of the bus device; if a non-delivered access request is received, generate a response message for the access request indicating that a failure of the bus device has occurred; and send the response message to the host machine to trigger the host machine to perform protection processing on the virtual machine running on the host machine in accordance with a preset protection policy. Embodiments of the present disclosure also provide a host machine equipped with a bus device and a proxy component disposed on the bus device. The host machine includes a memory, a processor, and a communication component. The memory is configured to store one or more computer instructions. The processor is coupled to the memory and the communication component and configured to execute the one or more computer instructions, which are configured to: initiate a non-posted access request to the bus device according to a bus protocol; and upon receiving a response message returned by the proxy component indicating a bus device failure, perform protection processing on the virtual machine according to a preset protection policy. The response message is issued by the proxy component on behalf of the bus device when a preset type of bus device failure occurs. Embodiments of the present disclosure also provide a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the aforementioned fault handling method. Embodiments of the present disclosure also provide a computer program. When executed in a computer, the computer program causes the computer to execute the aforementioned fault handling method. In an embodiment of the present disclosure, a proxy component can be provided on a bus device. When a bus device fails and is unable to respond to non-posted access requests, the proxy component can identify a predefined type of failure, promptly receive non-posted access requests initiated by the host on behalf of the bus device, and generate a response message for such access requests, indicating that the bus device has failed. The proxy component can then send the response message to the host. In this way, when a bus device fails and is unable to respond to non-posted access requests, the proxy function provided by the proxy component can prevent host machine downtime. Triggered by this response message, the host can then protect the virtual machines running on it according to a pre-set protection policy. This prevents the virtual machines from directly downtime due to bus device failures, thereby reducing the impact of bus device failures on the virtual machines. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are provided to provide a further understanding of the present disclosure and constitute a part of this disclosure. The illustrative embodiments of this disclosure and their description are provided to explain the present disclosure and are not intended to unduly limit it.In the accompanying drawings: Figure 1 is a schematic diagram illustrating the assembly relationship between a bus device and a host machine according to an exemplary embodiment of the present disclosure; Figure 2 is a flowchart illustrating a fault handling method according to an exemplary embodiment of the present disclosure; Figure 3 is a logical diagram illustrating a fault handling solution for an exemplary bus device according to an exemplary embodiment of the present disclosure; Figure 4 is a flowchart illustrating a fault handling method according to another exemplary embodiment of the present disclosure; Figure 5 is a logical diagram illustrating a fault solution for an exemplary bus device on the host machine side according to another exemplary embodiment of the present disclosure; Figure 6 is a schematic diagram illustrating the structure of a bus device according to another exemplary embodiment of the present disclosure; and Figure 7 is a schematic diagram illustrating the structure of a host machine according to another exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS To further clarify the objectives, technical solutions, and advantages of the present disclosure, the technical solutions of the present disclosure will be described clearly and completely below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. It should be understood that the described embodiments are only a portion of the embodiments of the present disclosure, and are not exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. Before describing in detail the technical solutions provided by the various embodiments of the present disclosure, several technical concepts involved in the present disclosure are briefly described as follows.
[0002] PCI (Peripheral Component Interconnect) is a high-performance local bus, and PCIE (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus. Both PCI and PCIE can be called buses. They are both proposed to meet the needs of high-speed data transmission between peripherals and between external devices and the host. Bus device: It can be understood as an external device assembled on the host via a bus. Bus devices and the host communicate in accordance with the bus protocol. Bus protocol: It defines the interaction rules between the various communication parties communicating based on the bus. The interaction rules may include but are not limited to the internal structure of the messages used in the interaction process, as well as the execution process within each communication party. Bus transaction: The sequence of operations from requesting the bus to completing the use of the bus is called a bus transaction. Bus transactions can include at least two types: non-posted and posted. oPosted bus transactions: After a request is issued over the bus, bus resources can be released step by step without waiting for a response message. Non-posted bus transactions: After a request is issued, bus resources can only be released after receiving a corresponding response message. Currently, for non-posted bus transactions, a host machine issues a non-posted access request to a bus device mounted on it. If the bus device does not respond promptly, a Completion Timeout (CTO) exception is triggered on the host machine, causing the host machine to experience a kernel lock or kernel panic, resulting in a crash. This also causes the virtual machines running on the host machine to crash, potentially leading to data loss. Therefore, this embodiment proposes a fault handling solution to reduce the impact of bus device failures on virtual machines. The following, combined with the accompanying drawings, details the technical solutions provided by various embodiments of the present disclosure. Figure 1 is a schematic diagram of the assembly relationship between a bus device and a host machine, provided by an exemplary embodiment of the present disclosure. As shown in Figure 1, a bus device is mounted on a host computer via a bus (PCI or PCIE). Referring to Figure 1, this embodiment proposes that a proxy component be provided on the bus device. In practical applications, the proxy component can be added to an existing bus device; of course, a new bus device with a built-in proxy component can also be provided. Furthermore, this embodiment does not limit other components within the bus device. Based on these other components, the bus device can provide a variety of device functions, and specific examples of these functions are not provided here. In this embodiment, the proxy component can provide a proxy function in the event of a pre-defined type of failure on the bus device. As mentioned above, after a host computer issues a non-delivered access request to a bus device mounted on it, it waits for the bus device to return a response message. A non-posted access request can be understood as a transaction layer message initiated by a host machine, carrying a non-posted transaction type. Under normal circumstances, upon receiving such a transaction layer message, a bus device can determine that it has received a non-posted access request based on the transaction type carried within the message. The bus device can then generate a response message and return it to the host machine in accordance with the processing rules for non-posted access requests defined in the bus protocol. Therefore, in this embodiment, the preset fault types may include, but are not limited to, kernel deadlocks, kernel errors, CPU exceptions, or memory exceptions that cause the bus device to be unable to respond to non-posted access requests. The list of preset fault types is not exhaustive.That is, in this embodiment, if a bus device is unable to respond to a non-delivered access request initiated by a host machine due to a fault, the proxy component configured on the bus device can be promptly activated. In this embodiment, the implementation of the proxy component is not limited. The proxy component can be implemented as software, such as an independent operating system, or as hardware, such as a field programmable gate array (FPGA) or an independent chip. Further implementation examples are not provided here; any implementation that enables the proxy component to function normally after a bus device experiences a predetermined type of fault is acceptable. Thus, in this embodiment, the independence of the proxy component ensures that the proxy component can function normally after a bus device fails. For different implementations, this embodiment can employ appropriate triggering methods to activate the proxy component. If the proxy component is implemented as an independent operating system, the proxy component can monitor the heartbeat of the bus device. If an abnormal heartbeat is detected on the bus device, the default type of fault is determined to have occurred on the bus device. The general principle of heartbeat monitoring is as follows: a heartbeat monitoring channel is established between the answering component and a bus device. The bus device proactively and regularly sends heartbeat packets to the answering component through this channel. If the answering component detects that the bus device has not sent a heartbeat packet on time, it can determine that the bus device's heartbeat is abnormal. Alternatively, the answering component can send a heartbeat detection packet to the bus device. If it does not receive an acknowledgment packet from the bus device on time, it can determine that the bus device's heartbeat is abnormal. For example, the answering component can define a return delay for the acknowledgment packet, such as 1 second. If the answering component does not receive an acknowledgment packet from the bus device within 1 second after sending the heartbeat detection packet, it can determine that the bus device's heartbeat is abnormal. The heartbeat monitoring principle described here is illustrative, and this embodiment is not limited to this. Under this triggering mode, if a bus device experiences a predetermined fault, it will typically be unable to send heartbeat packets or return acknowledgment packets on time. Therefore, the answering component can proactively detect whether a predetermined fault has occurred on the bus device by monitoring its heartbeat. If the proxy component is implemented as independent hardware, the bus device can activate the proxy component when a preset fault type occurs. In this case, the activation of the proxy component automatically indicates that the bus device has experienced a preset fault type. The inventors discovered during their research that a preset fault type typically triggers the bus device to initiate fault handling logic, such as MCE (Machine Check Except) processing logic.To this end, this embodiment proposes that, under this triggering method, the proxy component can be promptly activated by modifying the fault handling logic in the bus device. In one exemplary modification scheme, proxy component activation logic can be added to the fault handling logic in the bus device. Thus, when a predetermined fault occurs on the bus device, the bus device will enter the modified fault handling logic. During execution of the modified fault handling logic, the proxy component can be activated by sending a startup signal to the proxy component according to the proxy component activation logic. The activation signal can be an electrical signal or a command signal. For example, if the proxy component is implemented as an FPGA, an electrical signal can be sent to a startup pin on the proxy component. Alternatively, if the proxy component is implemented as a standalone chip with its own microprocessor, a command signal can be sent to the microprocessor in the proxy component to trigger the proxy component to activate. These triggering methods are merely exemplary and are not limited to these methods in this embodiment. Further examples of triggering methods will not be provided here. Based on this, when a predetermined fault occurs on the bus device, the proxy component will be promptly activated. In practical applications, the startup time of the proxy component can be sufficiently short. The startup time of the waiting component plus its response time to non-posted access requests can be at least shorter than the response waiting time specified in the bus protocol for non-posted access requests. This effectively ensures that any non-posted access requests issued by the host machine to the bus device are not missed, thereby preventing host machine crashes caused by missed requests. Furthermore, referring to Figure 1 , in this embodiment, a virtual machine runs on the host machine. In virtualization technology, a virtual machine manager (VMM) is typically used to create and manage virtual machines. A virtual machine manager typically includes user-mode components and kernel-mode components. User-state components are typically used to handle user-state-related tasks, such as the Quick Emulator (QEMU) component; while kernel-state components are typically used to handle kernel-state-related tasks, such as the Kernel-based Virtual Machine (KVM) component. The tasks handled by user-state components may include, but are not limited to, creating a virtual machine, allocating addresses from the virtual address space occupied by the virtual machine manager as the virtual machine's physical address, and simulating required virtual devices for the virtual machine, etc., and these are not exhaustive here.The kernel-mode component provides a series of interfaces to the user-mode component, allowing the user-mode component to control various aspects of the virtual machine, such as the number of CPUs, memory layout, and operation and shutdown. The kernel-mode component also processes privileged instructions issued by the virtual machine. Privileged instructions are typically those that may affect the entire host machine, such as various I / O requests issued by the virtual machine. Further background on the host machine and virtual machines will not be provided here. FIG2 is a flowchart illustrating a fault handling method provided in an exemplary embodiment of the present disclosure. Referring to FIG2 , this method can be applied to a proxy component provided in a bus device. The method may include: Step 100: When a bus device experiences a fault of a preset type, receiving an access request initiated by the host machine on behalf of the bus device; Step 101: If a non-posted access request is received, generating a response message for the access request indicating that the bus device has experienced a fault; Step 102: Sending the response message to the host machine to trigger the host machine to perform protection processing on the virtual machine running on it according to a preset protection policy. In step 100, the proxy component can be activated promptly when a bus device experiences a predetermined type of failure, and receive access requests initiated by the host on behalf of the bus device. It should be understood that bus device failures can be of various types. In this embodiment, the proxy component can be activated promptly when a bus device experiences a failure that prevents the bus device from responding to non-delivered access requests. For other types of failures, which typically do not cause host downtime, activation of the proxy component is unnecessary in this embodiment. For details on how to activate the proxy component, please refer to the previous description. As mentioned above, in this embodiment, activation of the proxy component can be achieved using a triggering method tailored to the implementation of the proxy component, and this description will not be repeated here. In this embodiment, the proxy component can monitor the interface for receiving access requests in the bus device's communication component to receive access requests initiated by the host. Of course, other methods can also be used in this embodiment to enable the proxy component to receive access requests initiated by the host to the bus device, and this is not limited to this method. Further examples are not provided here. Furthermore, because the proxy component takes time to start up, in some special cases, the host machine may initiate an access request to the bus device between the time a predetermined type of bus device failure occurs and the time the proxy component starts up. To address this issue, in this embodiment, the proxy component can determine whether there are any unprocessed access requests by examining the access request description information cached under the bus device's interface for receiving access requests. If so, it treats these as access requests received by the proxy component. This effectively avoids missed access requests.Continuing with Figure 2, in step 101, if a non-posted access request is received, a response message is generated for the access request, indicating a bus device failure. As mentioned above, both non-posted and posted access requests may occur between the host and the bus device. Given that posted access requests are not affected by bus device failures, in this embodiment, the proxy component does not need to process captured posted access requests. However, for non-posted access requests, this embodiment requires the proxy component to generate a response message for each captured non-posted access request, indicating a bus device failure. During research, the inventors discovered that the request header for an access request from a host machine to a bus device typically includes a request type field (Type field), which indicates the request type (Non-posted or Posted). Preferably, in this embodiment, the proxy-answerable component can generate a response message according to the bus protocol to indicate the completion of the access request, and mark the bus device's fault status in the response message. In the bus protocol, a response message indicating the completion of the access request is also called a completion message (CPL). Its message header may include a completion status field (i.e., Completion Status field), a requester ID field (i.e., Requster ID field), and a completer ID field (i.e., Completer ID field). The completion status field can be used to indicate whether the corresponding access request has been successfully completed; the requester ID field can be used to indicate the identity of the initiator of the access request; and the completer ID field can be used to indicate the identity of the responder that responded to the access request. It should be understood that the aforementioned fields are merely exemplary, and the response message may also include other fields to indicate other aspects of information. In this embodiment, the proxy module may define the completion status field in the response message as SC (i.e., successful completion) to indicate that the corresponding access request has been completed. In this embodiment, the response message generated by the proxy component is used not only to indicate that a non-delivered access request has been completed but also to indicate that a bus device has failed. In this embodiment, various implementations can be used to indicate a bus device failure based on the response message. In one optional implementation, the failure status of the bus device can be indicated in the response message.A fault state refers to an abnormal state of a bus device. This optional implementation proposes that the fault state can be marked using a marking dimension such as the error type. During research, the inventors discovered that, in addition to the aforementioned completion status field and requester identification field, the response message also contains a field for marking the error type. Based on this, this optional implementation proposes that the fault state of a bus device can be marked by setting the target field in the response message to a preset indicator. The target field can be any field in the response message that can be used to mark the error type. In one exemplary solution, the field in the response message used to mark a data error (i.e., Error Poisoned (EP)) can be set to a preset indicator, such as 1, to mark the fault state of the bus device. The EP field is a field in the response message used to mark the error type. The error type identified by the EP field is Error Forwarding / Data Poisoning. Of course, this is merely exemplary. The response message also includes other fields for identifying error types, though different fields may identify different error types. Based on this, the proxy component can also assign values to other fields to indicate other error types. Further examples are not provided here. Thus, in this embodiment, by assigning a value to the target field in the response message, the error type can be marked as the fault status of the bus device, thereby indirectly indicating that a bus device failure has occurred. It should be understood that this is merely a preferred embodiment. This embodiment supports the proxy component constructing a response message using other custom formats, ensuring that the response message can indicate that the corresponding access request has been responded to and that a bus device failure has occurred. If a message format that does not conform to the bus protocol is used, the host and the proxy component must pre-agreed on the response message format so that the host can parse and recognize the response message according to the agreement. This embodiment does not limit the format of the response message. Referring to Figure 2, in step 102, the proxy component may send the response message to the host. In practical applications, the proxy component can reuse the bus channel between the bus device and the host to send the response message to the host. First, in this embodiment, the proxy component generates a response message for the non-delivered access request it receives and returns it to the host. This ensures that the non-delivered access request it issued has been responded to, thus preventing host downtime.Secondly, in this embodiment, the response message generated by the proxy component also indicates that a bus device failure has occurred. This allows the host machine to promptly detect the bus device failure through the response message. Thus, in this embodiment, after a bus device failure of a preset type occurs, all non-posted access requests sent by the host machine to the bus device can be received by the proxy component. As mentioned above, the proxy component's startup time plus its response time to non-posted access requests is shorter than the response wait time specified in the bus protocol for non-posted access requests. Therefore, non-posted access requests arriving at the bus device after a preset type of failure occurs can be received and responded to by the proxy component without being missed. Furthermore, the response message generated by the proxy component can reach the host machine before the aforementioned response wait time expires, thereby preventing host machine downtime. During their research, the inventors discovered that most non-posted access requests initiated by a host machine to a bus device are caused by the virtual machines running on it. Although the proxy component responds to non-posted access requests on behalf of the bus device, this response can prevent host machine downtime. However, for the virtual machine, the original request initiated by the virtual machine is not properly responded to, and the virtual machine may still experience access anomalies. Therefore, when a bus device experiences a predetermined type of failure, while the proxy component prevents host machine downtime, it may still cause anomalies in the associated virtual machine. To address this, this embodiment also proposes pre-setting a protection policy in the host machine. Based on this, after receiving a response message, the host machine can perform protection processing on the virtual machine running on it according to the pre-set protection policy to prevent further anomalies or errors. This embodiment supports the on-demand setting of relevant protection policies in the host machine to better protect the virtual machine. The pre-set protection policies in the host machine are not limited here and will be discussed in detail in the host-side embodiments. In summary, this embodiment proposes that a proxy component can be set up on a bus device. When a bus device fails and is unable to respond to non-posted access requests, the proxy component can promptly receive non-posted access requests initiated by the host machine on behalf of the bus device based on a preset type of failure. The proxy component generates a response message for this type of access request, indicating that the bus device has failed. The proxy component can then send the response message to the host machine.In this way, if a bus device fails and is unable to respond to non-posted access requests, the proxy function provided by the proxy component can prevent host machine crashes. Triggered by this response message, the host machine can protect the virtual machines running on it according to a pre-set protection policy, preventing the virtual machines from directly crashing due to bus device failures. This reduces the impact of bus device failures on the virtual machines. As mentioned above, in the above-mentioned or following embodiments, the device functions provided by the bus device are not limited. That is, the fault handling method provided in this embodiment is applicable to various types of bus devices, particularly bus devices that may need to process non-posted access requests issued by the host. For example, various data processing units (DPUs) are examples. Figure 3 is a logical diagram of a fault handling solution for an exemplary bus device, provided in an exemplary embodiment of the present disclosure. Referring to Figure 3, the exemplary bus device runs an operating system, and within the operating system runs a user-mode component of a virtual machine manager, such as the QEMU component mentioned above. The kernel-mode components of the virtual machine manager, such as the KVM component mentioned above, run on the host machine. In addition, the bus device also has shared memory. Shared memory is the most efficient inter-process communication method. Processes can directly read and write shared memory without requiring any data copying between processes. In this embodiment, the shared memory in the bus device can be mapped to the host machine's memory address space. This allows the kernel-mode component in the host machine to initiate communication requests (essentially read / write requests to the shared memory) to the user-mode component in the bus device based on the host machine's memory address. Since the shared memory is mapped to the host machine's memory address space, information related to the communication request can be written to the shared memory in the bus device, allowing the user-mode component in the bus device to receive the communication request. Similarly, the user-mode component in the bus device can also initiate communication requests to the kernel-mode component in the host machine based on the shared memory. Therefore, in this embodiment, the interaction between the kernel-mode component in the host machine and the user-mode component in the bus device can be viewed as two processes, which can communicate with each other through the shared memory on the bus device. The exemplary bus device shown in Figure 3 can also be referred to as a Cloud Infrastructure Processing Unit (CIPU). This device is a chip designed specifically for cloud computing, enabling efficient management and coordinated acceleration of multi-dimensional resources such as computing, storage, networking, and security.The CI PU not only connects to physical computing, storage, and network resources, enabling rapid cloudification of hardware resources, but also connects to distributed platforms, enabling flexible management, scheduling, and orchestration of cloud resources. The CI PU can be combined with CPUs, GPUs, FPGAs, and other chips in the host machine to form a virtualized architecture that integrates hardware and software, achieving deep system-level integration. It should be understood that CI PU is merely an exemplary term for this type of bus device. Other terms for this type of bus device exist in the cloud computing field, such as Infrastructure Processing Unit (IPU), and more examples are not provided here. Referring to Figure 3, this exemplary bus device hosts the user-mode components of the virtual machine manager. Thus, the user-mode components of the bus device and the kernel-mode components of the host machine can cooperate to create and manage virtual machines on the host machine. Based on this, referring to Figure 3, after a virtual machine running on a host machine issues an I / O request, it will first be captured by the kernel-mode component of the host machine. Since I / O requests issued by the host machine require cooperation between the kernel-mode and user-mode components, the kernel-mode component needs to issue a read / write request for shared memory to the bus device in response to the captured I / O request. This read / write request is a typical non-posted access request occurring on the host machine. I / O requests issued by a virtual machine are typically used to access external devices on the host machine or other virtual machines, and will not be explained in detail here. I / O requests issued by a virtual machine may include, but are not limited to, port-mapped I / O requests (programmed I / O, PI0) and memory-mapped I / O requests (MM0). For example, when a virtual machine needs to use the GPU installed on the host machine, it can initiate an MM IO request. This MM IO request will be received by the kernel-mode component and forwarded to the user-mode component. The user-mode component can then perform processing logic such as address translation to ensure that the request reaches the hardware resource of the GPU installed on the host machine, thereby completing the hardware processing of the request. IO requests issued by virtual machines are common requests in virtualization technology and will not be further explained here. It should be understood that the read / write request for shared memory in the bus device issued by the kernel-mode component in Figure 3 is merely an exemplary non-posted access request, and the initiator of this non-posted access request is the kernel-mode component of the virtual machine manager in the host machine.Referring to Table 3, from the perspective of this exemplary bus device, it can manage hardware resources and virtual machines on the host machine. From the perspective of the host machine, however, this exemplary bus device is essentially still a bus device, and communication between the host machine and the bus device continues according to the bus protocol. In traditional fault handling solutions (i.e., when the bus device does not include the proxy module provided in this embodiment), a fault of a predetermined type in the bus device can cause the host machine to crash due to the aforementioned completion timeout (CTO) issue. Continuing with the description of the fault handling solution in the proxy component in the previous embodiment, when the exemplary bus device in Figure 3 encounters a predetermined type of fault, the proxy component therein can receive the aforementioned read / write requests for shared memory in the bus device issued by the kernel-mode component. Furthermore, the proxy component can generate response messages for these non-posted access requests, ensuring that the host machine receives the corresponding response messages, thereby preventing these non-posted access requests from causing host machine crashes due to bus device failures. Accordingly, for the exemplary bus devices shown in FIG3 , the fault handling solution provided in this embodiment can be used, through the proxy component provided on such exemplary bus devices, to effectively prevent host machine downtime caused by failures of such exemplary bus devices, thereby reducing the impact of such exemplary bus device failures on virtual machines running on the host machines. It is worth noting that the bus devices shown in FIG3 are merely exemplary. This embodiment does not limit the type of bus devices. This embodiment supports the application of the fault handling solution provided in this embodiment to various bus devices to minimize host machine downtime caused by bus device failures, thereby improving host machine stability and protecting virtual machines. FIG4 is a flowchart illustrating a fault handling method provided in another exemplary embodiment of the present disclosure. Referring to FIG4 , this method is applicable to a host machine running virtual machines, equipped with a bus device, and equipped with a proxy component. Referring to Figure 4 , the method may include: Step 400: Initiating a non-delivered access request to a bus device according to a bus protocol; Step 401: Upon receiving a response message from a proxy component indicating a bus device failure, performing protection processing on the virtual machine according to a preset protection policy. The response message is issued by the proxy component on behalf of the bus device when a preset type of bus device failure occurs. In this embodiment, the host machine may initiate a non-delivered access request to the bus device according to the bus protocol.During their research, the inventors discovered that there are many reasons why a host machine may need to initiate a non-posted access request to a bus device, including but not limited to the need to read or write memory, I / O, or configuration data to the bus device. These reasons are not exhaustive here. Regardless of the reason for the host machine's need to initiate a non-posted access request to the bus device, as mentioned above, if the bus device experiences a predetermined type of failure, preventing it from promptly responding to the host machine's non-posted access request, a completion timeout (CTO) exception will be triggered in the host machine, causing the host machine to experience a kernel lock or kernel error, leading to a crash. This will also cause the virtual machine running on the host machine to crash, potentially leading to data loss. Referring to Figure 4 , in this embodiment, after the host machine initiates a non-posted access request to the bus device, it will wait for a corresponding response message. Based on the fault handling solution provided in this embodiment, a proxy component provided on the bus device can, in the event of a predetermined type of bus device failure, generate a response message on behalf of the bus device for the non-posted access request initiated by the host machine. In this way, in this embodiment, after a bus device failure occurs, the host machine can receive the response message generated by the proxy component configured on the bus device in response to a non-posted access request. Because the host machine receives the response message returned to the non-posted access request it initiated, a CTO exception is no longer triggered in the host machine, preventing a kernel lockup or kernel error, and thus preventing the host machine from crashing. In other words, when a bus device fails and is unable to respond to a non-posted access request initiated by the host machine, the reply function provided by the proxy component effectively prevents host machine crashes. In this embodiment, the host machine does not crash in the event of a bus device failure, thus providing the host machine with an opportunity to protect the virtual machines running on it, preventing further exceptions or errors. To this end, this embodiment proposes pre-setting a protection policy in the host machine. Based on this, referring to Figure 4 , in step 401, after receiving a response message from the proxy component indicating a bus device failure, the host machine can protect the virtual machine according to the pre-set protection policy. In this embodiment, the pre-set protection policy can be flexibly set and updated as needed. The policy content within the pre-set protection policy is not limited; any effective policy that can reasonably prevent further anomalies or errors in the virtual machine can be used. In one exemplary pre-set protection policy, after receiving a response message from the proxy component to a non-delivered access request, the host machine can determine the initiator of the non-delivered access request. Based on this, the host machine can then transmit the response message to the initiator, triggering the initiator to protect the virtual machine associated with the non-delivered access request.The virtual machines associated with non-posted access requests can be understood as virtual machines that may be affected by a bus device failure. Therefore, this embodiment proposes that logic for processing such response messages be implemented in each initiator that may initiate a non-posted access request on a host machine. This processing logic can be used to identify the virtual machines that may be affected by a bus device failure and provide timely protection for these virtual machines. During research, the inventors discovered that the initiators of non-posted access requests on a host machine can be diverse, including, for example, kernel-mode components within a virtual machine manager, virtual machines, or processes within the host operating system. This embodiment does not exhaustively list these initiators. Furthermore, the processing logic for response messages returned by the proxy component within these initiators is not limited. The processing logic within an exemplary initiator will be described in detail later. In summary, this embodiment proposes that a proxy component be provided on a bus device. If a bus device experiences a predetermined type of failure, the proxy component can receive non-delivered access requests initiated by the host on behalf of the bus device and generate a response message for such access requests, indicating that the bus device has failed. The proxy component can then send the response message to the host. In this way, if a bus device fails and is unable to respond to non-delivered access requests, the proxy function provided by the proxy component can prevent host machine downtime. Triggered by this response message, the host can then protect the virtual machines running on it according to a pre-set protection policy, preventing the virtual machines from directly crashing due to bus device failures and thus reducing the impact of bus device failures on the virtual machines. Figure 5 is a logical diagram of a host-side fault resolution solution for an exemplary bus device, provided by another exemplary embodiment of the present disclosure. Referring to Figure 5 , the exemplary bus device runs an operating system, which includes a user-mode component of a virtual machine manager, such as the QEMU component mentioned above. The kernel-mode component of the virtual machine manager, such as the KVM component mentioned above, runs on the host. Based on this, kernel-mode components can interact with user-mode components via shared memory on a bus device. For the exemplary principles of communication between kernel-mode and user-mode components based on shared memory, please refer to the previous description and will not be repeated here. With reference to FIG5 , if the exemplary bus device experiences a predetermined type of failure and is unable to respond to a non-posted access request, the kernel-mode component in the host machine can initiate a communication request to the user-mode component in the bus device. Specifically, this can be a read / write request to the shared memory in the bus device. This read / write request is a non-posted access request.After the non-posted access request is sent to the bus device, the proxy component will provide a response message because the bus device has failed. Based on this, if the host machine determines that the initiator of the non-posted access request is a kernel-mode component, the response message will be passed to the kernel-mode component. The kernel-mode component will then be used to shut down the virtual machine running on the host machine. In other words, in this embodiment, an exemplary initiator of the non-posted access request to the bus device is a kernel-mode component. In this case, the host machine can pass the response message returned by the proxy component to the kernel-mode component. For a kernel-mode component, the virtual machines associated with the non-posted access request can be all virtual machines running on the host machine. As mentioned above, in this embodiment, processing logic for such response messages can be configured in the kernel-mode component. In this case, an exemplary processing logic for this type of response message in the kernel-mode component may be: after receiving the response message, parse the response message; if it is determined that the target field in the response message is set to a preset indicator, invoke a function for controlling virtual machine shutdown to shut down the virtual machine running on the host machine. As mentioned in the previous introduction to the kernel-mode component's functions, the kernel-mode component can be used to control the operation and shutdown of virtual machines. The kernel-mode component includes a function (API) for controlling virtual machine shutdown, which the kernel-mode component can call to shut down the virtual machine. This function is an existing function in the kernel-mode component and will not be described in detail here. During research, the inventors discovered that, for the exemplary bus device shown in Figure 5, the reason why the kernel-mode component in the host machine needs to initiate a non-posted access request to the bus device is typically due to an I0 request issued by the virtual machine on the host machine. The process of cooperation between the kernel-mode component and the user-mode component after the virtual machine issues an I0 request can be found in the previous description and will not be repeated here. Thus, on one hand, as described in the aforementioned embodiment, although the response message may indicate that the read / write request for the shared memory has been responded to, it also indicates that the exemplary bus device has failed. For example, the message header of the exemplary response message in the aforementioned embodiment defines SC+EPo. Accordingly, the result data required by the virtual machine's 10 request is not carried in the response message.Therefore, although the proxy component returns a response message on behalf of the bus device to the kernel-mode component's non-delivered access request, this response message does not correctly address the I / O request initiated by the virtual machine. Therefore, from the perspective of the virtual machine initiating the I / O request, its I / O request will still experience a response exception. However, this response exception will not cause a system crash and may only be an anomaly at the data or configuration level. Furthermore, the non-delivered access request responded to by the proxy component may only cause I / O request exceptions for a single or limited number of virtual machines, while other virtual machines will be unaware of the bus device anomaly. If other virtual machines subsequently initiate I / O requests, the kernel-mode component will initiate the same non-delivered access request to the bus device, causing I / O exceptions in more virtual machines, and the scope of the anomaly will continue to expand. Therefore, in the aforementioned response message processing logic implemented in the kernel-mode component, after determining that the target field in the response message is set to a preset indicator, the host machine running on the host machine is shut down. This effectively protects virtual machines that have not experienced any exceptions, while preventing further exceptions from occurring in virtual machines that have already experienced any exceptions and thus effectively protecting them. This effectively prevents further exceptions or errors in the virtual machines. Furthermore, in this case, in addition to using the kernel-mode component to shut down the virtual machines running on it, the host machine can also use the kernel-mode component to check whether any virtual machines running on the host machine have any pending I / O requests, and classify the virtual machines based on the presence of pending I / O requests. The kernel-mode component records the processing status information of each I / O request initiated by the virtual machine. Based on this information, the kernel-mode component can seamlessly detect whether any virtual machines running on the host machine have any pending I / O requests. Here, a pending I / O request can be understood as an I / O request that the virtual machine has issued but has not received any response data. It is understandable that the IO request corresponding to the non-delivered request responded to by the proxy component is an incomplete IO request. This is because the response message returned by the proxy component does not carry the response data required to respond to the IO request. Of course, IO requests issued by a virtual machine but not reaching the proxy component will also be determined to be incomplete IO requests. This processing status information is recorded in detail by the kernel-mode component. Therefore, the kernel-mode component can accurately identify IO requests issued by virtual machines with incomplete responses and classify virtual machines accordingly. An exemplary classification marking scheme may be to mark virtual machines with incomplete IO requests as incomplete, and to mark virtual machines without incomplete IO requests as complete.This further optimization solution proposes classifying virtual machines. It should be understood that this classification primarily serves as a reference for subsequent operations and maintenance. Based on the classification and labeling of virtual machines, this provides more operational basis for subsequent operations and maintenance, enabling differentiated operations and maintenance for virtual machines with different classification labels. This embodiment does not impose any restrictions on the subsequent operations and maintenance steps. For example, only virtual machines marked as complete may be migrated in subsequent operations and maintenance steps, which will not be further explained here. In summary, in this embodiment, the host machine can provide a fault handling solution for the exemplary bus device in Figure 5. This solution utilizes the proxy function provided by the proxy component in this exemplary bus device to avoid triggering a CTO exception, thereby preventing host machine downtime. Furthermore, by configuring processing logic for response messages returned by the proxy component in the kernel-mode component of the host machine, the host machine running on the host machine can be promptly shut down. This prevents other virtual machines on the host machine from receiving 10 requests, effectively preventing other virtual machines from experiencing exceptions due to the exemplary bus device failure. This prevents the failure of this exemplary bus device from affecting more virtual machines and preventing the scope of the exception from expanding, thereby better protecting the virtual machines. Furthermore, the classification tags configured for each virtual machine can be used as a reference in subsequent operations and maintenance, facilitating differentiated operations and maintenance for different types of virtual machines. In this embodiment, in addition to the aforementioned virtual machine initiating the request, another reason for a kernel-mode component to initiate a non-posted access request to a bus device may be that the kernel-mode component in the host machine needs to configure the user-mode component in the bus device. Such non-posted access requests may affect all virtual machines running on the host machine. Therefore, in response to such non-posted access requests, the kernel-mode component may also shut down the host machine after receiving the response message. Furthermore, in this embodiment, the kernel-mode component is used as the initiator of the non-posted access request. However, it should be understood that in this embodiment, the initiator of the non-posted access request is not limited to this. Different types of response message processing logic may also be configured for other initiators. For example, when accessing a pass-through device on a host machine, a virtual machine may initiate a direct memory access (DMA) request to the pass-through device. This is also a non-posted request. In this case, the initiator of the non-posted request is the virtual machine. An exemplary processing logic within the virtual machine for the response message may be to, upon determining that the target field in the response message is set to a preset indicator, request a shutdown from the kernel-mode component, so that the kernel-mode component can complete the shutdown. In this case, only the virtual machine can be shut down without shutting down other virtual machines.The processing logic for the response message within the initiator is not further described here; it is merely an example, and this embodiment is not limited thereto. It should be noted that some of the processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be understood that these operations may be executed in a different order than the order in which they appear herein or in parallel. Operation numbers, such as 101 and 102, are merely used to distinguish between different operations and do not represent any specific execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. Figure 6 is a schematic diagram of the structure of a bus device provided in another exemplary embodiment of the present disclosure. As shown in Figure 6, the bus device is provided with a proxy component 60, which is independent of the bus device. The bus device may be installed on a host computer, which in turn runs a virtual machine. The proxy component may execute one or more computer instructions to: receive access requests initiated by the host computer on behalf of the bus device when a preset type of fault occurs on the bus device; if a non-delivered access request is received, generate a response message indicating a bus device fault; and send the response message to the host computer to trigger the host computer to perform protection processing on the virtual machine running on the bus device according to a preset protection policy. In an optional embodiment, when generating a response message indicating a bus device fault for an access request, the proxy component 60 may specifically: generate a response message indicating that a response to the access request has been completed according to the bus protocol; and mark the bus device fault status in the response message. In an optional embodiment, when marking the bus device fault status in the response message, the proxy component 60 may specifically: set the target field in the response message to a preset indicator to indicate the bus device fault status; wherein the target field is any field in the response message used to mark the error type. In an optional embodiment, the proxy component 60 is implemented as an independent operating system and can also be used to: monitor the heartbeat of a bus device; if an abnormal heartbeat is detected on the bus device, determine that a predetermined fault type has occurred on the bus device; or, alternatively, the proxy component 60 is implemented as independent hardware, and the bus device activates the proxy component 60 when a predetermined fault type occurs; after activation, the proxy component 60 determines that a predetermined fault type has occurred on the bus device. In an optional embodiment, proxy component activation logic is added to the fault handling logic triggered by the bus device when a predetermined fault type occurs. During the execution of the fault handling logic, the bus device sends a startup signal to the proxy component 60 according to the proxy component activation logic to activate the proxy component 60.In an optional embodiment, a user-mode component of a virtual machine manager runs on the bus device, while the kernel-mode component of the virtual machine manager runs on the host machine. The kernel-mode component interacts with the user-mode component through shared memory on the bus device, and non-posted access requests are read / write requests to the shared memory initiated by the kernel-mode component. It is worth noting that the technical details of the aforementioned bus device embodiments can be found in the description of the aforementioned method embodiments on the proxy component side. To save space, these details are not repeated here, but this should not compromise the scope of protection of the present disclosure. Figure 7 is a schematic diagram of the structure of a host machine provided by another exemplary embodiment of the present disclosure. Referring to Figure 7, the host machine includes a memory 70, a processor 71, and a communication component 72. Referring to Figure 7, the host machine is equipped with a bus device 73, which is provided with a proxy component. Processor 71 is coupled to memory 70, communication component 72, and bus device 73 via a bus, and is configured to execute a computer program in memory 70 to: initiate a non-delivered access request to a bus device according to a bus protocol; and upon receiving a response message from a proxy component indicating a bus device failure, perform protection processing on the virtual machine according to a preset protection policy. The response message is sent by the proxy component on behalf of the bus device when a preset type of bus device failure occurs. In an optional embodiment, when performing protection processing on the virtual machine according to the preset protection policy, processor 71 may specifically: determine, within the host machine, the initiator of the non-delivered access request; and transmit the response message to the initiator, thereby triggering the initiator to perform protection on the virtual machine associated with the non-delivered access request. In an optional embodiment, a user-mode component of a virtual machine manager runs on the bus device, and the user-mode component interacts with a kernel-mode component of the virtual machine manager running on the host machine through shared memory on the bus device. When the processor 71 transmits a response message to the initiator to trigger the initiator to protect the virtual machine associated with the non-posted access request, the processor 71 may be specifically configured to: if the initiator corresponding to the non-posted access request is a kernel-mode component, transmit the response message to the kernel-mode component; and use the kernel-mode component to shut down the virtual machine running on the host machine. In an optional embodiment, when the processor 71 uses the kernel-mode component to shut down the virtual machine running on the host machine, the processor 71 may be specifically configured to: upon receiving the response message, the kernel-mode component parses the response message; and if it is determined that the target field in the response message is set to a preset indicator, the kernel-mode component invokes a function for controlling virtual machine shutdown to shut down the virtual machine running on the host machine.In an optional embodiment, processor 71 may also be configured to: utilize kernel-mode components to check whether any virtual machines running on the host machine have any uncompleted I / O requests; and classify the virtual machines based on whether any uncompleted I / O requests exist. In an optional embodiment, when classifying the virtual machines based on whether any uncompleted I / O requests exist, processor 71 may specifically be configured to: mark virtual machines with uncompleted I / O requests as being in an incomplete state; and mark virtual machines without uncompleted I / O requests as being in a complete state. Furthermore, as shown in FIG7 , the host machine also includes other components, such as a power supply component 74. FIG7 only schematically illustrates some components and does not imply that the host machine only includes the components shown in FIG7 . It is worth noting that the technical details of the aforementioned host machine embodiments can be found in the relevant descriptions of the aforementioned method embodiments. To save space, these details will not be repeated here, but this should not compromise the scope of protection of the present disclosure. Accordingly, embodiments of the present disclosure also provide a computer-readable storage medium storing computer instructions. When executed by one or more processors, the computer instructions cause the one or more processors to perform the steps of the aforementioned method embodiments. Accordingly, embodiments of the present disclosure also provide a computer program. When executed on a computer, the computer program causes the computer to perform the steps of the aforementioned method embodiments. The memory in FIG. 7 is used to store the computer program and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, images, videos, and the like. The memory can be implemented using any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. The communication component in Figure 7 is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access a wireless network based on a communication standard, such as Wi-Fi, 2G, 3G, 4G / LTE, 5G, or other mobile communication networks, or a combination thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a near-field communication (NFC) module to facilitate short-range communication.For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies. The power supply assembly in Figure 7 provides power to various components of the device in which the power supply assembly resides. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device in which the power supply assembly resides. Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code. The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing device, produce a device for implementing the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions can also be stored in a computer-readable memory capable of directing the computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, product, or device. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, product, or device comprising the element. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) referred to in this disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of such data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or deny. The foregoing description is merely an example of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations of the present disclosure are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this disclosure should be included in the scope of protection of this disclosure.
Claims
Claims 1. A fault handling method, applicable to a proxy answering component provided on a bus device, the bus device being assembled on a host, the method comprising: In the case that a preset type of fault occurs in the bus device, receive the access request initiated by the host instead of the bus device; If a non-delivery access request is received, generate a response message for the access request to characterize that a fault has occurred in the bus device; Send the response message to the host to trigger the host to perform protection processing on the virtual machines running thereon according to a preset protection policy.
2. The method according to claim 1, generating a response message for characterizing a fault occurring in the bus device for the access request, including: Generate a response message for characterizing that the response to the access request has been completed according to the bus protocol; Mark the fault status of the bus device in the response message.
3. The method according to claim 2, wherein marking the fault status of the bus device in the response message comprises: Set the target field in the response message to a preset identifier to mark the fault status of the bus device; wherein, the target field is any field in the response message for marking the error type.
4. The method according to any one of claims 1-3, wherein the answering component is implemented as an independent operating system, and the method further comprises: Perform heartbeat monitoring on the bus device; If an abnormal heartbeat of the bus device is detected, determine that a preset type of fault has occurred in the bus device; or, the proxy response component is implemented as independent hardware, and the bus device starts the proxy response component in the case of a preset type of fault; after the proxy response component is started, determine that a preset type of fault has occurred in the bus device.
5. According to the method described in claim 4, in the fault handling logic triggered when a preset type of fault occurs in the bus device, a proxy response component startup logic is added; if the proxy response component is implemented as independent hardware, then during the process of the bus device running the fault handling logic, according to the proxy response component startup logic, send a startup signal to the proxy response component to start the proxy response component.
6. According to the method described in any one of claims 1-5, a user-mode component of a virtual machine manager runs in the bus device, a kernel-mode component of the virtual machine manager runs in the host, the kernel-mode component interacts with the user-mode component through a shared memory on the bus device, and the read / write request of the kernel-mode component for the shared memory is used as the non-delivery access request.
7. A fault handling method, applicable to a host computer on which a virtual machine is running, and a bus device is installed. A proxy component is provided on the bus device. The method includes: Initiate a non-delivery access request to the bus device according to the bus protocol; If a response message for characterizing that a fault has occurred in the bus device returned by the proxy response component is received, perform protection processing on the virtual machine according to a preset protection policy; wherein, the response message is sent by the proxy response component on behalf of the bus device in the case of a preset type of fault occurring in the bus device.
8. According to the method described in claim 7, performing protection processing on the virtual machine according to a preset protection policy includes: in the host, determine the initiator corresponding to the non-delivery access request; transfer the response message to the initiator to trigger the initiator to protect the virtual machine related to the non-delivery access request. 9. The method according to claim 8, wherein a user-mode component of a virtual machine manager runs on the bus device, a kernel-mode component of the virtual machine manager runs in the host, and the kernel-mode component initiates a non-delivery access request to communicate with the user-mode component; Transfer the response message to the initiator to trigger the initiator to protect the virtual machine associated with the non-delivery access request, including: If the initiator corresponding to the non-delivery access request is the kernel-mode component, transfer the response message to the kernel-mode component; Use the kernel-mode component to power off the virtual machines running on the host.
10. The method according to claim 9, wherein shutting down the virtual machine running on the host by using the kernel-mode component comprises: After receiving the response message, the kernel-mode component parses the response message; If it is determined that the target field in the response message is set to a preset identifier, the kernel-mode component invokes a function for controlling the power-off of the virtual machine to power off the virtual machines running on the host.
11. The method according to claim 9 or 10, further comprising: Use the kernel-mode component to check whether there are any outstanding I / O requests in the virtual machines running on the host; Classify the virtual machines according to whether there are any outstanding I / O requests.
12. The method according to claim 11, classifying the virtual machine according to whether there is an I / O request with an uncompleted response, including: Mark the virtual machines with outstanding I / O requests as the incomplete-status class; Mark the virtual machines without outstanding I / O requests as the complete-status class.
13. A bus device is provided with a proxy responder component for executing one or more computer instructions for: in the case where a preset type of fault occurs in the bus device, receiving on behalf of the bus device an access request initiated by a host; if a non-delivery access request is received, generating for the access request a response message indicating that a fault has occurred in the bus device; and sending the response message to the host to trigger the host to perform protection processing on the virtual machines running thereon according to a preset protection policy.
14. A host is equipped with a bus device on which a proxy responder component is provided. The host includes a memory, a processor, and a communication component; the memory is used for storing one or more computer instructions; the processor is coupled to the memory and the communication component and is used for executing the one or more computer instructions for: initiating a non-delivery access request to the bus device according to a bus protocol; if a response message indicating that a fault has occurred in the bus device returned by the proxy responder component is received, performing protection processing on the virtual machines running thereon according to a preset protection policy; wherein the response message is sent on behalf of the bus device by the proxy responder component in the case where a preset type of fault occurs in the bus device.
15. The host according to claim 14, wherein when the processor performs protection processing on the virtual machine according to a preset protection policy, it is specifically used for: in the host, determining the initiator corresponding to the non-delivery access request; Transfer the response message to the initiator to trigger the initiator to protect the virtual machine associated with the non-delivery access request.
16. A computer-readable storage medium storing computer instructions, which cause one or more processors to execute the fault handling method according to any one of claims 1 to 2 when the computer instructions are executed by the one or more processors.
17. A computer program that causes a computer to execute the fault handling method according to any one of claims 1 to 12 when the computer program is executed on the computer. 17
Citation Information
Patent Citations
Control method and device for sharing FPGA by multiple virtual machines and electronic equipment
CN109656676A
PCIe device and operating method thereof
CN115203101A
Non-delivery write transactions for computer bus
CN115481071A
Sand timer algorithm for tracking in-flight data storage requests for data replication
US20200226097A1
Systems and methods for continuous data protection comprising storage of completed I / O requests intercepted from an I / O stream using touch points
US20230125719A1
Cited By
Bus communication inspection method based on master-slave communication mode and related equipment
CN120675833A