Fault processing method and device and storage medium
By setting up the pickup component on the bus device, the host and virtual machine downtime caused by bus device failure is solved, and the protection of the failure is achieved to ensure the stable operation of the virtual machine.
Patent Information
- Application Number
- CN202311873214.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In cloud computing, bus equipment failure causes the host to go down, which leads to all virtual machines going down, causing difficulties in operation and maintenance work.
Set up a pickup component on the bus device, instead of the bus device receiving a non-delivery access request and generating a response message in the event of a failure, notifying the host for protection processing.
It avoids host downtime, reduces the impact of bus device failure on virtual machines, and ensures the stable operation of virtual machines.
Smart Images

Figure CN120234170A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to a fault handling method, device, and storage medium. Background Art
[0002] In the field of cloud computing, virtualization technology is widely used. Based on virtualization technology, multiple virtual machines can be created on a host, and the virtual machines can be managed and controlled through a Virtual Machine Monitor (VMM) and other means.
[0003] Various bus devices (PCI / PCIE devices) often need to be installed on the host. Moreover, in many cases, the host needs to send a non-posted access request to the bus device, and the bus resources in the host can be normally released only after receiving the response message returned by the bus device.
[0004] Currently, when a bus device fails and cannot respond to a non-posted access request, the host will crash because the non-posted access request sent times out without being responded to. The virtual machines on the host will all crash because the host crashes, resulting in all virtual machines being abnormal, which brings great difficulties to subsequent operation and maintenance work. Summary of the Invention
[0005] Multiple aspects of this application provide a fault handling method, device, and storage medium to reduce the impact of bus device failures on virtual machines.
[0006] An embodiment of this application provides a fault handling method, which is adapted to a proxy component set on a bus device, and the bus device is installed on a host. The method includes:
[0007] When a preset type of fault occurs in the bus device, receive the access request initiated by the host on behalf of the bus device;
[0008] If a non-posted access request is received, generate a response message for characterizing that the bus device has a fault for the access request;
[0009] Send the response message to the host to trigger the host to perform protection processing on the virtual machines running thereon according to a preset protection policy.
[0010] An embodiment of this application also provides a fault handling method, which is adapted to a host on which virtual machines are running, and a bus device is installed on the host, and a proxy component is set on the bus device. The method includes:
[0011] In accordance with the bus protocol, send a non-posted access request to the bus device;
[0012] If a response message indicating that a bus device has failed is received from the proxy component, perform protection processing on the virtual machine according to a preset protection policy;
[0013] Wherein, the response message is sent by the proxy component on behalf of the bus device when the bus device has a preset type of failure.
[0014] An embodiment of the present application further provides a bus device, which is provided with a proxy component, and the proxy component is used to execute one or more computer instructions for:
[0015] When the bus device has a preset type of failure, receive an access request initiated by the host on behalf of the bus device;
[0016] If a non-delivery access request is received, generate a response message indicating that the bus device has failed for the access request;
[0017] Send the response message to the host to trigger the host to perform protection processing on the virtual machine running thereon according to a preset protection policy.
[0018] An embodiment of the present application further provides a host, which is equipped with a bus device, and the bus device is provided with a proxy component. The host includes a memory, a processor, and a communication component;
[0019] The memory is used to store one or more computer instructions;
[0020] The processor is coupled to the memory and the communication component, and is used to execute one or more computer instructions for:
[0021] In accordance with the bus protocol, initiate a non-delivery access request to the bus device;
[0022] If a response message indicating that the bus device has failed is received from the proxy component, perform protection processing on the virtual machine according to a preset protection policy;
[0023] Wherein, the response message is sent by the proxy component on behalf of the bus device when the bus device has a preset type of failure.
[0024] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to execute the foregoing fault processing method.
[0025] In an embodiment of the present application, it is proposed that a proxy component can be set on a bus device. The proxy component can preset a type of fault in the case where the bus device fails to respond to a non-delivery access request, and in a timely manner, replace the bus device to receive the non-delivery access request initiated by the host computer, and generate a response message for characterizing that the bus device has failed for such an access request. The proxy component can send the response message to the host computer. In this way, in the case where the bus device fails to respond to the non-delivery access request, based on the proxy function provided by the proxy component, the host computer can be prevented from crashing; and the host computer can, under the trigger of such a response message, perform protection processing on the virtual machines running thereon according to a preset protection policy, and the virtual machines will no longer directly crash due to the bus device failure, thereby reducing the impact of the bus device failure on the virtual machines. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The exemplary embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0027] Figure 1 FIG. 9 is a schematic diagram of an assembly relationship between a bus device and a host computer provided by an exemplary embodiment of the present application;
[0028] Figure 2 FIG. 13 is a schematic flowchart of a fault handling method provided by an exemplary embodiment of the present application;
[0029] Figure 3 FIG. 17 is a schematic logical diagram of a fault handling solution under an exemplary bus device provided by an exemplary embodiment of the present application;
[0030] Figure 4 FIG. 21 is a schematic flowchart of a fault handling method provided by another exemplary embodiment of the present application;
[0031] Figure 5 FIG. 25 is a schematic logical diagram of a fault solution on the host computer side under an exemplary bus device provided by another exemplary embodiment of the present application;
[0032] Figure 6 FIG. 29 is a schematic structural diagram of a bus device provided by still another exemplary embodiment of the present application;
[0033] Figure 7 FIG. 33 is a schematic structural diagram of a host computer provided by still another exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0035] Before starting to elaborate on the technical solutions provided by the embodiments of this application, several technical concepts related to this application are briefly described as follows.
[0036] PCI (Peripheral Component Interconnect) is a high-performance local bus, and PCIE (peripheral component interconnect express) is a high-speed serial computer expansion bus. Both PCI and PCIE can be referred to as buses, which are proposed to meet the high-speed data transmission between peripherals and between external devices and the host.
[0037] Bus device: It can be understood as an external device assembled on the host in a bus manner. The bus device and the host communicate in accordance with the bus protocol.
[0038] Bus protocol: It defines the interaction rules between the communication parties based on the bus. The interaction rules may include but are not limited to the internal structure of the messages used in the interaction process and the execution processes within each communication party, etc.
[0039] Bus transaction: The operation sequence from requesting the bus to completing the use of the bus is called a bus transaction. Bus transactions can include at least two categories: non-posted and posted.
[0040] Posted bus transaction: After the request is sent through the bus, the bus resources can be released step by step without waiting for the response message.
[0041] Non-posted bus transaction: After the request is sent, the bus resources can be released only after receiving the corresponding response message.
[0042] Currently, for non-posted bus transactions, the host will send a non-posted Non-Posted access request to the bus devices assembled on it. If it does not receive a timely response from the bus devices, a Completion Time Out (CTO) exception will be triggered in the host, resulting in the host crashing due to a kernel deadlock or a kernel panic, and the virtual machines running on the host will also crash, and the virtual machines may lose data due to the crash.
[0043] Therefore, this embodiment proposes a fault handling solution to reduce the impact of bus device failures on virtual machines.
[0044] The following will detail the technical solutions provided by the embodiments of the present application with reference to the accompanying drawings.
[0045] Figure 1 FIG. is a schematic diagram of the assembly relationship between a bus device and a host provided for an exemplary embodiment of the present application. As Figure 1 shown, the bus device is assembled on the host through a bus (PCI or PCIE) method.
[0046] Referring to Figure 1 , in this embodiment, it is proposed that a proxy response component can be set on the bus device. In practical applications, a proxy response component can be added to the existing bus devices; of course, new bus devices with built-in proxy response components can also be provided. In addition, in this embodiment, other components in the bus device are not limited. Based on other components, the bus device can provide various device functions, and specific examples of the device functions are not given here.
[0047] In this embodiment, the proxy response component can provide a proxy response function when the bus device has a preset type of failure.
[0048] As mentioned above, after the host sends a non-posted access request to the bus devices assembled on it, it will wait for the bus devices to return a response message. Among them, the non-posted Non-Posted access request can be understood as a transaction layer message initiated by the host. The transaction type carried in this message is non-posted Non-Posted. In this way, under normal circumstances, after receiving this transaction layer message, the bus device can determine that the received is a non-posted access request through the transaction type carried in it. Then, the bus device can generate a response message according to the processing rules for non-posted access requests defined in the bus protocol and return it to the host.
[0049] To this end, in this embodiment, the preset type of faults may include, but are not limited to, faults such as kernel deadlocks, kernel errors, CPU exceptions, or memory exceptions that cause the bus device to be unable to respond to non-posted access requests. The types of the preset type of faults are not exhaustively listed here. That is, in this embodiment, when the bus device is unable to respond to the non-posted access request initiated by the host due to a fault, the proxy component set on the bus device can be started in a timely manner.
[0050] In this embodiment, the implementation form of the proxy component is not limited. The proxy component can be implemented as software, for example, an independent operating system; the proxy component can also be implemented as hardware, for example, a field programmable gate array (FPGA), or an independent chip, etc. No more examples of implementation forms are given here, and any implementation form that can support the proxy component to still work properly after the bus device has a preset type of fault is acceptable. In this way, in this embodiment, based on the independence of the proxy component, it can be ensured that the proxy component can still work properly after the bus device fails.
[0051] For different implementation forms, in this embodiment, an adapted triggering method can be used to start the proxy component:
[0052] If the proxy component is implemented as an independent operating system, the proxy component can perform heartbeat monitoring on the bus device; if the heartbeat of the bus device is detected to be abnormal, it is determined that the bus device has a preset type of fault. The general principle of heartbeat monitoring is as follows: A heartbeat monitoring channel is established between the proxy component and the bus device, and the bus device actively sends heartbeat packets to the proxy component through this channel on time; if the proxy component detects that the bus device does not send heartbeat packets on time, it can be determined that the heartbeat of the bus device is abnormal. Of course, the proxy component can also send a heartbeat detection packet to the bus device. If the confirmation packet returned by the bus device is not received on time, the proxy component can determine that the heartbeat of the bus device is abnormal. For example, the return delay of the confirmation packet can be defined in the proxy component, such as 1s. Then, after the proxy component sends the heartbeat detection packet, if the confirmation packet returned by the bus device is not received within 1s, it can be determined that the heartbeat of the bus device is abnormal. The principle of heartbeat monitoring here is exemplary, and this embodiment is not limited to this. In this triggering method, if the bus device has the aforementioned preset type of fault, it usually cannot send heartbeat packets or return confirmation packets on time. Therefore, the proxy component can actively sense whether the bus device has a preset type of fault by performing heartbeat monitoring on the bus device.
[0053] If the proxy answering component is implemented as an independent hardware, it can be started by the bus device in the event of a preset type of fault. In this case, after the proxy answering component is started, it is assumed by default that the bus device has a preset type of fault. During the research process, the inventor found that after the bus device has a preset type of fault, it usually triggers the bus device to start the fault handling logic. For example, the MCE (Machine Check Exception) handling logic, etc. Therefore, in this embodiment, it is proposed that in this triggering method, the proxy answering component can be started in a timely manner by modifying the fault handling logic in the bus device. In an exemplary modification scheme: a proxy answering component start logic can be added to the fault handling logic in the bus device. In this way, when the bus device has a preset type of fault, the bus device will enter the modified fault handling logic, and during the execution of the modified fault handling logic, a start signal can be sent to the proxy answering component according to the proxy answering component start logic to actively start the proxy answering component. Among them, the start signal can be an electrical signal or an instruction signal, etc. For example, if the proxy answering component is implemented as an FPGA, an electrical signal can be transmitted to the start pin on the proxy answering component. Another example is that if the proxy answering component is implemented as an independent chip with its own microprocessor, an instruction signal can be sent to the microprocessor in the proxy answering component to trigger the start of the proxy answering component, which is not specifically limited here.
[0054] The above several triggering methods for the proxy answering component are only exemplary, and this embodiment is not limited to this, and no more examples of the triggering method are given here.
[0055] Based on this, when the bus device has a preset type of fault, the proxy answering component will be started in a timely manner. In practical applications, the startup time of the proxy answering component can be short enough. The startup time of the answering component plus its response time to non-delivery access requests can be at least shorter than the response waiting time specified in the bus protocol for non-delivery access requests. In this way, it can be effectively ensured that any non-delivery access requests sent by the host to the bus device will not be missed, thus avoiding the host from crashing due to missed requests.
[0056] In addition, referring to Figure 1, in this embodiment, virtual machines are running on the host. In virtualization technology, virtual machines are usually created and managed by a Virtual Machine Monitor (VMM). The virtual machine manager usually includes user-mode components and kernel-mode components. Among them, user-mode components are usually used to handle work related to the user mode. For example, components such as the Quick Emulator (QEMU); while kernel-mode components are usually used to handle work related to the kernel mode. For example, components such as the Kernel-based Virtual Machine (KVM). The work processed by user-mode components may include, but is not limited to, creating virtual machines, allocating addresses from the virtual address space occupied by the virtual machine manager as the physical addresses of the virtual machines, and simulating the required virtual devices for the virtual machines, etc., which will not be enumerated here. The kernel-mode components provide a series of interfaces to the user-mode components. The user-mode components can control various aspects of the virtual machines through these interfaces, such as the number of CPUs, memory layout, running and halting, etc.; the kernel-mode components are also used to process privileged instructions issued in the virtual machines. Privileged instructions are usually those instructions that may affect the entire host, such as various IO requests that occur in the virtual machines. No more introduction to the knowledge background related to virtual machines on the host will be given here.
[0057] On this basis, Figure 2 is a schematic flowchart of a fault handling method provided by an exemplary embodiment of the present application. Refer to Figure 2 , this method can be applied to the proxy component set in the bus device. The method may include:
[0058] Step 100, in the case where a preset type of fault occurs in the bus device, receive the access request initiated by the host on behalf of the bus device;
[0059] Step 101, if a non-delivery access request is received, generate a response message for the access request to indicate that a fault has occurred in the bus device;
[0060] Step 102, send the response message to the host to trigger the host to perform protection processing on the virtual machines running on it according to the preset protection policy.
[0061] Among them, in step 100, the proxy answering component can be started in time when a bus device has a preset type of fault, and replace the bus device to receive the access request initiated by the host computer. It should be understood here that the faults of the bus device can be various. In this embodiment, when the bus device has a fault that causes the bus device to be unable to respond to non-posted access requests, the proxy answering component can be started in time; for other types of faults, it usually does not cause the host computer to crash, and in this embodiment, the proxy answering component does not need to be started. The startup scheme of the proxy answering component can refer to the description in the previous text. As mentioned above, in this embodiment, the proxy answering component can be started by an adapted triggering method according to different implementation forms of the proxy answering component, and will not be repeated here.
[0062] In this embodiment, the proxy answering component can replace the bus device to receive the access request initiated by the host computer by monitoring the interface used to receive the access request in the communication component of the bus device. Of course, in this embodiment, other methods can also be used to support the proxy answering component to be able to receive the access request initiated by the host computer to the bus device, and it is not limited to this, and no more examples will be given here. In addition, since there is a startup time consumption for the proxy answering component, therefore, in some special cases, after the bus device has a preset type of fault and before the proxy answering component is started, the host computer may initiate an access request to the bus device during this period. In this regard, in this embodiment, the proxy answering component can determine whether there is an unprocessed access request by the access request description information cached under the interface used to receive the access request in the bus device. If so, it is used as the access request received by the proxy answering component. This can effectively avoid the problem of missing access requests.
[0063] Continue to refer to Figure 2 , in step 101, if a non-posted access request is received, a response message is generated for the access request to indicate that the bus device has a fault. As mentioned above, there may be non-posted access requests or posted access requests between the host computer and the bus device. Considering that the posted access requests are not affected by the bus device fault, in this embodiment, the proxy answering component does not need to handle the captured posted access requests.
[0064] For non-posted access requests, in this embodiment, it is proposed that the proxy answering component needs to generate a response message for the captured non-posted access request to indicate that the bus device has a fault.
[0065] The inventor found during the research process that the request header of the access request sent by the host computer to the bus device usually contains a request type field (Type field), which is marked with the request type (Non-posted or posted)
[0066] Preferably, in this embodiment, the proxy component can generate a response message for indicating the completion response of the access request according to the bus protocol; and mark the fault status of the bus device in the response message. In the bus protocol, the response message for indicating the completion response of the access request is also called a completion message (CPL), and its message header may include a completion status field (i.e., the Compl.Status field), a requester identification field (i.e., the Requster ID field), a completer identification field (i.e., the Completer ID field), etc. Among them, the completion status field can be used to mark whether the corresponding access request has been successfully completed; the requester identification field can be used to mark the identity of the initiator of the access request; and the completer identification field can be used to mark the identity of the responder who responds to the access request. It should be understood that the above several fields are only exemplary, and the response message may also include other fields for marking other aspects of information. In this embodiment, the proxy module can define the completion status field in the response message as SC (that is, successful completion) to indicate the completion response of the corresponding access request.
[0067] In this embodiment, the response message generated by the proxy component is used not only to indicate the completion response of the non-delivery access request, but also to indicate that the bus device has failed. In this embodiment, multiple implementation methods can be used to indicate that the bus device has failed based on the response message. In an alternative implementation: the fault status of the bus device can be marked in the response message. Among them, the fault status refers to the abnormal state of the bus device. In this alternative implementation, it is proposed that the fault status can be marked through marking dimensions such as error types.
[0068] During the research process, the inventor found that in addition to the aforementioned completion status field and requester identification field, etc., the response message also contains a field for marking the error type. Based on this, in this optional implementation, it is proposed that the fault status of the bus device can be marked by setting the target field in the response message to a preset identifier. Among them, the target field can be any field in the response message that can be used to mark the error type. In an exemplary solution: the field in the response message used to identify a data error (i.e., Error Poisoned, EP) can be set to a preset identifier, such as 1, to mark the fault status of the bus device. Among them, the EP field is a field in the response message used to mark the error type, and the error type identified by the EP field is error forwarding / data poisoning. Of course, this is only exemplary. The response message also contains other fields for identifying the error type, but the error types that different fields can identify may be different. Based on this, the proxy answering component can also assign values under other fields to mark other error types, and no more examples are given here. In this way, in this embodiment, the error type can be marked as the fault status of the bus device by assigning a value to the target field in the response message, thereby indirectly indicating that the bus device has failed.
[0069] It should be understood that this is only a preference. In this embodiment, the proxy answering component is supported to use other custom formats to construct the response message, as long as the response message can indicate that the corresponding access request has been completed and can indicate that the bus device has failed. If a message format that does not conform to the bus protocol is used, it is necessary to pre-agree on the message format of the response message between the host and the proxy answering component so that the host can parse and understand the response message according to the agreement. This embodiment does not limit the format of the response message.
[0070] Reference Figure 2 , in step 102, the proxy answering component can send the response message to the host. In actual applications, the proxy answering component can reuse the bus channel between the bus device and the host to send the response message to the host.
[0071] First of all, in this embodiment, the proxy answering component generates a response message for the non-delivery access request it receives and returns it to the host. In this way, from the perspective of the host, the non-delivery access request it issues has been responded to, so the host can be prevented from crashing. Secondly, in this embodiment, the response message generated by the proxy answering component also indicates that the bus device has failed. In this way, from the perspective of the host, it can timely sense that the bus device has failed through the response message.
[0072] In this way, in this embodiment, after a bus device has a preset type of fault, all non-posted access requests sent by the host to the bus device can be received by the proxy responder component. Moreover, as mentioned above, the startup time of the proxy responder component plus its response time to non-posted access requests is shorter than the response waiting time specified in the bus protocol for non-posted access requests. Therefore, it can be ensured that all non-posted access requests reaching the bus device after the bus device has a preset type of fault can be received and responded to by the proxy responder component without omission. Furthermore, the response message generated by the proxy responder component for this purpose can reach the host before the aforementioned response waiting time, thereby preventing the host from crashing.
[0073] During the research process, the inventors found that most of the non-posted access requests initiated by the host to the bus device are caused by the virtual machines running on it. Although the proxy responder component responds to the Non-posted access requests on behalf of the bus device, although this response can prevent the host from crashing, for the virtual machines, the original requests initiated by them are not normally responded to, and the virtual machines will still have access exceptions. Therefore, after the bus device has a preset type of fault, although the proxy responder component prevents the host from crashing, it may still cause related virtual machines to have exceptions.
[0074] For this reason, this embodiment also proposes to preset a protection policy in the host. On this basis, after the host receives the response message, it can perform protection processing on the virtual machines running on it according to the preset protection policy to prevent more exceptions or errors from occurring. In this embodiment, it is supported to set relevant protection policies in the host as needed to better protect the virtual machines. The preset protection policy in the host is not limited here and will be elaborated in the embodiments on the host side.
[0075] In summary, in this embodiment, it is proposed that a proxy responder component can be set on the bus device. The proxy responder component can, in the case of a bus device having a fault and being unable to respond to non-posted access requests (preset type of fault), timely receive the Non-posted access requests initiated by the host on behalf of the bus device and generate a response message indicating that the bus device has a fault for such access requests. The proxy responder component can send the response message to the host. In this way, in the case of a bus device having a fault and being unable to respond to non-posted access requests, based on the proxy function provided by the proxy responder component, the host can be prevented from crashing; and the host can, under the trigger of this response message, perform protection processing on the virtual machines running on it according to the preset protection policy, and the virtual machines will no longer directly crash due to the bus device fault, thereby reducing the impact of the bus device fault on the virtual machines.
[0076] In the above or following embodiments, as mentioned above, the device functions provided by the bus device are not limited. That is, the fault handling method provided in this embodiment is applicable to various types of bus devices, especially bus devices that may need to process Non-posted access requests issued by the host. For example, various Data Processing Units (DPUs), etc.
[0077] Figure 3 It is a logical schematic diagram of a fault handling solution under an exemplary bus device provided by an exemplary embodiment of the present application. Refer to Figure 3 , an operating system runs on the exemplary bus device, and a user-mode component of the virtual machine manager runs in the operating system. For example, the QEMU component mentioned above. And the kernel-mode component of the virtual machine manager, such as the KVM component mentioned above, runs in the host. In addition, a shared memory is also set in the bus device. Among them, the shared memory is the most efficient inter-process communication method. Processes can directly read and write the shared memory, and there is no need for any data copying between processes. In this embodiment, the shared memory in the bus device can be mapped to the memory address space on the host side. In this way, the kernel-mode component in the host can initiate a communication request to the user-mode component in the bus device based on the memory address on the host side (essentially a read / write request for the shared memory). And since the shared memory has been mapped to the memory address space based on the host side, the information related to the communication request can be written into the shared memory in the bus device, and the user-mode component in the bus device can receive the communication request. Similarly, the user-mode component in the bus device can also initiate a communication request to the kernel-mode component in the host based on the shared memory. Based on this, in this embodiment, the interaction between the kernel-mode component in the host and the user-mode component in the bus device can be regarded as two processes, and the two can perform inter-process communication through the shared memory on the bus device.
[0078] Figure 3The exemplary bus device provided herein may also be referred to as a Cloud Infrastructure Processing Unit (CIPU). It can be understood as a chip specifically designed for cloud computing scenarios, which can efficiently manage and synergistically accelerate multi-dimensional resources such as computing, storage, network, and security. The CIPU can not only access physical computing, storage, and network resources downward to quickly cloudify the hardware resources, but also access the distributed platform upward to flexibly manage, schedule, and orchestrate cloud resources. The CIPU can be combined with chips such as the CPU, GPU, and FPGA in the host to form a software-hardware integrated virtualization architecture, achieving deep integration at the system level. It should be understood that the CIPU is only an exemplary name for such bus devices, and there are other names for such bus devices in the cloud computing field, such as Infrastructure Processing Unit (IPU), etc. No more examples of names will be given here.
[0079] Reference Figure 3 , this exemplary bus device hosts the user-mode components in the virtual machine manager. In this way, the user-mode components in the bus device and the kernel-mode components in the host can cooperate with each other to create and manage virtual machines on the host.
[0080] Based on this, reference Figure 3 , after a virtual machine running on the host issues an IO request, it will first be captured by the kernel-mode components in the host. Since the IO requests issued by the host need to be completed through the cooperation of the kernel-mode components and the user-mode components, the kernel-mode components need to send read / write requests for shared memory to the bus device for the captured IO requests. Such read / write requests are typical Non-posted access requests that occur in the host. Among them, the IO requests issued by the virtual machine are usually used to access external devices on the host or other virtual machines, etc., which will not be elaborated here. The IO requests issued by the virtual machine may include, but are not limited to, programmed I / O (PIO) or Memory-mapped I / O (MMIO). For example, when a virtual machine needs to use the GPU installed on the host, it can initiate an MMIO request. The MMIO request will be received by the kernel-mode components and transferred to the user-mode components. The user-mode components can execute processing logics such as address translation to make the request reach the hardware resource of the GPU installed on the host, and then complete the processing of the request in the hardware. The IO requests issued by the virtual machine are common requests in virtualization technology and will not be elaborated here.
[0081] It should be understood that by Figure 3The read / write request for the shared memory in the bus device issued by the kernel-mode component in is only an exemplary type of Non-posted access request, and the initiator of this type of Non-posted access request is the kernel-mode component of the virtual machine manager in the host machine.
[0082] Reference Table Figure 3 From the perspective of the exemplary bus device, it can manage the hardware resources and virtual machines on the host machine. From the perspective of the host machine, the essence of the exemplary bus device is still a bus device, and the host machine and the bus device still communicate in accordance with the bus protocol. In the traditional fault handling solution (that is, when the answering module provided by this embodiment is not provided in the bus device), when the bus device has a preset type of fault, it will cause the host machine to crash due to the aforementioned completion timeout CTO problem.
[0083] Following the description of the fault handling solution in the answering component in the above embodiment, Figure 3 In the event that a preset type of failure occurs in the exemplary bus device, the proxy component provided therein can receive the aforementioned read / write request for the shared memory in the bus device issued by the kernel-mode component, and the proxy component can generate a response message for such Non-posted access request, thereby ensuring that the host machine can receive the response message corresponding to such Non-posted access request, thereby avoiding such Non-posted access request from causing the host machine to crash due to the bus device failure.
[0084] Accordingly, for Figure 3 The exemplary bus device provided can be based on the fault handling solution provided in this embodiment. By setting a proxy component on such exemplary bus device, it can effectively avoid the problem of host machine crash caused by failure of such exemplary bus device, thereby reducing the impact of failure of such exemplary bus device on virtual machines running on the host machine.
[0085] It is worth mentioning that Figure 3 The bus device provided is only exemplary. In this embodiment, the type of bus device is not limited. This embodiment supports the application of the fault handling solution provided in this embodiment in various bus devices to minimize the host machine downtime that may be caused by bus device failure, thereby improving the stability of the host machine and achieving protection of the virtual machine.
[0086] Figure 4 A flowchart of a fault handling method provided by another exemplary embodiment of the present application. Figure 4 This method is applicable to a host machine, a virtual machine is running on the host machine, a bus device is installed on the host machine, and a proxy component is set on the bus device. Figure 4, the method may include:
[0087] Step 400, initiate a non-posted access request to the bus device according to the bus protocol;
[0088] Step 401, if a response message for indicating that the bus device has a fault is received from the proxy responder component, perform protection processing on the virtual machine according to a preset protection policy;
[0089] Wherein, the response message is sent by the proxy responder component on behalf of the bus device when the bus device has a preset type of fault.
[0090] In this embodiment, the host can initiate a non-posted access request to the bus device according to the bus protocol. The inventor found during the research process that there are many reasons for the host to initiate a non-posted access request to the bus device, including but not limited to memory read / write, I / O read / write, and configuration read / write of the bus device, etc., which will not be enumerated here. However, no matter what the reason is for the host to initiate a non-posted access request to the bus device, as mentioned above, if the bus device has a preset type of fault and cannot respond to the non-posted access request initiated by the host in a timely manner, a completion timeout (CTO) exception will be triggered in the host, resulting in the host having a kernel lockup or kernel error and crashing. The virtual machines running on the host will also crash, and the virtual machines may lose data due to the crash.
[0091] Reference Figure 4 , in this embodiment, after the host initiates a non-posted access request to the bus device, it will wait for the corresponding response message. Based on the fault handling solution provided in this embodiment, the proxy responder component set on the bus device can generate a response message for the non-posted access request initiated by the host on behalf of the bus device when the bus device has a preset type of fault. In this way, in this embodiment, after the bus device has a fault, the host can receive the response message generated by the proxy responder component set on the bus device for the non-posted access request.
[0092] Since the host has received the response message returned for the non-posted access request it initiated, therefore, the CTO exception will no longer be triggered in the host, and it will no longer cause the host to have a kernel lockup or kernel error, and the host will not crash. That is to say, when the bus device has a fault and cannot respond to the Non-posted access request initiated by the host, through the answering function provided by the proxy responder component, the host crashing can be effectively avoided.
[0093] In this embodiment, when the bus device has a fault, the host does not crash. Therefore, an opportunity can be provided for the host to protect the virtual machines running on it to prevent more exceptions or errors from occurring.
[0094] To this end, in this embodiment, it is proposed that a protection policy can be preset in the host. Based on this, referring to Figure 4 , in step 401, after the host receives the response message returned by the proxy component for characterizing that a bus device has failed, the host can perform protection processing on the virtual machine according to the preset protection policy.
[0095] In this embodiment, the preset protection policy can be flexibly set and updated as needed, and the policy content in the preset protection policy is not limited. Any effective policy that can reasonably avoid more exceptions or errors in the virtual machine is acceptable.
[0096] In an exemplary preset protection policy, after the host receives the response message returned by the proxy component for a non-delivery access request, the host can determine the initiator corresponding to the non-delivery access request; on this basis, the host can pass the response message to the initiator to trigger the initiator to protect the virtual machine related to the non-delivery access request. Among them, the virtual machine related to the non-delivery access request can be understood as the virtual machine that can be affected by the bus device failure. To this end, in this embodiment, it is proposed that the processing logic for such response messages can be set in each initiator on the host that may initiate a non-delivery access request, so as to determine the virtual machines that can be affected by the bus device failure through these processing logics and protect these virtual machines in a timely manner. The inventor found during the research process that in the host, the initiators corresponding to non-delivery access requests are diverse. For example, it may be a kernel-mode component in the virtual machine manager, or it may be a virtual machine, or it may be a process in the host operating system in the host, etc. This embodiment does not list all these initiators exhaustively. And the processing logic for the response messages returned by the proxy component in each initiator is not limited either. The processing logic in the exemplary initiators will be elaborated in detail later.
[0097] In summary, in this embodiment, it is proposed that a proxy component can be set on the bus device. In the case where the bus device has a preset type of failure, the proxy component can replace the bus device to receive the non-delivery access request initiated by the host and generate a response message for characterizing that the bus device has failed for such an access request. The proxy component can send the response message to the host. In this way, in the case where the bus device fails and cannot respond to the non-delivery access request, based on the proxy function provided by the proxy component, the host can be prevented from crashing; and the host can, under the trigger of this response message, perform protection processing on the virtual machines running on it according to the preset protection policy, and the virtual machines will no longer crash directly due to the bus device failure, thereby reducing the impact of the bus device failure on the virtual machines.
[0098] Figure 5A logical schematic diagram of a fault solution on the host side provided for another exemplary embodiment of this application under an exemplary bus device. Refer to Figure 5 , an operating system runs on the exemplary bus device, and a user-mode component of a virtual machine manager runs in the operating system, such as the QEMU component mentioned above. The kernel-mode component of the virtual machine manager, such as the KVM component mentioned above, runs on the host. Based on this, the kernel-mode component can interact with the user-mode component through the shared memory on the bus device. For an exemplary principle of communication between the kernel-mode component and the user-mode component based on shared memory, refer to the description in the previous text and will not be repeated here.
[0099] Based on this, refer to Figure 5 , in the case where the exemplary bus device has a preset type of fault and cannot respond to a non-delivery access request, the kernel-mode component in the host can initiate a communication request to the user-mode component in the bus device, specifically, a read / write request to the shared memory in the bus device, and this read / write request is a non-delivery access request. After this non-delivery access request is sent to the bus device, since the bus device has failed, a response message will be provided by the proxy component.
[0100] On this basis, for the host, if it is determined that the initiator of the non-delivery access request is the kernel-mode component, the response message will be passed to the kernel-mode component; and the kernel-mode component will be used to shut down the virtual machines running on the host. That is, in this embodiment, an exemplary initiator of the non-delivery access request to the bus device is the kernel-mode component. In this case, the host can pass the response message returned by the proxy component to this kernel-mode component. Among them, for the kernel-mode component as the initiator, the virtual machines determined to be related to the non-delivery access request can be all the virtual machines running on the host.
[0101] As mentioned above, in this embodiment, a processing logic for such response messages can be set in the kernel-mode component. In this case, an exemplary processing logic for such response messages in the kernel-mode component can be: after receiving the response message, parse the response message; if it is determined that the target field in the response message is set to a preset identifier, call the function used to control the shutdown of the virtual machine to shut down the virtual machines running on the host.
[0102] During the process of introducing the functions of the kernel-mode component in the previous text, it has been mentioned that the kernel-mode component can be used to control the operation and shutdown of virtual machines. The kernel-mode component contains a function (API) for controlling the shutdown of virtual machines, and the kernel-mode component can call this function to shut down the virtual machines. This function is an existing function in the kernel-mode component and will not be elaborated here.
[0103] During the research process, the inventors found that for Figure 5 the exemplary bus device shown, the reason why the kernel-mode component in the host needs to initiate a non-delivery access request to this bus device is usually that the virtual machine on the host issues an IO request. Regarding the process of cooperation between the kernel-mode component and the user-mode component after the virtual machine issues an IO request, reference can be made to the description above, and it will not be repeated here.
[0104] In this way, on the one hand, as described in the foregoing embodiments, although the response message can indicate that the read / write request for the shared memory has been completed, it also indicates that the exemplary bus device has failed. For example, SC+EP is defined in the message header of the foregoing exemplary response message. Accordingly, the result data required for the IO request of the virtual machine is not carried in this response message. Therefore, although the proxy response component returns a response message to the non-delivery access request initiated by the kernel-mode component on behalf of the bus device, this response message cannot correctly respond to the IO request initiated by the virtual machine. Therefore, from the perspective of the virtual machine that initiates the IO request, the IO request it initiates will still have a response exception, but this response exception will not cause the system to crash, but may only be an exception at the data level or the configuration level.
[0105] On the other hand, the non-delivery access request responded to by the proxy response component may only cause an IO request exception for a single or part of the virtual machines, but other virtual machines are not aware of the exception of the bus device. If other virtual machines initiate an IO request subsequently, it will cause the kernel-mode component to initiate the same non-delivery access request to this bus device as this time, which will cause more virtual machines to have IO exceptions, and the scope of the exception will continue to expand.
[0106] Therefore, in the processing logic for the response message set in the foregoing kernel-mode component, after determining that the target field in the response message is set to the preset identifier, the host running on the host is shut down, which can effectively protect the virtual machines that have not had an exception, and the virtual machines that have had an exception will not have more exceptions and can also be effectively protected. Therefore, it can effectively prevent more exceptions or errors from occurring in the virtual machines.
[0107] Furthermore, in this case, in addition to using the kernel-mode component to shut down the virtual machines running on it, the host can also use the kernel-mode component to check whether there are any IO requests with uncompleted responses in the virtual machines running on the host; classify the virtual machines according to whether there are IO requests with uncompleted responses.
[0108] Among them, the kernel-mode component records the processing status information of each IO request initiated by the virtual machine. Based on this, the kernel-mode component can check without hindrance whether there are IO requests with uncompleted responses in the virtual machines running on the host. Here, an IO request with an uncompleted response can be understood as an IO request that has been sent by the virtual machine but has not received response data. It can be understood that the IO request corresponding to the non-delivery request responded by the proxy response component is an IO request with an uncompleted response, because the response message returned by the proxy response component does not carry the response data required for the IO request response. Of course, the IO requests that have been sent by the virtual machine but have not reached the proxy response component will also definitely be determined as IO requests with uncompleted responses. All of this processing status information is recorded in detail by the kernel-mode component. Therefore, the kernel-mode component can accurately check the IO requests with uncompleted responses sent by the virtual machine, and thus classify the virtual machines accordingly.
[0109] An exemplary classification and marking scheme can be:
[0110] Mark the virtual machines with IO requests having uncompleted responses as the class with incomplete status;
[0111] Mark the virtual machines without IO requests having uncompleted responses as the class with complete status.
[0112] In this further optimization scheme, a classification of virtual machines is proposed. It should be understood that the classification of virtual machines here is mainly used to provide a reference for subsequent operation and maintenance processes. Accordingly, based on the classification and marking of virtual machines, more operation and maintenance bases can be provided for subsequent operation and maintenance processes, so as to perform differential operation and maintenance on virtual machines with different classification marks in subsequent operation and maintenance processes. In this embodiment, no solution is defined for subsequent operation and maintenance processes. For example, only the virtual machines marked as the class with complete status can be migrated in subsequent operation and maintenance processes, etc., which will not be elaborated here.
[0113] In summary, in this embodiment, the host can target Figure 5The exemplary bus device in [description] provides a fault handling solution to avoid triggering a CTO exception through the answering function provided by the answering component in this type of exemplary bus device, thereby avoiding the host crashing; moreover, on the basis of avoiding the host crashing, the processing logic for the response message returned by the answering component can also be set in the kernel state component in the host, so as to stop the host running on the host in a timely manner. In this way, other virtual machines on the host will no longer generate IO requests, thus effectively preventing other virtual machines from experiencing anomalies due to the failure of this exemplary bus device, avoiding the failure of this type of exemplary bus device from affecting more virtual machines, preventing the expansion of the anomaly range, and thus better protecting the virtual machines. In addition, the classification tags configured for each virtual machine can be used as a reference basis in subsequent operation and maintenance processes to facilitate differentiated operation and maintenance of different types of virtual machines.
[0114] In this embodiment, in addition to the aforementioned virtual machine initiating an IO request, the reason for the kernel state component to initiate a non-delivery access request to the bus device may also be that the kernel state component in the host needs to configure the user state component in the bus device, etc. Such non-delivery access requests may affect all virtual machines running on the host. Therefore, for such non-delivery access requests, after receiving the response message, the kernel state component can also stop the host running on the host.
[0115] In addition, in this embodiment, the example of the initiator corresponding to the non-delivery access request is the kernel state component. However, it should be understood that in this embodiment, the initiator corresponding to the non-delivery access request is not limited to this. For other initiators, the processing logic for the response message of this type can also be set. For example, when a virtual machine accesses a direct access device on the host, it can initiate a direct memory access (DMA) request to the direct access device, which is also a non-delivery request. In this case, the initiator of the non-delivery request is the virtual machine. An exemplary processing logic for the response message in the virtual machine can be to request the kernel state component to stop when it is determined that the target field in the response message is set to a preset identifier, so as to complete the stop through the kernel state component. In this case, only this virtual machine needs to be stopped without stopping other virtual machines. The processing logic for the response message within more initiators is only an example here, and this embodiment is not limited to this.
[0116] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel.
[0117] Figure 6 A schematic structural diagram of a bus device provided for another exemplary embodiment of the present application. As Figure 6 shown, a proxy component 60 is provided on the bus device, and the proxy component 60 is independent of the bus device. The bus device can be assembled on a host computer, and a virtual machine runs on the host computer. The proxy component can execute one or more computer instructions for:
[0118] In the case where a preset type of fault occurs in the bus device, receiving the access request initiated by the host computer on behalf of the bus device;
[0119] If a non-delivery access request is received, generating a response message for characterizing that a fault has occurred in the bus device for the access request;
[0120] Sending the response message to the host computer to trigger the host computer to perform protection processing on the virtual machine running thereon according to a preset protection policy.
[0121] In an optional embodiment, when the proxy component 60 generates a response message for characterizing that a fault has occurred in the bus device for the access request, it can specifically be used for:
[0122] Generating a response message for characterizing that the access request has been completed in response according to the bus protocol;
[0123] Marking the fault status of the bus device in the response message.
[0124] In an optional embodiment, when the proxy component 60 marks the fault status of the bus device in the response message, it can specifically be used for:
[0125] Setting a target field in the response message to a preset identifier to mark the fault status of the bus device;
[0126] Wherein, the target field is any field in the response message for marking the error type.
[0127] In an optional embodiment, the proxy component 60 is implemented as an independent operating system, and the proxy component 60 can also be used for: performing heartbeat monitoring on the bus device; if the heartbeat of the bus device is monitored to be abnormal, determining that a preset type of fault has occurred in the bus device;
[0128] Alternatively, the proxy component 60 is implemented as independent hardware, and the bus device starts the proxy component 60 in the case where a preset type of fault occurs; after the proxy component 60 is started, it determines that a preset type of fault has occurred in the bus device.
[0129] In an alternative embodiment, in the fault handling logic triggered in the bus device when a preset type of fault occurs in the bus device, a proxy component startup logic is added; during the process of the bus device running the fault handling logic, according to the proxy component startup logic, a startup signal is sent to the proxy component 60 to start the proxy component 60.
[0130] In an alternative embodiment, a user-mode component of a virtual machine manager runs in the bus device, and a kernel-mode component of the virtual machine manager runs in the host. The kernel-mode component interacts with the user-mode component through shared memory on the bus device, and the non-delivery access request is a read / write request initiated by the kernel-mode component for the shared memory.
[0131] It should be noted that for the technical details in the above embodiments of the bus device, reference can be made to the relevant descriptions in the method embodiments on the proxy component side. To save space, they will not be elaborated here, but this should not cause a loss of the protection scope of this application.
[0132] Figure 7 The structure diagram of a host provided for another exemplary embodiment of this application. Refer to Figure 7 , the host includes: a memory 70, a processor 71, and a communication component 72. Refer to Figure 7 , a bus device 73 is assembled on the host, and a proxy component is provided on the bus device 73.
[0133] The processor 71 is coupled to the memory 70, the communication component 72, and the bus device 73 in a bus manner, and is configured to execute a computer program in the memory 70 for:
[0134] In accordance with the bus protocol, initiate a non-delivery access request to the bus device;
[0135] If a response message for indicating that the bus device has a fault returned by the proxy component is received, perform protection processing on the virtual machine according to a preset protection policy;
[0136] Among them, the response message is sent by the proxy component on behalf of the bus device when a preset type of fault occurs in the bus device.
[0137] In an alternative embodiment, when the processor 71 performs protection processing on the virtual machine according to the preset protection policy, it is specifically configured to:
[0138] In the host, determine the initiator corresponding to the non-delivery access request;
[0139] Transmit the response message to the initiator to trigger the initiator to protect the virtual machine related to the non-delivery access request.
[0140] In an alternative embodiment, a user-mode component of a virtual machine manager runs on a bus device, and the user-mode component interacts with a kernel-mode component of the virtual machine manager running in the host through shared memory on the bus device; when the processor 71 delivers a response message to the initiator to trigger the initiator to protect the virtual machine associated with the non-delivery access request, it may specifically be used for:
[0141] If the initiator corresponding to the non-delivery access request is the kernel-mode component, deliver the response message to the kernel-mode component;
[0142] Use the kernel-mode component to power off the virtual machine running on the host.
[0143] In an alternative embodiment, when the processor 71 uses the kernel-mode component to power off the virtual machine running on the host, it may specifically be used for:
[0144] After receiving the response message, the kernel-mode component parses the response message;
[0145] If it is determined that the target field in the response message is set to a preset identifier, the kernel-mode component invokes a function for controlling the power-off of the virtual machine to power off the virtual machine running on the host.
[0146] In an alternative embodiment, the processor 71 may also be used for:
[0147] Use the kernel-mode component to check whether there are any outstanding I / O requests in the virtual machine running on the host;
[0148] Classify the virtual machines according to whether there are any outstanding I / O requests.
[0149] In an alternative embodiment, when the processor 71 classifies the virtual machines according to whether there are any outstanding I / O requests, it may specifically be used for:
[0150] Mark the virtual machines with outstanding I / O requests as the incomplete status class;
[0151] Mark the virtual machines without outstanding I / O requests as the complete status class.
[0152] Furthermore, as Figure 7 shown, the host further includes: other components such as a power supply component 74. Figure 7 Only some components are schematically shown in Figure 7 and it does not mean that the host only includes
[0153] shown components. It should be noted that for the technical details in the above embodiments of the host, reference may be made to the relevant descriptions in the foregoing method embodiments. For the sake of brevity, they are not repeated here, but this should not cause any loss to the protection scope of this application.
[0154] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps in the above method embodiments.
[0155] The above Figure 7 The memory therein is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method for operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks or optical disks.
[0156] The above Figure 7 The communication component therein is configured to facilitate communication in a wired or wireless manner between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0157] The above Figure 7 The power component therein provides power for various components of the device where the power component is located. The power component can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power component is located.
[0158] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0159] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0162] It should also be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element.
[0163] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0164] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A fault handling method, characterized in that, Adapted to a call - answering component provided on a bus device, the bus device being assembled on a host machine, the method includes: In the case where a preset type of fault occurs in the bus device, receiving, on behalf of the bus device, an access request initiated by the host machine; If a non - delivery access request is received, generating, for the access request, a response message for characterizing that a fault has occurred in the bus device; Sending the response message to the host machine to trigger the host machine to perform protection processing on the virtual machines running thereon according to a preset protection policy.
2. The method according to claim 1, wherein Generating, for the access request, a response message for characterizing that a fault has occurred in the bus device, includes: Generating, according to the bus protocol, a response message for characterizing that the access request has been completed; Marking the fault state of the bus device in the response message.
3. The method according to claim 2, wherein Marking the fault state of the bus device in the response message, includes: Setting a target field in the response message to a preset identifier to mark the fault state of the bus device; Wherein, the target field is any field in the response message for marking the error type.
4. The method according to claim 1, characterized in that The call - answering component is implemented as an independent operating system, and the method further includes: performing heartbeat monitoring on the bus device; if the heartbeat of the bus device is monitored to be abnormal, determining that a preset type of fault has occurred in the bus device; Alternatively, the call - answering component is implemented as independent hardware, and the bus device starts the call - answering component in the case where a preset type of fault occurs; after the call - answering component is started, it determines that a preset type of fault has occurred in the bus device.
5. The method according to claim 4, wherein In the fault - handling logic triggered when a preset type of fault occurs in the bus device, a call - answering component startup logic is added; if the call - answering component is implemented as independent hardware, during the process of the bus device running the fault - handling logic, according to the call - answering component startup logic, a startup signal is sent to the call - answering component to start the call - answering component.
6. The method according to claim 1, wherein A user - mode component of a virtual machine manager runs in the bus device, a kernel - mode component of the virtual machine manager runs in the host machine, the kernel - mode component interacts with the user - mode component through a shared memory on the bus device, and a read / write request of the kernel - mode component to the shared memory serves as the non - delivery access request.
7. A fault handling method, characterized in that, Adapted to a host machine, virtual machines run on the host machine, a bus device is assembled on the host machine, and a call - answering component is provided on the bus device, the method includes: Initiating a non - delivery access request to the bus device according to the bus protocol; If a response message for characterizing that a fault has occurred in the bus device returned by the call - answering component is received, performing protection processing on the virtual machine according to a preset protection policy; Wherein, the response message is sent by the call - answering component on behalf of the bus device in the case where a preset type of fault occurs in the bus device.
8. The method according to claim 7, wherein Performing protection processing on the virtual machine according to a preset protection policy, includes: In the host machine, determining the initiator corresponding to the non - delivery access request; Transfer the response message to the initiator to trigger the initiator to protect the virtual machine associated with the non-delivery access request.
9. The method according to claim 8, characterized in that, A user-mode component of a virtual machine manager runs on the bus device, and a kernel-mode component of the virtual machine manager runs in the host. The kernel-mode component initiates a non-delivery access request to communicate with the user-mode component; Transferring the response message to the initiator to trigger the initiator to protect the virtual machine associated with the non-delivery access request includes: If the initiator corresponding to the non-delivery access request is the kernel-mode component, transfer the response message to the kernel-mode component; Use the kernel-mode component to power off the virtual machines running on the host.
10. The method according to claim 9, characterized in that, Using the kernel-mode component to power off the virtual machines running on the host includes: After receiving the response message, the kernel-mode component parses the response message; If it is determined that the target field in the response message is set to a preset identifier, the kernel-mode component invokes a function for controlling the power-off of the virtual machine to power off the virtual machines running on the host.
11. The method according to claim 9 or 10, characterized in that, It further includes: Use the kernel-mode component to check whether there are any uncompleted IO requests in the virtual machines running on the host; Classify the virtual machines according to whether there are uncompleted IO requests.
12. The method according to claim 11, wherein Classifying the virtual machines according to whether there are uncompleted IO requests includes: Mark the virtual machines with uncompleted IO requests as the status-incomplete class; Mark the virtual machines without uncompleted IO requests as the status-complete class.
13. A bus device, characterized in that, There is a proxy response component, and the proxy response component is used to execute one or more computer instructions for: In the case of a preset type of failure occurring in the bus device, receive the access request initiated by the host on behalf of the bus device; If a non-delivery access request is received, generate a response message for characterizing the failure of the bus device for the access request; Send the response message to the host to trigger the host to perform protection processing on the virtual machines running thereon according to a preset protection policy.
14. A host computer, characterized in that, The host is equipped with a bus device, and a proxy response component is set on the bus device. The host includes a memory, a processor, and a communication component; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component and is used to execute the one or more computer instructions for: Initiate a non-delivery access request to the bus device according to the bus protocol; If a response message for characterizing the failure of the bus device returned by the proxy response component is received, perform protection processing on the virtual machines running thereon according to a preset protection policy; Wherein, the response message is sent by the proxy response component on behalf of the bus device in the case of a preset type of failure occurring in the bus device.
15. The host computer according to claim 14, wherein When the processor performs protection processing on the virtual machine according to a preset protection policy, it is specifically used for: In the host, determine the initiator corresponding to the non-delivery access request; Transfer the response message to the initiator to trigger the initiator to protect the virtual machine associated with the non-delivery access request.
16. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, cause the one or more processors to execute the fault handling method according to any one of claims 1-12.