Server fault processing method and system, electronic equipment, computer storage medium and computer program product

By monitoring the status of processors and storage devices in the cloud server system, migrating virtual machines and switching storage device mounts, the problem of poor server fault handling reliability is solved, and rapid recovery and reduced impact scope is achieved.

CN120508454APending Publication Date: 2025-08-19HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410181979.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-18
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In cloud server systems, server fault handling is poor and the maintenance time is expensive, especially when the processor system or local storage device fails, resulting in the overall server being unavailable.

Method used

By monitoring the working status of the processor system and local storage devices, if a failure occurs, migrate the virtual machine and switch the storage device to a healthy processor system to ensure that the virtual machine can still access the original storage resources.

Benefits of technology

It reduces the impact of server failure on the entire system, improves the reliability of fault handling, avoids the server being completely unavailable, and improves maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508454A_ABST
    Figure CN120508454A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a server fault processing method and system, electronic equipment, a computer storage medium and a computer program product. The server fault processing method comprises the following steps: monitoring the working states of at least two processor systems in a target server when accessing respective mounted local storage devices, each processor system comprising computing resources deployed to a virtual machine of the processor system, the local storage device mounted on each processor system comprises storage resources deployed to the virtual machine of the processor system; if the working state of the first processor system in the at least two processor systems indicates that the first processor system has a system fault, migrating a virtual machine in the first processor system to a second processor system in the at least two processor systems, and switching the local storage device mounted on the first processor system to be mounted on the second processor system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular to a server fault handling method, system, electronic device, computer storage medium, and computer program product. Background Art

[0002] Generally speaking, a cloud server system includes multiple servers that are connected to each other in communication. A virtual machine can be configured in the physical machine of each server, and distributed computing is achieved through data communication between virtual machines in different servers.

[0003] For example, a processor system in a server may include computing resources such as memory and a processor. A local storage device is mounted on the processor system as a storage device. A virtual machine running in the memory accesses the local storage device through the processor to perform access to the local storage device. For example, the processor reads data from the local storage device to the virtual machine, or writes data from the virtual machine to the local storage device.

[0004] Since the virtual machines in each server are interconnected and have data communication, when the processor system or local storage device in a server fails, the required repair time and cost are high, resulting in poor reliability of server fault handling. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a server fault handling method, system, electronic device, computer storage medium, and computer program product to solve the above problems.

[0006] According to a first aspect of an embodiment of the present invention, a server failure handling method is provided, comprising: monitoring the working status of at least two processor systems in a target server when accessing their respective mounted local storage devices, wherein each processor system includes computing resources of a virtual machine deployed to the processor system, and the local storage device mounted to each processor system includes storage resources of the virtual machine deployed to the processor system; if the working status of a first processor system among the at least two processor systems indicates that a system failure has occurred in the first processor system, migrating the virtual machine in the first processor system to a second processor system among the at least two processor systems, and switching the local storage device mounted on the first processor system to be mounted on the second processor system.

[0007] According to a second aspect of an embodiment of the present invention, a server failure handling method is provided, comprising: monitoring the working status of a first local storage device among at least two local storage devices in a target server when accessed by a mounted processor system, the processor system including computing resources of a virtual machine deployed to the processor system, and the at least two local storage devices including storage resources of the virtual machine deployed to the processor system; if the working status indicates that a device failure has occurred in the first local storage device, switching the processor system from mounting the first local storage device to mounting a second local storage device among the at least two local storage devices.

[0008] According to a third aspect of an embodiment of the present invention, a server fault handling device is provided, comprising: a monitoring module, which monitors the working status of at least two processor systems in a target server when accessing their respective mounted local storage devices, wherein each processor system includes computing resources of a virtual machine deployed to the processor system, and the local storage device mounted to each processor system includes storage resources of the virtual machine deployed to the processor system; and a switching module, which migrates the virtual machines in the first processor system to a second processor system among the at least two processor systems, and switches the local storage device mounted on the first processor system to be mounted on the second processor system, if the working status of a first processor system among the at least two processor systems indicates that a system fault has occurred in the first processor system.

[0009] According to a fourth aspect of an embodiment of the present invention, a server fault handling device is provided, comprising: a monitoring module, monitoring the working status of a first local storage device among at least two local storage devices in a target server when accessed by a mounted processor system, the processor system including computing resources of a virtual machine deployed to the processor system, and the at least two local storage devices including storage resources of the virtual machine deployed to the processor system; a switching module, switching the processor system from mounting the first local storage device to mounting a second local storage device among the at least two local storage devices if the working status indicates that a device failure has occurred in the first local storage device.

[0010] According to a fifth aspect of an embodiment of the present invention, a server fault handling system is provided, comprising: at least one server and a monitoring device, wherein the monitoring device is configured to execute the method according to the first aspect or the second aspect.

[0011] According to the sixth aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect or the second aspect.

[0012] According to a seventh aspect of an embodiment of the present invention, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect or the second aspect is implemented.

[0013] According to an eighth aspect of an embodiment of the present invention, a computer program product is provided, comprising a computer program / instruction, which implements the method described in the first aspect or the second aspect when executed by a processor.

[0014] In the solution of the embodiment of the present invention, at least two processor systems in the target server can migrate the virtual machine in the first processor system to the second processor system when the working status of the first processor system indicates that the first processor system has a system failure. Accordingly, the local storage device mounted on the first processor system is switched to be mounted on the second processor system, so that the virtual machine can still access the storage resources of the local storage device previously mounted on the first processor system through the second processor system, avoiding the situation where the target server is completely unavailable, reducing the impact range of the failure of the target server on the entire server system, and improving the reliability of server failure handling. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0016] Figure 1 Schematic diagram of the physical machine configuration of some example servers.

[0017] Figure 2 The present invention is a flowchart of the steps of a server failure handling method according to some embodiments of the present invention.

[0018] Figure 3 for Figure 2 A schematic diagram of a physical machine configuration of a server of an embodiment.

[0019] Figure 4A and Figure 4B for Figure 2Schematic diagram of states of a processor system before and after switching according to an embodiment.

[0020] Figure 5A and Figure 5B for Figure 2 A schematic diagram of the state of the local storage device before and after switching of an embodiment.

[0021] Figure 6 The present invention is a flowchart of the steps of a server failure handling method according to some embodiments of the present invention.

[0022] Figure 7 Schematic block diagram of a server fault handling device according to some other embodiments of the present invention.

[0023] Figure 8 Schematic block diagram of a server fault handling device according to some other embodiments of the present invention.

[0024] Figure 9 2 is a structural block diagram of a server fault handling system according to some other embodiments of the present invention.

[0025] Figure 10 Schematic diagram of the structure of electronic devices according to other embodiments of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention.

[0027] The specific implementation of the embodiment of the present invention is further described below with reference to the accompanying drawings of the embodiment of the present invention.

[0028] Figure 1 The following is a diagram of the physical machine configuration of some example servers. The server system includes multiple servers. Figure 1 As shown, in some examples, a processor system in a server can implement computing resources such as memory and processors, and a local storage device, such as a hard disk, can be mounted on the processor system as a storage device. In the case of multiple processors, the multiple processors can communicate through a unified point-to-point interconnect (UPI) to improve computing efficiency.

[0029] In addition, the virtual machine running in the memory accesses the local storage device through the processor and the local storage device via a communication bus such as PCIe. For example, the processor reads data from the local storage device to the virtual machine via the communication bus, or the processor writes data from the virtual machine to the local storage device via the communication bus. In other words, the processor system can access the local storage device mounted to the processor system.

[0030] A local storage device can be an instance of a cloud storage system built from a hard disk such as an SSD or HDD. For example, a physical machine can be configured with a local disk instance in a cloud storage system.

[0031] For processor systems based on most processor architectures (for example, the X86 architecture), when a failure occurs in the processor system itself or when a failure occurs in the local storage device mounted on the processor system, the entire server may fail. Since the virtual machines in the processor systems of each server in the server system are uniformly deployed and configured, in order to minimize the impact on the entire server system, the data on the local storage device mounted on the failed processor system cannot be quickly migrated, and redeploying the local storage device may cause data loss in the local storage device. In addition, the local storage devices mounted on the processor systems of each server are uniformly deployed and configured as cloud disk instances. If a local storage device fails, it will also cause a server failure. In this case, in order to minimize the impact on the entire server system, the failed local storage device needs to be isolated, and then the physical machine needs to be restarted to restore to normal.

[0032] That is, when a processor system or local storage device in a server fails, the maintenance time and cost of such a server are high, and the reliability of server failure handling is poor. To this end, the embodiment of the present invention provides the following series of solutions to solve the above problems.

[0033] Figure 2 The present invention is a flowchart of the steps of a server failure handling method according to some embodiments of the present invention. Figure 2 The server fault handling method can be performed by the monitoring device, including:

[0034] S210: Monitor the working status of at least two processor systems in the target server when accessing their respective mounted local storage devices, wherein each processor system includes computing resources of a virtual machine deployed to the processor system, and the local storage device mounted to each processor system includes storage resources of the virtual machine deployed to the processor system.

[0035] It should be understood that the processor system can be implemented as a host motherboard, which includes the computing resources required by the virtual machines, such as processors such as a CPU and GPU, and memory such as RAM. Local storage devices can provide storage resources for the virtual machines, such as non-volatile storage media such as SSDs or HDDs. In other words, a target server including at least two processor systems is configured as a multi-host system.

[0036] It should also be understood that the working status of the processor system includes but is not limited to: processor load, processor frequency, memory usage status, etc. Specifically, processor load refers to the number of tasks and workload being processed by the processor system. The processor load can be evaluated by monitoring indicators such as processor usage and average load. Processor frequency refers to the operating frequency of the processor. The working status of the processor system can be used to determine whether the processor is working normally by monitoring the frequency changes of the processor. Memory usage status refers to the usage of memory in the processor system, including various memory indicators such as used memory, available memory, cache, etc., and memory usage can be monitored to understand whether the memory is sufficient to meet the needs of the current task. That is to say, when the working status of the first processor system does not exceed the normal range, the first processor system has a normal system; when the working status of the first processor system exceeds the normal range, the first processor system has a system failure.

[0037] S220: If the working status of the first processor system among the at least two processor systems indicates that a system failure occurs in the first processor system, the virtual machine in the first processor system is migrated to the second processor system among the at least two processor systems, and the local storage device mounted on the first processor system is switched to be mounted on the second processor system.

[0038] It should be understood that the operating status of the processor system may indicate the status of the processor system's access path to the mounted local storage device. When a device along the access path fails, the operating status of the processor system may be abnormal. For example, the operating status of the processor system may indicate a device failure of the local storage device and / or a system failure of the processor system.

[0039] It should also be understood that when executing the above-mentioned switching process, the operating system installed in the second processor system (e.g., motherboard) can detect the newly connected local storage device through a hardware detection program such as BIOS or UEFI. Then, the operating system can read the relevant information of the local storage device (e.g., device model, serial number, etc.) and assign an identifier to the local storage device. Then, the operating system of the second processor system establishes data communication between the second processor system and the local storage device by loading the local storage device driver. Then, the operating system can map the storage space of the local storage device to the file system structure of the operating system, so that the virtual machine can read and write data on the local storage device through the operating system.

[0040] It should also be understood that each processor system can independently install its own operating system, and the processor system that installs the operating system realizes the mounting of the local storage device by associating the identified storage partition of the mounted local storage device with the file system.

[0041] In the solution of the embodiment of the present invention, at least two processor systems in the target server can migrate the virtual machine in the first processor system to the second processor system when the working status of the first processor system indicates that the first processor system has a system failure. Accordingly, the local storage device mounted on the first processor system is switched to be mounted on the second processor system, so that the virtual machine can still access the storage resources of the local storage device previously mounted on the first processor system through the second processor system, avoiding the situation where the target server is completely unavailable, reducing the impact range of the failure of the target server on the entire server system, and improving the reliability of server failure handling.

[0042] In some embodiments, a cloud service management system can configure a virtual machine in a physical server and a local storage device of a cloud storage system. Figure 3 As shown, the physical machine of the server 10 includes at least two processor systems 20. The processor system 20 can be implemented as a motherboard in the physical machine. The processor system 20 includes memory and a processor such as a CPU, providing computing resources for the virtual machine, and the local storage device 40 provides storage resources for the virtual machine.

[0043] Furthermore, the cloud service management system may configure a virtual machine management agent such as a hypervisor in the processor system 20 to implement creation, destruction, and maintenance of virtual machines.

[0044] Furthermore, the monitoring device 50 is disposed outside each server to manage each server. For example, a monitoring agent 500 of the monitoring device 50 in the processor system 20 may be configured in the processor system 20 of the physical machine.

[0045] It should be understood that the monitoring device 50 can be part of a cloud service management system or a monitoring system independent of the cloud service management system. The monitoring agent 500 can be part of a virtual machine management agent or an agent independent of the virtual machine management agent.

[0046] Further, in Figure 4A and Figure 4B In the embodiment, the physical machine 10 includes a first processor system 21 and a second processor system 22, each of which is provided with a memory and a processor such as a CPU. The monitoring device 50 is configured with a monitoring agent 510 in the first processor system 21 and a monitoring agent 520 in the second processor system 22.

[0047] In addition, at least two processor systems (for example, the first processor system 21 or the second processor system 22) and their local storage devices are connected to a bus switch 30 such as a PCIe switch via a bus, and each local storage device 40 can be mounted to the corresponding processor system through the bus switch 30 so that the virtual machine in the processor system can access the local storage device 40 mounted on the processor system.

[0048] exist Figure 4A In the example of , the virtual machine A previously deployed in the first processor system 21 can access the local storage device 40 mounted to the first processor system 21. Figure 4B In the example, virtual machine A is deployed to the second processor system 22 through virtual machine migration. Since the previously accessed local storage device 40 is mounted to the second processor system 22 through switching, virtual machine A can still access the storage resources of the previous local storage device 40.

[0049] Without loss of generality, switching the local storage device 40 mounted on the first processor system 21 to be mounted on the second processor system 22 includes: switching the local storage device 40 mounted on the first processor system 21 to be mounted on the second processor system 22 by switching the port connection state of the bus switch. Based on the above processing method, the bus switch can quickly implement the switching of the local storage device from the first processor system to the second processor system.

[0050] In some examples, as an example of switching a local storage device mounted on a first processor system to be mounted on a second processor system by switching the port connection state of a bus switch, an association table entry between the bus port between the local storage device mounted on the first processor system and the bus switch and the bus port of the first processor system can be set to disabled, and an association table entry between the local storage device mounted on the first processor system and the bus switch and the bus port of the second processor system can be set to enabled.

[0051] Alternatively, as an example of switching a local storage device mounted on a first processor system to be mounted on a second processor system by switching the port connection state of a bus switch, the association entry between the bus port between the local storage device mounted on the first processor system and the bus switch and the bus port of the first processor system can be deleted from a port mapping table indicating the port connection state of the bus switch, and an association entry between the bus port between the local storage device mounted on the first processor system and the bus switch and the bus port of the second processor system can be established. Based on this processing approach, soft switching can be achieved between bus ports, thereby improving the switching efficiency of the processor system through the bus switch.

[0052] In other embodiments, migrating a virtual machine from a first processor system to a second processor system among at least two processor systems includes: obtaining configuration information and state information of the virtual machine from the memory of the first processor system; and restoring the virtual machine to the memory of the second processor system among the at least two processor systems based on the configuration information and state information of the virtual machine. Based on the above processing approach, migration of the virtual machine from the first processor system to the second processor system is reliably achieved.

[0053] For example, the configuration information of a virtual machine includes but is not limited to: the hardware configuration of the virtual machine, the network configuration, etc. The status information of a virtual machine includes but is not limited to: the memory status of the virtual machine, register values, hard disk status, network status, etc.

[0054] Specifically, when restoring a virtual machine to the memory of a second processor system among the at least two processor systems, a new virtual machine instance can be created in the second processor system using virtual machine management software such as VMware or VirtualBox, and configured with the same hardware configuration as the virtual machine in the first processor system. The virtual machine configuration file can then be imported into the virtual machine management software of the second processor system, ensuring that the configuration file correctly matches the virtual machine instance.

[0055] Furthermore, the state information of the virtual machine can be imported into the virtual machine instance in the second processor system. For example, information such as memory state, register value, hard disk state, etc. can be imported into the corresponding virtual machine instance.

[0056] In other embodiments, the server fault handling system may be implemented by a cloud service management system. Monitoring devices in the cloud service management system may configure monitoring agents for at least two processor systems in the server. For example, a first monitoring agent may be configured for the first processor system, and a second monitoring agent may be configured for the second processor system. The first monitoring agent runs in the memory of the first processor system, and the second monitoring agent runs in the memory of the second processor system.

[0057] Furthermore, the monitoring device can obtain the working status of each processor system when accessing the mounted local storage device through the monitoring agent of the processor system, and obtain the configuration information and status information of the virtual machine deployed to the processor system from the memory of the processor system.

[0058] It should be understood that each processor system and each mounted local storage device can be connected to a bus switch through a bus port. The cloud service management system can also configure the port configuration agent of the bus switch. The port configuration agent can manage the bus switch through a port mapping table between the bus port of the processor system and the bus port of the local storage device. For example, the monitoring device can send a port mapping table to the bus switch (for example, the port configuration agent) to enable the bus switch to configure each bus port to satisfy the port connection status in each associated table entry of the port mapping table.

[0059] Furthermore, obtaining the configuration information and status information of the virtual machine from the memory of the first processor system includes: obtaining the configuration information and status information of the virtual machine from the memory of the first processor system to a first monitoring agent running in the memory of the first processor system, and recording the configuration information and status information of the virtual machine by the first monitoring agent. Based on the above processing method, the first monitoring agent is conducive to achieving reliable monitoring while being compatible with the configuration of the processor system.

[0060] For example, the monitoring agent of the processor system can be loaded into the memory of the processor system when the processor system is started. The monitoring agent and the virtual machine transmit data through inter-process communication. For example, the virtual machine can periodically send its configuration information and status information to the monitoring agent, and the monitoring agent can then transmit the configuration information and status information of the virtual machine to the monitoring device. Alternatively, the monitoring agent can actively send a request to obtain configuration information and status information to the virtual machine (for example, in response to a monitoring request from the monitoring device), and after obtaining the configuration information and status information of the virtual machine, return it to the monitoring device.

[0061] In other embodiments, restoring a virtual machine to the memory of a second processor system of at least two processor systems based on the virtual machine's startup configuration information and state information includes: sending the virtual machine's configuration information and state information to a second monitoring agent in the memory of the second processor system, and loading the virtual machine's configuration information and state information into the memory of the second processor system via the second monitoring agent; and creating the virtual machine in the memory of the second processor system based on the virtual machine's configuration information and state information. Based on the above processing approach, the software agent further achieves compatibility with the processor system's configuration.

[0062] In other embodiments, the virtual machine's state information includes communication state information with virtual machines in other servers. Creating the virtual machine in the memory of the second processor system based on the virtual machine's configuration information and state information includes: creating the virtual machine in the memory of the second processor system such that the virtual machine is in a communication state with the virtual machines in other servers as indicated by the communication state information. Based on the above processing approach, in the server system, the communication state between the virtual machine in the target server and the other servers is maintained before a failure occurs in the local storage device or processor system, thereby reducing the scope of the target server's impact on other servers.

[0063] Further, in Figure 5A and Figure 5B In the embodiment, the physical machine 10 includes a first processor system 21 and a second processor system 22, each of which is provided with a memory and a processor such as a CPU. The monitoring device 50 is configured with a monitoring agent 510 in the first processor system 21 and a monitoring agent 520 in the second processor system 22.

[0064] In this example, the working status of the first local storage device 41 mounted to the first processor system 21 indicates that a device failure occurs in the first local storage device 41 among at least two local storage devices. Accordingly, the virtual machine in the first processor system 21 switches from accessing the storage resources of the local storage device 41 to accessing the storage resources of the local storage device 42.

[0065] Furthermore, by switching the port connection state of the bus switch 30, the first processor system or the second processor system switches from mounting the first local storage device to mounting the second local storage device of the at least two local storage devices. Based on the above processing method, the bus switch 30 can realize a quick switch from the first local storage device to the second local storage device.

[0066] Furthermore, in a port mapping table indicating the port connection status of the bus switch, the association entry between the bus port of the first processor system or the second processor system and the bus port of the first local storage device is deleted, and an association entry between the bus port of the first processor system or the second processor system and the bus port of the second local storage device is established. Based on this processing approach, soft switching can be achieved between the bus ports, thereby improving the switching efficiency of the local storage device through the bus switch.

[0067] Figure 6 The present invention is a flowchart of the steps of a server failure handling method according to some embodiments of the present invention. Figure 6 The server fault handling method can be performed by the monitoring device, including:

[0068] S610: Monitor the working status of a first local storage device among at least two local storage devices in a target server when accessed by a mounted processor system, where the processor system includes computing resources of a virtual machine deployed to the processor system, and at least two local storage devices include storage resources of the virtual machine deployed to the processor system.

[0069] It should be understood that the operating status of the first local storage device includes, but is not limited to, device space, device read / write speed, IOPS (input / output operations per second), device response time, device failure rate, etc. Specifically, device space refers to the remaining available space on the local storage device. The remaining capacity of the storage device can be monitored by monitoring disk usage to ensure that the storage space is sufficient to meet current data storage needs. Device read / write speed refers to the read and write speed of the local storage device. The performance of the storage device can be evaluated by monitoring the disk read and write speed to ensure that it can meet the data reading and writing requirements of the processor system. IOPS refers to the number of input and output operations that the local storage device can complete per second. By monitoring the IOPS metric, the performance and response speed of the storage device can be evaluated. Device response time refers to the response time of the local storage device to a request from the processor system. By monitoring the storage device's response time, the access latency of the processor system to the storage device can be understood. The device failure rate refers to the probability of failure of the local storage device. By monitoring the failure rate of the storage device, the reliability of the storage device can be evaluated and timely measures can be taken to prevent data loss. That is, when the working state of the first local storage device does not exceed the normal range, the system of the first local storage device is normal; when the working state of the first local storage device exceeds the normal range, the system of the first local storage device is faulty.

[0070] S620: If the working status indicates that a device failure occurs in the first local storage device, switch the processor system from mounting the first local storage device to mounting a second local storage device among the at least two local storage devices.

[0071] In the solution of the embodiment of the present invention, the processor system of the target server mounts at least two local storage devices. When the working status of the first local storage device indicates that a device failure has occurred in the first local storage device, the mounted local storage device of the processor system can be switched from the first local storage device to the second local storage device, thereby avoiding the situation where the target server is completely unavailable, reducing the impact range of the failure of the target server on the server system, and improving the reliability of server fault handling.

[0072] Specifically, if Figure 5A and Figure 5B As shown, the working status of the first local storage device 41 mounted to the first processor system 21 indicates that a device failure occurs in the first local storage device 41 among at least two local storage devices. Accordingly, the virtual machine in the first processor system 21 switches from accessing the storage resources of the local storage device 41 to accessing the storage resources of the local storage device 42.

[0073] It should be understood that the server fault handling system can be implemented by a cloud service management system. The monitoring device in the cloud service management system can configure monitoring agents for each local storage device mounted on the processor system (i.e., the first local storage device), with the monitoring agents running in the memory of the processor system. Furthermore, the monitoring device can obtain the operating status of the first local storage device when the processor system accesses the mounted first local storage device through the monitoring agent of the first local storage device.

[0074] It should also be understood that each processor system and each mounted local storage device can be connected to a bus switch through a bus port. The cloud service management system can also configure the port configuration agent of the bus switch. The port configuration agent can manage the bus switch through a port mapping table between the bus port of the processor system and the bus port of the local storage device. For example, the monitoring device can send a port mapping table to the bus switch (for example, the port configuration agent) to enable the bus switch to configure each bus port to satisfy the port connection status in each associated table entry of the port mapping table.

[0075] In other embodiments, the processor system and at least two local storage devices are connected to a bus switch via a bus. Switching the processor system from mounting the first local storage device to mounting the second local storage device among the at least two local storage devices includes: switching the processor system from mounting the first local storage device to mounting the second local storage device among the at least two local storage devices by switching the port connection state of the bus switch 30. Based on the above processing approach, the bus switch can achieve quick switching from the first local storage device to the second local storage device.

[0076] It should be understood that the drivers for at least two local storage devices may be pre-installed in the operating system of the processor system. Then, before the processor system switches from mounting the first local storage device to mounting the second local storage device, the operating system may associate the identified storage partition of the first local storage device with the file system to achieve mounting.

[0077] In other embodiments, as an example of switching the processor system from mounting a first local storage device to mounting a second local storage device among at least two local storage devices, in a port mapping table indicating the port connection status of a bus switch, an association table entry between the bus port of the processor system and the bus port of the first local storage device can be set to disabled, and an association table entry between the bus port of the processor system and the bus port of the second local storage device can be set to enabled.

[0078] Alternatively, as another example of switching the processor system from mounting the first local storage device to mounting the second local storage device among the at least two local storage devices, the association entry between the bus port of the processor system and the bus port of the first local storage device can be deleted from a port mapping table indicating the port connection status of the bus switch, and an association entry between the bus port of the processor system and the bus port of the second local storage device can be established. Based on this processing approach, soft switching can be implemented between the bus ports, thereby improving the switching efficiency of local storage devices through the bus switch.

[0079] Figure 7 Schematic block diagram of a server fault handling device according to some other embodiments of the present invention. Figure 7 Server fault handling device and Figure 2 The corresponding server fault handling methods include:

[0080] Monitoring module 710 monitors the working status of at least two processor systems in the target server when accessing their respective mounted local storage devices, wherein each processor system includes the computing resources of the virtual machine deployed to the processor system, and the local storage device mounted to each processor system includes the storage resources of the virtual machine deployed to the processor system.

[0081] The switching module 720 migrates the virtual machine in the first processor system to the second processor system among the at least two processor systems if the working status of the first processor system among the at least two processor systems indicates that the first processor system has a system failure, and switches the local storage device mounted on the first processor system to be mounted on the second processor system.

[0082] In the solution of the embodiment of the present invention, at least two processor systems in the target server can migrate the virtual machine in the first processor system to the second processor system when the working status of the first processor system indicates that the first processor system has a system failure. Accordingly, the local storage device mounted on the first processor system is switched to be mounted on the second processor system, so that the virtual machine can still access the storage resources of the local storage device previously mounted on the first processor system through the second processor system, avoiding the situation where the target server is completely unavailable, reducing the impact range of the failure of the target server on the entire server system, and improving the reliability of server failure handling.

[0083] In other embodiments, the at least two processor systems and their local storage devices are connected to a bus switch via a bus, and each processor system accesses the local storage device mounted on that processor system through the bus switch. The switching module is specifically configured to switch the local storage device mounted on the first processor system to be mounted on the second processor system by switching the port connection state of the bus switch.

[0084] In other embodiments, the switching module is specifically used to: delete the association table entry between the bus port between the local storage device mounted on the first processor system and the bus switch and the bus port of the first processor system in the port mapping table indicating the port connection status of the bus switch, and establish the association table entry between the bus port between the local storage device mounted on the first processor system and the bus switch and the bus port of the second processor system.

[0085] In other embodiments, the switching module is specifically used to: obtain the configuration information and status information of the virtual machine from the memory of the first processor system; and restore the virtual machine to the memory of the second processor system among the at least two processor systems based on the configuration information and status information of the virtual machine.

[0086] In other embodiments, the switching module is specifically used to: obtain the configuration information and status information of the virtual machine from the memory of the first processor system to the first monitoring agent running in the memory of the first processor system, and record the configuration information and status information of the virtual machine through the first monitoring agent.

[0087] In other embodiments, the switching module is specifically used to: send the configuration information and status information of the virtual machine to the second monitoring agent in the memory of the second processor system, and load the configuration information and status information of the virtual machine into the memory of the second processor system through the second monitoring agent; and create the virtual machine in the memory of the second processor system based on the configuration information and status information of the virtual machine.

[0088] In some other embodiments, the state information of the virtual machine includes communication state information with virtual machines in other servers. The switching module is specifically configured to: create the virtual machine in the memory of the second processor system, so that the virtual machine is in a communication state indicated by the communication state information with the virtual machines in other servers.

[0089] Figure 8 Schematic block diagram of a server fault handling device according to some other embodiments of the present invention. Figure 8 Server fault handling device and Figure 6 The corresponding server fault handling methods include:

[0090] a monitoring module 810 for monitoring an operating status of a first local storage device among at least two local storage devices in a target server when accessed by a mounted processor system, the processor system including computing resources of a virtual machine deployed to the processor system, the at least two local storage devices including storage resources of the virtual machine deployed to the processor system;

[0091] The switching module 820 switches the processor system from mounting the first local storage device to mounting a second local storage device among the at least two local storage devices if the working status indicates that a device failure occurs in the first local storage device.

[0092] In the solution of the embodiment of the present invention, the processor system of the target server mounts at least two local storage devices. When the working status of the first local storage device indicates that a device failure has occurred in the first local storage device, the mounted local storage device of the processor system can be switched from the first local storage device to the second local storage device, thereby avoiding the situation where the target server is completely unavailable, reducing the impact range of the failure of the target server on the server system, and improving the reliability of server fault handling.

[0093] In some other embodiments, the processor system and the at least two local storage devices are both connected to a bus switch via a bus. The switching module is specifically configured to switch the processor system from mounting the first local storage device to mounting the second local storage device of the at least two local storage devices by switching a port connection state of the bus switch.

[0094] In other embodiments, the switching module is specifically used to: delete the association table entry between the bus port of the processor system and the bus port of the first local storage device in the port mapping table indicating the port connection status of the bus switch, and establish the association table entry between the bus port of the processor system and the bus port of the second local storage device.

[0095] The specific implementation of the modules in each device can refer to the corresponding descriptions of the corresponding steps and units in the above method embodiments, and have corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the above method embodiments, and will not be repeated here.

[0096] Figure 9 2 is a structural block diagram of a server fault handling system according to some other embodiments of the present invention. Figure 9 The server fault processing system includes at least one server 10 and a monitoring device 50.

[0097] Reference Figure 10 , shows a schematic structural diagram of an electronic device according to another embodiment of the present invention. The specific embodiment of the present invention does not limit the specific implementation of the electronic device.

[0098] like Figure 10 As shown, the electronic device may include: a processor 1002 for executing a program 1010 , a communication interface 1004 , a memory 1006 , and a communication bus 1008 .

[0099] The processor, the communication interface, and the memory communicate with each other via a communication bus.

[0100] Communication interface, used to communicate with other electronic devices or servers.

[0101] The processor is used to execute a program, and specifically can execute the server failure handling method of any of the above embodiments.

[0102] Specifically, the program may include program codes including computer operation instructions.

[0103] The processor may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or different types of processors, such as one or more CPUs and one or more ASICs.

[0104] Memory is used to store programs. The memory may include high-speed RAM memory and may also include non-volatile memory (non-volatile memory), such as at least one disk storage.

[0105] The program may include multiple computer instructions. Specifically, the program may enable the processor to execute operations corresponding to each method described in any of the aforementioned method embodiments through the multiple computer instructions.

[0106] The specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, and will not be repeated here.

[0107] An embodiment of the present invention further provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments. The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk.

[0108] An embodiment of the present invention further provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to each method in the above-mentioned multiple method embodiments.

[0109] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0110] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0111] The method according to the embodiment of the present invention described above can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.

[0112] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present invention.

[0113] The above implementation methods are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Ordinary technicians in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the scope of patent protection of the embodiments of the present invention should be defined by the claims.

Claims

1. A server failure handling method, comprising: Monitoring the operating status of at least two processor systems in the target server when accessing local storage devices mounted thereto, wherein each processor system includes computing resources of a virtual machine deployed thereto, and the local storage devices mounted thereto include storage resources of the virtual machine deployed thereto; If the working status of the first processor system among the at least two processor systems indicates that a system failure occurs in the first processor system, the virtual machine in the first processor system is migrated to the second processor system among the at least two processor systems, and the local storage device mounted on the first processor system is switched to be mounted on the second processor system.

2. The method according to claim 1, wherein The at least two processor systems and their local storage devices are connected to a bus switch via a bus, and each processor system accesses the local storage device mounted on the processor system through the bus switch; Switching the local storage device mounted on the first processor system to be mounted on the second processor system includes: By switching the port connection state of the bus switch, the local storage device mounted on the first processor system is switched to be mounted on the second processor system.

3. The method according to claim 2, wherein: Switching the local storage device mounted on the first processor system to be mounted on the second processor system by switching the port connection state of the bus switch includes: In the port mapping table indicating the port connection status of the bus switch, an association table entry between the bus port between the local storage device mounted on the first processor system and the bus switch and the bus port of the first processor system is deleted, and an association table entry between the bus port between the local storage device mounted on the first processor system and the bus switch and the bus port of the second processor system is established.

4. The method according to claim 1, wherein Migrating the virtual machine in the first processor system to the second processor system in the at least two processor systems includes: Obtaining configuration information and status information of the virtual machine from the memory of the first processor system; The virtual machine is restored to a memory of a second processor system among the at least two processor systems based on the configuration information and state information of the virtual machine.

5. The method according to claim 4, wherein Acquiring configuration information and status information of the virtual machine from the memory of the first processor system includes: The configuration information and status information of the virtual machine are acquired from the memory of the first processor system to a first monitoring agent running in the memory of the first processor system, and the configuration information and status information of the virtual machine are recorded by the first monitoring agent.

6. The method according to claim 5, wherein: Restoring the virtual machine to a memory of a second processor system of the at least two processor systems based on startup configuration information and state information of the virtual machine includes: Sending the configuration information and status information of the virtual machine to a second monitoring agent in the memory of the second processor system, and loading the configuration information and status information of the virtual machine into the memory of the second processor system through the second monitoring agent; The virtual machine is created in the memory of the second processor system based on the configuration information and state information of the virtual machine.

7. The method according to claim 6, wherein: The state information of the virtual machine includes communication state information with virtual machines in other servers; Creating the virtual machine in the memory of the second processor system based on the configuration information and state information of the virtual machine includes: The virtual machine is created in the memory of the second processor system, so that the virtual machine and the virtual machines in other servers are in a communication state indicated by the communication state information.

8. A method for handling a server failure, comprising: monitoring an operating status of a first local storage device of at least two local storage devices in a target server when accessed by the mounted processor system, the processor system including computing resources of a virtual machine deployed to the processor system, the at least two local storage devices including storage resources of the virtual machine deployed to the processor system; If the working status indicates that a device failure occurs in the first local storage device, the processor system is switched from mounting the first local storage device to mounting a second local storage device of the at least two local storage devices.

9. The method according to claim 8, wherein The processor system and the at least two local storage devices are connected to a bus switch via a bus; Switching the processor system from mounting the first local storage device to mounting the second local storage device among the at least two local storage devices includes: By switching the port connection state of the bus switch, the processor system is switched from mounting the first local storage device to mounting the second local storage device of the at least two local storage devices.

10. The method according to claim 9, wherein: Switching the processor system from mounting the first local storage device to mounting the second local storage device of the at least two local storage devices by switching the port connection state of the bus switch includes: In the port mapping table indicating the port connection status of the bus switch, an association table entry between the bus port of the processor system and the bus port of the first local storage device is deleted, and an association table entry between the bus port of the processor system and the bus port of the second local storage device is established.

11. A server fault handling system, comprising: at least one server; A monitoring device, wherein the monitoring device is configured to execute the method according to any one of claims 1 to 10.

12. An electronic device comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, where the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 10.

13. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising a computer program / instruction, which implements the method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Cited By

  • Fault processing method of server and electronic equipment

    CN120994453A

  • Server troubleshooting methods and electronic equipment

    CN120994453B