Fault processing method of server and electronic equipment

By monitoring processor operating data in real time and converting physical address and memory data in the server, and building virtual configuration components, the problem of failure of the processor in traditional dual-processor servers cannot be taken over, and seamless migration and resource optimization are achieved.

CN120994453AActive Publication Date: 2025-11-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511519723.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In traditional dual-processor server architectures, when one processor fails, the other cannot take over, causing server operation to be interrupted, and backup devices remain idle for a long time, resulting in wasted resources.

Method used

By monitoring processor operating data in real time, disconnecting the communication link of the faulty processor, and transferring its physical address and memory data to another processor, a virtual configuration component is built to take over the faulty processor, achieving seamless migration.

Benefits of technology

It enables rapid takeover of faulty processors, avoids server outage risks, prevents resource waste of backup devices, and improves hardware utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994453A_ABST
    Figure CN120994453A_ABST
Patent Text Reader

Abstract

The invention discloses a fault processing method of a server and electronic equipment, and relates to the technical field of servers, the server comprises a first processor and a second processor, and operation data of the first processor is monitored in real time; when it is judged that the first processor is in an abnormal state according to the operation data, communication links of all ports of the first processor are disconnected; converting the physical address of each port of a first processor in the server to a corresponding physical address in a second processor, and converting the storage address of the memory data of the first processor in the server to a corresponding storage address in the second processor, so as to construct a communication link and a memory space of the ports in the second processor; constructing a virtual configuration component in the second processor, and configuring the virtual configuration component according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor; and taking over the first processor through the virtual configuration component in the second processor.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and in particular to a server fault processing method and an electronic device. BACKGROUND

[0002] In a traditional dual-processor server architecture, data isolation exists between processors, cross-domain access cannot be achieved, so when one of the processors fails, the associated firmware and hardware will fail, and the devices managed by the processor cannot be taken over by other processors due to the strong binding relationship between the physical link and the configuration space of the processor, resulting in server operation interruption. SUMMARY

[0003] The present application provides a server fault processing method and an electronic device to at least solve the problem that when one of the processors fails, it cannot be taken over by other processors, resulting in server operation interruption.

[0004] The present application provides a server fault processing method, the server comprising a first processor and a second processor, comprising: monitoring the running data of the first processor in real time; when it is determined according to the running data that the first processor is in an abnormal state, disconnecting the communication links of each port of the first processor; converting the physical addresses of each port of the first processor in the server to the corresponding physical addresses in the second processor, and converting the storage addresses of the memory data of the first processor in the server to the corresponding storage addresses in the second processor, to build the communication links of the ports and the memory space in the second processor; building a virtual configuration component in the second processor, and configuring the virtual configuration component according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor; taking over the first processor through the virtual configuration component in the second processor.

[0005] The present application also provides a server fault processing device, the server comprising a first processor and a second processor, comprising: an acquisition module for monitoring the running data of the first processor in real time; a processing module for disconnecting the communication links of each port of the first processor when it is determined according to the running data that the first processor is in an abnormal state; the processing module is also used for converting the physical addresses of each port of the first processor in the server to the corresponding physical addresses in the second processor, and converting the storage addresses of the memory data of the first processor in the server to the corresponding storage addresses in the second processor, to build the communication links of the ports and the memory space in the second processor; The processing module is further configured to construct a virtual configuration component in the second processor, and configure the virtual configuration component according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor. The processing module is further configured to take over the first processor by the virtual configuration component in the second processor.

[0006] The application further provides an electronic device, which comprises a memory configured to store a computer program, and a processor configured to execute the computer program to implement the steps of the fault processing method of the server.

[0007] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the fault processing method of the server.

[0008] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the fault processing method of the server.

[0009] In the application, the server comprises a first processor and a second processor, the running data of the first processor is monitored in real time, when it is judged according to the running data that the first processor is in an abnormal state, the communication links of each port of the first processor are disconnected, the physical addresses of each port of the first processor in the server are converted to the corresponding physical addresses in the second processor, and the storage addresses of the memory data of the first processor in the server are converted to the corresponding storage addresses in the second processor, so as to construct the communication links of the ports and the memory space in the second processor, a virtual configuration component is constructed in the second processor, and the virtual configuration component is configured according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor, and the first processor is taken over by the virtual configuration component in the second processor. In this scheme, the fast takeover of the fault CPU and the seamless migration of the device resources are realized in the server architecture of the dual CPU, the physical binding relationship between the CPU and the PCIe device in the traditional architecture is broken, so that the other CPU can take over the device resources when the single CPU fails, the service interruption risk caused by the hardware failure is completely eliminated, and the problem of resource waste caused by the long-term idling of the backup device in the traditional dual machine redundancy scheme is avoided. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 A flow of a server fault processing method provided by an embodiment of the present application Figure 1 ; Figure 2 A flow of a server fault processing method provided by an embodiment of the present application Figure 2 ; Figure 3 A structural diagram of a server fault processing device provided by an embodiment of the present application Figure 4 A structural diagram of an electronic device provided by an embodiment of the present application DETAILED DESCRIPTION

[0012] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0013] It should be noted that, in the description of the present application, the terms “comprising”, “containing” or any other variants thereof are intended to cover the non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes the elements inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0014] It should be noted that, in the embodiments of the present application, the words “exemplary” or “for example” are used to represent an example, illustration or description. Any embodiment or design scheme described as “exemplary” or “for example” in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words “exemplary” or “for example” are intended to present the relevant concept in a specific manner.

[0015] In traditional dual-CPU server architecture, the binding relationship between CPU and PCIe device has significant fault tolerance defects. PCIe (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus standard used to connect various hardware components on the motherboard, such as graphics cards, storage devices, and network adapters. When one CPU (CPU0) fails, the PCIe devices managed by it cannot be directly taken over by another normally running CPU (CPU1) due to the strong binding between the physical link and the configuration space of the Root Complex (RC) of the failed CPU, resulting in service interruption or the need to rely on inefficient software redirection solutions (such as network storage mapping or virtualization migration). Root Complex (RC) is the core component of PCIe bus architecture, responsible for connecting CPU / memory subsystems and PCIe devices, managing the initialization, data communication, address mapping, and device configuration of the entire PCIe subsystem. Current processor fault takeover technologies face the following core problems: PCIe device access isolation: PCIe bus domains and memory domains are naturally isolated, and cross-domain access requires address mapping through an Address Translation Unit (ATU). In traditional solutions, CPU1 cannot dynamically take over the PCIe bus domain address space of CPU0, causing the device access path to be interrupted. For example, the PCIe device BAR space of CPU0 (such as 0xA0000000-0xAFFFFFFF) cannot be directly accessed by CPU1 through the original physical link after CPU0 fails.

[0016] Redundancy architecture resource waste: Traditional dual-machine hot standby or Dual Root architecture requires independent PCIe devices for each CPU, resulting in low hardware utilization and high cost. For example, financial transaction systems need to deploy redundant Graphics Processing Units (GPUs) for each CPU, resulting in resource idle rates exceeding 40%.

[0017] Switching delay and reliability bottleneck: Existing fault switching relies on operating system-level device reenumeration, which takes up to several seconds to tens of seconds, and cannot meet the needs of high real-time scenarios (such as autonomous driving and high-frequency trading). In addition, PCIe link state detection only relies on a single signal (such as a heartbeat packet), which has a high risk of misjudgment and may cause cascading failures.

[0018] Interrupt and DMA consistency problem: the interrupt vector (MSI-X) and DMA channel of CPU0 device are bound to the APIC and memory space of CPU0, so when the fault switches, the interrupt routing and Input-Output Memory Management Unit (IOMMU) mapping need to be reconfigured, which is easy to cause data loss or system crash.

[0019] At present, in the related art, the processor fault takeover can be realized by the multiplexing function of the PCIe switch to share the device. The master and standby servers periodically synchronize the device configuration information to the shared storage, and the standby server needs to re-scan the PCIe bus and load the device driver when the fault switches. The core problem is that the device switching depends on the operating system level bus enumeration and driver initialization process, which causes the interrupt time to exceed 2 seconds, and does not solve the cross-CPU domain memory access conflict problem, which requires an additional software layer to intercept the DMA request. In addition, this method needs to reserve backup resources for each device, and the hardware utilization is low.

[0020] In addition, the related art also discloses that the device can be mounted to two independent servers at the same time through the PCIe switch. When the master server fails, the standby server needs to reset the device and reload the driver through the software layer to take over the device. The defect is that the operating system needs to restart the device driver during the switching process, which causes the service interruption time to exceed 5 seconds, and does not solve the synchronization problem of the device state (such as DMA context and interrupt binding). This scheme relies on the device reset operation of the software layer and cannot achieve seamless takeover, and the hardware resources need to be configured with independent Root Complex for double servers, which increases the complexity and cost of device management.

[0021] In addition, the related art also discloses that the baseboard management controller is in communication connection with the IO expander and a plurality of interface modules, each PCIE port of the processor is in communication connection with a corresponding switching module, and two PCIE ports of the switching module are in communication connection with a first type interface and a second type interface in the interface module respectively; the baseboard management controller acquires the in-place signal of the first type interface in the target interface module through the IO expander, and judges whether the first type interface in the target interface module is connected with a device and generates a corresponding control signal; according to the control signal, one PCIE port in the target switching module connected with the target interface module is controlled to be in communication with the corresponding PCIE port of the processor through the IO expander. This method needs all CPUs to share the same PCIe domain, and the hardware expandability is limited.

[0022] In summary, in the conventional dual-path architecture, the system startup process relies on CPU0 (main CPU0) to initialize hardware resources. If CPU0 is physically damaged (such as power short circuit, kernel breakdown), the firmware code (such as BIOS module) and hardware associated with CPU0 will be completely disabled, resulting in the failure of CPU1 to take over the startup process, and the overall system downtime. Traditional PCIe devices are strongly bound to the CPU, and when CPU0 fails, the PCIe devices managed by CPU0 cannot be directly accessed by CPU1 due to physical link isolation and loss of configuration space. When the connection is re-established, the server needs to interrupt the service and then connect again, which cannot achieve seamless connection.

[0023] To solve the above technical problems, the embodiments of the present application provide a server fault processing method, the server comprising a first processor and a second processor, real-time monitoring of the running data of the first processor; when it is judged that the first processor is in an abnormal state according to the running data, the communication link of each port of the first processor is disconnected; the physical address of each port of the first processor in the server is converted to the corresponding physical address in the second processor, and the storage address of the memory data of the first processor in the server is converted to the corresponding storage address in the second processor, to construct the communication link of the port and the memory space in the second processor; a virtual configuration component is constructed in the second processor, and the virtual configuration component is configured according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor; the first processor is taken over by the virtual configuration component in the second processor. In this scheme, the fast takeover of the faulty CPU and the seamless migration of the device resources are realized in the dual-CPU server architecture, breaking the physical binding relationship between the CPU and the PCIe device in the traditional architecture, so that another CPU can take over the device resources when a single CPU fails, completely eliminating the service interruption risk caused by hardware failure, and avoiding the problem of resource waste caused by the long-term idling of the backup device in the traditional dual-redundancy scheme.

[0024] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] As Figure 1 shown, Figure 1 The flowchart of the server fault processing method provided by the embodiments of the present application can include the following steps: 101、Real-time monitoring of the running data of the first processor.

[0026] In the embodiments of the present application, the server comprises a first processor and a second processor, the first processor and the second processor can be the same processor, when neither of the first processor and the second processor is faulty, only one of the processors can be used to process services, and when one of the processors is faulty, the other processor can be used to take over, so that the running data of the first processor, i.e., the processor currently running, can be monitored in real time, and the running data can at least comprise temperature data (such as a PROCHOT signal) of the first processor and link connection state data (PCIe device link).

[0027] In some embodiments, when the running data of the first processor is monitored, the running data of the first processor can be collected by a BMC management controller or a special monitoring circuit at a certain time interval.

[0028] 102. When it is judged according to the running data that the first processor is in an abnormal state, the communication links of the ports of the first processor are disconnected.

[0029] In the embodiments of the present application, in the process of monitoring the running data in real time, if it is judged according to the running data that the first processor is in an abnormal state, it indicates that the first processor cannot normally run at this time, and the server can have a service interruption, so the second processor is needed to take over the work of the first processor, but first the communication links of the first processor need to be disconnected immediately, which is equivalent to cutting off the command of the first processor, equivalent to taking away its official seal and intercom, so that it cannot order the subordinate departments (hardware devices) any more, and nanosecond-level electrical isolation is achieved.

[0030] 103. The physical addresses of the ports of the first processor in the server are converted to the corresponding physical addresses in the second processor, and the storage addresses of the memory data of the first processor in the server are converted to the corresponding storage addresses in the second processor, so as to construct the communication links of the ports and the memory space in the second processor.

[0031] In the embodiment of the present application, when the first processor is faulty, the ports of the first processor can no longer be used, and the second processor takes over the work of the first processor, that is, subsequent data transmission and service processing in the server are performed through the ports on the second processor, so it is necessary to convert the port address of the first processor in the server to the address corresponding to the port in the second processor, so that subsequent data transmission and service processing can be directly performed according to the address of the port in the second processor; in addition, the processor is provided with a memory space, and the memory space stores data, which can include system data generated when the processor is running, can include received and transmitted data, and can include service data, so it is also necessary to transfer the memory data of the first processor, which can be specifically achieved by converting the storage address of the memory data of the first processor to the corresponding storage address in the second processor, that is, after the storage address of the memory data of the first processor is converted to the corresponding storage address in the second processor, the second processor can take over the memory data of the first processor according to the storage address, so as to build the memory space in the second processor, which is equivalent to copying the memory space of the first processor.

[0032] 104. A virtual configuration component is built in the second processor, and the virtual configuration component is configured according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor.

[0033] In the embodiment of the present application, in order to enable the second processor to take over according to the configuration of the first processor and achieve the non-susceptible takeover of the processor, a virtual configuration component can be built in the second processor, and then the virtual configuration component is configured according to the configuration information of the first processor, so that the configuration information of the configured virtual configuration component is the same as the configuration information of the first processor, so that when the second server is running, the user or the external device thinks that the processor is still intact and cannot perceive that the original first processor has been taken over.

[0034] It should be noted that the virtual configuration component can be a virtual Root Complex, and the Root Complex (RC) is a core component of the PCIe bus architecture, which is responsible for connecting the CPU / memory subsystem and the PCIe device, and managing the initialization, data communication, address mapping and device configuration of the entire PCIe subsystem. In short, the virtual Root Complex is equivalent to a "virtual general manager's office", and the layout, extension number (ECAM configuration space, Bus / Device / Function number) of the "virtual general manager's office" and the first processor are exactly the same, so that the other employees (external devices, operating systems and application programs) of the company completely feel that the general manager has been replaced, thinking that everything is normal, so as to achieve the non-susceptible takeover of the processor.

[0035] 105、through the virtual configuration component in the second processor to take over the first processor.

[0036] In the embodiment of the present application, after the physical addresses of each port in the second processor, the storage addresses of the memory data, and the virtual configuration component are all set, the second processor is equivalent to a copy processor of the first processor at this time, and thus the first processor can be taken over by the virtual configuration component in the second processor. Subsequent data transmission and service processing in the server can be implemented by the second processor.

[0037] In the embodiment of the present application, the server includes the first processor and the second processor, and the running data of the first processor is monitored in real time. When it is determined according to the running data that the first processor is in an abnormal state, the communication link of each port of the first processor is disconnected. The physical address of each port of the first processor in the server is converted to the corresponding physical address in the second processor, and the storage address of the memory data of the first processor in the server is converted to the corresponding storage address in the second processor, so as to construct the communication link of the port and the memory space in the second processor. The virtual configuration component is constructed in the second processor, and the virtual configuration component is configured according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor. The first processor is taken over by the virtual configuration component in the second processor. In this scheme, the rapid takeover of the faulty CPU and the seamless migration of the device resources are implemented in the server architecture of the dual CPU, the physical binding relationship between the CPU and the PCIe device in the traditional architecture is broken, so that the single CPU failure can be taken over by the other CPU, the service interruption risk caused by the hardware failure is completely eliminated, and the problem of resource waste caused by the long-term idling of the backup device in the traditional dual-machine redundancy scheme is avoided.

[0038] As shown in Figure 2 , another flowchart of the server fault processing method provided by the embodiment of the present application is shown. The method can include the following steps: Figure 2 201, the running data of the first processor is monitored in real time.

[0039] In the embodiment of the present application, for the description of step 201, please refer to the detailed description of step 101 in the above embodiment, and the embodiment of the present application will not be described again.

[0040] 202, when it is detected that the duration that the temperature of the first processor is less than the preset temperature value reaches the preset duration, and / or, it is detected that the link connection state of the first processor is in an abnormal state, it is determined that the first processor is in an abnormal state.

[0041] ​In the embodiment of the present application, the temperature of the processor will rise during normal operation and be obviously higher than the room temperature, and if the processor is not running due to failure, the temperature of the processor will drop and approach the room temperature, so whether the processor is faulty can be determined according to the temperature of the processor. If the temperature of the first processor is less than the preset temperature value for a preset time length, that is, the first processor has been at a lower temperature for a period of time, it can be considered that the first processor is currently faulty. In addition, during normal operation of the processor, the link can normally communicate, and if the link cannot communicate, that is, the link connection is faulty, it can also indicate that the processor is faulty, so whether the processor is faulty can be determined according to the link connection state of the processor. If the link connection state of the first processor is in an abnormal state, it can be considered that the first processor is currently faulty.

[0042] It should be noted that the above temperature and link connection state are two ways of judging the failure of the processor. As long as at least one of the above judging methods is met, it can be considered that the processor is faulty. That is, when the temperature of the first processor is less than the preset temperature value for a preset time length, or when the link connection state of the first processor is in an abnormal state, or when the temperature of the first processor is less than the preset temperature value for a preset time length and the link connection state of the first processor is in an abnormal state, it can be considered that the first processor is in an abnormal state.

[0043] In the embodiment of the present application, whether the processor is faulty is determined by two parameters of the temperature and the link connection state of the processor, which improves the accuracy of fault diagnosis and enables another processor to take over in time when the processor is determined to be faulty.

[0044] In some embodiments, dual independent power supply and signal isolation are required at the hardware level. Independent VRM (voltage regulation module) is configured for the first processor and the second processor, a super capacitor (such as Maxwell 2.7V3000F) is added to the PCIe device slot, and the device power supply is maintained ≥60 seconds after the first processor is powered off. The BMC monitors the PROCHOT signal of the first processor through the SMBus, and polls the PCIe link LTSSM state (once every 10 ms), and continuously abnormal determines that the first processor is invalid.

[0045] 203、When it is determined according to the running data that the first processor is in an abnormal state, the communication link of each port of the first processor is disconnected by adjusting the global reset signal value of the first processor.

[0046] In the embodiment of the present application, when the communication link of the port of the first processor is disconnected, the global reset signal, that is, the PCIe_RST# signal, can be used for setting, the PCIe_RST# signal is a global low active signal in the PCIe bus for initializing the internal logic of the device, and the reset of the device register and the state machine is triggered by a hardware means, so as to ensure that the device enters a controllable initial state when the system starts or is abnormal. The PCIe_RST# signal has the characteristic of low active, that is, the reset is triggered when the signal is pulled low, and is released after being pulled high. The reset can restore the internal registers (such as configuration registers, state registers), timing logic and state machine of the device to the initial value, and clear the temporary state in operation. When the first processor is normally running, the value of the global reset signal can be 1, and when the first processor is in an abnormal state, the value of the global reset signal can be adjusted from 1 to 0, so as to disconnect the communication link of each port of the first processor.

[0047] In some embodiments, the FPGA-based arbitration module immediately pulls down the PCIe_RST# signal of the first processor after detecting that the first processor is invalid, and cuts off the PERST#, CLKREQ# and other key control signals through an optocoupler relay, to realize nanosecond-level electrical isolation.

[0048] 204、updating the physical addresses of the ports of the first processor stored in the routing table of the server to the physical addresses of the corresponding ports in the second processor, to construct the communication link of the ports in the second processor.

[0049] In the embodiment of the present application, in the process of converting the physical addresses of the ports of the first processor, the physical addresses of the ports can be stored in the routing table of the server, so that the physical addresses of the ports stored in the routing table can be updated, that is, the physical addresses of the ports of the first processor originally stored are updated to the physical addresses of the corresponding ports in the second processor, the routing table of the server is dynamically modified through the PCIe switch, the PCIe ports (Port 0-7) originally belonging to the first processor are redirected to the second processor domain, and the address space conversion (for example, the Outbound ATU maps the device BAR address 0xA0000000 of the first processor to 0xD0000000 of the second processor domain, and the Inbound reversely analyzes) is performed, to ensure that the device physical layer is transparently accessed. Simply, it can be understood that the internal telephone switchboard (PCIe switch) of the company will transfer all internal telephone calls (data requests) originally made to the first processor to the online line of the second processor. The process can be completed within 10 milliseconds (ms).

[0050] 205. In response to a memory access request, determine the storage address of the memory data from the first processor in the server, and redirect the storage address of the memory data to the pre-stored memory pool corresponding to the second processor, so as to construct memory space in the second processor.

[0051] In this embodiment, when converting memory data, it can be done based on a memory access request. This memory access request can be equivalent to a "key" to access the memory space of the first processor. The storage address of the memory data is determined from the first processor and then redirected to the corresponding new address in the second processor. This new address can be understood as a shared memory space, specifically a CMA memory pool or CXL memory pool that can be accessed at high speed by the CPU. This allows the data originally stored in the first processor to be found in the pre-stored memory pool, and new data will also be stored in the pre-stored memory pool.

[0052] In simple terms, the files being processed by the first processor are locked in their own "file cabinet." The system uses the IOMMU (a type of memory management unit) to assign a "master key" to the second processor, dynamically mapping the address of the first processor's file cabinet (e.g., 0xA0000000) to a new address (e.g., 0xD0000000) that the second processor can recognize. This allows the second processor to seamlessly continue processing this data, achieving "zero-copy" transfer—the data itself doesn't need to be physically moved; only a different key is used to open it. This shared file space can be understood as a CMA or CXL memory pool that the CPU can access at high speed.

[0053] In some embodiments, PCIe device takeover and virtualization reconstruction achieve cross-domain transparent access based on a non-transparent bridge (NTB). Core technologies include address isolation mapping, cross-domain interrupt forwarding, and virtual root complex emulation. The NTB maps the physical address of a device in the first processor domain (e.g., 0xA0000000) to the virtual address of a device in the second processor domain (0xD0000000) by configuring an Outbound / Inbound Address Translation Unit (ATU), and isolates the TLP transmission paths between the two domains to avoid address conflicts. At the interrupt level, the NTB intercepts MSI-X messages from the first processor device, rewrites the target address to the APIC ID of the second processor (0xFEE00000→0xFEE01000), and dynamically allocates vector numbers using the kernel interrupt remapping table (IRT) to achieve cross-domain interrupt response. DMA operations redirect device requests to the CMA memory pool reserved in the CPU1 domain through the IOMMU dynamic page table, and achieve zero-copy transfer using NTB reverse address resolution.

[0054] In the embodiment of the present application, the continuity of the service process and the data consistency are ensured by updating and redirecting the physical addresses of each port in the first processor and the storage addresses of the memory data, in combination with the memory cache synchronization and IO address remapping technology.

[0055] 206. constructing a virtual configuration component in the second processor.

[0056] In the embodiment of the present application, for the description of step 206, please refer to the detailed description of step 104 in the above embodiment, and the embodiment of the present application will not be repeated here.

[0057] 207. starting the firmware image in the second processor.

[0058] In the embodiment of the present application, through the redundant design mechanism of the BMC, when the fault of the firmware storage area (SPI Flash) of the first processor is detected, the system is automatically started by switching to the independent firmware image of the second processor to complete the system startup, ensuring the high availability under hardware failure.

[0059] It should be noted that the second processor is equipped with an independent SPI Flash storage area, which stores a UEFI firmware image completely isolated from the first processor. The image contains a complete startup program (such as U-Boot, kernel, file system), which can independently perform hardware initialization, driver loading and operating system boot. The firmware image is isolated from the firmware in the first processor, avoiding single point failure propagation. The firmware image contains a complete UEFI startup process (such as SEC, PEI, DXE stage), which can independently perform hardware initialization, driver loading and operating system boot without relying on the firmware state of the first processor. The BMC monitors the fault state of the first processor through a hardware timeout mechanism (such as a Watchdog timer), and if the first processor has a fault, the BMC automatically switches the SPI interface to the Flash area of the second processor to trigger the backup image to start.

[0060] 208. retrieving the configuration information of the first processor from the pre-stored global device list.

[0061] 209. configuring the virtual configuration component according to the configuration information of the first processor based on the firmware image.

[0062] In the embodiment of the present application, when the virtual configuration component is configured, the configuration information of the first processor can be used, and the configuration information of the first processor can be stored in the global device list, so the configuration information of the first processor can be retrieved from the global device list for configuration, so that the configuration information of the virtual configuration component is the same as that of the first processor.

[0063] In some embodiments, the global device list (GDT) should actually be a global descriptor table (GDT), which is a core data structure for memory segmentation management in the protected mode of the x86 architecture. The GDT is a table stored in memory that defines the properties of each memory segment in the system, including the base address, segment limit (i.e., the size of the segment), access permissions, and the like. Each memory segment can be understood as a contiguous region in memory for storing code, data, or stack information, and the like.

[0064] In some embodiments, the BIOS pre-scans the dual-domain PCIe device information and writes it into the NVRAM (such as Intel Optane PMem) to form a global device list (GDT), and the second processor copies the ECAM configuration space (Bus / Device / Function number and Class Code) of the first processor to the virtual Root Complex according to the GDT, so that the operating system mistakenly believes that the device is still mounted at the original location.

[0065] In simple terms, the virtual Root Complex created is like a fake "hardware management center" for another "brain" (such as the second processor) in a server, allowing it to take over and control hardware devices (such as network cards, GPUs) that do not belong to it. When a major "brain" (such as the first processor) in the server suddenly fails, in order to prevent the hardware devices it manages from being paralyzed, the system quickly copies a virtual hardware configuration that is identical to the original environment on another healthy "brain" (the second processor) according to a pre-prepared "hardware map" (global device list GDT). In this way, the operating system running on the second processor will mistakenly believe that all hardware is still mounted at the original location, allowing it to continue using these devices without modifying any software, achieving nearly seamless failover.

[0066] In some embodiments, the firmware layer initialization and configuration achieve seamless takeover through dual Boot ROM switching and virtualization context reconstruction, and the second processor can take over and control the hardware originally managed by the first processor through a virtual configuration component with the same configuration as the first processor, achieving nearly seamless failover of the processor.

[0067] 210、delete the configuration data of the first processor stored in the configuration management list, and update the configuration management list according to the configuration data of the second processor.

[0068] In the embodiments of the present application, since the first processor is faulty, the second processor takes over the first processor subsequently, and therefore the configuration data of the processor can be updated in the configuration management list, that is, the originally stored configuration data of the first processor is updated to the configuration data of the second processor.

[0069] In some embodiments, the configuration management list can be an Advanced Configuration and Power Interface (ACPI) table. ACPI is a standard interface for operating systems to interact with hardware. The ACPI table is used for system hardware configuration and power management, and contains sub-tables such as DSDT (Differential System Description Table), MADT (Multiple APIC Description Table), SRAT (System Resource Affinity Table), etc.

[0070] MADT (Multiple APIC Description Table) is a table in ACPI, which describes the configuration of Advanced Programmable Interrupt Controller (APIC) in the system, including Local APIC entries. Each CPU core usually has a Local APIC entry for interrupt handling and inter-processor communication. When the first processor fails, the Local APIC entry of the first processor in the MADT can be deleted, which means that the interrupt controller description of the first processor is removed in the ACPI table.

[0071] SRAT (System Resource Affinity Table) is another table in ACPI, which describes the affinity between memory and processor, that is, which memory region is associated with which processor core or node. When the first processor fails, the memory topology in SRAT can be reconfigured, and the association between memory and processor is reorganized. The memory node originally associated with the first processor is associated with the second processor domain, which means that the memory originally allocated to the first processor is now associated with the domain of the second processor, and the memory node is re-allocated to avoid conflicts when the operating system identifies the CPU and memory.

[0072] 211、In response to the hot plug trigger signal, loading the system driver from the second processor.

[0073] 212、Based on the system driver, taking over the first processor through the virtual configuration component.

[0074] In this embodiment, operating system recovery and driver loading achieve seamless service recovery through kernel hot-plug event triggering and dynamic resource redirection. After the hardware layer switch is complete, the kernel simulates a PCIe device hot-plugging via an ACPI HotPlug event (e.g., _OST0x103), triggering the driver manager (e.g., udev) to load pre-registered backup drivers (e.g., NVIDIA GPU drivers). The kernel receives a simulated "new device inserted" signal (ACPI HotPlug event), which triggers it to rescan the hardware. However, since the virtual configuration components ("virtual office") have already been configured, the kernel scan will find all hardware intact and running in place. Therefore, it will automatically load the drivers for these hardware components (e.g., GPU, network card), allowing services to resume normal operation.

[0075] 213. Monitor the first data sent by the first processor and the second data received in real time.

[0076] 214. Intercept the first data and the second data, and forward both the first data and the second data to the second processor.

[0077] In this embodiment of the application, after the first processor fails, all services of the first processor will be implemented through the second processor. Therefore, the second data sent to the first processor needs to be forwarded to the second processor. Similarly, the data sent by the first processor may be corrupted due to the failure of the first processor, so it is necessary to intercept the first data sent by the first processor.

[0078] In some embodiments, the NTB intercepts MSI-X messages from the first processor's device, rewrites the target address to the APIC ID of the second processor (0xFEE00000→0xFEE01000), and dynamically allocates vector numbers using the kernel interrupt remapping table (IRT) to achieve cross-domain interrupt response. This can be understood as follows: previously, when the hardware department encountered a problem, it needed to request instructions from the first processor (sending an interrupt signal, such as MSI-X). Now, these requests are automatically transferred to the second processor (APIC controller) by the "Interrupt Routing Table (IRT)," ensuring unimpeded communication and timely response.

[0079] In this embodiment, when the first processor malfunctions and cannot continue to work, the data sent by the first processor may be abnormal and therefore needs to be intercepted. At the same time, the data sent to the first processor will not be processed further. Therefore, the data sent and received by the first processor can be forwarded to the second processor. This can ensure data integrity, prevent data loss, and enable the second processor to process the data in a timely manner.

[0080] In the embodiments of the present application, a full-process high availability solution is realized by a hardware-firmware-operating system cooperative mechanism, in which the second processor takes over and accesses the PCIe device of the first processor when the first processor fails. First, the BMC reads the PROCHOT (temperature signal) of the first processor through the SMBus, and if the abnormality is detected for three times in succession, the BMC triggers the system management interrupt to notify the second processor to start taking over and to pull down the PCIe_RST# signal of the first processor, so that the optocoupler relay cuts off the control link and triggers the FPGA to perform hardware isolation. Then, the BMC switches the SPI Flash chip selection signal to force the UEFI image of the second processor to start. After the UEFI image of the second processor starts, the PCIe switch dynamically switches the device port to the second processor domain, and the non-transparent bridge (NTB) synchronously updates the address mapping table (0xA0000000→0xD0000000), to realize the transparent migration at the physical layer. The operating system activates the backup driver instance through the hot plug event, dynamically remaps the IOMMU page table and migrates the GPU memory context, and finally restores the service.

[0081] In some embodiments, the following is a relatively popular description to explain the server failure processing method provided by the present application. Under normal circumstances, CPU0 of a high-end server is responsible for processing all daily business, and CPU1 is in standby state and synchronizes the work progress of CPU0 at all times. To ensure that nothing goes wrong, the system is designed with an extremely rapid automatic fault switching solution.

[0082] Step 1: “Health monitor” (fault detection) that is always vigilant: The “health monitor” (usually a BMC management controller or a special monitoring circuit) checks the “pulse” (such as the PROCHOT signal) of CPU0 and whether the PCIe device link managed by it is in smooth communication every extremely short time (such as 10 milliseconds).

[0083] Quick determination: Once CPU0 abnormality or PCIe device link disconnection is found for several times in succession, the monitor will determine that CPU0 loses the working ability within a few milliseconds.

[0084] Start emergency: The monitor will immediately notify the “emergency command center” (usually a FPGA chip) to start the emergency plan.

[0085] Step 2: Isolation and takeover (hardware layer switching): After receiving the alarm, the emergency command center (FPGA) will immediately perform two key operations: Forced Rest: First, it will cut off the "command power" of CPU0, such as pulling down its PCIe reset signal, which is equivalent to taking away its seal and intercom, so that it can no longer command subordinate departments (hardware devices), achieving nanosecond-level electrical isolation.

[0086] Seamless takeover: The command center will start a "virtual general manager's office" (virtual Root Complex) for CPU1. The layout, phone extension number (ECAM configuration space, Bus / Device / Function number) of this virtual office is exactly the same as that of CPU0. At the same time, the company's "internal telephone switchboard" (PCIe switch) will instantly transfer all internal calls (data requests) originally made to CPU0 to CPU1 online. In this way, the company's other employees (operating system and application) feel that the general manager has been replaced, thinking that everything is normal.

[0087] Step 3: Update the address book and permissions (firmware and memory switching): Now CPU1 has taken over CPU0 to support the server to boot normally, but it needs to be able to handle business normally.

[0088] Update the employee handbook: The company's "basic administrative system" (BIOS / UEFI firmware) will start a backup plan (independent UEFI image) specifically for CPU1. It will modify the company's "organizational chart" (ACPI table) to remove CPU0's name and clearly mark CPU1 as the new leader, and tell the system that all resources (such as memory) now report to CPU1 to avoid confusion.

[0089] Authorized access to shared file cabinet: The most critical step is to handle the "shared file cabinet" (memory data). The files being processed by CPU0 are locked in its own file cabinet. The system will configure a "universal key" for CPU1 through IOMMU (a memory management unit) to dynamically map the address of CPU0's file cabinet (such as 0xA0000000) to a new address that CPU1 can recognize (such as 0xD0000000). In this way, CPU1 can seamlessly continue to process these data, achieving "zero-copy" transmission, that is, the data itself does not need to be physically moved, only a new key is used to open it. This shared file space can be understood as a CMA memory pool or CXL memory pool that all CPUs can access at high speed.

[0090] Step 4: Resume normal work (operating system and driver recovery): When the hardware and firmware level switching is complete, the company's "management" (operating system kernel) needs to be notified.

[0091] Simulated hot plug: the kernel receives a simulated "new device insertion" signal (ACPI HotPlug event), which triggers it to rescan the hardware. But since the "virtual office" has already been set up, the kernel finds all the hardware intact and hanging in place after the scan, so it automatically loads the drivers for the hardware (such as GPU, network card) and restores normal business operation.

[0092] Redirected work report: previously, the hardware department encountered problems and needed to report to the general manager (send an interrupt signal, such as MSI-X). Now these reports will be automatically forwarded to the "secretariat" (APIC controller) of CPU1 by the "interrupt routing table (IRT)", ensuring smooth communication and timely response.

[0093] In some embodiments, through the synergistic innovation of hardware redundancy design, firmware-level fault isolation, and operating system dynamic resource management, the rapid takeover of the failed CPU and the seamless migration of device resources are realized in the dual-CPU server architecture, bringing multiple substantial improvements to high-availability computing scenarios. The core value lies in breaking the physical binding relationship between CPU and PCIe devices in the traditional architecture through non-transparent bridge technology and virtualization layer reconstruction, so that when a single CPU fails, the other CPU can take over the device resources, completely eliminating the risk of service interruption caused by hardware failure, while avoiding the waste of resources caused by the long-term idling of backup devices in traditional dual-redundancy schemes. This scheme uses global device table pre-registration and interrupt redirection mechanisms to ensure that the operating system can dynamically identify and load cross-domain device drivers without restarting, and combines memory cache synchronization and IO address remapping technology to ensure the continuity of business processes and data consistency, making it particularly suitable for real-time transaction systems and high-throughput data processing scenarios that have extremely low tolerance for service interruption. The hardware layer realizes precise isolation of the fault domain through intelligent power management and signal isolation mechanisms, and cooperates with PCIe exchange network dynamic routing switching technology to avoid the risk of fault propagation at the physical level, while supporting on-demand scheduling of heterogeneous computing devices and parallel processing of mixed loads, significantly improving hardware resource utilization and system flexibility.

[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and the necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better implementation.

[0095] As shown in Figure 3 The embodiments of the present application also provide a server fault processing device, which can include: The acquisition module 301 is configured to monitor the running data of the first processor in real time. The processing module 302 is configured to disconnect the communication links of the ports of the first processor when it is determined that the first processor is in an abnormal state according to the running data. The processing module 302 is further configured to convert the physical addresses of the ports of the first processor in the server to the corresponding physical addresses in the second processor, and convert the storage addresses of the memory data of the first processor in the server to the corresponding storage addresses in the second processor, so as to construct the communication links of the ports and the memory space in the second processor. The processing module 302 is further configured to construct a virtual configuration component in the second processor, and configure the virtual configuration component according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor. The processing module 302 is further configured to take over the first processor by the virtual configuration component in the second processor.

[0096] In some embodiments, the processing module 302 is further configured to determine that the first processor is in an abnormal state when it is detected that the temperature of the first processor is less than a preset temperature value for a preset time length, and / or it is detected that the link connection state of the first processor is in an abnormal state.

[0097] In some embodiments, the processing module 302 is specifically configured to disconnect the communication links of the ports of the first processor by adjusting the value of the global reset signal of the first processor when it is determined that the first processor is in an abnormal state according to the running data.

[0098] In some embodiments, the processing module 302 is specifically configured to update the physical addresses of the ports of the first processor stored in the routing table of the server to the physical addresses of the corresponding ports in the second processor, so as to construct the communication links of the ports in the second processor.

[0099] In some embodiments, the processing module 302 is specifically configured to determine the storage address of the memory data from the first processor in the server in response to the memory access request, and redirect the storage address of the memory data to the corresponding pre-stored memory pool in the second processor, so as to construct the memory space in the second processor.

[0100] In some embodiments, the processing module 302 is specifically configured to start a firmware image in the second processor, and the firmware image is isolated from the firmware in the first processor. The processing module 302 is specifically configured to call the configuration information of the first processor from a pre-stored global device list. The processing module 302 is specifically configured to configure the virtual configuration component according to the configuration information of the first processor based on the firmware image.

[0101] In some embodiments, the processing module 302 is further configured to delete the configuration data of the first processor stored in the configuration management list, and update the configuration management list according to the configuration data of the second processor.

[0102] In some embodiments, the processing module 302 is specifically configured to load the system driver from the second processor in response to the hot plug trigger signal. The processing module 302 is specifically configured to take over the first processor through the virtual configuration component based on the system driver.

[0103] In some embodiments, the obtaining module 301 is further configured to monitor the first data sent by the first processor and the second data received in real time. The processing module 302 is further configured to intercept the first data and the second data, and forward the first data and the second data to the second processor.

[0104] In the embodiments of the present application, the features of the embodiments corresponding to the fault handling device of the server can be referred to the related descriptions of the embodiments corresponding to the fault handling method of the server, which will not be repeated here.

[0105] As shown in Figure 4 The embodiments of the present application further provide an electronic device, which comprises a memory 401 and a processor 402, the memory 401 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any of the above-mentioned server fault handling method embodiments.

[0106] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above-mentioned server fault handling method embodiments when running.

[0107] In an exemplary embodiment, the above-mentioned computer readable storage medium can include but is not limited to: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0108] The embodiments of the present application further provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned server fault handling method embodiments.

[0109] The embodiment of the present application further provides another computer program product, comprising a nonvolatile computer readable storage medium, the nonvolatile computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps in any of the server fault processing method embodiments.

[0110] Those skilled in the art will further appreciate that the functions of the examples described herein-based units and algorithm steps can be implemented using electronic hardware, computer software, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their general functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0111] The above describes in detail the process monitoring of the storage system provided by the present application. The principles and implementation manners of the present application are described herein by applying specific examples, and the above description of the examples is only applicable to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for handling server faults, characterized in that, The server includes a first processor and a second processor, and the method includes: Real-time monitoring of the operating data of the first processor; When the first processor is determined to be in an abnormal state based on the operating data, the communication links of each port of the first processor are disconnected. The physical addresses of each port of the first processor in the server are translated to the corresponding physical addresses in the second processor, and the storage addresses of the memory data of the first processor in the server are translated to the corresponding storage addresses in the second processor, so as to construct the communication links and memory space of the ports in the second processor. A virtual configuration component is constructed in the second processor, and the virtual configuration component is configured according to the configuration information of the first processor, so that the configuration information of the virtual configuration component is the same as the configuration information of the first processor; The first processor is taken over by the virtual configuration component in the second processor.

2. The method according to claim 1, characterized in that, After real-time monitoring of the operating data of the first processor, the method further includes: When the temperature of the first processor is detected to be lower than the preset temperature value for a preset duration, and / or when the link connection status of the first processor is detected to be abnormal, it is determined that the first processor is in an abnormal state.

3. The method according to claim 1, characterized in that, The step of disconnecting the communication links of each port of the first processor when it is determined from the operating data that the first processor is in an abnormal state includes: When the first processor is determined to be in an abnormal state based on the operating data, the communication links of each port of the first processor are disconnected by adjusting the global reset signal value of the first processor.

4. The method according to claim 1, characterized in that, The step of translating the physical addresses of each port of the first processor in the server to the corresponding physical addresses in the second processor includes: The physical addresses of each port of the first processor stored in the routing table of the server are updated to the physical addresses of the corresponding ports in the second processor, so as to establish a communication link for the ports in the second processor.

5. The method according to claim 1, characterized in that, The step of converting the storage address of the memory data in the first processor of the server to the corresponding storage address in the second processor includes: In response to a memory access request, the storage address of the memory data is determined from the first processor in the server, and the storage address of the memory data is redirected to the pre-stored memory pool corresponding to the second processor, so as to construct the memory space in the second processor.

6. The method according to claim 1, characterized in that, The step of configuring the virtual configuration component according to the configuration information of the first processor includes: The firmware image in the second processor is started, and the firmware image is isolated from the firmware in the first processor. Retrieve the configuration information of the first processor from the pre-stored global device list; Based on the firmware image, the virtual configuration component is configured according to the configuration information of the first processor.

7. The method according to claim 1, characterized in that, The method further includes: Delete the configuration data of the first processor stored in the configuration management list, and update the configuration management list according to the configuration data of the second processor.

8. The method according to claim 1, characterized in that, The step of taking over the first processor through the virtual configuration component in the second processor includes: In response to a hot-plug trigger signal, load the system driver from the second processor; Based on the system driver, the first processor is taken over by the virtual configuration component.

9. The method according to claim 1, characterized in that, The method further includes: Real-time monitoring of the first data sent by the first processor and the second data received; The first data and the second data are intercepted, and both the first data and the second data are forwarded to the second processor.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the server fault handling method as described in any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Hot switching method for dual-mode redundant microprocessor

    CN105630732A

  • Server data processing method and device and storage medium

    CN109445995A

  • Fault processing method, device and system

    CN113360325A

  • Fault processing system and method, electronic equipment and storage medium

    CN117112317A

  • Server fault processing method and system, electronic equipment, computer storage medium and computer program product

    CN120508454A