Fault handling system, method, electronic device, and storage medium
Patent Information
- Application Number
- CN202310919850.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-25
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-07-25
AI Technical Summary
[0004]然而,BMC获取到各种错误报告后,并不会对错误进行处理,只会记录显示出来,服务器中出现的故障错误需要人为操作处理,故障处理效率低下,影响服务器的稳定性
[0025]本发明实施例提供的一种故障处理系统,包括多个处理器,处理器包括互为冗余的第一处理器和第二处理器,还包括逻辑器件和基板管理控制器;第一处理器与第二处理器上分别设置一对中断接口,第一处理器与第二处理器通过中断接口与基板管理控制器相连,中断接口用于在监测到所在部件出现运行错误时上报错误信息至逻辑器件,基板管理控制器用于记录错误日志,逻辑器件用于若接收到第一处理器的错误信息,将第一处理器上的运行数据切换至第二处理器,并控制第一处理器重启为冗余状态。本发明实施例中利用多个处理器互为冗余,在当前运行的处理器出现故障时,通过中断机制及时上报错误,使得逻辑器件及时切换故障处理器,实现故障快速上报处理,整个系统能够自动快速上报并处理各模块出现的致命性错误,实现故障处理的智能化,避免无法及时处理故障出现整个系统宕机的风险,增加平台的竞争力提高系统可用性,进而提高服务器的可靠性。
Smart Images

Figure CN117112317B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server technology, and in particular to a fault handling system, method, electronic device, and storage medium. Background Technology
[0002] With the rapid development of complex computing scenarios such as artificial intelligence, machine learning, and high-performance computing, new requirements have been put forward for data center architecture. In order to meet the different needs in different scenarios, data centers are accelerating their transformation from a computing-centric architecture to a data-centric converged architecture.
[0003] In converged architectures, server processors (CPUs) are often configured with out-of-band interfaces. The baseboard management controller (BMC) within the server monitors and records the CPU's status based on these out-of-band interfaces. Currently, servers obtain sensor information from various components through the BMC, determine server malfunctions based on thresholds, or obtain error reports via out-of-band IPMI.
[0004] However, after receiving various error reports, BMC does not process the errors, but only records and displays them. Faults and errors in the server require manual handling, which is inefficient and affects the stability of the server. Summary of the Invention
[0005] In view of this, the present invention aims to provide a fault handling system, method, electronic device and storage medium to solve the problem of enabling rapid fault reporting and timely processing, improving fault handling efficiency and ensuring stable server operation.
[0006] According to a first aspect of the present invention, a fault handling system is provided, the system including a plurality of processors, the processors including a first processor and a second processor, the first processor and the second processor being redundant to each other, the system further including logic devices and a baseboard management controller;
[0007] The first processor and the second processor are each provided with a pair of interrupt interfaces, and the first processor and the second processor are connected to the baseboard management controller through the interrupt interfaces;
[0008] The interrupt interface is used to report error information to the logic device when an operational error is detected in the component, and the baseboard management controller is used to record error logs;
[0009] The logic device is used to switch the running data on the first processor to the second processor and control the first processor to restart into a redundant state if it receives an error message from the first processor.
[0010] Furthermore, the baseboard management controller is also provided with a pair of interrupt interfaces for receiving interrupt signals from the first processor and the second processor, and sending a square wave of a preset frequency to the first processor and the second processor through the interrupt interfaces, so that the first processor and the second processor can monitor the operating status of the baseboard management controller through the square wave.
[0011] Furthermore, the first processor and the second processor are used to record an error log when an error is detected in the operation of the baseboard management controller and the interrupt interface interrupts the transmission of square waves.
[0012] Furthermore, the first processor and the second processor are also configured to report the error log to the logic device when the baseboard management controller interrupt interface resumes transmitting square waves, so as to control the baseboard management controller to restart.
[0013] Furthermore, the system also includes a high-speed peripheral component interconnect standard, which connects multiple devices and is used to monitor the operating status of the multiple devices.
[0014] Furthermore, the high-speed peripheral component interconnect standard is also used to control the interrupt interface of the device to report error information to the first processor when an error is detected in the operation of the device; the first processor controls the device to reset through the logic device; wherein, before the device is reset, the data of the device is switched to the normal device.
[0015] Furthermore, the high-speed peripheral component interconnect standard is also used to report to the first processor through an interrupt interface when an error occurs in the operation of the high-speed peripheral component interconnect standard itself, switch the data of the high-speed peripheral component interconnect standard to the normal high-speed peripheral component interconnect standard, and reset the high-speed peripheral component interconnect standard through the interrupt interface controlled by the first processor.
[0016] According to a second aspect of the present invention, a fault handling method is provided, applied to any of the fault handling systems described above, the method comprising:
[0017] Monitor the operating status of the first processor; wherein the first processor is the currently running main processor;
[0018] If an operational error is detected in the first processor, the error information is reported to the logic device;
[0019] The running data on the first processor is switched to the second processor, and the first processor is restarted to a redundant state.
[0020] According to another aspect of the present invention, an electronic device is also provided, comprising:
[0021] processor;
[0022] Memory used to store the processor's executable instructions;
[0023] The processor is configured to execute the instructions to implement the fault handling method described above.
[0024] According to another aspect of the present invention, a readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the fault handling method described above.
[0025] This invention provides a fault handling system comprising multiple processors, including a redundant first processor and a second processor, as well as logic devices and a baseboard management controller. Each of the first and second processors has a pair of interrupt interfaces, which are connected to the baseboard management controller. The interrupt interfaces report error information to the logic devices when a malfunction is detected in the component. The baseboard management controller records an error log. Upon receiving an error message from the first processor, the logic devices switch the running data from the first processor to the second processor and control the first processor to restart in a redundant state. This invention utilizes multiple redundant processors. When a currently running processor malfunctions, an interrupt mechanism promptly reports the error, allowing the logic devices to switch the faulty processor quickly. This enables rapid fault reporting and processing, allowing the entire system to automatically and quickly report and handle fatal errors in each module. This intelligent fault handling avoids the risk of system downtime due to untimely fault handling, increases platform competitiveness, improves system availability, and ultimately enhances server reliability.
[0026] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0027] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0028] Figure 1This is one of the structural schematic diagrams of a fault handling system provided in an embodiment of the present invention;
[0029] Figure 2 This is a second schematic diagram of the structure of a fault handling system provided in an embodiment of the present invention;
[0030] Figure 3 yes Figure 1 A schematic diagram of processor connection for a fault handling system according to an embodiment of the present invention is provided.
[0031] Figure 4 This is a flowchart of the steps of a fault handling method provided in an embodiment of the present invention;
[0032] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the various embodiments of the present invention to facilitate a better understanding of this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for ease of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with and referenced by each other without contradiction.
[0034] In existing server architectures, due to the involvement of multiple CPUs and other chips, troubleshooting is difficult when operational failures occur. Developers conduct system testing during server product development to promptly identify and resolve design flaws, ensuring the quality of new server products. However, current fault handling during the development phase relies heavily on developers manually searching test logs and diagnostic data, or on their own experience, which is labor-intensive and time-consuming. To address these issues, this invention, in a server architecture with redundant processors, uses an interrupt mechanism to promptly report errors when the currently running processor fails. This allows logic devices to switch to the faulty processor in a timely manner, enabling rapid fault reporting and processing. The entire system can automatically and quickly report and handle fatal errors in various modules, achieving intelligent fault handling.
[0035] Reference Figure 1 The diagram shows a structural schematic of a fault handling system provided in an embodiment of the present invention, such as... Figure 1As shown, the system includes multiple processors, including a first processor and a second processor, which are redundant with each other. The system also includes logic devices and a baseboard management controller.
[0036] The first processor and the second processor are each provided with a pair of interrupt interfaces, and the first processor and the second processor are connected to the baseboard management controller through the interrupt interfaces;
[0037] The interrupt interface is used to report error information to the logic device when an operational error is detected in the component it is located in, and the board management controller is used to record the error log;
[0038] The logic device is used to switch the running data on the first processor to the second processor and control the first processor to restart into a redundant state if it receives an error message from the first processor.
[0039] In this embodiment of the invention, the fault reporting system can be understood as a server system. The server includes multiple processors, and the system also includes logic devices and a baseboard management controller. This embodiment takes the first processor and the second processor as examples for illustration. The first processor and the second processor are redundant CPUs. By default, the first processor is the main CPU when the system is powered on and is connected to various devices through logic devices. The second processor is the secondary CPU. When a fatal and unrecoverable error occurs during the operation of the first processor, the various programs running on the first processor are switched to the second processor to achieve rapid processing of CPU faults.
[0040] In this embodiment, the first processor and the second processor are each provided with a pair of interrupt interfaces, and the first processor and the second processor are connected to the baseboard management controller through the interrupt interfaces. For example, see [reference needed]. Figure 2 There are two interrupt GPIOs between the first processor, the second processor, and the baseboard management controller, which are CPU INT GPIO X2 in the figure. They are respectively INT GPIO X1 of the first processor, INT GPIO X2 of the first processor, INT GPIO X2 of the second processor, and INT GPIO X2 of the second processor. INT GPIO X1 of the first processor is connected to INT GPIO X2 of the second processor. INT GPIO X2 of the second processor is connected to INT GPIO X1 of the first processor. The baseboard management controller receives interrupt signals from the two processors.
[0041] It should be noted that the interrupt interface in this embodiment is GPIO (General-purpose input / output). GPIO is a type of GPIO that is equipped with interrupt functionality. The interrupt mechanism means that when unexpected situations occur during computer operation requiring host intervention, the machine can automatically stop the currently running program and switch to a program to handle the new situation. After handling the situation, it can return to the originally suspended program and continue running. It should also be noted that the interrupt interface has the highest priority during operation. When an error or fault is detected, it can notify the host as quickly as possible to execute the corresponding instructions. The interrupt interface is used to report error information to the logic device when an operational error is detected in the component it is connected to. The board management controller is used to record the error log.
[0042] Specifically, in this embodiment, the server's Baseboard Management Controller (BMC) is primarily responsible for out-of-band management, system status monitoring, and restart, power-on / power-off control. In this embodiment, the BMC mainly functions as the refresher of error logs. The logic device can be a CPLD, which, as the server's logic device, is responsible for various logic switching of each board and power-on / off management of each chip and module. Specifically, it is responsible for monitoring the operating status of all components and executing various fault handling operations according to processing commands. In this embodiment, if the logic device receives an error message from the first processor, it switches the operating data on the first processor to the second processor and controls the first processor to restart into a redundant state.
[0043] It should be noted that under normal circumstances, the first processor operates normally, while the second processor is in a redundant state. When the first processor encounters a fatal, unrecoverable error, the BIOS controls the first processor's GPIO port to notify the second processor to switch operating states via an interrupt interface. The first processor stores its running data and sends it to the second processor, which then takes over operation. Furthermore, the first processor notifies the baseboard management controller via the interrupt interface that when a fatal, unrecoverable error occurs, this error information is recorded in the BMC log. After the second processor has switched over, the baseboard management controller notifies the CPLD logic device to restart the first processor via IIC. The CPLD logic device then notifies the CPLD logic device on the processor board via its CPLD GPIO pin, putting the first processor in a redundant state, ready to take over the running data from the second processor.
[0044] For example, when a system failure occurs, specifically when the currently running primary processor fails, the logic devices within the redundant secondary processor CPU automatically switch the control function to the redundant secondary processor. In other words, when the primary CPU fails, the logic devices, upon receiving the primary CPU's fault information, switch the control function to the backup CPU within 1 to 2 preset cycles. The backup CPU's output is enabled, and the backup CPU controls the system's operation, while the primary CPU's output is disabled. This ensures that when one processor fails, other components continue to operate, guaranteeing server stability. It should be noted that the above is merely a specific example; the duration of the preset cycle is pre-set based on the server's processing needs and can be adjusted according to actual fault reporting requirements. No specific limitation is made here.
[0045] In addition, after the first processor recovers from the fault, the recovered main CPU sends a signal to the backup CPU. After receiving the signal from the main CPU, the logic device switches the control function to the main CPU within 1 to 2 preset cycles. The main CPU controls the server to work, and the output of the backup CPU is disabled.
[0046] This invention provides a fault handling system comprising multiple processors, including a redundant first processor and a second processor, as well as logic devices and a baseboard management controller. Each of the first and second processors has a pair of interrupt interfaces, which are connected to the baseboard management controller. The interrupt interfaces report error information to the logic devices when a malfunction is detected in the component. The baseboard management controller records an error log. Upon receiving an error message from the first processor, the logic devices switch the running data from the first processor to the second processor and control the first processor to restart in a redundant state. This invention utilizes multiple redundant processors. When a currently running processor malfunctions, an interrupt mechanism promptly reports the error, allowing the logic devices to switch the faulty processor quickly. This enables rapid fault reporting and processing, allowing the entire system to automatically and quickly report and handle fatal errors in each module. This intelligent fault handling avoids the risk of system downtime due to untimely fault handling, increases platform competitiveness, improves system availability, and ultimately enhances server reliability.
[0047] Reference Figure 2 , Figure 2 This is a second structural schematic diagram of a fault handling system provided in an embodiment of the present invention. Further, in this embodiment, according to the connection relationship of each board in the server, the error types are classified according to the board module: CPU error, BMC error, PCIe SW error, and Device error. Each error corresponds to its own fault reporting and handling method. (Refer to...) Figure 3 ,yes Figure 1 One of the processor connection diagrams of a fault handling system provided in this embodiment of the invention includes a pair of interrupt interfaces on the baseboard management controller for receiving interrupt signals from the first processor and the second processor, and sending a square wave of a preset frequency to the first processor and the second processor through the interrupt interfaces, so that the first processor and the second processor can monitor the operating status of the baseboard management controller through the square wave.
[0048] Specifically, the Baseboard Management Controller (BMC), as an out-of-band management module, will not affect the normal operation of the server system when an error occurs. In this embodiment, the BMC is also equipped with a pair of interrupt interfaces, which serve as the BMC's HeartBeat. This HeartBeat represents a signal when the BMC has a fatal error due to a code defect that is unrecoverable. Through the HeartBeat, resources (such as IP and program services) can be quickly transferred from a faulty device to another normally operating machine to continue providing services.
[0049] It should be noted that under normal conditions, the interrupt interface GPIO sends a square wave of a preset frequency, that is, it sends a square wave of a preset frequency to the first processor and the second processor through the interrupt interface. The first processor, the second processor, and the logic devices simultaneously monitor the operating status of the board management controller through the square wave.
[0050] Furthermore, the first and second processors are used to record error logs when an error is detected in the operation of the baseboard management controller and the interrupt interface interrupts the transmission of square waves.
[0051] Furthermore, the first and second processors are also used to report error logs to the logic devices when the board management controller resumes transmitting square waves after the interrupt interface, so as to control the board management controller to restart.
[0052] Specifically, if the first processor and the second processor detect that the GPIO state of the interrupt interface of the baseboard management controller remains unchanged, that is, when the square wave is resumed, the logic device performs a reset operation on the baseboard management controller to restart the baseboard management controller. The first processor and the second processor then report the error log. In other words, during this fault process, the error log recording function originally performed by the baseboard management controller is transferred to the first processor and the second processor for execution. After the BMC is working normally, the error log is sent to the BMC via IPMI and the BMC re-records it.
[0053] In this embodiment of the invention, when the currently running baseboard management controller malfunctions, the error is reported in a timely manner through an interrupt mechanism, enabling the logic device to switch the running data of the faulty baseboard management controller to the processor in a timely manner, thereby realizing rapid fault reporting and processing. The entire system can automatically and quickly report and process fatal errors that occur in each module, realizing intelligent fault handling, avoiding the risk of the entire system crashing due to failure to handle faults in a timely manner, increasing the competitiveness of the platform, improving system availability, and thus improving the reliability of the server.
[0054] Furthermore, refer to Figure 3 The system also includes a high-speed peripheral component interconnect standard, which connects multiple devices and is used to monitor the operating status of multiple devices.
[0055] Specifically, the high-speed peripheral component interconnect standard PCIe Switch (SW) is used to expand the number of PCIe lanes, allowing more PCIe devices to be mounted on the processor and enabling in-band monitoring of device status. Furthermore, because it embeds a small ARM core, its SW has some data processing capabilities. The PCIe SW expands the CPU's PCIe lanes, allowing more PCIe devices to be connected via connectors.
[0056] Among them, the connector (Mini Cool Edge IO, MCIO) is a flexible, robust, and cost-effective connector that helps product designers improve flexibility, reduce overall space requirements, and expand the coverage of high-speed signals, without being limited to any specific type.
[0057] Furthermore, the high-speed peripheral component interconnect standard is also used to control the interrupt interface of the device to report error information to the first processor when an error is detected in the device operation; the first processor controls the device to reset through the logic device; wherein, before the device is reset, the data of the device is switched to the normal device.
[0058] Specifically, in this embodiment, the high-speed peripheral component interconnect standard is also used to monitor the device's operating status. When the device encounters an error, the error information will be communicated to the high-speed peripheral component interconnect standard via in-band PCIe. When the high-speed peripheral component interconnect standard detects an error in the device's operation, it controls the device's interrupt interface to report the error information to the first processor. The first processor then controls the device to reset via logic devices.
[0059] For example, the first processor is notified via the interrupt GPIO of the corresponding device 1, and this is synchronized to the second processor and the baseboard management controller. When the first processor determines that the error is unrecoverable and causes the device to malfunction, it instructs the CPLD to reset the corresponding faulty device via GPIO X5. It should be noted that GPIO X5 represents 5 GPIO bits, which, when combined, can represent 32 states, corresponding to the 8 high-speed peripheral component interconnect standard PCIe SWs and the 24 connected devices. Preferably, before resetting the faulty device, the data on the faulty device needs to be seamlessly switched to another normal device before the first processor controls the device reset via logic devices.
[0060] Furthermore, the high-speed peripheral component interconnect standard is also used to report to the first processor through an interrupt interface when an error occurs in the operation of the high-speed peripheral component interconnect standard itself, switch the data of the high-speed peripheral component interconnect standard to the normal high-speed peripheral component interconnect standard, and reset the high-speed peripheral component interconnect standard through the interrupt interface controlled by the first processor.
[0061] For example, when a high-speed peripheral component interconnect standard PCIe switch itself experiences a fatal error, the PCIe switch notifies the processor via GPIO. The PCIe switch can then transfer its running data to another PCIe switch via the Fabric PCIe channel. The CPU then resets the faulty switch by controlling GPIO X5. It's important to note that all fatal errors in PCIe switches and devices are sent to the BMC and CPU via SWINT GPIO. The faulty PCIe switch automatically switches to the working device or transfers its own data to another healthy PCIe switch via Fabric PCIe.
[0062] In this embodiment of the invention, when the currently running high-speed peripheral component interconnect standard PCIe SW itself or the device malfunctions, the error is reported in a timely manner through an interrupt mechanism, enabling the logic device to switch the operating data in the faulty component to the normal component in a timely manner, realizing rapid fault reporting and processing. The entire system can automatically and quickly report and process fatal errors that occur in each module, realizing intelligent fault handling, avoiding the risk of the entire system crashing due to failure to handle faults in a timely manner, increasing the competitiveness of the platform, improving system availability, and thus improving the reliability of the server.
[0063] This invention provides a fault handling system comprising multiple processors, including a redundant first processor and a second processor, as well as logic devices and a baseboard management controller. Each of the first and second processors has a pair of interrupt interfaces, which are connected to the baseboard management controller. The interrupt interfaces report error information to the logic devices when a malfunction is detected in the component. The baseboard management controller records an error log. Upon receiving an error message from the first processor, the logic devices switch the running data from the first processor to the second processor and control the first processor to restart in a redundant state. This invention utilizes multiple redundant processors. When a currently running processor malfunctions, an interrupt mechanism promptly reports the error, allowing the logic devices to switch the faulty processor quickly. This enables rapid fault reporting and processing, allowing the entire system to automatically and quickly report and handle fatal errors in each module. This intelligent fault handling avoids the risk of system downtime due to untimely fault handling, increases platform competitiveness, improves system availability, and ultimately enhances server reliability.
[0064] Reference Figure 4 The flowchart illustrates the steps of a fault handling method provided in an embodiment of the present invention, which is applied to... Figures 1 to 3 In any of the fault handling systems shown, the method may include:
[0065] Step 101: Monitor the operating status of the first processor; wherein, the first processor is the currently running main processor.
[0066] In this embodiment of the invention, the error types are classified according to the board module based on the connection relationship of each board in the server, namely CPU error, BMC error, PCIe SW error, and Device error. Each type of error corresponds to its own fault reporting and handling method. In this embodiment, CPU error is used as an example for explanation.
[0067] Specifically, the first processor is the currently running main processor. The server's baseboard management controller is mainly responsible for out-of-band management of the server, monitoring of system status, and restart, power-on, and power-off control. In this embodiment, the baseboard management controller monitors the operating status of the first processor.
[0068] It should be noted that the first processor is the currently running main processor. However, it should also be clarified that the first and second processors described in this embodiment are only used to distinguish between the main processor and the backup processor in operation; there is no specific order between them. This embodiment is only used as an example where the first processor is the currently running main processor, and the second processor is a backup processor waiting to take over the operating data of the faulty first processor.
[0069] Step 102: If an operational error is detected in the first processor, the error information is reported to the logic device.
[0070] In this embodiment of the invention, if an operational error is detected in the first processor, the error information is reported to the logic device. The logic device is responsible for various logic switching of each board and power-on / off management of each chip and module. Specifically, it is responsible for monitoring the operating status of all components and executing various fault handling operations according to the processing command.
[0071] In this embodiment, error information is reported through an interrupt interface, specifically a general-purpose input / output (GPIO). All GPIOs are interrupt-enabled. The interrupt mechanism means that when unexpected situations arise during computer operation requiring host intervention, the machine can automatically stop the currently running program and switch to a program to handle the new situation. After handling the situation, it returns to the originally suspended program to continue running. It should be noted that the interrupt interface has the highest priority during operation, allowing it to quickly notify the host and execute corresponding instructions when an error or fault is detected.
[0072] In this embodiment, the baseboard management controller records error logs and uses the interrupt interface to report error information to the logic device when an operational error is detected in the component. This allows the logic device to promptly process the faulty component based on the error information, perform data migration, and switch devices.
[0073] Step 103: Switch the running data on the first processor to the second processor, and control the first processor to restart into a redundant state.
[0074] In this embodiment of the invention, when a running error occurs in the first processor, the logic device switches the running data on the first processor to the second processor and controls the first processor to restart into a redundant state.
[0075] The logic device is responsible for various logic switching of each board and power-on / off management of each chip and module. Specifically, it is responsible for monitoring the operating status of all components and executing various fault handling operations according to processing commands. In this embodiment, if the logic device receives an error message from the first processor, it switches the operating data on the first processor to the second processor and controls the first processor to restart into a redundant state.
[0076] Specifically, in this embodiment, the first processor operates normally, while the second processor is in a redundant state. When the first processor experiences a fatal, unrecoverable error, the BIOS controls the first processor's GPIO port to notify the second processor to switch operating states via an interrupt interface. The first processor stores its running data and sends it to the second processor, which then takes over operation. Furthermore, the first processor notifies the baseboard management controller via the interrupt interface that when a fatal, unrecoverable error occurs, this error information is recorded in the BMC log. After the second processor has switched over, the baseboard management controller notifies the CPLD logic device to restart the first processor via IIC. The CPLD logic device then notifies the CPLD logic device on the processor board via its CPLD GPIO pin, putting the first processor in a redundant state, ready to take over the second processor's running data.
[0077] This invention provides a fault handling method. Compared to existing technologies, this invention, based on the beneficial effects of the first embodiment, monitors the operating status of a first processor. The first processor is the currently running main processor. If an error is detected in the first processor, the error information is reported to the logic device. The running data on the first processor is switched to a second processor, and the first processor is restarted into a redundant state. This invention utilizes multiple processors for redundancy. When a fault occurs in the currently running processor, an interrupt mechanism promptly reports the error, allowing the logic device to switch the faulty processor in a timely manner. This enables rapid fault reporting and processing, allowing the entire system to automatically and quickly report and handle fatal errors in each module. This intelligent fault handling avoids the risk of system downtime due to untimely fault handling, increases platform competitiveness, improves system availability, and ultimately enhances server reliability.
[0078] This invention also provides an electronic device, such as... Figure 5 As shown, it includes a processor 201, a communication interface 202, a memory 203, and a communication bus 204. The processor 201, communication interface 202, and memory 203 communicate with each other through the communication bus 204.
[0079] Memory 203 is used to store computer programs;
[0080] When processor 201 executes a program stored in memory 203, it performs the following steps:
[0081] Monitor the operating status of the first processor; where the first processor is the currently running main processor.
[0082] If an operational error is detected in the first processor, the error information is reported to the logic device.
[0083] The running data on the first processor is switched to the second processor, and the first processor is restarted to a redundant state.
[0084] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0085] The communication interface is used for communication between the aforementioned terminal and other devices.
[0086] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0087] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0088] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the fault handling methods described in the above embodiments.
[0089] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0090] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0091] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0092] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A fault handling system, characterized in that, The system includes multiple processors, including a first processor and a second processor, which are redundant with each other. The system also includes logic devices and a baseboard management controller. The first processor and the second processor are each provided with a pair of interrupt interfaces, and the first processor and the second processor are connected to the baseboard management controller through the interrupt interfaces; The interrupt interface is used to report error information to the logic device when an operational error is detected in the component, and the baseboard management controller is used to record error logs; The logic device is used to switch the running data on the first processor to the second processor and control the first processor to restart into a redundant state if it receives an error message from the first processor.
2. The fault handling system according to claim 1, characterized in that, The baseboard management controller is also provided with a pair of interrupt interfaces for receiving interrupt signals from the first processor and the second processor, and sending a square wave of a preset frequency to the first processor and the second processor through the interrupt interfaces, so that the first processor and the second processor can monitor the operating status of the baseboard management controller through the square wave.
3. The fault handling system according to claim 2, characterized in that, The first processor and the second processor are used to record an error log when an error is detected in the operation of the baseboard management controller and the interrupt interface interrupts the transmission of square waves.
4. The fault handling system according to claim 2, characterized in that, The first processor and the second processor are further configured to report the error log to the logic device when the baseboard management controller interrupt interface resumes transmitting square waves, so as to control the baseboard management controller to restart.
5. The fault handling system according to claim 1, characterized in that, The system also includes a high-speed peripheral component interconnect standard, which connects multiple devices and is used to monitor the operating status of the multiple devices.
6. The fault handling system according to claim 5, characterized in that, The high-speed peripheral component interconnect standard is also used to control the interrupt interface of the device to report error information to the first processor when an error is detected in the operation of the device; the first processor controls the device to reset through the logic device; wherein, before the device is reset, the data of the device is switched to the normal device.
7. The fault handling system according to claim 5, characterized in that, The high-speed peripheral component interconnect standard is also used to report to the first processor through an interrupt interface when an error occurs in the operation of the high-speed peripheral component interconnect standard itself, switch the data of the high-speed peripheral component interconnect standard to the normal high-speed peripheral component interconnect standard, and reset the high-speed peripheral component interconnect standard through the interrupt interface controlled by the first processor.
8. A fault handling method, characterized in that, The method, applied to the fault handling system according to any one of claims 1 to 7, comprises: Monitor the operating status of the first processor; wherein the first processor is the currently running main processor; If an operational error is detected in the first processor, the error information is reported to the logic device; The running data on the first processor is switched to the second processor, and the first processor is restarted to a redundant state.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to implement the fault handling method as described in claim 8.
10. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, which, when executed by a processor, implements the fault handling method as described in claim 8.
Citation Information
Patent Citations
System and method for processing error
CN104424041A
Server dual-redundancy CPU device and switching method
CN115454730A