Troubleshooting methods, related devices, computer equipment, media and programs

By isolating and recovering embedded processor faults in the data processor through programmable logic devices, the host interruption problem caused by embedded processor faults is solved, fault handling without restarting is achieved, and the normal operation and efficient recovery of the computer host are ensured.

CN115454705BActive Publication Date: 2025-09-26SHENZHEN XINGYUN ZHILIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211212545.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-01
Publication Date
2025-09-26
Estimated Expiration
2042-07-01

AI Technical Summary

Technical Problem

A failure of an embedded processor in a data processor may cause a host user service interruption or a host crash. Existing technologies require restarting the host to recover, affecting user service processing.

Method used

Entering the answering mode through the programmable logic device, a hot plug interrupt signal is sent to the computer host to isolate the embedded processor fault. After the fault is detected and repaired, a hot plug signal is sent to restore communication.

Benefits of technology

Fault isolation and recovery are achieved without restarting the computer host, minimizing the impact on the host, ensuring normal operation of the host, and improving fault handling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115454705B_ABST
    Figure CN115454705B_ABST
Patent Text Reader

Abstract

The present application relates to the field of electronic digital data processing in the Internet industry, and discloses a method, apparatus, computer equipment, and storage medium for fault handling. The method includes: in response to detecting that an embedded processor in a data processor has a fault, entering a proxy mode, the proxy mode including sending a hot plug interrupt signal to a computer host to isolate the computer host from the embedded processor fault, the hot plug interrupt signal being used to indicate that the embedded processor has performed a hot plug operation; in response to detecting that the embedded processor has been repaired, sending a hot plug signal to the computer host, exiting the proxy mode to complete fault recovery, the hot plug signal being used to indicate that the embedded processor has performed a hot plug operation. By implementing the embodiments of the present application, fault isolation can be effectively achieved without restarting the computer host, and the impact on the computer host can be minimized, thereby ensuring the normal operation of the computer host.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic digital data processing in the Internet industry, and in particular to a fault handling method, related devices, computer equipment, storage medium and program. Background Art

[0002] With the rapid development of data centers, communication and computing capabilities have become two mutually reinforcing and important development directions for data center infrastructure. If data centers focus solely on improving computing power without keeping pace with improvements in communication infrastructure, the overall system performance of the data center will remain limited and its true potential will not be realized. To cope with the increasingly large and complex data volumes, data processing units (DPUs) have emerged.

[0003] The data processor is positioned as a collaborative processing unit and is an implementation of the idea of ​​separating the data plane from the control plane. It cooperates with the central processing unit (CPU), with the latter responsible for general control and the former focusing on data processing. In other words, the data processor can offload data processing / preprocessing from the CPU and distribute computing power closer to where the data occurs, thereby reducing communication volume. Since the data processor needs to move the calculation close to the data, this also means that higher requirements are placed on the reliability and availability of the data processor. In order to meet the network's demand for large data transmission, the data processor needs to use a high-speed serial computer expansion bus standard (Peripheral Component Interconnect Express, PCIe) interface for data transmission. Therefore, the data processor is often involved in the failure of the PCIe device, and the data processor system is required to assist the PCIe device in fault recovery.

[0004] Currently, the data processor simulates a PCIe device on the embedded central processing unit (ECPU). The host's PCIe-related transaction layer pockets (TLPs) are forwarded to the ECPU for processing. This means that if the embedded processor's PCIe emulation program or the system itself fails, two scenarios may occur: one is that the host user's services may be affected, even causing the host to crash; the other is that restoring the data processor requires restarting the host, interrupting all running programs. Therefore, how to isolate the fault without affecting user services is a problem that technicians in this field need to solve. Summary of the Invention

[0005] The embodiments of the present application provide a method, related apparatus, computer equipment, storage medium and program for fault handling, which can effectively implement fault isolation without restarting the computer host, can minimize the impact on the computer host, and thus can ensure the normal operation of the computer host.

[0006] In a first aspect, an embodiment of the present application provides a fault handling method, which is applied to a programmable logic device, wherein:

[0007] In response to detecting a fault in an embedded processor in a data processor, the programmable logic device enters a proxy mode, the proxy mode including sending a hot plug interrupt signal to a computer host to isolate the computer host from the fault of the embedded processor, the hot plug interrupt signal being used to indicate that the embedded processor has performed a hot plug operation;

[0008] In response to detecting that the embedded processor is repaired, the programmable logic device sends a hot plug signal to the computer host to exit the answering mode to complete fault recovery. The hot plug signal is used to indicate that the embedded processor has performed a hot plug operation.

[0009] In a second aspect, an embodiment of the present application provides a fault handling device, which is applied to a programmable logic device, wherein:

[0010] a fault isolation unit configured to, in response to detecting a fault in an embedded processor in a data processor, cause the programmable logic device to enter a proxy mode, the proxy mode comprising sending a hot-swap interrupt signal to a computer host to isolate the computer host from the embedded processor fault, the hot-swap interrupt signal being used to indicate that the embedded processor has performed a hot-swap operation;

[0011] A fault recovery unit is used to respond to detecting that the repair of the embedded processor is completed, and the programmable logic device sends a hot plug signal to the computer host to exit the answering mode to complete fault recovery. The hot plug signal is used to indicate that the embedded processor has performed a hot plug operation.

[0012] In a third aspect, an embodiment of the present application provides a computer device comprising a processor, a memory, and a communication interface, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for some or all of the steps described in the first aspect of the embodiment of the present application.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.

[0014] Implementing the embodiments of this application will have the following beneficial effects:

[0015] By using the above-mentioned fault handling method, device, computer equipment and storage medium, after detecting that the embedded processor in the data processor has a fault, the programmable logic device can enter the proxy mode, and send a hot plug interrupt signal to the computer host, which is used to indicate that the embedded processor has performed a hot plug operation to disconnect the communication between the computer host and the embedded processor, thereby completing fault isolation simply and efficiently. In this way, it is possible to avoid the problem in the traditional technical solution that once the embedded processor fails, the computer host must be restarted, resulting in the interruption of all running programs, thereby minimizing the impact on the computer host and ensuring the normal operation of the computer host. After detecting that the embedded processor has been repaired, the programmable logic device sends a hot plug signal to the computer host, which is used to indicate that the embedded processor has performed a hot plug operation and exits the proxy mode, so that the computer host and the embedded processor can re-communicate, thereby quickly completing fault recovery and improving the efficiency of fault handling. In addition, the embodiment of the present application does not require the participation of external tools such as the BMC system and the control platform, which can reduce the degree of dependence and have higher reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. Among them:

[0017] Figure 1 A schematic diagram of a system architecture provided in an embodiment of the present application;

[0018] Figure 2 A flowchart of a method for troubleshooting provided in an embodiment of the present application;

[0019] Figure 3 A schematic diagram of the structure of a fault handling device provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] The terms "first," "second," "predetermined," and "fourth" in the specification, claims, and drawings of this application are used to distinguish between different objects, not to describe a specific order. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0023] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0024] It should also be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " as used herein generally indicates that the associated objects are in an "or" relationship.

[0025] To facilitate understanding, several basic concepts involved in the embodiments of the present application are first introduced below.

[0026] The data processing unit (DPU) is a newly developed category of specialized processors. Following the central processing unit (CPU) and graphics processing unit (GPU), it is the third most important computing chip in data center scenarios, providing a computing engine for high-bandwidth, low-latency, and data-intensive computing scenarios. Data processors have three primary characteristics: offloading, acceleration, and isolation. Accordingly, the three main application scenarios for data processors are networking, storage, and security. In terms of offloading, data processors can serve as offloading engines for the CPU, freeing up the CPU's computing power for upper-layer applications. For example, data processors can offload data center network services (virtual switching, virtual routing, etc.), data center storage services, and data center security services (firewalls, encryption and decryption, etc.). In terms of acceleration, data processors will become sandboxes for algorithm acceleration and the most flexible accelerator carrier. Data processors are more than just fixed application-specific integrated circuits (ASICs). Data consistency access protocols for CPUs, graphics processors, and data processors, promoted by standards organizations like Compute Express Link (CXL), will further reduce data processor programming barriers. Combined with programmable devices like field-programmable gate arrays (FPGAs), customizable hardware will have greater potential. "Software-to-hardware" will become the norm, and the widespread adoption of various data processors will fully unleash the potential of heterogeneous computing. Regarding isolation, data processors will become new data gateways, elevating security and privacy to a new level. Asymmetric encryption algorithms such as SM2, hashing algorithms such as SM3, and symmetric block ciphers such as SM4 can all be implemented by integrating them into data processors.

[0027] The peripheral component interconnect express (PCIe) standard is a high-speed serial point-to-point dual-channel, high-bandwidth transmission standard. Connected devices are allocated exclusive channel bandwidth and do not share bus bandwidth. It defines slots and connectors with multiple widths: x1, x4, x8, x12, x16, and x32. Typically, low-speed peripherals (such as WiFi cards) use single-channel (x1) links, while graphics adapters more often use faster and wider x16 channel links.

[0028] The field programmable gate array (FPGA) is a further development of programmable devices such as programmable logic array (PAL), effectively addressing the limited gate count of existing devices. The basic structure of an FPGA includes programmable input and output units, configurable logic blocks, a digital clock management module, routing resources, embedded dedicated hard cores, and underlying embedded functional units. Due to its rich routing resources, reprogrammability, high integration, and low investment, FPGAs have been widely used in digital circuit design.

[0029] Complex programmable logic devices (CPLDs) utilize programming technologies such as electrically erasable programmable read-only memory (EEPROM), flash memory, and static random access memory (SRAM) to create high-density, high-speed, and low-power programmable logic devices. CPLDs are digital integrated circuits in which users customize their logic functions based on their needs. The basic design approach utilizes an integrated development software platform, using schematics and hardware description languages ​​to generate target files. The code is then transferred to the target chip via a download cable to implement the designed digital system.

[0030] See Figure 1 , Figure 1 This is a schematic diagram of a system architecture provided by an embodiment of the present application. Figure 1 As shown, the system architecture may include a computer host 100 and a data processor 200. The computer host 100 may include a computer device with a processor having various architectures as a computing core, such as an Intel x86 architecture CPU, an ARM architecture CPU, or a MIPS architecture CPU, including but not limited to industrial computers, servers, vehicle-mounted computers, mobile workstations, and other application forms. The present embodiment does not impose any restrictions on the type of computer host.

[0031] In an embodiment of the present application, the data processor 200 may include a programmable logic device 201 and an embedded processor 202. The data processor 200 may use the PCIe standard protocol for data transmission, and therefore, the data processor 200 may be a PCIe device. The programmable logic device 201 may be an FPGA, a system on chip (SOC), an ASIC, or a multi-core processor, or other programmable logic devices, which are not limited in the embodiment of the present application. The programmable logic device 201 may communicate with the computer host 100 via a PCIe interface. The programmable logic device 201 may communicate with the embedded processor 202 via a first interface. The first interface may be a PCIe interface, a common flash interface (CFI), a serial peripheral interface (SPI), a peripheral component interconnect (PCI) interface, a local bus (LocalBus) interface, etc., which are not limited in the embodiment of the present application.

[0032] The programmable logic device 201 may include a status register, which may be used to record the operating status of the embedded processor 202 or other devices in the data processor 200. The programmable logic device 201 reads the information in the status register to determine whether the embedded processor 202 has failed. Optionally, a first-in-first-out (FIFO) channel may be established between the computer host 100, the programmable logic device 201, and the embedded processor 202, so that the computer host 100, the programmable logic device 201, and the embedded processor 202 can exchange data through the FIFO channel to increase data transmission speed.

[0033] The data processor 200 may include Figure 1 The complex programmable logic device (CPLD) not shown in the figure can communicate with the programmable logic device 201 through the first interface, and can also communicate with the embedded processor 202 through the first interface. That is, the CPLD can communicate with the programmable logic device 201 and the embedded processor 202 respectively. The data processor 200 may also include Figure 1 The memory (such as Flash memory) not shown in the figure can be used to store the program to be run, and the memory can communicate with the programmable logic device 201 through the CFI interface. In addition, the data processor 200 may also include Figure 1The PCIe switch, GPU, digital signal processor (DSP), redundant arrays of independent disks (RAID), etc. not shown in the figure are not limited in the embodiments of the present application.

[0034] With the rapid development of data centers, communication and computing capabilities have become two mutually reinforcing and important development directions for data center infrastructure. If data centers focus solely on improving computing power without keeping pace with improvements in communication infrastructure, the overall system performance of the data center will remain limited and its true potential will not be realized. To cope with the increasingly large and complex data volumes, data processing units (DPUs) have emerged.

[0035] The data processor is positioned as a collaborative processing unit and is an implementation of the idea of ​​separating the data plane from the control plane. It cooperates with the central processing unit (CPU), with the latter responsible for general control and the former focusing on data processing. In other words, the data processor can offload data processing / preprocessing from the CPU and distribute computing power closer to where the data occurs, thereby reducing communication volume. Since the data processor needs to move the calculation close to the data, this also means that higher requirements are placed on the reliability and availability of the data processor. In order to meet the network's demand for large data transmission, the data processor needs to use a high-speed serial computer expansion bus standard (Peripheral Component Interconnect Express, PCIe) interface for data transmission. Therefore, the data processor is often involved in the failure of the PCIe device, and the data processor system is required to assist the PCIe device in fault recovery.

[0036] Currently, the data processor simulates a PCIe device on the embedded central processing unit (ECPU). The host's PCIe-related transaction layer pockets (TLPs) are forwarded to the ECPU for processing. This means that if the embedded processor's PCIe emulation program or the system itself fails, two scenarios may occur: one is that the host user's services may be affected, even causing the host to crash; the other is that restoring the data processor requires restarting the host, interrupting all running programs. Therefore, how to isolate the fault without affecting user services is a problem that technicians in this field need to solve.

[0037] In order to solve the above problems, an embodiment of the present application provides a fault handling method. By implementing this method, fault isolation can be effectively achieved without restarting the computer host, which can minimize the impact on the computer host and thus ensure the normal operation of the computer host.

[0038] Please refer to Figure 2 , Figure 2 This is a flowchart of a method for troubleshooting provided by an embodiment of the present application. It can be understood that this method can be used to Figure 1 In the system architecture shown, it can be specifically Figure 1 The programmable logic device shown in FIG. 1 is executed, and the method may include the following steps S201-S202, wherein:

[0039] Step S201: In response to detecting a fault in an embedded processor in a data processor, the programmable logic device enters a proxy mode, which includes sending a hot plug interrupt signal to a computer host to isolate the computer host from the embedded processor fault.

[0040] As the number of PCIe devices connected to a computer host increases, the probability of PCIe device failure also increases. PCIe device failure may affect the normal operation of the computer host, and in severe cases, may even cause the computer host to hang. Therefore, handling failed PCIe devices is an important part of maintaining the normal operation of the computer host. Since the data processor can use the PCIe standard protocol for data transmission, the data processor can be a PCIe device. Figure 1 As shown, the data processor may include a programmable logic device and an embedded processor. The programmable logic device may be an FPGA, a SOC, an ASIC, or a multi-core processor, or other programmable logic devices, which are not limited in the embodiments of the present application. For the data processor, the data processor may simulate a PCIe device on the embedded processor end, and the embedded processor may be used to process PCIe-related TLP packets of the computer host forwarded thereto. This means that when the PCIe simulation program of the embedded processor or the system itself fails, the failure of the embedded processor may affect the normal operation of the computer host.

[0041] In order to prevent the fault of the embedded processor from affecting the normal operation of the computer host, it is necessary to isolate the fault of the embedded processor from the computer host. In an embodiment of the present application, a programmable logic device can be used to isolate the fault of the embedded processor from the computer host. Specifically, after detecting that the embedded processor has a fault, the programmable logic device enters the proxy mode to complete the fault isolation. Among them, a possible implementation method for the programmable logic device to enter the proxy mode can be: the programmable logic device sends a hot plug interrupt signal to the computer host, and the hot plug interrupt signal is used to indicate that the embedded processor has performed a hot plug operation. When the computer host receives the hot plug interrupt signal, it will think that the embedded processor has been hot unplugged, and the computer host will no longer interact with the embedded processor, thereby disconnecting the communication connection between the computer host and the embedded processor to complete the fault isolation.

[0042] As can be seen, after detecting an embedded processor fault, the programmable logic device sends a hot-swap interrupt signal to the host computer, isolating the fault from the host computer. This fault isolation method is more thorough, preventing the fault from spreading to the entire host computer. Furthermore, this fault isolation method eliminates the need to restart the host computer, making it simple and efficient, minimizing the impact on the host computer and ensuring that other host services continue to operate without interruption.

[0043] In a possible implementation, before step S201, a fault detection phase may be further included, which may specifically include the following steps:

[0044] In response to reading a preset flag bit from the computer host, the programmable logic device determines that the embedded processor fails. The preset flag bit is used to indicate that the embedded processor fails.

[0045] In an embodiment of the present application, a FIFO channel can be established between the computer host, the programmable logic device and the embedded processor, and the computer host, the programmable logic device and the embedded processor can exchange data through the FIFO channel. Specifically, the computer host can send a heartbeat packet to the embedded processor according to a predetermined time period, and send the heartbeat packet to the embedded processor through the FIFO channel. After receiving the heartbeat packet, the embedded processor replies a heartbeat packet to the computer host to maintain mutual communication between the computer host and the embedded processor. If the computer host receives the heartbeat packet sent by the embedded processor within a preset time length, it can be determined that the embedded processor is in normal working state, that is, the embedded processor has not failed. If the computer host does not receive the heartbeat packet sent by the embedded processor when exceeding the preset time length, it can be determined that the embedded processor has failed, that is, the embedded processor may have been hung.

[0046] In one possible implementation, the computer host may be provided with a preset program. Upon detecting a fault in the embedded processor, a preset flag may be set in the preset program to indicate that the embedded processor has failed. Upon detecting the preset flag, the programmable logic device may determine that the embedded processor has failed and enter a fault isolation phase, such as executing step S201 to isolate the fault between the computer host and the embedded processor.

[0047] In one possible embodiment, the computer host may be provided with a first status register, which may be used to store the operating status of the embedded processor. The status value of the first status register may be set to: "00" indicates that the embedded processor is operating normally, and "10" indicates that the embedded processor has failed. The preset flag bit may be an identifier used to indicate that the embedded processor has failed. In this embodiment of the present application, the preset flag bit may be understood as the status value "10" of the first status register. When the computer host detects that the embedded processor is operating normally, the status value of the first status register remains unchanged at "00". After detecting that the embedded processor has failed, the computer host may change the status value of the first status register from "00" to "10", indicating that the embedded processor has failed. The programmable logic device may poll and read the status value of the first status register. When the status value of the first status register is "10", i.e., when the preset flag bit is read, it determines that the embedded processor has failed. After detecting that the embedded processor has failed, the programmable logic device enters a fault isolation phase, for example, executing step S201 to isolate the computer host from the embedded processor.

[0048] As can be seen, a FIFO channel is established between the host computer, the programmable logic device, and the embedded processor, enabling data exchange between the host computer, the programmable logic device, and the embedded processor. Upon detecting a fault in the embedded processor, the host computer can set a preset flag to indicate the fault. Upon reading this preset flag, the programmable logic device can determine that the embedded processor has failed. This fault detection method places little burden on the programmable logic device and is simple and efficient.

[0049] In a possible implementation, the specific implementation of fault detection can also be implemented through steps A1-A2:

[0050] Step A1: The programmable logic device obtains register information of a status register, where the register information is used to record the operating status of the embedded processor.

[0051] Step A2: The programmable logic device determines whether the embedded processor fails according to the register information.

[0052] In an embodiment of the present application, the programmable logic device may be provided with a status register, which is used to store the operating status of the embedded processor. The register information stored in the status register is used to record the operating status of the embedded processor.

[0053] Taking a programmable logic device including a 32-bit status register as an example, the meaning of each bit of the status register can be shown in Table 1:

[0054] Table 1

[0055] Bit0 Bit1 Bit2-Bit3 Bit4-Bit8 Bit9-Bit31 First flag Second flag reserve Fault information code reserve

[0056] The specific instructions are as follows:

[0057] Bit 0: The programmable logic device determines whether the embedded processor is working properly. If the embedded processor is working properly, it sets the first flag bit in the status register within a preset time. Here, setting can be understood as setting the corresponding flag bit to 0 or 1.

[0058] Bit 1: The second flag bit is used to store the operating status of the embedded processor fed back by the complex programmable logic device. The programmable logic device can store the information fed back by the complex programmable logic device in Bit 1.

[0059] Bit2-Bit3: reserved bits.

[0060] Bit4-Bit8: If the programmable logic device detects that the embedded processor has a fault, it can store a corresponding fault information code in Bit4-Bit8 according to the fault source and fault type of the embedded processor.

[0061] Bit9-Bit31: reserved bits.

[0062] After obtaining the register information of the status register, it is possible to determine whether the embedded processor is faulty based on the register information. For specific implementation methods, please refer to the following description and will not be repeated here.

[0063] As can be seen, by obtaining the register information of the status register and then judging whether the embedded processor has a fault based on the register information, this method of implementing self-test of the embedded processor through the status register provided by the programmable logic device is simple and efficient, with fewer devices involved in the test and higher fault detection accuracy.

[0064] In a possible implementation, step A2 may specifically include the following steps:

[0065] The programmable logic device reads a status value of a first flag bit in a status register from register information according to a preset cycle; the programmable logic device determines whether the first flag bit is set based on the read status value of the first flag bit; in response to the first flag bit not being set at a preset time, the programmable logic device determines that a fault has occurred in the embedded processor.

[0066] In an embodiment of the present application, the self-detection of the embedded processor can be implemented through the status register provided by the programmable logic device. Specifically, the operating status of the embedded processor can be associated with the first flag bit of the status register (for example, Bit 0 shown in Table 1). The programmable logic device periodically (for example, 1 minute) sends a discovery request to the embedded processor. If the embedded processor that receives the discovery request is operating normally, it will reply a discovery response to the programmable logic device, that is, respond to the discovery request sent by the programmable logic device, and at this time the first flag bit can be set. Wherein, setting can be understood as setting the corresponding flag bit to 0 or 1. For example, the programmable logic device responds to the discovery request sent by the embedded processor. If the status value of the first flag bit in the previous cycle is "0", the status value of the first flag bit is set to 1; or if the status value of the first flag bit in the previous cycle is "1", the status value of the first flag bit is set to 0. In other words, if the embedded processor is operating normally, the status value of the first flag bit of the status register in the programmable logic device will change periodically. If the embedded processor fails, it will not be able to respond to the discovery request sent by the embedded processor, and the first flag bit cannot be set. For example, if the programmable logic device is unable to respond to the discovery request sent by the embedded processor, if the status value of the first flag bit in the previous cycle is "0", the status value of the first flag bit this time is still "0"; or if the status value of the first flag bit in the previous cycle is "1", the status value of the first flag bit this time is still "1". In other words, if the embedded processor fails, the status value of the first flag bit of the status register in the programmable logic device will no longer change periodically. Therefore, in an embodiment of the present application, the programmable logic device can read the status value of the first flag bit in the status register from the register information according to a preset period, and then determine whether the first flag bit is set based on the status value of the read first flag bit. If the first flag bit is not set when the preset time comes, it can be determined that the embedded processor has failed. The preset period and preset time can be determined according to the actual application scenario, and the embodiment of the present application does not limit this.

[0067] As can be seen, the programmable logic device can read the status value of the first flag bit in the status register from the register information at a preset period. Then, based on the read status value of the first flag bit, it determines whether the first flag bit is set. If the first flag bit is not set after the preset time, it is determined that the embedded processor has failed. This method of implementing embedded processor self-testing through the status register provided by the programmable logic device is simple and efficient, involves fewer devices in the test, and has high fault detection accuracy.

[0068] In a possible implementation, step A2 may further include the following steps:

[0069] The programmable logic device reads the status value of the second flag bit in the status register from the register information. The status value of the second flag bit is associated with the first signal fed back by the complex programmable logic device. The complex programmable logic device is used to detect the operating status of the embedded processor and feed back the first signal to the programmable logic device based on the operating status of the embedded processor. In response to the status value of the second flag bit being a preset value, the programmable logic device determines that a fault has occurred in the embedded processor.

[0070] In an embodiment of the present application, a programmable logic device can detect whether an embedded processor fails through a complex programmable logic device connected to the embedded processor. Among them, the complex programmable logic device can communicate with the embedded processor and the programmable logic device respectively, and the complex programmable logic device can have a built-in watchdog module for monitoring the operation of the embedded processor. When the embedded processor is working normally, it will send a feedback signal to the timer of the watchdog module within a preset time (for example, 1 second or 0.5 seconds, etc.) to clear the timer to realize the dog feeding function. If the embedded processor does not feed the timer within the preset time, the complex programmable logic device can feedback a first signal to the programmable logic device to remind the embedded processor that it has not responded within the preset time. When the programmable logic device receives the first signal, it can modify the status value of the bit corresponding to the complex programmable logic device in the status register (for example, Bit1 shown in Table 1) to a preset value to indicate that the embedded processor has failed.

[0071] For example, the status value of the second flag bit in the status register (see Bit 1 shown in Table 1) can be set to: "00" indicates that the embedded processor is operating normally, and "10" indicates that the embedded processor has failed. The preset value can be an identifier used to indicate that the embedded processor has failed. In the embodiment of the present application, the preset value can be understood as the status value "10" of the second flag bit. When the programmable logic device does not receive the first signal for reporting that the embedded processor has failed, the status value of the second flag bit remains unchanged at "00". After receiving the first signal, the programmable logic device can change the status value of the second flag bit from "00" to "10", indicating that the embedded processor has failed. When the programmable logic device reads the status value of the second flag bit as "10", that is, when the preset flag bit is read, it determines that the embedded processor has failed. In one possible embodiment, the second flag bit and the first flag bit can also be the same flag bit, that is, the same bit bit is used to implement the functions of the first flag bit and the second flag bit. For example, Bit 1 can be used to store information about the operating status of the embedded processor reported by the complex programmable logic device, and can also be used to store information that the programmable logic device determines whether the embedded processor is operating normally. After detecting that the embedded processor has failed, the programmable logic device enters a fault isolation phase, for example, step S201 may be executed to complete fault isolation between the computer host and the embedded processor.

[0072] As can be seen, a complex programmable logic device (CPLD) is used to detect faults in an embedded processor. The CPLD's detection results are then correlated with the PLD's status register. If the second flag in the status register reads a preset value, the embedded processor is determined to have a fault. Using a CPLD to detect faults in an embedded processor expands the range of methods available for detecting faults in embedded processors, while placing a relatively small burden on the PLD and improving the accuracy of fault detection.

[0073] In a possible implementation, the proxy mode further includes the programmable logic device sending a first processing layer data packet for reporting an error to the computer host, so as to isolate the computer host from the embedded processor fault.

[0074] In an embodiment of the present application, when a computer host performs business interaction with an embedded processor, the computer host may first send the business data to a programmable logic device, which then forwards the business data to the embedded processor for processing. If the embedded processor fails, it will be unable to process the business data. At this time, after the computer host sends the business data to the programmable logic device, the programmable logic device may enter a proxy mode to prompt the computer host that the embedded processor has failed, thereby isolating the computer host from the embedded processor. Specifically, the programmable logic device may send a first processing layer data packet to the computer host, and the first processing layer data packet is used to report an error to the computer host to prompt the embedded processor that a fault has occurred, thereby isolating the computer host from the embedded processor.

[0075] As can be seen, after detecting an embedded processor fault, the programmable logic device sends a first-layer processing packet reporting the error to the computer host, isolating the fault between the computer host and the embedded processor. This fault isolation method eliminates the need to restart the computer host, minimizing the impact on the computer host and ensuring that other computer services can operate normally without interruption.

[0076] In a possible implementation, the programmable logic device sending a first processing layer data packet for reporting an error to a computer host may include the following steps:

[0077] Acquire register information of a status register, where the register information is used to record the operating status of the embedded processor; read the fault source and fault type of the embedded processor from the register information; generate a first processing layer data packet for reporting an error according to the fault source and fault type; and send the first processing layer data packet to a computer host.

[0078] The PCIe standard classifies PCIe device faults into three categories based on severity: correctable error (CE), non-fatal uncorrectable error (NFE), and fatal uncorrectable error (FE). Correctable errors are automatically identified and corrected or recovered by the hardware; non-fatal uncorrectable errors are typically handled directly by the device driver software; and fatal uncorrectable errors are typically handled by system software and typically require a reset. The source of the fault can be understood as the specific device experiencing the fault.

[0079] In an embodiment of the present application, multiple embedded processors may exist. For example, embedded processor A and embedded processor B may be associated with respective bits of a status register in a programmable logic device. As shown in Table 1, the operating status of embedded processor A may be stored in Bit 4, and the operating status of embedded processor B may be stored in Bit 5. Accordingly, the status values ​​of Bits 4 and 5 in the status register may be set to "00" to indicate normal operation, "01" to indicate a correctable error, "10" to indicate a non-fatal uncorrectable error, and "11" to indicate a fatal uncorrectable error. For example, when the status value of Bit 4 is "00" read from the programmable logic device, embedded processor A can be determined to be operating normally. When the status value of Bit 5 is "01" read from the programmable logic device, the fault source can be determined to be embedded processor B, and the fault type is a correctable error. Alternatively, when the status value of Bit 5 is "10" read from the programmable logic device, the fault source can be determined to be embedded processor B, and the fault type is a non-fatal uncorrectable error. Then, a first processing layer data packet can be generated based on the processor fault source and fault type in the register information, and the first processing layer data packet can be fed back to the computer host for error reporting, thereby completing fault isolation. The first processing layer data packet can also include a PCIe device identifier of the faulty embedded processor. The PCIe device identifier can be an identifier that identifies the PCIe device in the PCIe bus system. Specifically, the PCIe device identifier can be a bus device function (BDF) of the PCIe device or address information of the PCIe device, etc., so as to quickly locate the faulty embedded processor.

[0080] It can be seen that the programmable logic device reads the fault source and fault type of the embedded processor from the register information, then generates a first processing layer data packet based on the fault source and fault type, and feeds back the first processing layer data packet to the computer host, which helps to quickly locate the faulty embedded processor and improve the accuracy of fault location.

[0081] Step S202: In response to detecting that the embedded processor is repaired, the programmable logic device sends a hot plug signal to the computer host and exits the answering mode to complete the fault recovery.

[0082] In an embodiment of the present application, the fault repair method may be self-repair or restart of the faulty embedded processor, etc., which is not limited in the embodiment of the present application. After the embedded processor fault repair is completed and can operate normally, a recovery request may be sent to the programmable logic device, and the recovery request is used to indicate that the embedded processor fault repair is complete. After receiving the recovery request, the programmable logic device determines that the embedded processor repair is complete based on the recovery request, and may send a hot plug signal to the computer host, and the hot plug signal may be used to indicate that the embedded processor has performed a hot plug operation. After receiving the hot plug signal, the computer host will believe that the embedded processor has performed a hot plug, and the computer host can interact with the embedded processor normally. At this point, the programmable logic device exits the answering mode, the computer host can communicate with the embedded processor again, and the fault recovery is completed.

[0083] exist Figure 2 As can be seen from the method shown, after detecting that the embedded processor in the data processor has a fault, the programmable logic device can enter the proxy mode and send a hot plug interrupt signal to the computer host. The hot plug interrupt signal is used to indicate that the embedded processor has performed a hot plug operation to disconnect the communication between the computer host and the embedded processor, thereby completing fault isolation simply and efficiently. In this way, the problem that in the traditional technical solution, once the embedded processor fails, the computer host must be restarted, resulting in the interruption of all running programs, can be avoided, thereby minimizing the impact on the computer host and ensuring the normal operation of the computer host. After detecting that the embedded processor has been repaired, the programmable logic device sends a hot plug signal to the computer host. The hot plug signal is used to indicate that the embedded processor has performed a hot plug operation and exits the proxy mode, allowing the computer host to re-communicate with the embedded processor, thereby quickly completing fault recovery and improving the efficiency of fault handling. In addition, the embodiment of the present application does not require the participation of external tools such as the BMC system and the control platform, which can reduce the degree of dependence and have higher reliability.

[0084] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.

[0085] Please refer to Figure 3 , Figure 3 This is a schematic diagram of a fault handling device provided in an embodiment of the present application. The device is applied to a programmable logic device. Figure 3 As shown, the fault handling apparatus 300 includes a fault isolation unit 301 and a fault recovery unit 302. The detailed description of each unit is as follows:

[0086] a fault isolation unit 301 configured to enter a proxy mode in response to detecting a fault in an embedded processor in a data processor, wherein the proxy mode includes sending a hot plug interrupt signal to a computer host to isolate the computer host from the embedded processor fault, wherein the hot plug interrupt signal is used to indicate that the embedded processor has performed a hot plug operation;

[0087] The fault recovery unit 302 is used to send a hot plug signal to the computer host in response to detecting that the embedded processor is repaired, exit the answering mode, and complete fault recovery. The hot plug signal is used to indicate that the embedded processor has performed a hot plug operation.

[0088] In a possible implementation, the fault handling apparatus 300 may further include: Figure 3 The fault detection unit not shown in the figure can be specifically used to obtain register information of the status register, the register information is used to record the operating status of the embedded processor; and determine whether the embedded processor has a fault according to the register information.

[0089] In a possible implementation, the fault handling apparatus 300 may further include: Figure 3 A fault detection unit not shown in the figure can be specifically used to read the status value of the first flag bit in the status register from the register information according to a preset period; determine whether the first flag bit is set according to the read status value of the first flag bit; in response to the first flag bit not being set when a preset time is reached, determine that the embedded processor has a fault.

[0090] In a possible implementation, the fault handling apparatus 300 may further include: Figure 3 A fault detection unit not shown in the figure can be specifically used to read the status value of the second flag bit in the status register from the register information, the status value of the second flag bit is associated with the first signal fed back by the complex programmable logic device, the complex programmable logic device is used to detect the operating status of the embedded processor and feed back the first signal to the programmable logic device according to the operating status of the embedded processor; in response to the status value of the second flag bit being read as a preset value, it is determined that the embedded processor has a fault.

[0091] In a possible implementation, the fault isolation unit 301 is further configured to send a first processing layer data packet for reporting an error to the computer host, so as to isolate the computer host from the embedded processor fault.

[0092] In one possible embodiment, the fault isolation unit 301 is specifically used to obtain register information of the status register, where the register information is used to record the operating status of the embedded processor; read the fault source and fault type of the embedded processor from the register information; generate a first processing layer data packet for reporting an error based on the fault source and the fault type; and send the first processing layer data packet to the computer host.

[0093] In a possible implementation, there is a first-in-first-out queue channel between the computer host, the programmable logic device, and the embedded processor, and the fault handling device 300 may further include: Figure 3 The fault detection unit not shown in the figure can be specifically used to determine that the embedded processor has a fault in response to reading a preset flag bit from the computer host, and the preset flag bit is used to indicate that the embedded processor has a fault.

[0094] It should be noted that the implementation of each unit can also refer to Figure 2 The corresponding description of the method embodiment shown.

[0095] Please refer to Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 4 As shown, the computer device 400 includes a processor 401, a memory 402, and a communication interface 403, wherein the memory 402 stores a computer program 404. The processor 401, the memory 402, the communication interface 403, and the computer program 404 may be connected via a bus 405.

[0096] When the computer device is a programmable logic device, the computer program 404 is used to execute instructions for the following steps:

[0097] In response to detecting a fault in an embedded processor in a data processor, entering a proxy mode, the proxy mode comprising sending a hot plug interrupt signal to a computer host to isolate the computer host from the fault in the embedded processor, the hot plug interrupt signal being used to indicate that the embedded processor has performed a hot plug operation;

[0098] In response to detecting that the embedded processor is repaired, a hot plug signal is sent to the computer host to exit the answering mode to complete fault recovery. The hot plug signal is used to indicate that the embedded processor has performed a hot plug operation.

[0099] In one possible implementation, the programmable logic device includes a status register, and before the programmable logic device enters the answering mode in response to detecting a fault in the embedded processor in the data processor, the computer program 404 is further configured to execute instructions for the following steps:

[0100] Acquiring register information of the status register, where the register information is used to record the operating status of the embedded processor;

[0101] Determine whether the embedded processor fails according to the register information.

[0102] In a possible implementation, in determining whether the embedded processor fails according to the register information, the computer program 404 is specifically configured to execute instructions for the following steps:

[0103] Reading a status value of a first flag bit in the status register from the register information according to a preset period;

[0104] Determining whether the first flag bit is set according to the read state value of the first flag bit;

[0105] In response to the first flag not being set when a preset time arrives, it is determined that the embedded processor fails.

[0106] In a possible implementation, in determining whether the embedded processor fails according to the register information, the computer program 404 is specifically configured to execute instructions for the following steps:

[0107] Reading a status value of a second flag bit in the status register from the register information, where the status value of the second flag bit is associated with a first signal fed back by a complex programmable logic device, the complex programmable logic device being configured to detect an operating status of the embedded processor and feed back the first signal to the programmable logic device based on the operating status of the embedded processor;

[0108] In response to reading that the status value of the second flag bit is a preset value, it is determined that the embedded processor fails.

[0109] In a possible implementation, the answering mode further includes:

[0110] A first processing layer data packet for reporting an error is sent to the computer host, so as to isolate the computer host from the embedded processor fault.

[0111] In a possible implementation, the programmable logic device includes a status register, and in terms of sending a first processing layer data packet for reporting an error to the computer host, the computer program 404 is specifically configured to execute instructions for the following steps:

[0112] Acquiring register information of the status register, where the register information is used to record the operating status of the embedded processor;

[0113] Reading the fault source and fault type of the embedded processor from the register information;

[0114] generating a first processing layer data packet for error reporting according to the fault source and the fault type;

[0115] The first processing layer data packet is sent to the computer host.

[0116] In one possible implementation, a first-in-first-out queue channel exists between the computer host, the programmable logic device, and the embedded processor. Before the programmable logic device enters the proxy mode in response to detecting a fault in the embedded processor in the data processor, the computer program 404 is further configured to execute instructions for the following steps:

[0117] In response to reading a preset flag bit from the computer host, it is determined that the embedded processor fails, and the preset flag bit is used to indicate that the embedded processor fails.

[0118] Those skilled in the art will understand that for ease of explanation, Figure 4 Only one memory and processor are shown. In an actual terminal or server, there may be multiple processors and memories. Memory 402 may also be referred to as a storage medium or storage device, etc., which is not limited in this embodiment of the present application.

[0119] It should be understood that in the embodiment of the present application, the processor 401 can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0120] It should also be understood that the memory 402 mentioned in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronize link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0121] It should be noted that when the processor 401 is a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) is integrated into the processor.

[0122] It should be noted that the memory 402 described herein is intended to comprise, but is not limited to, these and any other suitable types of memory.

[0123] The bus 405 may include, in addition to the data bus, a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are all labeled as buses in the figure.

[0124] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.

[0125] In various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0126] Those skilled in the art will appreciate that the various illustrative logical blocks (ILBs) and steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0127] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0128] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0129] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0130] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).

[0131] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement part or all of the steps of any fault handling method recorded in the above method embodiments.

[0132] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute part or all of the steps of any fault handling method recorded in the above method embodiments.

[0133] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A fault handling method, applied to a programmable logic device in a data processor, characterized in that: The data processor includes the programmable logic device and an embedded processor, the programmable logic device includes a field programmable gate array or a dedicated integrated circuit; a first-in first-out queue channel is established between the computer host, the programmable logic device and the embedded processor; The programmable logic device communicates with the computer host via a PCIe interface; The method includes: in response to a preset flag being read from the computer host based on the first-in-first-out queue channel, the programmable logic device determining that the embedded processor has failed, the preset flag being used to indicate that the embedded processor has failed; the programmable logic device is provided with a status register, the status register being used to store an operating status of the embedded processor; In response to detecting a fault in an embedded processor in a data processor, the programmable logic device enters a proxy mode, the proxy mode including sending a hot plug interrupt signal to a computer host to isolate the computer host from the fault of the embedded processor, the hot plug interrupt signal being used to indicate that the embedded processor has performed a hot plug operation; In response to detecting that the embedded processor is repaired, the programmable logic device sends a hot plug signal to the computer host to exit the answering mode to complete fault recovery. The hot plug signal is used to indicate that the embedded processor has performed a hot plug operation.

2. The method according to claim 1, characterized in that The programmable logic device includes a status register, and before the programmable logic device enters the answering mode in response to detecting that the embedded processor in the data processor has failed, further includes: The programmable logic device obtains register information of the status register; The programmable logic device determines whether a fault occurs in the embedded processor according to the register information.

3. The method according to claim 2, characterized in that The programmable logic device determines whether the embedded processor fails according to the register information, including: The programmable logic device reads the state value of the first flag bit in the state register from the register information according to a preset cycle; The programmable logic device determines whether the first flag bit is set according to the state value of the read first flag bit; In response to the first flag not being set at the preset time, the programmable logic device determines that the embedded processor fails.

4. The method according to claim 2, characterized in that The programmable logic device determines whether the embedded processor fails according to the register information, including: The programmable logic device reads a status value of a second flag bit in the status register from the register information, where the status value of the second flag bit is associated with a first signal fed back by a complex programmable logic device, and the complex programmable logic device is used to detect an operating status of the embedded processor and feed back the first signal to the programmable logic device based on the operating status of the embedded processor; In response to reading that the status value of the second flag bit is a preset value, the programmable logic device determines that the embedded processor fails.

5. The method according to any one of claims 1 to 4, characterized in that The answering mode also includes: The programmable logic device sends a first processing layer data packet for reporting an error to the computer host, so as to isolate the computer host from the embedded processor fault.

6. The method according to claim 5, characterized in that The programmable logic device includes a status register, and the programmable logic device sends a first processing layer data packet for reporting an error to the computer host, including: The programmable logic device obtains register information of the status register, where the register information is used to record the operating status of the embedded processor; The programmable logic device reads the fault source and fault type of the embedded processor from the register information; The programmable logic device generates a first processing layer data packet for error reporting according to the fault source and the fault type; The programmable logic device sends the first processing layer data packet to the computer host.

7. A fault handling device, applied to a programmable logic device in a data processor, characterized in that: The data processor includes the programmable logic device and an embedded processor, the programmable logic device includes a field programmable gate array or a dedicated integrated circuit; a first-in first-out queue channel is established between the computer host, the programmable logic device and the embedded processor; The programmable logic device communicates with the computer host via a PCIe interface; The fault handling device includes: a fault isolation unit, configured to, in response to a preset flag being read from the computer host based on the first-in-first-out queue channel, cause the programmable logic device to determine that the embedded processor has failed, the preset flag being used to indicate that the embedded processor has failed; the programmable logic device being provided with a status register for storing an operating status of the embedded processor; in response to detecting that the embedded processor in the data processor has failed, the programmable logic device entering a proxy mode, the proxy mode including sending a hot-swap interrupt signal to the computer host to isolate the computer host from the embedded processor failure, the hot-swap interrupt signal being used to indicate that the embedded processor has performed a hot-swap operation; A fault recovery unit is used to respond to detecting that the repair of the embedded processor is completed, and the programmable logic device sends a hot plug signal to the computer host to exit the answering mode to complete fault recovery. The hot plug signal is used to indicate that the embedded processor has performed a hot plug operation.

8. A computer device, characterized in that: The method comprises a processor, a memory and a communication interface, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which enables a computer to execute the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product is used to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Main / standby switching system and method of dual-wan PORT network apparatus

    CN104753710A

  • Equipment management method and device and server

    CN110457164A