PCIe device fault alarm method, system, device, medium and product
By obtaining and judging the prefetchable memory value of PCIe devices, disabling abnormal devices and recording log trigger alarms, it solves the server downtime caused by abnormal memory configuration space integration of PCIe devices under the Switch chip, and improves server stability and operation and maintenance efficiency.
Patent Information
- Application Number
- CN202510668324.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-22
AI Technical Summary
When there is an integration abnormality in the prefetchable memory configuration space of the PCIe device hanging under the Switch chip, there is no log or alarm and the server is down on the boot logo interface, which makes it difficult to quickly locate the problem.
By obtaining the current prefetchable memory values of multiple PCIe devices, it is determined whether they are all greater than or equal to the preset value or are all less than or equal to the preset value. If it is uneven, the PCIe device with the current prefetchable memory value greater than or equal to the preset value is disabled, and the hardware signal of the PCIe device with the current prefetchable memory value less than the preset value is obtained based on the substrate management controller, identify the target PCIe device with inconsistent signals, record the exception log and trigger an alarm.
The situation of no logs and alarms has been solved, which has significantly improved the stability and reliability of the server system, optimized the fault alarm and logging mechanism, reduced operation and maintenance costs, and improved overall operation and maintenance efficiency.
Smart Images

Figure CN120196519A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a fault alarm method, system, device medium and product for a PCIe device. Background Art
[0002] In the current AI server architecture, the Switch chip, as a key component for expanding PCIe (Peripheral Component Interconnect Express) devices, plays an indispensable role in connection and exchange, providing a powerful expansion capability for the CPU (Central Processing Unit), allowing the CPU to easily connect to more devices and achieve efficient interconnection between PCIe devices. As server configurations become increasingly complex, PCIe Switch chips provide strong support for the connection and expansion of multiple devices with their rich port and channel resources. In actual application scenarios, PCIe Switch chips need to connect various PCIe devices with different functions, and uniformly integrate and manage the memory resource space of the PCIe devices expanded by their downstream ports. However, the memory resource space of these PCIe devices requires more than 4GB of memory space, while others only require less than 4GB of memory space.
[0003] However, because the PCIe Switch can only allocate a unified address range when integrating the pre-fetchable memory resource space of downstream devices, and the address range is either above 4GB or below 4GB, when the PCIe devices under the PCIe Switch chip include both devices that require more than 4GB of memory resource space and devices that require less than 4GB of memory resource space, the PCIe Switch chip prioritizes reserving the memory resource space below 4GB, resulting in the PCIe devices that require more than 4GB of memory resource space being unable to obtain the required memory resources and unable to complete the initialization process, causing the server to crash on the startup logo interface. During this process, the system will not generate any relevant logs and alarm information, which undoubtedly makes it extremely difficult to quickly locate the problem. Summary of the invention
[0004] The present invention provides a fault alarm method, system, device, medium and product for PCIe devices, so as to at least solve the problem that when an integration abnormality occurs in the prefetchable memory configuration space of the PCIe device hung under the Switch chip, there is no log and alarm and the server crashes on the boot logo interface, which is not conducive to rapid problem location.
[0005] The present invention provides a method for fault warning of PCIe devices, including the following steps: obtaining the current prefetchable memory values of multiple PCIe devices, and determining whether the current prefetchable memory values of the multiple PCIe devices are all greater than or equal to a preset value, or whether the current prefetchable memory values of the multiple PCIe devices are all less than or equal to the preset value; if the current prefetchable memory values of the multiple PCIe devices are not all greater than or equal to the preset value, and the current prefetchable memory values of the multiple PCIe devices are not all less than or equal to the preset value, then disabling at least one first PCIe device among the multiple PCIe devices whose current prefetchable memory value is greater than or equal to the preset value, and obtaining a first hardware signal of at least one second PCIe device among the multiple PCIe devices whose current prefetchable memory value is less than the preset value based on a baseboard management controller, and identifying a second hardware signal of the at least one second PCIe device based on a host system; determining at least one target PCIe device whose first hardware signal and second hardware signal are inconsistent among the at least one second PCIe device, and recording an exception log of the at least one target PCIe device, and triggering an exception warning.
[0006] The present invention also provides a fault warning system for PCIe devices, including: a judgment module, configured to obtain the current prefetchable memory values of multiple PCIe devices, and determine whether the current prefetchable memory values of the multiple PCIe devices are all greater than or equal to a preset value, or whether the current prefetchable memory values of the multiple PCIe devices are all less than or equal to the preset value; an obtaining module, configured to, if the current prefetchable memory values of the multiple PCIe devices are not all greater than or equal to the preset value, and the current prefetchable memory values of the multiple PCIe devices are not all less than or equal to the preset value, then disable at least one first PCIe device among the multiple PCIe devices whose current prefetchable memory value is greater than or equal to the preset value, and obtain a first hardware signal of at least one second PCIe device among the multiple PCIe devices whose current prefetchable memory value is less than the preset value based on a baseboard management controller, and identify a second hardware signal of the at least one second PCIe device based on a host system; an alarm module, configured to determine at least one target PCIe device whose first hardware signal and second hardware signal are inconsistent among the at least one second PCIe device, and record an exception log of the at least one target PCIe device, and trigger an exception warning.
[0007] The present invention also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the above-mentioned method for fault warning of PCIe devices when executing the computer program.
[0008] The present invention also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the above-mentioned fault warning method for PCIe devices are implemented.
[0009] The present invention also provides a computer program product including a computer program, wherein when the computer program is executed by a processor, the above-mentioned fault warning method for PCIe devices is implemented.
[0010] Through the present invention, if the current prefetch memory values of multiple PCIe devices are not all greater than or equal to a preset value and are not all less than or equal to the preset value, at least one first PCIe device among the multiple PCIe devices with a current prefetch memory value greater than or equal to the preset value is disabled, and a first hardware signal of at least one second PCIe device among the multiple PCIe devices with a current prefetch memory value less than the preset value is obtained based on a baseboard management controller, and a second hardware signal of at least one second PCIe device is identified based on a host system; at least one target PCIe device with inconsistent first and second hardware signals of at least one second PCIe device is determined, an exception log of at least one target PCIe device is recorded, and an exception warning is triggered. Thereby, the problem that when the prefetch memory configuration space of PCIe devices under a Switch chip has an integration anomaly, there is no log or warning and the server crashes at the boot logo interface, which is not conducive to quickly locating problems, is solved. This not only significantly improves the stability and reliability of the server system, but also optimizes the fault warning and log recording mechanisms, reduces the operation and maintenance costs, and improves the overall operation and maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0012] Figure 1 FIG. is a flowchart of a fault warning method for a PCIe device according to an embodiment of the present invention; Figure 2 FIG. is a connection schematic diagram of a hardware design according to an embodiment of the present invention; Figure 3 FIG. is a schematic flow diagram of a fault warning method for a PCIe device according to an embodiment of the present invention; Figure 4 FIG. is a schematic diagram of a fault warning system for a PCIe device according to an embodiment of the present invention; Figure 5Schematic diagram of the structure of an electronic device according to an embodiment of the present invention.
[0013] Reference numerals: 10 - Fault warning system of PCIe device, 100 - Judgment module, 200 - Acquisition module, 300 - Warning module, 503 - Communication interface, 501 - Memory, 502 - Processor. Detailed implementation manners
[0014] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and not to describe a specific order or sequence.
[0016] Before specifically introducing the embodiments of the present invention, a brief introduction to PCIe devices is given.
[0017] A brief introduction to the configuration space of PCIe devices is as follows: The PCIe configuration space is 4KB in total and is divided into multiple parts, each with a specific function; it is mainly divided into the Header and the Device-Specific Register + PCIe Optional Configuration Space. 0-3Fh (64 bytes) is the PCI-compatible configuration space header, which can be divided into Type 0 configuration space header and Type 1 configuration space header according to the type. 40h-FFh (192 bytes) is the configuration space for PCI / PCI-X and PCIe extensions of the peripheral device interconnect, mainly storing some capability structures related to the MSI (Message Signaled Interrupts, a way to generate interrupts by writing information in memory) or MSI-X (an interrupt method extended based on MSI) interrupt mechanism and power management. 100h-FFFh (3840 bytes) is the optional configuration space extended by the PCIe protocol, mainly storing capability structures such as AER, virtual channels, and device serial numbers. Among them, the Prefetchable Memory Base (used to save the low 32-bit information of the starting address of the prefetchable memory area), Prefetchable Memory Limit (used to save the low 32-bit information of the ending address of the prefetchable memory area), Prefetchable Base Upper 32 Bits (used to save the high 32-bit information of the ending address of the prefetchable memory area), and Prefetchable Limit Upper 32 Bits (used to save the high 32-bit information of the ending address of the prefetchable memory area) jointly determine the range of the prefetchable address space.
[0018] In the first related technology, a PCIe component identification device is provided, including a CPLD (Complex Programmable Logic Device). The CPLD is connected to a BMC chip through an I2C link. At the same time, the CPLD is also communicatively connected to a PCA (Programmable Counter Array) chip. The PCA chip, as an expansion unit, branches out multiple I2C links, each link corresponding to an I2C port, and these I2C ports are successively connected to multiple PCIe connectors. The BMC (Baseboard Management Controller) chip is connected to the flash memory part of the BIOS chip (Basic Input Output System Chip) through a KCS channel. The BMC chip can read the PCIe component asset information identified by the BIOS chip from this flash memory, and then perform a legality check on this information. After the check is completed, the BMC chip sends the check result back to the CPLD through the previous I2C link. To intuitively display the status of PCIe components, the product structure also includes an indicator light electrically connected to the CPLD. This device can identify and manage illegal PCIe components.
[0019] The second related technology mainly discloses a method for a PCIE hot-plug device, a status warning method, and a readable storage medium. The method is applied to a PCIE hot-plug device and includes: when a PCIE device removal instruction is received by a processing module, determining whether the PCIE device is occupied; in response to the PCIE device not being occupied, unloading the driver of the PCIE device through the processing module, controlling an alarm to give a first alarm, and controlling the power supply to stop supplying power to the PCIE slot; or, in response to the PCIE device being occupied, controlling the alarm to give a second alarm through the processing module; where: the first alarm is used to notify the user that it is allowed to remove the PCIE device, and the second alarm is used to notify the user that it is not allowed to remove the PCIE device. The present invention can enable the user to perceive the device status of the PCIE hot-plug device, reducing the risk of damage to the PCIE device during hot plugging due to the user's inability to timely perceive the device status.
[0020] However, the first related technology mainly focuses on obtaining PCIe component asset information through the I2C link after the PCIe device enumeration is completed, without comparing the hardware presence status of the PCIe device with the system-recognized PCIe device status. It mainly realizes the identification and restriction of illegal PCIe components in the server, avoiding bugs caused by illegal components, but does not solve the complex scenario where the component is physically present but logically invisible; The second related technology mainly focuses on solving the problem that in the process of PCIe hot plugging, due to the lack of a notification mechanism, operators cannot timely perceive the device status, which increases the risk of damage. Through an alarm mechanism, operators can timely perceive the device status, clarify whether they can safely remove or connect a PCIE device, reduce the risk of device damage, and improve the reliability and security of the hot plugging process.
[0021] To sum up, the current PCIe error reports of servers are mainly divided into Baseline Error Reporting and Advanced Error Reporting (AER, Advanced Error Reporting). When the prefetchable memory configuration space of PCIe devices under a Switch cannot be integrated, there is no log or alarm, and the server directly crashes at the startup log interface. It is difficult for development and maintenance personnel to determine the source of the problem and locate whether it is a hardware problem, a component problem, or a related software problem, which brings great difficulties to the development and maintenance personnel in locating the problem.
[0022] To solve the above problems, an embodiment of the present invention provides a method for fault alarm of PCIe devices.
[0023] As Figure 1 shown, the method for fault alarm of the PCIe device includes the following steps: Step S101, obtain the current prefetchable memory values of multiple PCIe devices, and determine whether the current prefetchable memory values of the multiple PCIe devices are all greater than or equal to a preset value, or whether the current prefetchable memory values of the multiple PCIe devices are all less than or equal to the preset value.
[0024] Among them, prefetchable memory refers to the memory area that allows the CPU to perform cache optimization, and the current prefetchable memory value can be understood as the memory value that the PCIe device requires for cache optimization.
[0025] Among them, the PCIe Switch chip connects multiple PCIe devices to the upstream port and configures the resource allocation of each port. In the embodiment of the present invention, the preset value is 4G (4G refers to the upper limit of the 32-bit address space, and 4GB is 0xFFFFFFFF).
[0026] It should be understood that the embodiment of the present invention adds a pre-check design to the Switch firmware, that is, during the enumeration stage of multiple PCIe devices, it is detected whether the current prefetchable memory value of the multiple PCIe devices is greater than or equal to the preset value.
[0027] Specifically, a memory resource predictor is integrated inside the Switch chip. When a PCIe device is enumerated, by reading the base address registers BAR0 to BAR5 of multiple PCIe devices attached to the switch chip, the memory resource space sizes of multiple PCIe devices connected to the downstream ports of the Switch are traversed, and ports without connected PCIe devices are automatically skipped to obtain the current prefetchable memory value of each PCIe device. Among them, the base address register BAR defines the location and size of the memory or I / O space required by the PCIe device.
[0028] Optionally, in some embodiments, obtaining the current prefetchable memory values of multiple PCIe devices includes: determining whether a preset flag bit in the base address registers of the multiple PCIe devices is a prefetchable memory flag bit; if the preset flag bit is a prefetchable memory flag bit, then obtaining the current prefetchable memory values of the multiple PCIe devices according to the received prefetchable memory request.
[0029] In the embodiments of the present invention, the preset flag bit of the base address register is bit3. If bit3 is 1, it is determined that the preset flag bit is a prefetchable memory flag bit.
[0030] Specifically, when the type of the base address register BAR of multiple PCIe devices is Memory Space (that is, bit0 of the base address register BAR is 0), and the preset flag bit bit3 is 1, it is determined that the preset flag bit is a prefetchable memory flag bit, and it is determined that the memory requests sent by the multiple PCIe devices are prefetchable memory requests. At this time, the current prefetchable memory values of the multiple PCIe devices can be obtained through the base address register BAR.
[0031] Through the above technical solution, based on the standard fields (bit0 and bit3) of the BAR register in the PCIe specification, the logic for detecting whether the memory requests of multiple PCIe devices are prefetchable memory requests is implemented. Different vendor devices can be adapted without modifying the hardware design, reducing the firmware development complexity.
[0032] Step S102, if the current prefetchable memory values of multiple PCIe devices are not all greater than or equal to the preset value, and the current prefetchable memory values of multiple PCIe devices are not all less than or equal to the preset value, then disable at least one first PCIe device among the multiple PCIe devices whose current prefetchable memory values are greater than or equal to the preset value, and obtain the first hardware signals of at least one second PCIe device among the multiple PCIe devices whose current prefetchable memory values are less than the preset value based on the baseboard management controller, and identify the second hardware signals of at least one second PCIe device based on the host system.
[0033] It is understandable that the current prefetchable memory values of multiple PCIe devices are not all greater than or equal to the preset value, and the current prefetchable memory values of multiple PCIe devices are not all less than or equal to the preset value. That is, among multiple PCIe devices, there are devices with current prefetchable memory values less than or equal to 4G and devices with current prefetchable memory values greater than or equal to 4G at the same time.
[0034] If there are devices with values greater than 4G and less than 4G among multiple PCIe devices at the same time, the PCIe Switch preferentially reserves the ports of the memory resource space of at least one second PCIe device with a current prefetchable memory value less than or equal to 4G, and disables the ports of at least one first PCIe device with a current prefetchable memory value greater than or equal to the preset value.
[0035] Furthermore, the host system BIOS (Basic Input / Output System) identifies the second hardware signals of multiple PCIe devices (the second hardware signals are obtained through PCIe enumeration or the ACPI table). Compare the first hardware signal of at least one second PCIe device with the second hardware signal identified by the host system for at least one second PCIe device to detect whether there are PCIe devices with signal conflicts among at least one second PCIe device.
[0036] Optionally, in some embodiments, after disabling at least one first PCIe device with a current prefetchable memory value greater than or equal to the preset value among multiple PCIe devices, it includes: sending a thermal reset signal to the port corresponding to at least one first PCIe device to reset at least one first PCIe device.
[0037] It is understandable that after disabling at least one first PCIe device with a current prefetchable memory value greater than or equal to the preset value, a PCIe thermal reset signal is sent to the port corresponding to at least one first PCIe device to reset at least one first PCIe device, and the complete configuration of the port of the memory resource space of at least one second PCIe device with a current prefetchable memory value less than or equal to 4G is retained.
[0038] Through the above technical solution, after disabling at least one first PCIe device, by sending a PCIe thermal reset signal to the corresponding port, the device state is forcibly reset to ensure that the residual configuration does not interfere with the operation of other devices, and at the same time, the complete configuration of the legal devices is retained to avoid system restart.
[0039] Optionally, in some embodiments, after sending a thermal reset signal to the ports corresponding to at least one first PCIe device to reset at least one first PCIe device, it includes: re-enumerating at least one first PCIe device, and detecting whether the current prefetchable memory value of at least one first PCIe device is less than a preset value; if the current prefetchable memory value of at least one first PCIe device is greater than the preset value, continue to disable the ports corresponding to at least one first PCIe device.
[0040] It should be understood that after a thermal reset, re-enumerate at least one first PCIe device that has been disabled, and detect the current prefetchable memory value of at least one first PCIe device again.
[0041] If it is detected that the current prefetchable memory value of at least one first PCIe device is still ≥ 4GB, continue to disable its port, block its occupation of system resources, and notify the operation and maintenance personnel by sending a reminder message that at least one first PCIe device is still unavailable after a thermal reset.
[0042] If it is detected that the current prefetchable memory value of at least one first PCIe device is < 4GB, re-enable the device.
[0043] Through the above technical solution, after a thermal reset, the host system re-enumerates the devices and reads the BAR registers of at least one first PCIe device to achieve secondary verification. If the device still does not meet the conditions, directly disable its port to avoid its occupation of system resources and avoid the recurrence of conflicts caused by the device restoring its original configuration after a thermal reset.
[0044] Optionally, in some embodiments, when obtaining the first hardware signal of at least one second PCIe device with a current prefetchable memory value less than a preset value among multiple PCIe devices based on the baseboard management controller, it includes: obtaining a first verification signal and a second verification signal respectively based on the first presence pin and the second presence pin of at least one second PCIe device; if the first verification signal and the second verification signal of at least one second PCIe device are consistent, generate the first hardware signal corresponding to the second PCIe device based on the first verification signal and the second verification signal of at least one second PCIe device.
[0045] It can be understood that at least one second PCIe device has presence pins, including a first presence pin (Present Pin 1) and a second presence pin (Present Pin 2). The baseboard management controller samples the first presence pin and the second presence pin multiple times within a fixed time window, respectively obtains a first verification signal and a second verification signal, and determines a first hardware signal based on the obtained first verification signal and second verification signal through majority voting.
[0046] Through the above technical solution, by collecting dual-channel signals (the first verification signal and the second verification signal) multiple times, the error caused by collecting single-channel signals is avoided. Through the majority voting mechanism, it is ensured that the generated hardware signal is consistent with the actual state of the device, avoiding resource allocation conflicts or device function abnormalities caused by signal errors.
[0047] Optionally, in some embodiments, after generating a first hardware signal corresponding to at least one second PCIe device based on the first verification signal and the second verification signal of the at least one second PCIe device, it includes: obtaining the first hardware signal generated by the at least one second PCIe device through the input / output pins of an input / output expander (i.e., an IO expander) using a preset interface of the baseboard management controller.
[0048] Specifically, as Figure 2 shown, the presence pin Present PIN of the PCIe device is connected to the input / output pin GPIO_IN (PIN X) of the IO expander (9555 chip). After power-on, the baseboard management controller BMC periodically reads the status of 9555 GPIO through I2C (Inter-Integrated Circuit) / SPI (Serial Peripheral Interface) (which is the preset interface), so as to obtain the first hardware signal generated by the PCIe device.
[0049] In the embodiment of the present invention, a dual-channel signal acquisition + cross-verification design is adopted, and it is connected to the input / output pin GPIO PIN of 9555. Specifically: The first presence pin Present Pin 1 of the PCIe device is connected to the input / output pin 9555_GPIO12 to collect the first verification signal as the main path, and the second presence pin Present Pin 2 is connected to the input / output pin 9555_GPIO13 to collect the second verification signal as the verification path.
[0050] The embodiments of the present invention also perform anti-jitter processing: a filtering circuit / software debounce logic can be added on the 9555 or BMC side. Moreover, the embodiments of the present invention have high scalability. The 9555 supports multiple GPIOs and is suitable for monitoring multiple PCIe slots.
[0051] Through the above technical solution, a reliable hardware signal is generated through dual-channel signal cross-checking, avoiding misjudgment of the device, ensuring signal stability. Through anti-jitter processing, the frequent alarms or resource reallocations of the BMC caused by signal jitter are reduced, improving the overall reliability of the system.
[0052] Optionally, in some embodiments, after obtaining a first check signal and a second check signal based on a first presence pin and a second presence pin of at least one second PCIe device respectively, it further includes: at least one check-failed PCIe device based on the inconsistent first check signal and second check signal of at least one second PCIe device; recording the signal conflict timestamp, device information, signal sampling data, and signal check result of at least one check-failed second PCIe device, and sending an alarm message to a preset terminal.
[0053] It can be understood that if there is at least one check-failed PCIe device in at least one second PCIe device where the first check signal obtained through the first presence pin is inconsistent with the second check signal obtained through the second presence pin, then record the error log of at least one check-failed PCIe device and send an alarm message to a preset terminal. Specifically, recording the error log of at least one check-failed PCIe device includes: conflict timestamp, PCIe device information involved (such as port number, device ID), signal sampling data, and check result.
[0054] Through the above technical solution, by recording the sampling signal differences of the two pins, the operation and maintenance personnel do not need to check the devices one by one, and can directly locate the faulty device through the device ID and port number in the log, reducing the fault location time. Through the real-time push of the preset terminal, the operation and maintenance personnel can quickly receive an alarm after the signal conflict occurs.
[0055] Step S103, determine at least one target PCIe device where the first hardware signal and the second hardware signal of at least one second PCIe device are inconsistent, record the exception log of at least one target PCIe device, and trigger an exception alarm.
[0056] Specifically, as Figure 3 shown, compare the first hardware signal of at least one second PCIe device with the second hardware signal by which the host system recognizes at least one second PCIe device, and detect whether there is a device with a signal conflict in at least one second PCIe device.
[0057] If at least one target PCIe device with inconsistent first and second hardware signals is detected among at least one second PCIe device, an exception log of at least one target PCIe device is recorded, including IDL logs (Inventory and Diagnostic Log) and Switch error logs, to enable root cause tracing of faults, and an exception alarm is triggered through BMC. The exception log includes: conflict timestamp, PCIe device information involved (such as port number, device ID), signal sampling data. The exception alarm includes lighting a fault indicator and sending an alarm message to a preset terminal (i.e., Figure 3 the management platform in
[0058] It should be noted that the exception alarm can not only be triggered by lighting a fault indicator, but also by pushing emails / sms alarms to the mobile phones or emails of operation and maintenance personnel, etc.
[0059] As an embodiment of the present invention, specifically: at least one second PCIe device is monitored in real time for the presence of first and second hardware signals, and an alarm is triggered when the first and second hardware signals are detected to be inconsistent.
[0060] The logic for the host system to detect the second hardware signal is: the host system enumerates at least one second PCIe device through the PCIe bus. If at least one second PCIe device responds, the output second hardware signal is an "in-position signal", and the "in-position signal" indicates that the PCIe device is in the in-position state. Otherwise, the output second hardware signal is an "out-of-position signal", and the "out-of-position signal" indicates that the PCIe device is in the out-of-position state.
[0061] The first hardware signal is used to store the hardware signal status of the corresponding ports of at least one second PCIe device detected, that is, BMC periodically reads the 9555 GPIO status to obtain the status of the Present PIN of at least one second PCIe device; The second hardware signal refers to the status of at least one second PCIe device reported by the host system. In this part, BMC passes at least one second PCIe device obtained after PCIe enumeration is completed by BIOS to the BMC system, and BMC uses an updated method to update the status of at least one second PCIe device with the information obtained regularly.
[0062] If the first hardware signal and the second hardware signal are inconsistent, it indicates that there may be a "phantom device" problem. A "phantom device" refers to at least one target PCIe device, and there may be two situations: the first hardware signal obtained by the BMC shows that the device is in the present state, but the host system does not recognize it; or the host system recognizes the device, but the first hardware signal obtained by the BMC shows that the device is in the absent state. Both of these situations belong to conflicts.
[0063] If a conflict is detected, trigger a BMC alarm and record detailed logs, including: conflict timestamp, PCIe device information involved (such as port number, device ID), signal sampling data, and verification results.
[0064] Optionally, in some embodiments, after identifying the second hardware signal of at least one second PCIe device based on the host system, it further includes: based on the first hardware signal and the second hardware signal of at least one second PCIe device, determining whether there is a second PCIe device in at least one second PCIe device where the first hardware signal and the second hardware signal are consistent; if there is a second PCIe device in at least one second PCIe device where the first hardware signal and the second hardware signal are consistent, it is determined that the second PCIe device in at least one second PCIe device where the first hardware signal and the second hardware signal are consistent is in a normal state.
[0065] It should be understood that when there is a second PCIe device in at least one second PCIe device where the first hardware signal and the second hardware signal are consistent, it indicates that the second PCIe device is in a normal state, and there is no need to perform an abnormal alarm action, and regular monitoring is sufficient.
[0066] Through the above technical solution, there is no need to perform an abnormal alarm action on normal PCIe devices, and regular monitoring is carried out to reduce resource occupancy.
[0067] Optionally, in some embodiments, after determining whether the current prefetch memory values of multiple PCIe devices are all greater than or equal to a preset value, or whether the current prefetch memory values of multiple PCIe devices are all less than or equal to a preset value, it includes: if the current prefetch memory values of multiple PCIe devices are all greater than or equal to the preset value, or the current prefetch memory values of multiple PCIe devices are all less than or equal to the preset value, then do not disable the ports corresponding to the multiple PCIe devices, and obtain the third hardware signal of the multiple PCIe devices based on the baseboard management controller, and identify the fourth hardware signal of the multiple PCIe devices based on the host system; determine at least one abnormal PCIe device where the third hardware signal and the fourth hardware signal of the multiple PCIe devices are inconsistent, and record the abnormal logs of the at least one abnormal PCIe device, and trigger an abnormal alarm.
[0068] It can be understood that if the current prefetchable memory values of multiple PCIe devices are all greater than or equal to 4G, or the current prefetchable memory values of multiple PCIe devices are all less than or equal to 4G, the ports corresponding to the multiple PCIe devices are not disabled, and the third hardware signals of the multiple PCIe devices are obtained through the baseboard management controller, the fourth hardware signals of the multiple PCIe devices are identified through the host system, and corresponding judgments are made based on the third hardware signals and the fourth hardware signals of the multiple PCIe devices to determine whether they are consistent.
[0069] It can be understood that the third hardware signals of multiple PCIe devices are obtained through the baseboard management controller (BMC), and the fourth hardware signals of multiple PCIe devices are identified through the host system, so as to determine at least one abnormal PCIe device whose third hardware signal and fourth hardware signal are inconsistent, and record the exception logs of the at least one abnormal PCIe device, including IDL logs (Inventory and Diagnostic Log) and Switch error logs, to achieve root cause tracing of faults, and trigger an exception alarm through the baseboard management controller. The exception logs include: conflict timestamps, PCIe device information involved (such as port numbers, device IDs), signal sampling data, and verification results. The exception alarm includes lighting a fault indicator light and sending an alarm message to a preset terminal.
[0070] Through the above technical solution, by judging the prefetchable memory values of multiple PCIe devices, the ports are enabled only when the memory values of the multiple PCIe devices are consistent, avoiding the degradation of system performance caused by resource conflicts. The dual-signal verification filters out false alarms caused by abnormal single signal channels, improves the accuracy of fault detection, records the dual-hardware signal data, timestamps, and context information of abnormal devices, and supports maintenance personnel to quickly trace the root cause of faults.
[0071] Optionally, in some embodiments, after determining that a second PCIe device in which the first hardware signal and the second hardware signal are consistent among at least one second PCIe device is in a normal state, it includes: using a hot-swap controller to poll and detect whether there is a newly inserted PCIe device; if there is a newly inserted PCIe device, read the current prefetchable memory value of the newly inserted PCIe device; when the current prefetchable memory value of the newly inserted PCIe device is greater than a preset value, disable the port corresponding to the newly inserted PCIe device, and record the port disable log of the newly inserted PCIe device.
[0072] Among them, in the embodiment of the present invention, a pull-up resistor is used to ensure the default high level. If the PCIe device is not inserted, the presence pin (Present Pin) of the corresponding PCIe device is at a high level. When the PCIe device is inserted, the corresponding presence pin (Present Pin) is at a low level. Thus, it can be determined whether there is a newly inserted PCIe device.
[0073] It should be understood that for the hot-plug situation of PCIe devices, in the embodiment of the present invention, the ACPI (Advanced Configuration and Power Interface) Hot Plug Controller (HPC) driver is modified. The hot-plug controller or its driver program will poll the status of the Switch-side chip every 500 milliseconds to detect whether there is a newly inserted PCIe device or a removed PCIe device on the Switch-side chip. Through the polling mechanism, it is ensured that the system can respond to hot-plug events in a timely manner.
[0074] When a hot-plug device on the downstream port of the PCIe Switch is inserted, it is determined that there is a newly inserted PCIe device. Before the newly inserted PCIe device is enumerated, a pre-check process is also performed. Specifically: Judge whether the current prefetchable memory value of the newly inserted PCIe device is consistent with the current prefetchable memory values of at least one second PCIe device on the existing Switch downstream port, that is, judge whether the current prefetchable memory value of the newly inserted PCIe device is less than or equal to 4G. If the current prefetchable memory value of the newly inserted PCIe device is greater than 4G, the port of the newly inserted PCIe device needs to be disabled. At the same time, the Switch chip records relevant information of the disabled port of the newly inserted PCIe device, such as port PORT ID, pre-check result, disable time and other information. If the current prefetchable memory value of the newly inserted PCIe device is less than or equal to 4G, the port corresponding to the newly inserted PCIe device is not disabled.
[0075] Through the above technical solution, the newly inserted PCIe device is intercepted before enumeration, avoiding system crashes caused by resource allocation failures. Through the closed-loop of logs and alarms, operation and maintenance personnel can quickly locate the root cause of hot-plug problems and shorten the fault recovery time.
[0076] Optionally, in some embodiments, when the current prefetchable memory value of the newly inserted PCIe device is less than a preset value, the following steps are included: obtaining a fifth hardware signal of the newly inserted PCIe device based on the baseboard management controller, identifying a sixth hardware signal of the newly inserted PCIe device based on the host system, and determining whether the fifth hardware signal of the newly inserted PCIe device is consistent with the sixth hardware signal; if the fifth hardware signal is inconsistent with the sixth hardware signal, it is determined that the newly inserted PCIe device is in an abnormal state, an abnormal log of the newly inserted PCIe device is recorded, and an abnormal alarm is triggered.
[0077] Specifically, when the current prefetchable memory value of the newly inserted PCIe device is less than 4G, obtain the fifth hardware signal of the newly inserted PCIe device through the baseboard management controller, and identify the sixth hardware signal of the newly inserted PCIe device based on the host system. If the fifth hardware signal and the sixth hardware signal of the newly inserted PCIe device are inconsistent, it is determined that the newly inserted PCIe device is in an abnormal state, indicating that the newly inserted PCIe device is a "phantom device".
[0078] At this time, the Switch chip records the abnormal log of the newly inserted PCIe device and triggers an abnormal alarm. The abnormal log includes data such as conflict timestamps, PCIe device information involved (such as port numbers, device IDs), signal sampling data, and verification results, and sends an alarm message to a preset terminal, and lights up a fault indicator to remind the operation and maintenance personnel that the newly inserted PCIe device is abnormal.
[0079] Through the above technical solution, false alarms caused by abnormal single signal channels are filtered through dual-signal cross-verification, the alarm accuracy is improved, and the fault location time is reduced through log recording and alarming.
[0080] To enable those skilled in the art to further understand the fault alarm method of the PCIe device in the embodiments of the present invention, the following is elaborated in detail with specific embodiments.
[0081] (1) The embodiments of the present invention set BIOS (BIOS (Basic Input / Output System, basic input and output system) / UEFI and UEFI (Unified Extensible Firmware Interface, unified extensible firmware interface) pre-check: Specifically: The Switch firmware adds a pre-check design. That is, during the device enumeration phase, the BAR information of the ports of each PCIe device is read to determine the current prefetchable memory value of the PCIe device. If the current prefetchable memory value of some PCIe devices exceeds 4G and that of some PCIe devices is less than 4G, the ports of the PCIe devices with a prefetchable memory value exceeding 4G are disabled; otherwise, no processing is required. In this way, when there are both PCIe devices with more than 4G and less than 4G in the PCIe devices, the Switch chip preferentially retains the PCIe devices with a current prefetchable memory value less than 4G to ensure reasonable resource allocation.
[0082] (2)Enhanced Hot Plugging Driver When a newly inserted PCIe device exists at the downstream port of the PCIe Switch, before enumerating the newly inserted PCIe device, a pre-check process is also performed. When the current prefetchable memory value of the newly inserted PCIe device is the same as that of the existing PCIe devices at the downstream port of the Switch, for example, both are above 4G or below 4G, no processing is done. When the current prefetchable memory value of the newly inserted PCIe device is inconsistent with that of the existing PCIe devices at the downstream port of the Switch, that is, when there are both PCIe devices with more than 4G and less than 4G at the downstream port of the Switch, the PCIe port of the newly inserted PCIe device needs to be disabled, and at the same time, the Switch chip records the relevant information of the port of the newly inserted PCIe device, such as the port PORT ID, pre-check result, disable time, etc.
[0083] (3)The Switch adds a port fusing mechanism, adding a gating circuit at the physical layer of the Switch to support dynamic port disabling: When all PCIe devices connected to the Switch chip include both PCIe devices with a prefetchable memory value above 4G and below 4G, the ports of all PCIe devices with a prefetchable memory value above 4G are immediately disabled, and a PCIe hot reset signal is sent to the ports of the PCIe devices with a prefetchable memory value above 4G, while retaining the complete configuration of the ports of the PCIe devices with a prefetchable memory value below 4G.
[0084] (4) Connect the presence pin of the PCIe device to the GPIO_IN PIN of the IO expander 9555. After the startup is completed, the BMC periodically reads the status of the 9555 GPIO through the I2C interface, obtains the first hardware signal of the PCIe device, and compares it with the second hardware signal of the PCIe device recognized under the host system. When the first hardware signal of the PCIe device is consistent with the second hardware signal, it is determined that the PCIe device is in a normal state, and no action is required, and only regular monitoring is needed. When the first hardware signal of the PCIe device is inconsistent with the second hardware signal, the BMC triggers an alarm, records the log of the abnormal PCIe device, lights up the fault indicator, and sends an alarm message to the preset terminal to remind the operation and maintenance personnel.
[0085] In summary, the technical effects brought by the embodiments of the present invention are as follows: (1) Through the 4GB memory threshold pre-check mechanism, potential resource conflicts are intercepted during the device initialization phase, avoiding system crashes caused by improper allocation of prefetchable memory, and fundamentally improving the reliability of the server.
[0086] When the memory resources requested by the PCIe device exceed the threshold, the system automatically masks the excess part and disables the corresponding Switch port to ensure the normal startup of the server and solve the problem of startup failure caused by resource conflicts.
[0087] (2) By obtaining the physical signal of the PCIe device through the BMC and comparing it with the logical state after the host system starts up, "phantom devices" are accurately identified, greatly enhancing the reliability of the system.
[0088] (3) Combining the BMC alarm, IDL log (device initialization log) and Switch error log constitutes a three-level log system, quickly locking the root cause of the problem, shortening the average fault diagnosis time by more than 80%, greatly improving the fault location efficiency, saving the labor cost of reproducing the problem, and improving the overall delivery efficiency and quality. Through multi-channel alarm push such as BMC alarm, email notification, and indicator light, the automation and standardization of fault handling are realized.
[0089] (4) After the hot reset of the PCIe device, the host system re-enumerates the device and reads the BAR register of the PCIe device to achieve secondary verification. If the device still does not meet the conditions, its port is directly disabled to avoid it occupying system resources and avoid the recurrence of conflicts caused by the device restoring the original configuration after the hot reset.
[0090] (5) Generate reliable hardware signals through dual-channel signal cross-checking to avoid misjudgment of the device, ensure signal stability, and reduce frequent BMC alarms or resource reallocations caused by signal jitter through anti-jitter processing, improving the overall reliability of the system.
[0091] (6) By collecting the dual-channel signals (the first verification signal and the second verification signal) multiple times, the error caused by collecting a single-channel signal is avoided, and through the majority voting mechanism, it is ensured that the generated hardware signal is consistent with the actual state of the device, avoiding resource allocation conflicts or device function abnormalities caused by signal errors.
[0092] According to the fault warning method of the PCIe device proposed by the embodiment of the present invention, if the current prefetch memory values of multiple PCIe devices are not all greater than or equal to the preset value, and the current prefetch memory values of multiple PCIe devices are not all less than or equal to the preset value, at least one first PCIe device among the multiple PCIe devices whose current prefetch memory value is greater than or equal to the preset value is disabled, and based on the baseboard management controller, a first hardware signal of at least one second PCIe device among the multiple PCIe devices whose current prefetch memory value is less than the preset value is obtained, and a second hardware signal of at least one second PCIe device is identified based on the host system; at least one target PCIe device whose first hardware signal and second hardware signal are inconsistent is determined, and an exception log of at least one target PCIe device is recorded, and an exception warning is triggered. Thereby, when the prefetch memory configuration space of the PCIe devices under the Switch chip has an integration exception, the problem of no log and warning and the server crashing at the startup logo interface, which is not conducive to quickly locating the problem, is solved. It not only significantly improves the stability and reliability of the server system, but also optimizes the fault warning and log recording mechanism, reduces the operation and maintenance cost, and improves the overall operation and maintenance efficiency.
[0093] Next, a fault warning system for a PCIe device according to an embodiment of the present invention is described with reference to the accompanying drawings.
[0094] Figure 4 It is a schematic diagram of a fault warning system for a PCIe device according to an embodiment of the present invention.
[0095] As Figure 4 shown, the fault warning system 10 of the PCIe device includes: a judgment module 100, an acquisition module 200, and an alarm module 300.
[0096] Among them, the judgment module 100 is used to obtain the current prefetch memory values of multiple PCIe devices, and judge whether the current prefetch memory values of multiple PCIe devices are all greater than or equal to the preset value, or whether the current prefetch memory values of multiple PCIe devices are all less than or equal to the preset value; An acquisition module 200, configured to, if the current prefetchable memory values of multiple PCIe devices are not all greater than or equal to a preset value and are not all less than or equal to the preset value, disable at least one first PCIe device among the multiple PCIe devices whose current prefetchable memory values are greater than or equal to the preset value, and obtain a first hardware signal of at least one second PCIe device among the multiple PCIe devices whose current prefetchable memory values are less than the preset value based on a baseboard management controller, and identify a second hardware signal of at least one second PCIe device based on a host system; An alarm module 300, configured to determine at least one target PCIe device for which the first hardware signal and the second hardware signal of at least one second PCIe device are inconsistent, record an exception log of at least one target PCIe device, and trigger an exception alarm.
[0097] Optionally, in some embodiments, after determining whether the current prefetchable memory values of multiple PCIe devices are all greater than or equal to a preset value, or whether the current prefetchable memory values of multiple PCIe devices are all less than or equal to the preset value, the determination module 100 is further configured to: if the current prefetchable memory values of multiple PCIe devices are all greater than or equal to the preset value, or the current prefetchable memory values of multiple PCIe devices are all less than or equal to the preset value, do not disable the ports corresponding to the multiple PCIe devices, and obtain a third hardware signal of the multiple PCIe devices based on a baseboard management controller, and identify a fourth hardware signal of the multiple PCIe devices based on a host system; determine at least one abnormal PCIe device for which the third hardware signal and the fourth hardware signal of the multiple PCIe devices are inconsistent, record an exception log of at least one abnormal PCIe device, and trigger an exception alarm.
[0098] Optionally, in some embodiments, the determination module 100 is further configured to: determine whether a preset flag bit of a base address register of multiple PCIe devices is a prefetchable memory flag bit; if the preset flag bit is a prefetchable memory flag bit, obtain the current prefetchable memory values of the multiple PCIe devices according to a received prefetchable memory request.
[0099] Optionally, in some embodiments, the acquisition module 200 is further configured to: respectively obtain a first check signal and a second check signal based on a first presence pin and a second presence pin of at least one second PCIe device; if the first check signal and the second check signal of at least one second PCIe device are consistent, generate a first hardware signal corresponding to the at least one second PCIe device based on the first check signal and the second check signal of the at least one second PCIe device.
[0100] Optionally, in some embodiments, after generating a first hardware signal corresponding to at least one second PCIe device based on a first check signal and a second check signal of the at least one second PCIe device, the obtaining module 200 is further configured to: obtain the first hardware signal generated by the at least one second PCIe device through input / output pins of an IO expander by using a preset interface of a baseboard management controller.
[0101] Optionally, in some embodiments, after respectively obtaining a first check signal and a second check signal based on a first presence pin and a second presence pin of at least one second PCIe device, the obtaining module 200 is further configured to: determine at least one check-failed PCIe device based on the first check signal and the second check signal of the at least one second PCIe device being inconsistent; record a signal conflict timestamp, device information, signal sampling data, and a signal check result of the at least one check-failed second PCIe device, and send an alarm message to a preset terminal.
[0102] Optionally, in some embodiments, after disabling at least one first PCIe device among a plurality of PCIe devices whose current prefetch memory value is greater than or equal to a preset value, the obtaining module 200 is further configured to: send a thermal reset signal to a port corresponding to the at least one first PCIe device to reset the at least one first PCIe device.
[0103] Optionally, in some embodiments, after sending a thermal reset signal to a port corresponding to the at least one first PCIe device to reset the at least one first PCIe device, the obtaining module 200 is further configured to: re-enumerate the at least one first PCIe device, and detect whether the current prefetch memory value of the at least one first PCIe device is less than the preset value; if the current prefetch memory value of the at least one first PCIe device is greater than the preset value, continue to disable the port corresponding to the at least one first PCIe device.
[0104] Optionally, in some embodiments, after identifying a second hardware signal of at least one second PCIe device based on a host system, the obtaining module 200 is further configured to: determine whether there is a second PCIe device in the at least one second PCIe device where the first hardware signal and the second hardware signal are consistent based on the first hardware signal and the second hardware signal of the at least one second PCIe device; if there is a second PCIe device in the at least one second PCIe device where the first hardware signal and the second hardware signal are consistent, determine that the second PCIe device in the at least one second PCIe device where the first hardware signal and the second hardware signal are consistent is in a normal state.
[0105] Optionally, in some embodiments, after determining that a second PCIe device in which the first hardware signal and the second hardware signal are consistent among at least one second PCIe device is in a normal state, the obtaining module 200 is further configured to: poll and detect whether there is a newly inserted PCIe device by using a hot plug controller; if there is a newly inserted PCIe device, read the current prefetch memory value of the newly inserted PCIe device; when the current prefetch memory value of the newly inserted PCIe device is greater than a preset value, disable the port corresponding to the newly inserted PCIe device, and record the port disable log of the newly inserted PCIe device.
[0106] Optionally, in some embodiments, when the current prefetch memory value of the newly inserted PCIe device is less than the preset value, the obtaining module 200 is further configured to: obtain a fifth hardware signal of the newly inserted PCIe device based on a baseboard management controller, and identify a sixth hardware signal of the newly inserted PCIe device based on a host system, and determine whether the fifth hardware signal of the newly inserted PCIe device is consistent with the sixth hardware signal; if the fifth hardware signal is not consistent with the sixth hardware signal, determine that the newly inserted PCIe device is in an abnormal state, record the abnormal log of the newly inserted PCIe device, and trigger an abnormal alarm.
[0107] It should be noted that for the description of the features in the corresponding embodiments of the PCIe device failure alarm system, reference can be made to the relevant descriptions in the corresponding embodiments of the above-mentioned PCIe device failure alarm method, which will not be elaborated here one by one.
[0108] Figure 5 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device may include: A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0109] When the processor 502 executes the program, it implements the PCIe device failure alarm method provided in the above embodiments.
[0110] Further, the electronic device further includes: A communication interface 503, configured for communication between the memory 501 and the processor 502.
[0111] The memory 501 is used to store a computer program executable on the processor 502.
[0112] The memory 501 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0113] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 only a thick line is used in Figure 5 , but it does not mean that there is only one bus or one type of bus.
[0114] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.
[0115] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0116] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the fault warning method for a PCIe device when running.
[0117] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0118] Embodiments of the present invention also provide a computer program product including a computer program, and the computer program implements the above fault warning method for a PCIe device when executed by a processor.
[0119] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0120] The above has introduced in detail a method, system, device, medium, and product for fault warning of a PCIe device provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for fault warning of a PCIe device, characterized in that, Including the following steps: Obtain the current prefetchable memory values of multiple PCIe devices, and determine whether the current prefetchable memory values of the multiple PCIe devices are all greater than or equal to a preset value, or whether the current prefetchable memory values of the multiple PCIe devices are all less than or equal to the preset value; If the current prefetchable memory values of the multiple PCIe devices are not all greater than or equal to the preset value, and the current prefetchable memory values of the multiple PCIe devices are not all less than or equal to the preset value, then disable at least one first PCIe device among the multiple PCIe devices whose current prefetchable memory values are greater than or equal to the preset value, and obtain the first hardware signal of at least one second PCIe device among the multiple PCIe devices whose current prefetchable memory values are less than the preset value based on the baseboard management controller, and identify the second hardware signal of the at least one second PCIe device based on the host system; Determine at least one target PCIe device whose first hardware signal and second hardware signal are inconsistent among the at least one second PCIe device, record the exception log of the at least one target PCIe device, and trigger an exception alarm.
2. The method for fault warning of a PCIe device according to claim 1, wherein, After determining whether the current prefetchable memory values of the multiple PCIe devices are all greater than or equal to the preset value, or whether the current prefetchable memory values of the multiple PCIe devices are all less than or equal to the preset value, it includes: If the current prefetchable memory values of the multiple PCIe devices are all greater than or equal to the preset value, or the current prefetchable memory values of the multiple PCIe devices are all less than or equal to the preset value, then do not disable the ports corresponding to the multiple PCIe devices, and obtain the third hardware signal of the multiple PCIe devices based on the baseboard management controller, and identify the fourth hardware signal of the multiple PCIe devices based on the host system; Determine at least one abnormal PCIe device whose third hardware signal and fourth hardware signal are inconsistent among the multiple PCIe devices, record the exception log of the at least one abnormal PCIe device, and trigger an exception alarm.
3. The fault warning method of the PCIe device according to claim 1, characterized in that, The obtaining of the current prefetchable memory values of the multiple PCIe devices includes: Judge whether the preset flag bits of the base address registers of the multiple PCIe devices are prefetchable memory flag bits; If the preset flag bits are all the prefetchable memory flag bits, then obtain the current prefetchable memory values of the multiple PCIe devices according to the received prefetchable memory request.
4. The method for fault warning of the PCIe device according to claim 1, wherein When obtaining the first hardware signal of at least one second PCIe device among the multiple PCIe devices whose current prefetchable memory values are less than the preset value based on the baseboard management controller, it includes: Obtain a first check signal and a second check signal respectively based on the first presence pin and the second presence pin of the at least one second PCIe device; If the first check signal and the second check signal of the at least one second PCIe device are consistent, then generate the first hardware signal corresponding to the second PCIe device based on the first check signal and the second check signal of the at least one second PCIe device.
5. The method for fault warning of a PCIe device according to claim 4, wherein, After generating a first hardware signal corresponding to the at least one second PCIe device based on the first check signal and the second check signal of the at least one second PCIe device, it includes: Obtain the first hardware signal generated by the at least one second PCIe device through the input / output pins of the IO expander using a preset interface of the baseboard management controller.
6. The method for fault warning of a PCIe device according to claim 4, wherein, After respectively obtaining a first check signal and a second check signal based on a first presence pin and a second presence pin of the at least one second PCIe device, it further includes: At least one check-failed PCIe device based on the first check signal and the second check signal of the at least one second PCIe device being inconsistent; Record the signal conflict timestamp, device information, signal sampling data, and signal check result of the at least one check-failed second PCIe device, and send an alarm message to a preset terminal.
7. The method for fault warning of the PCIe device according to claim 1, characterized in that, After disabling at least one first PCIe device among the multiple PCIe devices whose current prefetch memory value is greater than or equal to the preset value, it includes: Send a thermal reset signal to the port corresponding to the at least one first PCIe device to reset the at least one first PCIe device.
8. The method for fault warning of a PCIe device according to claim 7, wherein After sending a thermal reset signal to the port corresponding to the at least one first PCIe device to reset the at least one first PCIe device, it includes: Re-enumerate the at least one first PCIe device, and detect whether the current prefetch memory value of the at least one first PCIe device is less than the preset value; If the current prefetch memory value of the at least one first PCIe device is greater than the preset value, continue to disable the port corresponding to the at least one first PCIe device.
9. The method for fault warning of a PCIe device according to claim 1, wherein After identifying a second hardware signal of the at least one second PCIe device based on the host system, it further includes: Based on the first hardware signal and the second hardware signal of the at least one second PCIe device, determine whether there is a second PCIe device among the at least one second PCIe device where the first hardware signal and the second hardware signal are consistent; If there is a second PCIe device among the at least one second PCIe device where the first hardware signal and the second hardware signal are consistent, determine that the second PCIe device among the at least one second PCIe device where the first hardware signal and the second hardware signal are consistent is in a normal state.
10. The method for fault warning of a PCIe device according to claim 9, wherein After determining that the second PCIe device among the at least one second PCIe device where the first hardware signal and the second hardware signal are consistent is in a normal state, it includes: Use the hot-swap controller to poll and detect whether there is a newly inserted PCIe device; If there is the newly inserted PCIe device, read the current prefetch memory value of the newly inserted PCIe device; When the current prefetch memory value of the newly inserted PCIe device is greater than the preset value, disable the port corresponding to the newly inserted PCIe device, and record the port disable log of the newly inserted PCIe device.
11. The method for fault warning of the PCIe device according to claim 10, wherein When the current prefetch memory value of the newly inserted PCIe device is less than the preset value, it includes: The baseboard management controller obtains the fifth hardware signal of the newly inserted PCIe device, and the host system identifies the sixth hardware signal of the newly inserted PCIe device, and determines whether the fifth hardware signal of the newly inserted PCIe device is consistent with the sixth hardware signal; If the fifth hardware signal is inconsistent with the sixth hardware signal, it is determined that the newly inserted PCIe device is in an abnormal state, an exception log of the newly inserted PCIe device is recorded, and an exception alarm is triggered.
12. A fault warning system for a PCIe device, characterized in that, Comprising: A judgment module, configured to obtain the current prefetchable memory values of multiple PCIe devices, and determine whether the current prefetchable memory values of the multiple PCIe devices are all greater than or equal to a preset value, or whether the current prefetchable memory values of the multiple PCIe devices are all less than or equal to the preset value; An obtaining module, configured to, if the current prefetchable memory values of the multiple PCIe devices are not all greater than or equal to the preset value, and the current prefetchable memory values of the multiple PCIe devices are not all less than or equal to the preset value, disable at least one first PCIe device among the multiple PCIe devices whose current prefetchable memory value is greater than or equal to the preset value, and obtain, based on the baseboard management controller, the first hardware signal of at least one second PCIe device among the multiple PCIe devices whose current prefetchable memory value is less than the preset value, and identify the second hardware signal of the at least one second PCIe device based on the host system; An alarm module, configured to determine at least one target PCIe device whose first hardware signal and second hardware signal are inconsistent among the at least one second PCIe device, record an exception log of the at least one target PCIe device, and trigger an exception alarm.
13. An electronic device, characterized in that, Comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executes the program to implement the fault alarm method for a PCIe device according to any one of claims 1-11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the fault alarm method for a PCIe device according to any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fault alarm method for a PCIe device according to any one of claims 1-11.
Citation Information
Patent Citations
Mobile terminal, memory allocation control method and storage medium
CN107656811A
PCIE equipment fault monitoring method and device, communication equipment and storage medium
CN115878430A
Server system and log capturing system of memory expansion card
CN118377738A
Device operation method, system and apparatus, and non-volatile readable storage medium and electronic device
WO2024183334A1
Cited By
Fault analysis method and device for switch chip
CN120474904A