Switch chip fault analysis method and device
By generating error logs and converting error codes into fault description information, and combining them with storage media for associated storage, the problem of low efficiency in switching chip fault analysis is solved, the automation and real-time nature of fault analysis is achieved, and operation and maintenance efficiency and system reliability are improved.
Patent Information
- Application Number
- CN202510957895.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-11
AI Technical Summary
In the existing technology, the fault analysis efficiency of switching chips is low, relying on manual operations and lacking a persistent storage mechanism, resulting in fault diagnosis delays and inefficient maintenance.
By generating error logs and converting error codes into fault description information, and combining them with storage media for associated storage, fault tracing and real-time alarms can be achieved, and automated fault analysis can be performed using preset fault detection mechanisms and error code mapping tables.
It achieves automation and real-time fault analysis, improves operation and maintenance efficiency, and enhances system reliability and fault tracing capabilities.
Smart Images

Figure CN120474904B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for analyzing faults of a switching chip. Background Art
[0002] With the advancement of information technology, data centers are increasingly required to store, compute, and interact with increasing amounts of data. PCIe, as a key hardware bus for data interaction within servers, plays an indispensable role. PCIe utilizes a point-to-point transmission mechanism, transmitting data directly from the CPU root port to the device end. To enhance system flexibility, PCIe switch chips were introduced to expand the number of signals and bandwidth, allowing a single root port to connect to multiple end devices, such as GPUs or storage drives. However, existing technologies present significant limitations when errors occur in the switch chip's extended signal link. Fault analysis relies on manual labor and requires log export using vendor-specific tools. System management modules, such as the BMC, only record basic error codes and cannot provide specific fault location or type analysis. Furthermore, link error detection requires physical access to analysis equipment, which is often impractical due to hardware space limitations. Error logs are easily lost after power failures and lack a persistent storage mechanism. These shortcomings lead to delayed fault diagnosis and inefficient maintenance. Consequently, the related art suffers from the low efficiency of fault analysis in switch chips. Summary of the Invention
[0003] The present application provides a method and apparatus for fault analysis of a switching chip, so as to at least solve the technical problem of low efficiency in fault analysis of a switching chip existing in the related art.
[0004] The present application provides a fault analysis method for a switching chip, comprising: in response to an error event occurring in a target link, generating an error log based on a preset fault detection mechanism, wherein the error log includes an error code and a port number, and the target link represents a link for the switching chip to extend a signal; when the error event is detected by a trigger, converting the error code into fault description information based on a pre-generated error code mapping table, wherein the trigger represents a trigger set by a target protocol analysis tool; and associating the fault description information with the error log and storing it in a storage partition corresponding to the port number in a storage medium.
[0005] The present application also provides a fault analysis device for a switching chip, comprising: a generation module for generating an error log based on a preset fault detection mechanism in response to an error event occurring in a target link, wherein the error log includes an error code and a port number, and the target link represents the link for the switching chip to extend the signal; a conversion module for converting the error code into fault description information based on a pre-generated error code mapping table when the error event is detected by a trigger, wherein the trigger represents a trigger set by a target protocol analysis tool; a storage module for associating the fault description information and the error log and storing them in a storage partition corresponding to the port number in a storage medium.
[0006] In an exemplary embodiment, the device is used to generate an error log based on a preset fault detection mechanism in response to an error event occurring in a target link in the following manner: detecting the connection status of each port in the target link through a general input and output interface; when an abnormality is detected in the connection status, determining that an error event has occurred in the target link, and generating the error log based on a preset fault detection mechanism; and exporting the error log through a serial debugging interface.
[0007] In an exemplary embodiment, the device is also used to: when the error event is detected by a trigger, before converting the error code into fault description information based on a pre-generated error code mapping table, use the target protocol tool to load a pre-configuration file; and set the trigger for each port in the target link based on the pre-configuration file.
[0008] In an exemplary embodiment, the device is used to set the trigger for each port in the target link based on the preconfiguration file in at least one of the following ways: selecting the trigger set for each port in the target link from a group of triggers configured according to preset rules based on the preconfiguration file; setting the trigger for each port in the target link in a customized manner based on the preconfiguration file.
[0009] In an exemplary embodiment, the device is used to convert the error code into fault description information based on a pre-generated error code mapping table when the error event is detected by a trigger in the following manner: when the error event is detected by a trigger, use a target protocol tool to obtain a trigger data packet corresponding to the error event; search the error log based on the trigger data packet to determine the error code; and convert the error code into fault description information based on a pre-generated error code mapping table.
[0010] In an exemplary embodiment, the device is also used to: in response to the switching chip being started, load the error code mapping table from the storage medium, wherein the error code mapping table is used to indicate the mapping relationship between the error code and the fault description information; when the error log is generated based on the preset fault detection mechanism, store the error log in a storage partition corresponding to the port number; when the error code is converted into the fault description information based on the error code mapping table, match the fault description information with the error code, store it in the storage partition corresponding to the port number, and record the timestamp and error level associated with the error log.
[0011] In an exemplary embodiment, the device is further used to: in response to the switching chip starting up, load the error code mapping table from the storage medium, and use the target protocol tool to load the pre-configuration file; set the trigger for each port in the target link based on the pre-configuration file; detect the connection status of each port in the target link through a general input and output interface; when an abnormality is detected in the connection status, determine that an error event has occurred in the target link, and generate the error log based on a preset fault detection mechanism; export the error log through a serial debugging interface and store it in a storage partition corresponding to the port number; when the error event is detected by the trigger, use the target protocol tool to obtain the a trigger data packet corresponding to the error event; searching the error log based on the trigger data packet to determine the error code; converting the error code into fault description information based on a pre-generated error code mapping table, and storing the information in a storage partition corresponding to the port number, while recording a timestamp and an error level associated with the error log; in the case of detecting continuous error events, calling a built-in script to export the register status of the switching chip, and updating the fault description information based on the register status; in response to the fault description information being stored in the storage partition corresponding to the port number, sending a signal to a baseboard management controller to instruct the baseboard management controller to read the fault description information and generate an alarm event based on the fault description information.
[0012] In an exemplary embodiment, the device is also used to: obtain the remote mapping table version through an out-of-band management channel in response to an update instruction issued by the baseboard management controller; compare the difference between the locally stored error code mapping table version number and the remote version number; when it is detected that the version difference exceeds a threshold, mount the non-volatile storage medium to the baseboard management controller; clear the historical mapping table partition through a secure erase operation, and write a new version of the error code mapping table.
[0013] In an exemplary embodiment, when the device is applied to a multi-switch board backplane system, it is used to execute the following when the baseboard management controller detects that the current switch board port error level is serious: broadcasting the port error event to the adjacent switch board through the backplane management bus; receiving associated link quality indicators returned by the adjacent board, the indicators including at least one of the following: signal crosstalk strength, reference clock offset value, power rail ripple coefficient; generating a joint fault analysis report based on multi-board indicators, and updating the fault description information to the cross-board interference type.
[0014] In an exemplary embodiment, the device is used to set triggers in the following manner: parsing the energy efficiency policy field in the pre-configuration file to identify the port activity level tag; configuring a real-time hardware error trigger for a high-activity level port; enabling a periodic polling mode for a low-activity level port, including: activating polling during system idle periods; setting a multi-level cache temporary error event; triggering batch parsing when the cumulative number of error events exceeds a threshold; the micro control unit dynamically switches the working mode according to the port activity status and enters a sleep state when there is no error event.
[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned switching chip fault analysis methods when executing the computer program.
[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned switching chip fault analysis methods are implemented.
[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned switching chip fault analysis methods when executed by a processor.
[0018] The present application also provides a switching chip fault analysis system, including: a switching chip, connected to a microcontroller unit via a universal input / output interface, for transmitting protocol layer error events; a microcontroller unit, integrating a target protocol analysis tool and mounting a non-volatile storage medium, configured to: capture the trigger data packet of the switching chip in real time via the universal input / output interface; convert the original error code into a fault description with a port identifier based on a pre-stored error code mapping table; a baseboard management controller, communicating with the microcontroller unit via an inter-integrated circuit bus, configured to: receive an interrupt signal from the microcontroller unit; dynamically switch control of the non-volatile storage medium to read fault description information; a physical layer monitoring circuit, including a universal input / output interface array, each interface being connected to an independent port connection status detection pin.
[0019] Through the embodiments of the present application, by monitoring error events, an error log containing an error code and port number is instantly generated when an abnormality occurs in the target link; then, through precise fault mapping, based on the error code mapping table pre-stored in non-volatile media, the original error code is dynamically converted into a readable fault description, and the physical location is associated with the data packet captured by the trigger to achieve fault tracing; then, through structured storage management, the fault description and the original log are partitioned and stored by port number to form a traceable historical database; finally, through system collaborative optimization, real-time alarms and in-depth diagnosis are triggered. This achieves the technical effects of a leap in operation and maintenance efficiency, enhanced fault tracing, and enhanced system reliability. Therefore, it can solve the technical problem of low fault analysis efficiency of switching chips in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A system architecture diagram of a switch chip fault analysis method provided in an embodiment of the present application;
[0022] Figure 2 A flowchart of a method for analyzing a fault in a switching chip provided in an embodiment of the present application;
[0023] Figure 3 A schematic diagram of a flow chart of a method for analyzing a fault of a switching chip provided in an embodiment of the present application;
[0024] Figure 4 A schematic diagram of the structure of a switching chip fault analysis device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0027] First, a brief explanation of the technical terms involved in this application:
[0028] BDF: Bus / Device / Function, bus number / device number / function number, used to uniquely identify the location of a PCIe device and assist in locating the source of a fault.
[0029] BMC: Baseboard Management Controller, responsible for overall server monitoring, receiving fault alarms and coordinating responses.
[0030] FW: Firmware, embedded software that controls the underlying operations and error handling logic of the switching chip.
[0031] GPIO: General Purpose Input / Output, general input and output interface, detects the physical connection status of the port (such as device presence or disconnection).
[0032] GPU: Graphics Processing Unit, as a PCIe terminal device, may trigger error events due to signal problems.
[0033] IB card: InfiniBand card, Infini Band card, high-speed network device, connected to the PCIe Switch downstream port.
[0034] MCU: Microcontroller Unit, microcontroller unit chip, core control unit, performs error log generation, parsing and storage.
[0035] NVMe: Non-Volatile Memory Express, a non-volatile memory fast interface, a storage device type, commonly seen in downstream port error detection.
[0036] OTC: On-Chip Trace, on-chip tracing function, refers to the tracing capability built into the PCIe switch chip, which is used to capture real-time data packets.
[0037] PCIe: Peripheral Component Interconnect Express, a fast bus for interconnecting peripheral devices, the core protocol for server data interaction, and the target link is built based on this protocol.
[0038] PCIe Switch: PCIe Switch chip, PCIe switching chip, core component for extending PCIe signal bandwidth (in this application, it can be understood to include but not be limited to Broadcom PCIe5.0 Switch chip).
[0039] SD card: Secure Digital card, secure digital memory card, non-volatile storage medium, stores error logs and mapping tables.
[0040] SDIO: Secure Digital Input Output, secure digital input and output interface, data transmission channel between MCU and SD card.
[0041] SDB: Serial Debug Bus, debug interface, used to export logs (such as connecting to SDBHeader).
[0042] SMBUS: System Management Bus, which transmits sensor data (such as temperature or voltage).
[0043] Switch board: PCIe Switch board, PCIe switch board, an independent board with an integrated PCIe Switch chip.
[0044] UART: Universal Asynchronous Receiver / Transmitter, serial communication interface, used for data transmission between MCU and PCIe Switch.
[0045] AER: Advanced Error Reporting, an error detection mechanism for the PCIe protocol that generates error codes and logs.
[0046] CE: Correctable Error, a transmission error that can be automatically corrected (such as packet retransmission).
[0047] CRC: Cyclic Redundancy Check, a data integrity check method that detects transmission errors.
[0048] DLLP: Data Link Layer Packet, data link layer packet, PCIe protocol control packet, errors may occur at this layer.
[0049] ECRC: End-to-End CRC, end-to-end cyclic redundancy check, end-to-end data integrity check, used for serious error detection.
[0050] ECC: Error Correction Code, error correction code, error correction mechanism in memory or data transmission.
[0051] FLIT: Flow Control Unit, flow control unit, data flow management unit, errors may cause bandwidth to drop.
[0052] I2C: Inter-Integrated Circuit, internal integrated circuit bus, a bidirectional synchronous serial bus used for communication between BMC and MCU.
[0053] IPMI: Intelligent Platform Management Interface, BMC remote management protocol, supports fault alarms.
[0054] TLP: Transaction Layer Packet, transport layer packet, PCIe protocol data packet, errors often involve packet header verification failure.
[0055] UCE: Uncorrectable Error, a serious error type that requires manual intervention (such as hardware failure).
[0056] CLI: Command Line Interface, the basis of PTraceCLI, executes protocol analysis commands.
[0057] ipalstruct.ini: Custom configuration file for the PCIe embedded analyzer, which sets detailed triggers and filter conditions and supports in-depth error analysis.
[0058] PEA: PCIe Embedded Analyzer, PCIe Embedded Analyzer, a built-in tool of PTraceCLI, used to capture data packets in real time.
[0059] precanned.ini: Predefined trigger configuration file that simplifies the setup of common error detection (such as hardware errors or packet loss).
[0060] PTraceCLI: PCIe trace command-line tool, integrated in the MCU, used to capture and analyze PCIe link error data.
[0061] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0062] like Figure 1 As shown, in combination with the specific application environment architecture or specific hardware architecture on which the execution of the fault analysis method of the switching chip depends, the specific application environment architecture or specific hardware architecture is described here.
[0063] S1: Error detection and log generation:
[0064] When an error event (such as a protocol layer violation or physical layer signal integrity failure) occurs on the target link (i.e., the physical link extended by a Broadcom Expander) on a PCIe (Peripheral Component Interconnect Express) switch board connected to a terminal or server, the switch board's MCU (Microcontroller Unit) instantly captures the event based on a pre-defined hardware fault detection mechanism (the Advanced Error Reporting (AER) feature integrated into the expander chip and the CLI tool). The MCU then generates an error log containing the original error code (derived from the protocol layer status register), the triggering port number, and the BDF (Bus / Device / Function) number of the device involved. Real-time error detection and information capture at the hardware level are crucial at this stage.
[0065] S2: Error code analysis and information conversion:
[0066] Once activated, the MCU's pre-built logic monitoring unit acts as a trigger (a hardware event configured by a target protocol analysis tool like PTraceCLI, such as the detection of a CE (correctable error) or UCE (uncorrectable error) interrupt) to initiate the analysis process. The MCU reads the error code mapping table pre-programmed on the SD card (which maps error codes to human-readable fault descriptions) and automatically converts the raw error code captured in step S1 into a specific fault description (for example, "Port 0x3 link timeout" or "Device configuration space exception"). This step automatically translates protocol-layer error information into operational semantics.
[0067] S3: Association information storage and partition management:
[0068] After parsing is complete, the MCU dynamically correlates and integrates the generated fault description information with the original error log (including message-level tracing data captured by the PTraceCLI tool). Subsequently, based on the port number contained in the log, the MCU stores the correlated complete diagnostic record in a specific pre-partitioned storage partition on the local SD card media via the SDIO (Secure Digital Input Output) interface (the partitioning strategy strictly corresponds to the physical port or slot number, such as Slot1 corresponding to Partition / P1). At the same time, the onboard I2C (Inter-Integrated Circuit) bus notifies the upstream BMC (Baseboard Management Controller) of critical alarm events and storage paths, enabling remote access to the log. This process ensures that fault information is accurately classified and stored by port and has the ability to persist even after power failure.
[0069] Based on a hardware-assisted fault diagnosis architecture, this framework enables closed-loop management of PCIe link failures in data center-level terminals or servers throughout their entire lifecycle. Through deep collaboration between the MCU and expander chips, millisecond-level lossless acquisition and analysis of abnormal events are automated, fundamentally eliminating the response delays associated with traditional manual export and debugging. A dual-path redundant design ensures complete logging of fault points (the SD card maintains independent partitions to store raw logs, while the BMC synchronizes key events). A strong mapping between physical ports and storage intervals ensures data isolation and retrieval efficiency under concurrent multi-device failures. Overall system resource optimization is significant. The MCU maintains a low-power polling mode in steady state, activating real-time processing only when high-risk error events are triggered, effectively balancing diagnostic accuracy with platform energy consumption. Ultimately, adaptive fault recovery and proactive early warning capabilities are established.
[0070] The embodiment of the present application provides a method for analyzing a fault of a switching chip. The method is described in detail in conjunction with the execution flow of the method for analyzing a fault of a switching chip. Figure 2 As shown, including but not limited to the following steps:
[0071] S202, in response to an error event occurring on a target link, generating an error log based on a preset fault detection mechanism, wherein the error log includes an error code and a port number, and the target link represents a link on which the switch chip extends a signal;
[0072] S204, when an error event is detected by a trigger, converting the error code into fault description information based on a pre-generated error code mapping table, wherein the trigger represents a trigger set by a target protocol analysis tool;
[0073] S206: Associate the fault description information and the error log and store them in a storage partition corresponding to the port number in the storage medium.
[0074] Optionally, in the embodiment of the present application, combined with Figure 1 The hardware architecture shown above, the target link can be understood as Figure 1 The PCIe physical connection channel (marked with a solid arrow in the figure) between the Broadcom Expander chip and downstream devices (such as GPU cards, NIC cards, and NVMe drives) includes, but is not limited to, the five PCIe interfaces extending from the right side of the Expander chip. When a GPU card fails, the target link transmits a TLP message containing an error status bit, achieving hardware-level error propagation and enabling the switch board to detect physical layer anomalies.
[0075] Optionally, in the embodiment of the present application, combined with Figure 1 The hardware architecture shown above, the error log can be understood as structured diagnostic data generated by the Broadcom Expander chip and transmitted to the MCU via the I2C bus ( Figure 1 The I2C connection between the Expander and the MCU carries this data stream. Specifically, when a DLLP (Data Link Layer Packet) sequence number error occurs on the NIC card, the Expander chip status register generates an error code of 0x8012. This, combined with the port number (corresponding to physical slot P3) and timestamp, forms a log, achieving a digital snapshot of the fault and providing original evidence for protocol analysis.
[0076] Optionally, in the embodiment of the present application, combined with Figure 1 The hardware architecture shown above can be understood as the PTraceCLI protocol analysis engine running inside the MCU ( Figure 1 The built-in logic unit of the MCU module in the NVMe storage system. Specifically, this includes but is not limited to hardware error monitoring strategies configured through pre-burned scripts (such as setting the CE / UCE interrupt enable bit of the PCIe AER register). When a FLIT (Flow Control Unit) check fails in NVMe disk transmission, a trigger activates the MCU's log collection thread, enabling zero-latency capture of critical errors.
[0077] Optionally, in the embodiment of the present application, combined with Figure 1 The hardware architecture shown above, the error code mapping table can be understood as a machine-readable fault dictionary stored in the SD card ( Figure 1The MCU is connected to the memory card via the SDIO interface. Specifically, this includes but is not limited to the CSV file stored in the / SYS / ERRMAP directory (for example, mapping the Expander error code 0x8012 to "Port 3 Data Link Layer Serial Number Timeout"). When a power management state conflict occurs on the GPU card, the MCU retrieves the table in real time to generate a description text, making the protocol knowledge explicit.
[0078] Optionally, in the embodiment of the present application, combined with Figure 1 The hardware architecture shown above can be understood as an SD card physical storage array partitioned by port ( Figure 1 The rectangular module marked with the SDIO interface in the figure). Specifically, this includes but is not limited to the pre-partitioned / P1- / P8 directories on the card (corresponding to eight physical ports). When a physical layer electrical error is detected on port 4 of the NIC card, the associated log and description text are written to the / P4 partition, achieving physical isolation of the fault data.
[0079] Optionally, in the embodiment of the present application, the above-mentioned error event can be understood as a hardware identifier of a protocol violation event on the PCIe link ( Figure 1 The error flag in the Expander chip's internal status register is set. Specifically, this includes, but is not limited to, cases where the NVMe drive's clock skew causes continuous ECRC check failures (exceeding the PCIe protocol threshold). In this case, the Expander chip's AER (Advanced Error Reporting) module will mark the error status, enabling automatic calibration of protocol layer anomalies.
[0080] Optionally, in the embodiment of the present application, the above-mentioned response to the error event occurring in the target link can be understood as the deterministic processing flow of the MCU for the hardware interrupt signal ( Figure 1 The MCU receives Expander interrupts via the GPIO_IN pin. Specifically, when the Expander chip pulls up the GPIO_IN:0 level (indicating a UCE uncorrectable error has been detected), the MCU immediately suspends the current task and calls the AER processing firmware stored in the FLASH, achieving a jitter-free effect of hardware and software coordinated response.
[0081] Optionally, in the embodiment of the present application, the above-mentioned generation of error logs can be understood as a multi-source data fusion operation of the MCU ( Figure 1 The MCU collects Expander register data via the I2C bus. Specifically, this includes but is not limited to reading the Expander's error status register (including the BDF device identifier) from the MCU, the 128-byte error message captured by PTraceCLI, and the RTC (real-time clock) timestamp, all combined into a JSON-formatted log file, enabling fault site reconstruction.
[0082] Optionally, in the embodiment of the present application, the error event detected by the trigger can be understood as the physical layer real-time signal analysis of the Expander chip ( Figure 1 Specifically, this includes, but is not limited to, monitoring the voltage swing of the signal sent by the GPU card through the built-in Rx equalizer (for example, detecting a voltage swing below 65mV for 3UI continuously), comparing it to the PCIe base specification threshold, and triggering an error counter, enabling nanosecond-level electrical anomaly detection.
[0083] Optionally, in the embodiment of the present application, the above-mentioned conversion of the error code into the fault description information can be understood as the machine semantic translation performed by the MCU ( Figure 1 The MCU reads the mapping table in the SD card through the SDIO interface. Specifically, including but not limited to, when error code 0x7005 is captured, the MCU retrieves the mapping table to obtain the text description of "Port 5 Receiver Equalization Coefficient Exceeds Limit" and associates it with Section 4.2.6 of the PCIe specification, thereby achieving lossless transmission of error knowledge.
[0084] Optionally, in the embodiment of the present application, the above-mentioned storage of the fault description information and the error log in association can be understood as intelligent archiving based on the physical topology ( Figure 1 SDIO data channel between the MCU and the SD card. Specifically, when parsing out a "link training timeout" fault for the NVMe drive on port 2, the MCU packages the fault description, original log (including PTrace messages), and BMC alarm ID, writes them to the / P2 partition of the SD card, and creates a cross-partition index file, enabling spatial and temporal correlation of fault factors.
[0085] For example, in the attached Figure 1 In the scenario shown, first, when a GPU card's physical layer signal amplitude exceeds the limit due to poor contact of the gold finger, the AER (Advanced Error Reporting) module embedded in the Expander chip will detect the error event in real time and send it to the I2C bus ( Figure 1 The MCU (connecting the expander to the MCU) sends an interrupt signal to the MCU. Upon responding to this interrupt, the MCU immediately reads the expander's error status register, extracts the hexadecimal error code (for example, 0x7001 indicates an abnormal signal amplitude at the receiving end), the associated physical port number (for example, Port3 corresponds to the third PCIe interface), and the occurrence timestamp, and encapsulates it into a JSON-formatted error log. This process relies on a robust fault detection mechanism (including hardware register polling and protocol-layer error counters) to ensure that error-site snapshots are captured within microseconds, providing the data foundation for accurately locating the faulty device slot.
[0086] Second, the trigger is attached Figure 1 It is reflected in the PTraceCLI protocol analysis engine running inside the MCU ( Figure 1 The MCU module has built-in logic units), and the hardware error monitoring strategy is configured through the pre-burned script. For example, when an uncorrectable error (UCE) occurs on port 4 of the NIC card (such as continuous DLLP sequence number timeout), PTraceCLI triggers the "Hardware Error" interrupt condition (preset in the script's CE / UCE interrupt enable bit) through the AER register of the Expander chip, activating the MCU's parsing thread. The MCU then accesses the SD card ( Figure 1 The system uses a pre-stored error code mapping table (CSV file, such as / SYS / ERRMAP / PCIE_MAP.csv) for storage media connected via the SDIO interface in the NVMe storage system. This table converts the original error code 0x8012 into a human-readable semantic description: "Port 4: Data link layer sequence number synchronization failed. Check link negotiation status." This mapping table covers hundreds of error codes defined by the PCIe protocol. For example, if an NVMe drive experiences a link training timeout (error code 0xA005) due to unstable power supply, the resulting output will be "Port 2: TS1 sequence lost during link training. Verify device power supply." This step translates protocol knowledge to achieve the transition from machine code to operational instructions.
[0087] Finally, the storage medium is attached Figure 1 Specifically refers to the SD card (rectangular module marked with SDIO) connected to the MCU through the SDIO interface, and its physical storage space is pre-divided into / P1- / P8 partitions according to the port number. When a transport layer protocol timeout error occurs on the GPU card of port 1, the MCU performs associated storage operations: first, the original error log generated in step 1 (including the 128-byte error message captured by PTraceCLI, the device BDF number and the timestamp) is dynamically bound to the fault description converted in step 2 (such as "Port 1: TLP transport layer ACK / NAK handshake timeout"); then, according to the port number "1", the / P1 storage partition of the SD card is indexed, and the associated data packet is written to the partition through the SDIO interface, and a file named "ERR_PORT1_20240326T142356.json" is generated (including the cross-log index mark). At the same time, the MCU notifies the BMC via the I2C bus ( Figure 1 A new log path has been added to the BMC module connected to the MCU in the system, enabling multi-system collaborative tracing. This design ensures that if multiple devices fail simultaneously (for example, the NVMe drive on port 2 and the NIC card on port 5 report errors simultaneously), the error data is isolated and stored in the / P2 and / P5 partitions, preventing data overwriting and enabling rapid retrieval of historical faults by physical slot.
[0088] It should be noted that the triggering scenarios of target link error events are highly diverse, and this application does not make specific restrictions on this. Specifically: in the fault detection dimension, in addition to detecting basic protocol violations (such as PCIeTLP packet header check errors), the preset fault detection mechanism can also be extended to chip-level health monitoring (such as capturing abnormal power supply voltage fluctuations exceeding ±5% through the Expander built-in sensor), physical connection quality analysis (such as detecting excessive contact resistance of gold fingers based on UART interface loopback testing), and environmental interference event capture (such as receiving external electromagnetic interference alarms through GPIO_IN: 1). In addition to the error code and port number, the error log generated by it can also contain additional diagnostic fields such as ambient temperature and power supply noise spectrum.
[0089] It should be noted that there are multiple implementation forms of trigger configuration strategies, which are not specifically limited in this application. Specifically, the trigger type dimension can include hardware event triggers (such as configuring GPIO_IN:0 to respond to the Expander chip temperature exceeding the threshold interrupt), protocol layer behavior triggers (such as setting it through PTraceCLI to activate when three consecutive DLLP retransmissions are detected), and cross-device association triggers (such as starting log collection based on the GPU card fan failure event reported by the BMC). In addition to the static table fixed on the SD card, the error code mapping table it calls can also support dynamic updates at runtime (such as the MCU downloading the new version of the mapping rules from the server via the UART interface) or a hierarchical mapping mechanism (such as only outputting a simple code for correctable errors and associating uncorrectable errors with a maintenance manual chapter).
[0090] It should be noted that the organization of storage media is strongly related to the actual application scenario, and this application does not make specific restrictions on this. Specifically: in addition to physical partition storage by port number (such as isolated storage of the / P1- / P8 directory of the SD card), the storage architecture dimension can adapt to redundant storage (such as synchronous caching of key logs in the MCU's built-in FLASH), encrypted storage (such as enabling AES encryption for NVMe disk error logs containing sensitive data), and distributed storage (such as uploading logs to the server storage pool through the PCIe Switch). In addition to basic data packaging, its associated storage operations can also achieve intelligent indexing (such as generating a multi-dimensional search tree by timestamp and error level), storage policy linkage (such as automatically backing up to a remote NAS when the BMC receives a serious error), or lifecycle management (such as setting differentiated log retention periods by port).
[0091] It's important to note that the various scenarios described above are all implemented based on the hardware architecture shown in the accompanying diagram: voltage fluctuations are captured by the Expander's SMBUS sensor and transmitted to the MCU via I2C; electromagnetic interference events are input by an external monitoring device via GPIO_IN:1; DLLP retransmission triggers rely on PTraceCLI parsing Expander registers; and distributed storage utilizes the PCIeSwitch's uplink port for data transmission. Each expansion solution maintains the core port number indexing mechanism, but the processing granularity and functional boundaries can be dynamically adjusted to meet operational needs.
[0092] Through the embodiments of the present application, by monitoring error events, an error log containing an error code and port number is instantly generated when an abnormality occurs in the target link; then, through precise fault mapping, based on the error code mapping table pre-stored in non-volatile media, the original error code is dynamically converted into a readable fault description, and the physical location is associated with the data packet captured by the trigger to achieve fault tracing; then, through structured storage management, the fault description and the original log are partitioned and stored by port number to form a traceable historical database; finally, through system collaborative optimization, real-time alarms and in-depth diagnosis are triggered. This achieves the technical effects of a leap in operation and maintenance efficiency, enhanced fault tracing, and enhanced system reliability. Therefore, it can solve the technical problem of low fault analysis efficiency of switching chips in related technologies.
[0093] As an optional solution, in response to an error event occurring on the target link, an error log is generated based on a preset fault detection mechanism, including:
[0094] Detect the connection status of each port in the target link through the general input and output interface;
[0095] When an abnormality in the connection state is detected, it is determined that an error event has occurred on the target link, and an error log is generated based on a preset fault detection mechanism;
[0096] Export error logs via the serial debug interface.
[0097] Optionally, in the embodiment of the present application, the general input and output interface may include but is not limited to programmable hardware pins, including but not limited to Figure 1 The GPIO_IN:0 / GPIO_IN:1 pin group is marked on the left side of the MCU module. For example, when the NIC card is suddenly physically disconnected on port 1, GPIO_IN:0 will instantly jump from a high level to a continuous low level, directly reporting the physical layer connection loss event, which can reduce status detection latency.
[0098] Optionally, in the embodiment of the present application, the connection status may include but is not limited to the electrical connectivity status of the port and the device, including but not limited to Figure 1The link handshake signal sequence fed back by the Expander chip via GPIO. For example, when an NVMe drive is inserted into port 5, if the slot is deformed and causes impedance mismatch, GPIO_IN:1 will output periodic level oscillation, indicating that the link is in an unstable negotiation state.
[0099] Optionally, in the embodiment of the present application, the above-mentioned preset fault detection mechanism may include but is not limited to a hierarchical diagnosis strategy, including but not limited to: Figure 1 The AER register polling mechanism is used by the MCU and the expander to communicate via the I2C bus. For example, for the NIC card on port 4, if the physical layer error counter exceeds the threshold five times within 10ms, PTraceCLI message capture and register snapshot storage are triggered simultaneously.
[0100] Optionally, in the embodiment of the present application, the serial debug interface may include but is not limited to an offline diagnostic data channel, including but not limited to Figure 1 The multiple serial ports labeled UART to the right of the MCU (connected to the SBD Header) are shown. For example, maintenance personnel can use a USB-to-TTL converter to connect to these ports and send historical logs to a diagnostic terminal.
[0101] For example, the above-mentioned detection of the connection status of each port in the target link through the universal input and output interface can be understood as follows: Figure 1 As shown in the MCU module, GPIO_IN:0 is used to detect the physical plug-in and unplug status of the port (high level = device is in place), and GPIO_IN:1 monitors the link handshake result (low level = training failure). For example, if the GPU card on port 2 is accidentally unplugged, the GPIO_IN:0 level drops sharply, and the MCU immediately marks the port as "physically disconnected."
[0102] For example, when an abnormality in the connection state is detected, it is determined that an error event has occurred in the target link. Generating an error log based on a preset fault detection mechanism can be understood as establishing a mapping mechanism from physical abnormalities to logical errors. Figure 1 In a typical scenario, the NVMe disk on port 5 repeatedly retrains the link due to power supply noise: GPIO_IN: 1 continuously outputs a signal indicating an abnormal connection status. After the MCU identifies this as an error event, it reads the Expander's PHY register through I2C, generates a log file containing the error code 0x7003 (signal amplitude exceeds the limit) and the port number, and writes it to the SD card / P5 partition via SDIO.
[0103] For example, the above-mentioned export of error logs via the serial debug interface can be understood as an offline diagnostic data output solution. Figure 1The UART port on the right side of the MCU is connected to the SBD Header. When an error occurs on port 1, the operator sends a command via the USB-TTL converter, and the MCU automatically transfers the logs from the SD card / P1 partition to the terminal. This design avoids server downtime and disassembly, saving time in fault location.
[0104] Through the embodiments of the present application, the real-time diagnostic capability is significantly improved. Relying on the fusion mechanism of direct monitoring of the physical layer and deep detection of the protocol layer, the faults that traditionally require long-term manual positioning are transformed into automatic capture at the millisecond level. For example, the change in port connection status can instantly generate a log with time and space tags; the operation and maintenance efficiency is fundamentally improved, and the risk of equipment disassembly in traditional operations is effectively avoided with the help of the non-disassembly log export interface, which supports operation and maintenance personnel to remotely retrieve historical fault records to reduce the probability of system interruption; the system resilience is significantly enhanced, and the layered error handling mechanism fully covers the compound fault scenarios of physical connection anomalies and data transmission errors. When physical looseness of the port and protocol verification errors occur concurrently, the system can automatically generate multiple independent logs and store them in the corresponding partitions of the port to ensure complete traceability of complex faults. The full-process closed-loop design achieves high reliability of the fault handling process through direct connection of hardware-level status, port-based storage isolation and standardized interface docking, while controlling system performance loss at an extremely low level.
[0105] As an optional solution, when an error event is detected by a trigger, before converting the error code into fault description information based on a pre-generated error code mapping table, the method further includes:
[0106] Use the target protocol tool to load the pre-configuration file;
[0107] Set triggers for each port in the target link based on a pre-configured file.
[0108] Optionally, in an embodiment of the present application, the above-mentioned trigger may include but is not limited to a logic mechanism for automatically detecting specific error conditions in a switching chip fault analysis scenario, which activates subsequent operations based on preset rules or hardware signal changes, including but not limited to capturing abnormal events by real-time monitoring of PCIe link status changes. For example, when a CRC check error or link interruption occurs during PCIe data transmission, the trigger can respond immediately and start the log capture process to ensure that the fault is identified in time without manual intervention. In a switching chip environment, the design of the trigger relies on configurable hardware or software parameters, such as setting a specific data packet type or error level as a trigger point to efficiently filter irrelevant noise and focus on key fault signals, thereby improving system response speed and reliability. For example, including but not limited to configuring a DLLPs packet loss-based trigger for each PCIe port in a server board, when packet loss is detected three times in a row, the system automatically records the error location and notifies the management unit to avoid fault amplification due to delayed response.
[0109] Optionally, in an embodiment of the present application, the target protocol tool may include but is not limited to a software or firmware component dedicated to performing specific protocol operations in a switching chip environment, including but not limited to loading configuration files and setting monitoring triggers. In a PCIe fault analysis scenario, the tool, such as the PTraceCLI function, interacts with the chip through a UART interface and supports automated capture of packet trace information without the need for external equipment intervention. Its core advantage lies in its high degree of integration and its ability to dynamically adapt to different link states based on predefined rules. For example, it includes but is not limited to using a built-in tool in a server board to load an initialization file, activate an error detection module for each port, and automatically capture logs when a hardware error occurs.
[0110] Optionally, in an embodiment of the present application, the above-mentioned pre-configuration file may include but is not limited to a text or binary file that stores preset parameters, which is used to initialize the behavior settings of the target protocol tool, including but not limited to defining trigger conditions, filtering rules or error capture ranges in switching chip fault analysis. The file is structured data in INI format, which supports dynamic loading to adapt to different port requirements and ensure the consistency and efficiency of trigger configuration. For example, it includes but is not limited to loading a configuration file at system startup, specifying that only transmission errors of specific data packet types such as TLPs are monitored, ignoring low-priority events, and optimizing resource utilization.
[0111] Optionally, in an embodiment of the present application, the above-mentioned port may include but is not limited to a physical or logical interface on a switching chip for connecting to an external device or link, including but not limited to a unique identifier such as a BDF number corresponding to each port in a PCIe environment. In fault analysis, the port serves as the source location identifier of the error event, and the system stores logs based on the port number classification to facilitate accurate problem location. For example, including but not limited to setting up an independent partition for each GPU connection port in the board design, recording the port number when an error occurs and mapping it to a specific slot to accelerate fault isolation.
[0112] Exemplarily, the above-mentioned use of the target protocol tool to load the pre-configuration file can be understood as initializing the system settings by loading the preset file through a dedicated tool during fault analysis. Target protocol tools, such as PTraceCLI, are responsible for reading and executing the parameters in the configuration files, which define detailed rules for error monitoring such as trigger types or filter conditions. In a switching chip environment, loading pre-configuration files ensures standardized tool behavior and supports automated operations without manual intervention. For example, including but not limited to when the PCIe board is started, the MCU runs a built-in tool to load an INI format file that specifies which error events to monitor, such as hardware errors or packet loss, and sets the length of the captured packet. For example, when the system detects an abnormality in the downstream port signal, the tool automatically activates the tracking function according to the configuration file, captures the original data packet and temporarily stores it on the SD card to provide complete context data for error analysis. This step optimizes resource allocation, focuses only on key events, and avoids invalid monitoring and waste of processing power.
[0113] Exemplarily, the above-mentioned setting of triggers for each port in the target link based on a pre-configuration file can be understood as that after loading the configuration file, the system uses the parameters in the file to configure a dedicated trigger for each port to achieve link-level error monitoring. In the switching chip scenario, the target link represents the data path, and each port corresponds to an independent interface. After setting the trigger, port-level events such as connection interruption or signal error can be detected in real time. For example, based on the configuration file definition, the system sets a trigger based on the order set data packet for the upstream CPU port and a CRC error trigger for the downstream device port, and automatically activates log capture when an error occurs. For example, including but not limited to server maintenance, the configuration file specifies that only high-priority ports are monitored and idle interfaces are ignored, and the system allocates resources accordingly; when a GPU port detects continuous data packet loss, the trigger immediately notifies the MCU to export the register status to assist in in-depth analysis. This method improves the accuracy of fault location through refined configuration and ensures that error events are handled in a targeted manner.
[0114] Through the embodiments of the present application, the automation level and processing efficiency in switching chip fault analysis are significantly improved. By preloading configuration files and setting port triggers, the system can respond to error events in real time and capture detailed logs, reducing the need for manual intervention. At the same time, the fault description is converted based on the error code mapping table to ensure that the output information is intuitive and easy to understand, making it easy to quickly locate the root cause of the problem. Its advantages include reducing operation and maintenance costs, avoiding the tedious process of relying on special tools or protocol analyzers in traditional methods, and retaining historical data with the help of non-volatile storage to support long-term analysis. In addition, refined port monitoring optimizes resource utilization, prevents misjudgments or delays caused by missing configurations, and ultimately enhances system reliability and ease of maintenance. It is suitable for continuous fault management in high-density data environments.
[0115] As an optional solution, setting a trigger for each port in the target link based on the pre-configuration file includes at least one of the following:
[0116] Selecting a trigger set for each port in the target link from a group of triggers configured according to preset rules based on a pre-configured configuration file;
[0117] Triggers are customized for each port in the target link based on a pre-configuration file.
[0118] Optionally, in an embodiment of the present application, the triggers configured by the above-mentioned preset rules may include, but are not limited to, a predefined standard screening mechanism for quickly deploying common monitoring conditions in switching chip fault analysis scenarios, including but not limited to selecting a configuration scheme that adapts to the current link requirements from a predefined trigger library provided by the manufacturer. This method automatically matches port types and error levels by loading standardized configuration files, avoids repeated writing of complex rules, and significantly reduces configuration complexity in resource-constrained environments. For example, but not limited to, during the initialization phase of the PCIe switch board, the system automatically assigns a link status monitoring trigger to the upstream CPU port and a data packet integrity check trigger to the downstream storage device port after loading the pre-stored file, ensuring priority coverage of critical paths and improving deployment efficiency.
[0119] Optionally, in an embodiment of the present application, the above-mentioned customization method may include but is not limited to a flexible configuration mechanism that allows complete customization of trigger parameters according to the specific environmental requirements of the switching chip, including but not limited to manually defining error detection thresholds, packet filtering conditions, or combinational logic rules. In the fault analysis scenario, the customization method is suitable for non-standard equipment or special debugging requirements, and supports deep adaptation to hardware differences. For example, but not limited to, in a heterogeneous server environment, the operation and maintenance personnel can set a signal jitter tolerance threshold for a customized acceleration card port by editing a configuration file, and automatically trigger logging when a clock offset exceeding the set value is detected, thereby solving special failure modes that cannot be covered by standardized solutions.
[0120] Exemplarily, the aforementioned setting of triggers for each port in the target link based on a pre-configuration file, including at least one of the following, can be understood as providing two complementary configuration paths within the switch chip fault management framework, with port monitoring policies uniformly managed through pre-configuration files. This design allows the system to select a configuration method that prioritizes efficiency or flexibility based on actual needs, avoiding the limitations of a single solution. In specific implementations, for example, when deploying standardized server boards, engineers can use preset rules to batch configure port groups; while when debugging customized hardware, custom mode can be enabled to fine-tune the error sensitivity of specific ports.
[0121] For example, the above-mentioned selection of triggers for each port in the target link from a set of triggers configured according to preset rules based on a pre-configured configuration file can be understood as an efficient mode for quickly deploying monitoring policies using a pre-stored template library. The system automatically selects the best practice solution by matching port attributes, significantly reducing the workload of manual configuration. For example, during the startup process of a data center switch card, after the target protocol tool reads the configuration file, it automatically assigns preset bandwidth utilization monitoring triggers to all PCIe Gen4 ports and delay detection triggers to Gen3 ports, achieving minute-level full-link monitoring coverage.
[0122] For example, the triggers set for each port in the target link through customization based on the pre-configured file can be understood as providing expansion capabilities that exceed standard limitations. Users can freely define trigger conditions, associated actions, and error levels by editing the configuration file to meet the needs of specific scenarios. For example, during the R&D and testing phase, engineers can customize a composite trigger for a port: when the CRC error rate exceeds the threshold and the temperature sensor alarm is detected simultaneously, a chip register snapshot is automatically exported and the port is isolated, enabling multi-dimensional fault correlation analysis.
[0123] Through the embodiments of the present application, the elastic expansion and precise control of configuration strategies are achieved in the analysis of switching chip faults, and two complementary trigger deployment modes are uniformly carried through pre-configured files: the preset rule selection provides efficient configuration capabilities out of the box, significantly reducing the complexity of conventional deployment; the customized method gives the flexibility to deeply adapt to special needs and supports refined monitoring of complex scenarios. The two work together to enable the system to dynamically adapt to different hardware environments, such as quickly applying standardized monitoring in batch server scenarios and customizing special diagnostic rules in heterogeneous computing scenarios. This design greatly improves the coverage and accuracy of fault capture, ensuring that everything from common protocol errors to marginal anomalies can be effectively detected. At the same time, resource consumption is optimized through fine-grained configuration at the port level to avoid performance loss caused by invalid monitoring. Ultimately, an extensible automated diagnostic framework is formed to shorten fault location time, enhance system maintenance efficiency, and provide underlying protection for high-reliability applications.
[0124] As an optional solution, when an error event is detected by a trigger, the error code is converted into fault description information based on a pre-generated error code mapping table, including:
[0125] When an error event is detected by a trigger, use the target protocol tool to obtain the trigger data packet corresponding to the error event;
[0126] Search the error log based on the triggering data packet and determine the error code;
[0127] The error code is converted into fault description information based on the pre-generated error code mapping table.
[0128] Optionally, in an embodiment of the present application, the above-mentioned trigger data packet may include but is not limited to the original protocol transmission unit captured by the monitoring mechanism during the switching chip fault analysis process, containing complete context information when the error event occurs, including but not limited to TLPs data packets, DLLPs control packets or specific sequence sets transmitted in the PCIe link. In the fault diagnosis scenario, the trigger data packet is used as the core analysis object, and the root cause of the error is located by parsing its header field, payload content and transmission timing. For example, including but not limited to when the port detects a CRC check failure, the target protocol tool automatically captures the 128-byte data packet transmitted before and after this moment, which contains the source address, transaction type and check code field of the error data frame, providing the original basis for subsequent error classification.
[0129] For example, when an error event is detected by a trigger, the use of the target protocol tool to obtain the trigger data packet corresponding to the error event can be understood as an operation that prioritizes capturing the original chain of evidence in the error response process. When the preset trigger condition is met, the target protocol tool immediately locks the data stream snapshot at the time the error occurs to ensure the integrity of the fault site. For example, in a storage server scenario, when the DLLP packet loss trigger of the downstream port of the PCIe switch chip is activated, the PTraceCLI tool automatically captures the 64 most recently transmitted data packets of the port, including the protocol interaction sequence before and after the error occurs, and stores them in the non-volatile storage area for in-depth analysis.
[0130] For example, the aforementioned search for error logs based on trigger packets and determination of error codes can be understood as a technical approach that correlates raw packet characteristics with structured error records. The system parses key fields in the trigger packet, matches predefined error patterns in chip firmware or registers, and extracts standardized error identifiers. For example, when a packet containing an abnormal TLP header is captured, the diagnostic algorithm compares its transaction type field with the record in the AER log, determines that the error corresponds to error code 0x102, indicating a data link layer protocol violation, and simultaneously correlates the generated port number and timestamp information.
[0131] Exemplarily, the conversion of error codes into fault descriptions based on a pre-generated error code mapping table can be understood as a semantic conversion process from machine-readable code to human-understandable descriptions. The system queries a pre-loaded mapping database and converts the abstract code into a complete description containing the fault location, type, and repair suggestions. For example, when the error code 0x301 is input, the mapping table outputs a text description indicating that a power management status conflict has occurred on downstream port 2 and recommends checking the device's power configuration. This information is pushed to the operation and maintenance terminal in real time via the management bus to guide on-site maintenance.
[0132] Through the embodiments of the present application, a closed-loop automated diagnostic system is constructed in the analysis of switching chip faults. The key data packets of the error event are actively captured through the trigger mechanism to ensure that the fault site is completely preserved, avoiding the problem of error information loss in traditional methods; based on the characteristics of the data packet, the structured error log is accurately associated and the standard error code is extracted to eliminate the uncertainty of manual analysis; finally, an intuitive fault description is generated through a preset mapping table to achieve end-to-end conversion from raw signals to actionable guidance. This method significantly shortens the fault location time. For example, in a complex link error scenario, the system automatically identifies the abnormal sequence in the data packet and associates it with the physical port. Operations and maintenance personnel can directly obtain repair suggestions without analyzing the binary log. At the same time, standardized processing procedures reduce dependence on professional tools, allowing non-professionals to efficiently handle switching chip-level faults and improve system maintainability. In addition, structured data storage supports historical failure mode analysis, provides a data basis for preventive maintenance, and overall enhances service continuity in high-density computing environments.
[0133] As an optional solution, the above method further includes:
[0134] In response to the switching chip being started, loading an error code mapping table from a storage medium, wherein the error code mapping table is used to indicate a mapping relationship between error codes and fault description information;
[0135] When an error log is generated based on a preset fault detection mechanism, the error log is stored in a storage partition corresponding to the port number;
[0136] When the error code is converted into fault description information based on the error code mapping table, the fault description information is matched with the error code and stored in the storage partition corresponding to the port number, and the timestamp and error level associated with the error log are recorded.
[0137] Optionally, in an embodiment of the present application, the above-mentioned storage medium may include but is not limited to a persistent data carrier integrated into the switch board for long-term storage of key configuration information and fault records, including but not limited to non-volatile storage devices such as SD cards, eMMC chips or NOR Flash. In the fault analysis scenario, the medium must meet the requirements of frequent reading and writing and power-off preservation, and support partition management to achieve port-level log isolation. For example, including but not limited to the use of industrial-grade SD cards in the design of the server PCIe switch board, divided into 32 independent storage areas corresponding to different physical ports, to ensure that when an error occurs in port 3, the relevant log is only written to the third partition to avoid data cross-contamination.
[0138] Optionally, in an embodiment of the present application, the above-mentioned error level may include, but is not limited to, a severity classification system based on the scope of the fault impact, including, but not limited to, recoverable errors, service degradation errors, and unrecoverable system errors. This classification guides the operation and maintenance response priority and log storage strategy, and realizes optimal resource allocation in switch chip management. For example, including but not limited to, when continuous CRC check errors are detected on a port, the system marks it as a recoverable error level and only records the log without interrupting service; if the clock signal is lost, it is marked as an unrecoverable error and immediately triggers a port isolation alarm.
[0139] Optionally, in embodiments of the present application, the timestamp may include, but is not limited to, millisecond-accurate timing data used to record the absolute time point at which the error log was generated, including but not limited to a time source synchronized with an onboard real-time clock chip or a network time protocol. In fault analysis, the timestamp supports multi-module event correlation and fault chain reconstruction. For example, including but not limited to, when a bandwidth drop occurs on a downlink port of a switching chip, the system logs the year / month / day / hour / minute / second / millisecond information of the event, facilitating correlation with historical data such as temperature sensors and power supply fluctuations to analyze the root cause.
[0140] Exemplarily, in response to the switching chip starting up, the error code mapping table is loaded from the storage medium, wherein the error code mapping table is used to indicate the mapping relationship between the error code and the fault description information, which can be understood as a key pre-configuration operation in the system initialization phase. After the switching chip completes the hardware reset, the control unit automatically reads the pre-compiled mapping relationship data set from the non-volatile storage to establish a fast conversion channel between the error code and the readable description. For example, when a certain model of PCIe switching board is powered on, the MCU immediately loads a file named error_map.csv from the SD card. The table maps the hexadecimal code 0x205 to an uplink negotiation timeout and 0x310 to a downlink device power management conflict, providing a semantic conversion basis for real-time fault analysis.
[0141] For example, when an error log is generated based on a preset fault detection mechanism, storing the error log in a storage partition corresponding to the port number can be understood as a specific implementation of a structured storage strategy. When the monitoring module captures a valid error event, the system stores its original log in a preset independent storage area based on the physical identifier of the port where the error originated. For example, when a trigger detects a TLP transmission error on port BDF_0x1B, the diagnostic engine automatically extracts the binary log containing the error code and register snapshot and writes it to the partition marked as Port_1B on the SD card, forming a strong association between the physical location and the storage space.
[0142] Exemplarily, in the case where the error code is converted into fault description information based on the error code mapping table, the fault description information is matched with the error code and stored in the storage partition corresponding to the port number, while recording the timestamp and error level associated with the error log, which can be understood as an enhanced log archiving mechanism. After the error code conversion is completed, the system not only saves the readable description text, but also associates the original error code to form a comparison relationship, and attaches the event time and severity mark to build a multi-dimensional fault archive. For example, when the mapping table parses the error code 0x415 as a port training failure, the system simultaneously records in the partition corresponding to the port: error code 0x415, description text signal impedance abnormality, timestamp, and serious error level, forming a complete event record that supports multi-dimensional retrieval.
[0143] Through the embodiments of the present application, refined data management of the entire life cycle of fault analysis is achieved: the error code mapping table is preloaded during the initialization phase, which significantly shortens the analysis response time after the fault occurs. For example, the switching chip can establish a complete code conversion capability in a short time after startup; a port partition storage mechanism is adopted to ensure efficient retrieval and isolated storage of massive logs. When a specific port reports an error, the historical record can be immediately located to avoid the waste of resources of a full disk scan; a multi-dimensional archiving strategy with additional timestamps and error levels not only records the fault phenomenon itself, but also builds a traceable event sequence and impact assessment framework. For example, by analyzing the transient errors of a port in the same time period for three consecutive days, the signal integrity problems caused by ambient temperature fluctuations can be accurately identified. This method fundamentally improves the availability of fault data, supports operation and maintenance personnel to quickly perform root cause analysis, and provides a structured data foundation for preventive maintenance, ultimately enhancing the continuous service capabilities of high-availability systems.
[0144] As an optional solution, the above method further includes:
[0145] In response to the switch chip starting up, the error code mapping table is loaded from the storage medium, and a pre-configuration file is loaded using a target protocol tool;
[0146] Set triggers for each port in the target link based on a pre-configured file;
[0147] Detect the connection status of each port in the target link through the general input and output interface;
[0148] When an abnormality in the connection state is detected, it is determined that an error event has occurred on the target link, and an error log is generated based on a preset fault detection mechanism;
[0149] Export error logs through the serial debug interface and store them in the storage partition corresponding to the port number;
[0150] When an error event is detected by a trigger, use the target protocol tool to obtain the trigger data packet corresponding to the error event;
[0151] Search the error log based on the triggering data packet and determine the error code;
[0152] Convert the error code into fault description information based on the pre-generated error code mapping table, store it in the storage partition corresponding to the port number, and record the timestamp and error level associated with the error log;
[0153] When continuous error events are detected, the built-in script is called to export the register status of the switch chip and update the fault description information based on the register status;
[0154] In response to the fault description information being stored in the storage partition corresponding to the port number, a signal is sent to the baseboard management controller to instruct the baseboard management controller to read the fault description information and generate an alarm event based on the fault description information.
[0155] Optionally, in an embodiment of the present application, the aforementioned continuous error events may include, but are not limited to, repeated occurrences of the same type of fault signal within a preset time window, including, but not limited to, persistent triggering of CRC check errors or repeated link training failures. In the fault escalation mechanism, this state triggers a deep diagnostic mode to avoid over-response to occasional errors. For example, including but not limited to, when a port records more than five packet retransmission failures within 10 seconds, the system automatically determines it as a persistent error and initiates the register snapshot export process.
[0156] Optionally, in embodiments of the present application, the register status described above may include, but is not limited to, real-time configuration snapshots of the switch chip's internal functional modules, including, but not limited to, link control registers, error counters, and power management status words. This data provides a snapshot of hardware-level operation, revealing underlying anomalies during complex fault location. For example, when a persistent bandwidth drop is detected, a script may derive the physical layer equalizer coefficient register values and identify an abnormal offset in the receiver's pre-emphasis configuration.
[0157] Optionally, in embodiments of the present application, the aforementioned built-in scripts may include, but are not limited to, automated diagnostic routines pre-installed in the control unit, including, but not limited to, register operation sequences written in Python or Shell. Such scripts orchestrate actions for in-depth fault analysis, extending the boundaries of basic monitoring capabilities. For example, when a persistent error is triggered, the script automatically performs three operations: freezing the error port status, exporting the associated register group, and generating a binary snapshot file and linking it to the original log.
[0158] Optionally, in embodiments of the present application, the aforementioned alarm event may include, but is not limited to, a structured fault notification medium, including, but not limited to, a system event log in IPMI format or an SNMP alarm message. This event integrates the fault description, severity, and location information to drive an operations and maintenance response. For example, upon analyzing a port hardware fault, the system generates an alarm event containing the slot number, error description, and a recommendation to replace the hardware component, which is then pushed to the management platform to trigger the work order system.
[0159] Exemplarily, the aforementioned steps of loading the error code mapping table from the storage medium in response to the switch chip starting up and loading the pre-configuration file using the target protocol tool can be understood as a dual pre-loading operation during the initialization phase, simultaneously establishing the error semantic conversion foundation and monitoring rule base. After the switch chip completes the power-on self-test, the control unit loads key configuration files from non-volatile storage in parallel: the error code mapping table provides code translation capabilities, and the pre-configuration file defines the subsequent monitoring behavior guidelines.
[0160] For example, setting triggers for each port on the target link based on a pre-configured file can be understood as a refined deployment process for distributing monitoring policies based on port characteristics. The system parses the parameter rules in the pre-configured file and dynamically assigns dedicated trigger combinations based on port type, speed level, and connected device attributes.
[0161] For example, the above-mentioned detection of the connection status of each port in the target link through the universal input and output interface can be understood as a real-time inspection mechanism for the physical layer status. The control unit periodically scans the port detection pin level status and determines the device connection stability through the binary signal flow.
[0162] For example, when a connection status anomaly is detected, an error event is determined to have occurred on the target link. Generating an error log based on the pre-set fault detection mechanism can be understood as the process of converting a physical layer anomaly into a manageable fault record. When the connection status detection module reports an anomaly, the system automatically classifies it as a specific error type and generates a standardized log entry with a location identifier.
[0163] For example, exporting error logs via the serial debug interface and storing them in the storage partition corresponding to the port number can be understood as an operation chain that uses the underlying interface to achieve targeted log archiving. The control unit reads the chip's internal error buffer through a dedicated debug channel and writes the original log to the partitioned storage structure based on the port identifier.
[0164] For example, when an error event is detected by a trigger, using the target protocol tool to obtain the triggering packet corresponding to the error event can be understood as a key step in capturing a snapshot of the protocol layer's error response. When the preset trigger conditions are met, the target tool immediately captures the complete protocol interaction sequence at the moment of the error, preserving the chain of evidence at the fault site.
[0165] For example, the aforementioned error code determination based on the triggering packet's error log search can be understood as an analysis path that reverse-correlates structured error records with packet features. The diagnostic engine parses key fields in the packet, matches corresponding entries in the firmware error log, and extracts standard error identifiers.
[0166] For example, the system converts error codes into fault descriptions based on a pre-generated error code mapping table and stores them in the storage partition corresponding to the port number. Simultaneously recording the timestamp and error severity associated with the error log can be understood as an archiving strategy for building a multi-dimensional fault archive. After the system completes the code conversion, it simultaneously saves the following: human-readable description text, original error code, millisecond-level timestamp, and severity tag in the port-specific storage area, forming a traceable event record.
[0167] For example, when continuous error events are detected, the system invokes a built-in script to derive the register status of the switch chip and updates the fault description based on the register status. This can be considered a deep diagnostic enhancement mechanism for persistent faults. When the system identifies frequent occurrences of errors from the same source, it automatically triggers a diagnostic script to obtain a snapshot of the hardware registers and, based on this snapshot, revise the initial fault conclusion.
[0168] Exemplarily, the above process, in response to the fault description information being stored in the storage partition corresponding to the port number, sends a signal to the baseboard management controller, instructing the controller to read the fault description information and generate an alarm event based on the fault description information, which can be understood as a closed-loop alarm triggering process. Once the log is archived, the control unit notifies the management system via a hardware interrupt to read the complete diagnostic report and convert it into a standard alarm format.
[0169] Through the embodiments of the present application, a full-lifecycle automated fault management system is constructed: key configuration files are loaded in parallel during the initialization phase, shortening the fault response preparation time to milliseconds; through a two-layer monitoring network of port-level trigger configuration and physical status monitoring, full-dimensional error coverage from protocol layer anomalies to physical connection interruptions is achieved; a directional log archiving mechanism based on a partitioned storage architecture ensures structured storage and millisecond-level retrieval capabilities for massive fault data; triggers intelligent correlation analysis of data packet capture and error logs to eliminate the uncertainty of manual analysis and improve the accuracy of root cause diagnosis; a register deep snapshot function for persistent faults provides a basis for hardware-level problem location; and finally, through an event-driven alarm generation mechanism, end-to-end automation from fault detection to operation and maintenance response is achieved. This method compresses the average fault location time to one percent of that of traditional methods in a complex switching environment, while reducing the monitoring load through refined resource management and control, providing continuous and reliable hardware assurance capabilities for high-density computing scenarios.
[0170] In an exemplary embodiment, the method further includes: in response to an update instruction issued by a baseboard management controller, obtaining a remote mapping table version through an out-of-band management channel; comparing the locally stored error code mapping table version number with the remote version number; when it is detected that the version difference exceeds a threshold, mounting the non-volatile storage medium to the baseboard management controller; clearing the historical mapping table partition through a secure erase operation, and writing a new version of the error code mapping table.
[0171] In an exemplary embodiment, when applied to a multi-switch board backplane system, when the baseboard management controller detects that the current switch board port error level is serious, it executes: broadcasting the port error event to the adjacent switch board through the backplane management bus; receiving the associated link quality indicators returned by the adjacent board, the indicators including at least one of the following: signal crosstalk strength, reference clock offset value, power rail ripple coefficient; generating a joint fault analysis report based on the multi-board indicators, and updating the fault description information to the cross-board interference type.
[0172] In an exemplary embodiment, the process of setting the trigger includes: parsing the energy efficiency policy field in the pre-configuration file to identify the port activity level tag; configuring a real-time hardware error trigger for a high-activity level port; enabling a periodic polling mode for a low-activity level port, including: activating polling during the system idle period; setting a multi-level cache temporary error event; triggering batch parsing when the cumulative number of error events exceeds a threshold; the micro control unit dynamically switches the working mode according to the port activity status and enters a dormant state when there is no error event.
[0173] The following is a further explanation of this application with reference to specific examples:
[0174] With the development of information technology, data centers need to store, compute, and interact with increasing amounts of data. PCIe, as a key hardware bus for data exchange within servers, is an essential component. PCIe itself operates on a point-to-point basis, with hardware circuits from the CPU to the device transmitting data point-to-point. To increase PCIe's flexibility, PCIe Bridges or Switches are incorporated into PCIe hardware designs. This expands the number of PCIe signals and bandwidth, allowing multiple PCIe devices to be connected to the same PCIe port, enabling simultaneous information exchange between multiple devices.
[0175] PCIe Switch is a PCIe switching chip widely used in servers. It is used to expand PCIe signals so that the system can support more PCIe devices, such as GPU cards, network cards, IB cards, and NVMe storage disks.
[0176] Switch boards with PCIe switches are generally independent boards with limited communication signals with the mainboard's BMC. Because the BMC is responsible for monitoring and managing the entire server, it cannot occupy resources for long periods of time to perform PCIe fault analysis. Therefore, its ability to analyze PCIe switch board errors is limited. If a PCIe switch board fails in a server, logs must be exported and analyzed using the manufacturer's tools; the switch board cannot automatically analyze the fault.
[0177] This application proposes a method for automatically analyzing PCIe Switch board faults. This method monitors the PCIe Switch's link status, data transmission errors, and configuration anomalies in real time. By combining a non-volatile storage SD card with the system management module (MCU / BMC), this method automatically analyzes, stores, and reports fault information, significantly improving server system maintenance efficiency. Specifically, it achieves the following:
[0178] (1) Automatically analyze fault information, convert PCIe switch error codes into specific fault types (such as link interruption, CRC error, device configuration conflict, etc.), and associate them with the corresponding physical port or device slot.
[0179] (2) Implement non-volatile storage to ensure that fault logs are retained after power failure and support historical fault tracing.
[0180] (3) Real-time reporting, automatically notifying operation and maintenance personnel through system management modules (such as BMC), reducing manual intervention.
[0181] like Figure 3As shown, this application proposes an automatic fault analysis method for PCIe Switch boards based on Broadcom PCIe 5.0 Switches. The core modules include the following: A monitoring module: This module uses the PTraceCLI function and the PCIe AER mechanism to capture PCIe Switch error events in real time and analyze the error type and source port. A parsing module (MCU): This module integrates the PTraceCLI tool and a fault analysis algorithm, converts raw error codes into readable information, and triggers log export and reporting. A storage module: This module uses an SD card or eMMC to store error mapping tables and historical logs, dividing the storage area by port.
[0182] In addition, it can also include a system management module (BMC): communicate with the MCU through the I2C or IPMI protocol to achieve real-time push of fault information.
[0183] The mainboard houses the BMC management chip, while the PCIe switch board houses the PCIe switch chip and MCU. The PCIe switch connects to the CPU upstream via PCIe signals and to the GPU, network interface card, and NVMe storage drive downstream. The PCIe switch integrates a protocol analyzer, and the MCU integrates the PCIe switch's PTraceCLI tool, which can be used to access the protocol analyzer. The MCU monitors PCIe switch data in real time and stores PCIe switch error information on an SD card. The BMC communicates with the MCU via I2C. Upon receiving error information, the MCU notifies the BMC to read the specific error content and location. The BMC can then read or switch control of the SD card via I2C and retrieve fault analysis logs from the SD card.
[0184] The PCIe switch chip serves as the core switching chip, connecting upstream to the CPU and downstream to multiple PCIe end devices (such as GPUs and NVMe network cards). The MCU chip communicates with the PCIe switch via UART or I2C and is responsible for error code analysis and log management. The non-volatile storage unit (NSM) is mounted on the MCU and stores error mapping tables and historical logs. The BMC chip interacts with the MCU via the I2C bus, receiving fault notifications and forwarding them to the management platform. The MCU's SD card pre-stores the switch board's port configuration and bandwidth rate. The MCU connects to the PCIe switch's SDB and UART interfaces via UART. After the server boots up and the PCIe switch is reset, it begins operation. The MCU first reads the PCIe switch's OTC firmware log to obtain the actual physical connection topology of the PCIe switch, the devices corresponding to each BDF, and the bandwidth rate. The MCU then compares this information with the information recorded on the MCU's SD card. The OTC log is also stored on the SD card for the BMC to read.
[0185] The MCU integrates the PTraceCLI tool. After power-up, the tool loads and runs in real time. It can be used to configure hardware errors or various CE / UCE trigger modes, such as TLPs, DLLPs, and order sets. When a CE / UCE error occurs on the PCIe link, PTraceCLI captures a packet segment via UART and stores it on an SD card. The MCU uses GPIO to obtain the connection status of each PCIe switch port. If a port problem is detected, the g4Xtool tool is used to export the PCIe switch chip error log via UART and store it on an SD card. A preset file named precanned.ini is provided to simplify the use of common triggers. This file contains three sections: filters, trigger 0, and triggers. The ipalstruct.ini file configures the PEA to capture specific packets on ingress and egress ports. This file contains four sections: capture settings, filter settings, condition settings, and trigger settings. The MCU uses the precanned.ini file or the ipalstruct.ini file to configure triggers for each PCIe port. The precanned.ini file allows selection from a set of commonly used triggers. The ipalstruct.ini file allows you to customize all available triggers and filters. By setting precanned.ini and ipalstruct.ini, when hardware or firmware issues are detected, the MCU will automatically capture logs through PTraceCLI and store them on the SD card.
[0186] The SD card stores a list of all PCIe error types. The MCU compiles the received PtraceCLI package and the error triggering cause into an error record. The MCU compares the error record with the error list stored on the SD card, parses the error type, and writes the error information to the designated device slot area on the SD card. The MCU communicates with the BMC via I2C, instructing it to read the error log and data packet for the corresponding PCIe switch port. The BMC can read error information from the MCU via the I2C bus. Alternatively, the BMC can switch control of the SD card, retrieve the corresponding error information from the SD card, and convert it into an error log.
[0187] Specifically, it may include but is not limited to the following embodiments:
[0188] (1) During the initialization phase, the MCU loads the error mapping table (e.g., the correspondence between PCIe error codes and fault descriptions) from the storage unit. The PCIe Switch configures the AER function and enables error event interrupt notification.
[0189] (2) Error capture and analysis. When an error occurs in the PCIe link, the PCIe Switch generates an error log through the AER mechanism, which includes the error type (such as Uncorrectable Internal Error), port number, and device ID. The MCU runs the PTraceCLI tool in real time, configures the precanned.ini file or ipalstruct.ini file, and sets various CE / UCE trigger methods such as hardware error or TLPs, DLLPs, and order set. When an error event occurs, the PCIe package trace data is automatically captured. The MCU reads the original log through the UART and parses it into a specific fault description based on the mapping table (for example, "Port 3 link interruption, device BDF 0x1A"). It also stores the corresponding AER error and the PCIeTrace data captured by PTraceCLI.
[0190] (3) Log storage and classification: The parsed logs are stored in the corresponding partition of the SD card by port number, and the timestamp and error level (serious / recoverable) are recorded. If continuous errors are detected (such as multiple CRC check failures), the MCU calls a built-in script to export the internal register status of the PCIe Switch to assist in in-depth analysis.
[0191] (4) Fault reporting: The MCU sends an interrupt signal to the BMC via GPIO or I2C. The BMC reads the fault information in the SD card and generates an alarm event. The BMC supports remote access, and operation and maintenance personnel can view detailed logs or download historical data through the IPMI protocol.
[0192] Through the embodiments of this application, PCIe Switch error information and raw PCIe PackageTrace data can be obtained in real time, and the fault type can be automatically parsed and PCIe Switch FW logs can be stored, making it easier for maintenance personnel to locate the problem, saving manpower and rework and repair costs. A high degree of automation is achieved, automating the entire process from error capture and parsing to reporting, reducing manual operations. Fault location is accurate, and faulty hardware can be quickly located by associating the port number with the device ID. Data persistence and non-volatile storage ensure long-term log retention, supporting fault reproduction and root cause analysis.
[0193] This effectively solves the problem in PCIe switch board applications where, when error messages appear on the upstream and downstream ports of the PCIe switch, including the CPU and GPU, only the PCIe error type is reported but no original PCIe trace data is available, making it impossible to analyze the specific error cause. This avoids the need to manually edit logs to reproduce the problem and capture logs using a PCIe protocol analyzer.
[0194] In summary, this application proposes a method for automatically analyzing PCIe switch board faults. The PCIe switch board integrates an MCU chip, which is connected to the PCIe switch via UART. The MCU monitors PCIe switch error information in real time by dynamically associating PCIe AER error codes with physical ports, reading the original packet trace information and corresponding locations of PCIe switch errors. The MCU compiles the read PtraceCLI package and the error triggering cause into an error record and writes the error information to the corresponding location on an SD card. The error information is then compared with the error list stored on the SD card, parsed into the specific error type, and recorded. The MCU integrates the scrutiny CLI tool, which can export the PCIe switch firmware log. The MCU and the baseboard management computer (BMC) collaborate in real-time fault reporting. After receiving error information, the MCU communicates with the BMC via I2C, instructing the BMC to read the error log and trace data for the corresponding PCIe switch port. The error mapping table is dynamically updated, supporting remote updating of error code parsing rules via the BMC, adapting to devices from different manufacturers. A partitioned log management mechanism based on non-volatile storage and a dual-storage redundancy design with SD card and eMMC backing up each other prevents data loss due to single points of failure. Low-power mode optimization: the MCU enters sleep mode when there are no errors and only periodically polls the PCIe switch status to reduce energy consumption.
[0195] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0196] The embodiment of the present application also provides a fault analysis device for a switching chip, such as Figure 4 As shown, the device includes:
[0197] A generating module 402 is configured to generate an error log based on a preset fault detection mechanism in response to an error event occurring on a target link, wherein the error log includes an error code and a port number, and the target link indicates a link on which the switch chip extends a signal;
[0198] a conversion module 404 for converting an error code into fault description information based on a pre-generated error code mapping table when an error event is detected by a trigger, wherein the trigger represents a trigger set by a target protocol analysis tool;
[0199] The storage module 406 is used to store the fault description information and the error log in association with each other in a storage partition corresponding to the port number in the storage medium.
[0200] As an optional solution, the above-mentioned device is used to respond to an error event occurring in the target link and generate an error log based on a preset fault detection mechanism in the following manner: detecting the connection status of each port in the target link through a general input and output interface; when an abnormality in the connection status is detected, determining that an error event has occurred in the target link, and generating an error log based on the preset fault detection mechanism; and exporting the error log through a serial debugging interface.
[0201] As an optional solution, the above-mentioned device is also used to: when an error event is detected by a trigger, before converting the error code into fault description information based on a pre-generated error code mapping table, use the target protocol tool to load a pre-configuration file; and set a trigger for each port in the target link based on the pre-configuration file.
[0202] As an optional solution, the above-mentioned device is used to set triggers for each port in the target link based on a pre-configuration file in at least one of the following ways: selecting triggers set for each port in the target link from a group of triggers configured according to preset rules based on the pre-configuration file; setting triggers for each port in the target link in a customized manner based on the pre-configuration file.
[0203] As an optional solution, the above-mentioned device is used to convert the error code into fault description information based on a pre-generated error code mapping table when an error event is detected by a trigger in the following manner: when an error event is detected by a trigger, use the target protocol tool to obtain the trigger data packet corresponding to the error event; search the error log based on the trigger data packet to determine the error code; convert the error code into fault description information based on the pre-generated error code mapping table.
[0204] As an optional solution, the above-mentioned device is also used to: in response to the startup of the switching chip, load an error code mapping table from a storage medium, wherein the error code mapping table is used to indicate the mapping relationship between the error code and the fault description information; when an error log is generated based on a preset fault detection mechanism, the error log is stored in a storage partition corresponding to the port number; when the error code is converted into fault description information based on the error code mapping table, the fault description information is matched with the error code and stored in a storage partition corresponding to the port number, and the timestamp and error level associated with the error log are recorded at the same time.
[0205] As an optional solution, the above-mentioned device is further configured to: in response to the startup of the switching chip, load an error code mapping table from a storage medium and load a pre-configuration file using a target protocol tool; set a trigger for each port in the target link based on the pre-configuration file; detect the connection status of each port in the target link through a general input / output interface; when an abnormality in the connection status is detected, determine that an error event has occurred in the target link, and generate an error log based on a preset fault detection mechanism; export the error log through a serial debug interface and store it in a storage partition corresponding to the port number; when an error event is detected through a trigger, obtain a trigger data packet corresponding to the error event using a target protocol tool; search the error log based on the trigger data packet and determine the error code; convert the error code into fault description information based on a pre-generated error code mapping table, store it in a storage partition corresponding to the port number, and record a timestamp and error level associated with the error log; when continuous error events are detected, call a built-in script to export the register status of the switching chip and update the fault description information based on the register status; in response to the fault description information being stored in the storage partition corresponding to the port number, send a signal to a baseboard management controller to instruct the baseboard management controller to read the fault description information and generate an alarm event based on the fault description information.
[0206] As an optional solution, the device is also used to: respond to an update instruction issued by the baseboard management controller, obtain the remote mapping table version through an out-of-band management channel; compare the difference between the locally stored error code mapping table version number and the remote version number; when it is detected that the version difference exceeds a threshold, mount the non-volatile storage medium to the baseboard management controller; clear the historical mapping table partition through a secure erase operation, and write a new version of the error code mapping table.
[0207] As an optional solution, when the device is applied to a multi-switch board backplane system, it is used to execute the following when the baseboard management controller detects that the error level of the current switch board port is serious: broadcasting the port error event to the adjacent switch board through the backplane management bus; receiving the associated link quality indicators returned by the adjacent board, the indicators including at least one of the following: signal crosstalk strength, reference clock offset value, power supply rail ripple coefficient; generating a joint fault analysis report based on the multi-board indicators, and updating the fault description information to the cross-board interference type.
[0208] As an optional solution, the device is used to set triggers in the following manner: parsing the energy efficiency policy field in the pre-configuration file to identify the port activity level tag; configuring a real-time hardware error trigger for a high-activity level port; enabling a periodic polling mode for a low-activity level port, including: activating polling during system idle periods; setting a multi-level cache to temporarily store error events; triggering batch parsing when the cumulative number of error events exceeds a threshold; the micro control unit dynamically switches the working mode according to the port activity status and enters a dormant state when there is no error event.
[0209] For the description of the features in the embodiment corresponding to the fault analysis device of the switching chip, reference can be made to the relevant description of the embodiment corresponding to the fault analysis method of the switching chip, which will not be repeated here.
[0210] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned switching chip fault analysis method embodiments.
[0211] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned switching chip fault analysis method embodiments when running.
[0212] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0213] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned switching chip fault analysis method embodiments are implemented.
[0214] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned switching chip fault analysis method embodiments are implemented.
[0215] The present application also provides a switching chip fault analysis system, including: a switching chip, connected to a microcontroller unit via a universal input / output interface, for transmitting protocol layer error events; a microcontroller unit, integrating a target protocol analysis tool and mounting a non-volatile storage medium, configured to: capture the trigger data packet of the switching chip in real time via the universal input / output interface; convert the original error code into a fault description with a port identifier based on a pre-stored error code mapping table; a baseboard management controller, communicating with the microcontroller unit via an inter-integrated circuit bus, configured to: receive an interrupt signal from the microcontroller unit; dynamically switch control of the non-volatile storage medium to read fault description information; a physical layer monitoring circuit, including a universal input / output interface array, each interface being connected to an independent port connection status detection pin.
[0216] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0217] The above describes in detail the switching chip fault analysis method and device provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. It should be noted that for ordinary technicians in this technical field, without departing from the principles of this application, various improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A method for analyzing a switch chip fault, characterized in that: include: In response to an error event occurring on a target link, generating an error log based on a preset fault detection mechanism, wherein the error log includes an error code and a port number, and the target link represents a link on which the switch chip extends a signal; When the error event is detected by a trigger, converting the error code into fault description information based on a pre-generated error code mapping table, wherein the trigger represents a trigger set by a target protocol analysis tool; storing the fault description information and the error log in association with each other in a storage partition corresponding to the port number in a storage medium; When the error event is detected by a trigger, before converting the error code into fault description information based on a pre-generated error code mapping table, the method also includes: using the target protocol analysis tool to load a pre-configuration file; setting the trigger for each port in the target link based on the pre-configuration file includes at least one of the following: selecting the trigger set for each port in the target link from a group of triggers configured according to preset rules based on the pre-configuration file; setting the trigger for each port in the target link in a customized manner based on the pre-configuration file.
2. The fault analysis method for a switching chip according to claim 1, characterized in that: The generating of an error log based on a preset fault detection mechanism in response to an error event occurring on the target link includes: Detecting the connection status of each port in the target link through a general input and output interface; When detecting that the connection state is abnormal, determining that an error event has occurred on the target link, and generating the error log based on a preset fault detection mechanism; The error log is exported via the serial debug interface.
3. The fault analysis method for a switching chip according to claim 1, characterized in that: When the error event is detected by the trigger, converting the error code into fault description information based on a pre-generated error code mapping table includes: When the error event is detected by the trigger, using the target protocol analysis tool to obtain a trigger data packet corresponding to the error event; Searching the error log based on the triggering data packet to determine the error code; The error code is converted into fault description information based on a pre-generated error code mapping table.
4. The switching chip fault analysis method according to claim 1, characterized in that: The method further comprises: In response to the switching chip being started, loading the error code mapping table from the storage medium, wherein the error code mapping table is used to indicate a mapping relationship between error codes and fault description information; In a case where the error log is generated based on the preset fault detection mechanism, storing the error log in a storage partition corresponding to the port number; When the error code is converted into the fault description information based on the error code mapping table, the fault description information is matched with the error code and stored in the storage partition corresponding to the port number, and the timestamp and error level associated with the error log are recorded.
5. The fault analysis method for a switching chip according to claim 1, characterized in that: The method further comprises: In response to the switching chip being started, loading the error code mapping table from the storage medium and loading a pre-configuration file using the target protocol analysis tool; Setting the trigger for each port in the target link based on the preconfigured file; Detecting the connection status of each port in the target link through a general input and output interface; When detecting that the connection state is abnormal, determining that an error event has occurred on the target link, and generating the error log based on a preset fault detection mechanism; Exporting the error log through a serial debug interface and storing it in a storage partition corresponding to the port number; When the error event is detected by the trigger, using the target protocol analysis tool to obtain a trigger data packet corresponding to the error event; Searching the error log based on the triggering data packet to determine the error code; Convert the error code into fault description information based on a pre-generated error code mapping table, store the information in a storage partition corresponding to the port number, and record the timestamp and error level associated with the error log; In the case of detecting continuous error events, calling a built-in script to derive the register status of the switching chip, and updating the fault description information based on the register status; In response to the fault description information being stored in the storage partition corresponding to the port number, a signal is sent to a baseboard management controller to instruct the baseboard management controller to read the fault description information and generate an alarm event based on the fault description information.
6. The method for analyzing a fault of a switching chip according to any one of claims 1 to 5, characterized in that: The method further comprises: In response to an update instruction issued by the baseboard management controller, obtaining a remote mapping table version through an out-of-band management channel; Compare the locally stored error code mapping table version number with the remote version number; When it is detected that the version difference exceeds a threshold, the non-volatile storage medium is mounted to the baseboard management controller; Clear the historical mapping table partition through the secure erase operation and write the new version of the error code mapping table.
7. The switch chip fault analysis method according to claim 5, when applied to a multi-switch card backplane system, is characterized in that: When the baseboard management controller detects that the error level of the current switch board port is serious, the baseboard management controller executes: Broadcast port error events to adjacent switch cards via the backplane management bus; Receive an associated link quality indicator returned by an adjacent board, the indicator including at least one of the following: signal crosstalk strength, reference clock offset value, and power rail ripple coefficient; A joint fault analysis report is generated based on multi-board indicators, and the fault description information is updated to a cross-board interference type.
8. The method for analyzing a fault of a switching chip according to claim 1, wherein: The process of setting the trigger includes: Parse the energy efficiency policy field in the pre-configuration file and identify the port activity level tag; Configure real-time hardware error triggers for high-activity ports; Enable periodic polling mode for low activity level ports, including: Activate polling during system idle periods; Set up multi-level cache temporary error events; When the cumulative number of error events exceeds a threshold, batch parsing is triggered; The microcontroller unit dynamically switches the working mode according to the port activity status and enters the sleep state when there is no error event.
9. A switching chip fault analysis system, characterized in that: include: The switch chip is connected to the microcontroller unit through a general-purpose input / output interface and is used to transmit protocol layer error events; A microcontrol unit, integrating a target protocol analysis tool and mounting a non-volatile storage medium, is configured to: capture trigger data packets of the switching chip in real time through a universal input and output interface; convert the original error code into a fault description with a port identifier based on a pre-stored error code mapping table; a baseboard management controller, communicating with the microcontroller unit via an inter-integrated circuit bus, and configured to: receive an interrupt signal from the microcontroller unit; dynamically switch control of the non-volatile storage medium to read fault description information, wherein the fault description information and the error log are stored in association with a storage partition corresponding to the port number in the non-volatile storage medium; The physical layer monitoring circuit includes a general input and output interface array, each interface is connected to an independent port connection status detection pin.
10. A fault analysis device for a switching chip, characterized in that: include: In response to an error event occurring on a target link, generating an error log based on a preset fault detection mechanism, wherein the error log includes an error code and a port number, and the target link represents a link on which the switch chip extends a signal; When the error event is detected by a trigger, converting the error code into fault description information based on a pre-generated error code mapping table, wherein the trigger represents a trigger set by a target protocol analysis tool; storing the fault description information and the error log in association with each other in a storage partition corresponding to the port number in a storage medium; When the error event is detected by the trigger, before converting the error code into fault description information based on the pre-generated error code mapping table, the device is also used to: use the target protocol analysis tool to load the pre-configuration file; set the trigger for each port in the target link based on the pre-configuration file, including at least one of the following: select the trigger set for each port in the target link from a group of triggers configured according to preset rules based on the pre-configuration file; set the trigger for each port in the target link in a customized manner based on the pre-configuration file.
11. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the switching chip fault analysis method according to any one of claims 1 to 8 when executing the computer program.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the fault analysis method for a switching chip according to any one of claims 1 to 8.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the fault analysis method for a switching chip according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
SAS expander error code analysis method and device, equipment and storage medium
CN115437833A