Server link fault positioning system and method, electronic equipment and storage medium

By setting up a probe array with multiple data collection points on the server link, collecting and analyzing physical layer and protocol layer data, the fault location and type are automatically determined, solving the low positioning efficiency problem of traditional methods and achieving efficient troubleshooting.

CN120602390AActive Publication Date: 2025-09-05INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511056818.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-05
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Traditional GPU fault location methods rely on a single stress test, which cannot accurately locate the fault location and makes it difficult to distinguish between hardware link failures and software errors. Manual hardware replacement and troubleshooting are required, resulting in low troubleshooting efficiency.

Method used

By setting up a probe array with multiple data collection points on the server link, physical layer and protocol layer data are collected. Combined with the control module and analysis module, the fault location and type are automatically determined, including high-speed differential probes, impedance test probes and power supply noise probes. The time synchronization clock is used to align the data and generate topology maps and eye diagrams for fault annotation.

Benefits of technology

It realizes automated link quality assessment, can accurately locate the fault location and type, improves troubleshooting efficiency, and reduces manual intervention and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602390A_ABST
    Figure CN120602390A_ABST
Patent Text Reader

Abstract

The invention discloses a server link fault positioning system and method, electronic equipment and a storage medium, and relates to the technical field of server link tests.Physical layer data and protocol layer data of a data collection point are collected through a collection assembly, a test assembly is controlled to conduct server link testing, the test assembly comprises a control module and an analysis module, and the control module is connected with the analysis module. The test instruction is issued to the server link through the control module, when the server link fault is monitored, the analysis module acquires the physical layer data and the protocol layer data of each data acquisition point from the acquisition assembly, and determines the fault position and / or the fault type of the server link according to the acquired data, so that the automatic link quality evaluation is realized, and the efficiency is improved. And the fault position and the fault type can be determined, so that the problems of difficulty in accurately positioning the fault position and distinguishing the hardware link fault and the software error through a single pressure test and low troubleshooting efficiency caused by manual fault positioning in the prior art are solved, and the technical effect of improving the troubleshooting efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of server link testing, and in particular to a server link fault locating system, method, electronic device, and storage medium. Background Art

[0002] AI (Artificial Intelligence) servers, as core devices for high-performance computing, are widely used in fields such as artificial intelligence, scientific computing, and graphics rendering. They consist of modules such as GPUs (Graphics Processing Units), CPUs (Central Processing Units), memory, storage, and networking, relying on the collaborative work of these components to achieve efficient parallel computing. GPU link stability directly impacts computing efficiency, distributed training stability, and data I / O (Input / Output) performance, and is crucial for ensuring business continuity. Therefore, the quality of AI server links plays a crucial role in business operations.

[0003] However, traditional GPU fault location methods rely on a single stress test result and cannot distinguish between hardware link failures and software errors. They cannot distinguish between abnormal links between the CPU, server motherboard, SW (Switch Board), and GPU board. Manual hardware replacement and troubleshooting are required, resulting in low troubleshooting efficiency and prone to misjudgment, making it difficult to meet the operation and maintenance requirements of high-reliability AI server systems. Summary of the Invention

[0004] The present application provides a server link fault location system, method, electronic device and storage medium to at least solve the technical problems in the related art that it is difficult to accurately locate the fault location and distinguish between hardware link failures and software errors through a single stress test, and manual fault location is required, resulting in low troubleshooting efficiency.

[0005] The present application provides a server link fault location system, wherein the server link includes a central processing unit, a server mainboard, a switching component and a graphics processing component, and multiple data collection points are set on the server link, wherein the system includes: an acquisition component, wherein the acquisition component includes multiple probe arrays, the probe arrays are set at the data collection points, the probe arrays include a physical layer and a protocol layer, the physical layer collects physical layer data of the data collection points, and the protocol layer collects protocol layer data of the data collection points; a test component, wherein the test component includes a control module and an analysis module, the control module sends a test instruction to the server link, and when a server link fault is detected, the analysis module obtains the physical layer data and protocol layer data of each data collection point from the acquisition component, and determines the fault location and / or fault type of the server link based on the physical layer data and protocol layer data of each data collection point.

[0006] The present application also provides a server link fault location method, which is applied to the test component of the above-mentioned server link fault location system, wherein the method includes: issuing a test instruction to the server link; when a server link fault is detected, obtaining physical layer data and protocol layer data of each data collection point from the collection component; and determining the fault location and / or fault type of the server link based on the target physical layer data and target protocol layer data of each data collection point.

[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned server link fault locating method when executing the computer program.

[0008] The present application also provides a non-volatile computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned server link fault locating method are implemented.

[0009] Through this application, the physical layer data and protocol layer data of the data collection point can be collected by the collection component, and the test component can be controlled to perform server link testing. The test component includes a control module and an analysis module. The control module sends a test instruction to the server link. When a server link failure is detected, the analysis module obtains the physical layer data and protocol layer data of each data collection point from the collection component, and determines the fault location and / or fault type of the server link based on the acquired data, thereby realizing automated link quality evaluation and being able to determine the fault location and fault type. Therefore, it can solve the problem that the related technology is difficult to accurately locate the fault location and distinguish between hardware link failures and software errors through a single stress test, and manual fault location is required, resulting in low troubleshooting efficiency, thereby achieving the technical effect of improving troubleshooting efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A schematic diagram of the structure of a server link fault location system provided in an embodiment of the present application; Figure 2 A structural diagram of the main hardware links of a server provided in one embodiment of the present application; Figure 3 A structural diagram of a server link fault location system provided in one embodiment of the present application; Figure 4 A flowchart of the server link fault location system processing process provided by one embodiment of the present application; Figure 5 A flow chart of a server link fault locating method provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0013] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0014] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0015] Figure 1 This is a schematic diagram of the structure of the server link fault location system provided in the embodiment of the present application. Figure 1As shown, the system 10 includes: an acquisition component 101, multiple probe arrays 1011, a test component 102, a control module 1021 and an analysis module 1022; the server link 20 includes a central processing unit 201, a server motherboard 202, a switching component 203 and a graphics processing component 204.

[0016] First, the server link 20, which is the target of the embodiment of the present application, is described in detail. In the server link 20, the central processing unit 201 is the processing core of the server, and its main functions are task scheduling and management, logical processing, data pre-processing and post-processing, load sharing and hybrid computing, and fault tolerance processing. The server motherboard 202 is the core platform connecting all hardware components. Its design directly affects the performance, scalability and stability of the system and needs to be optimized for high-performance computing, large-scale parallel processing and multi-accelerator collaboration. The switching component 203 is a SW board used to achieve high-speed interconnection and communication, responsible for managing data flow, optimizing communication efficiency, and reducing latency. The graphics processing component 204 is a GPU board, which is a core hardware module that carries the graphics processor and is responsible for executing large-scale parallel computing tasks. The server link 20 includes links connecting the central processing unit 201, the server motherboard 202, the switching component 203 and the graphics processing component 204. The components are connected to each other through standards such as PCIe (Peripheral Component Interconnect Express) and OCulink (Optical Copper Link) according to actual needs.

[0017] Secondly, an embodiment of the present application is described in detail. In the embodiment of the present application, the acquisition component 101 includes multiple probe arrays 1011, and the probe arrays 1011 are arranged at data acquisition points. The probe arrays 1011 include a physical layer and a protocol layer. The physical layer collects physical layer data of the data acquisition point, and the protocol layer collects protocol layer data of the data acquisition point; the test component 102 includes a control module 1021 and an analysis module 1022. The control module 1021 sends a test instruction to the server link 20. When a fault in the server link 20 is detected, the analysis module 1022 obtains the physical layer data and protocol layer data of each data acquisition point from the acquisition component 101, and determines the fault location and / or fault type of the server link 20 based on the physical layer data and protocol layer data of each data acquisition point.

[0018] The physical layer data of the probe array 1011 and the data collection points will be described in detail below and will not be repeated here. The protocol layer data of the data collection points include the data packet type (such as TLP (Transaction Layer Packet), DLLP (Data Link Layer Packet)), packet sequence number, number of retransmissions, etc. The test instruction issued by the control module 1021 is a graphics processor stress test instruction. After receiving the test instruction, the server link 20 triggers a stress test on the graphics processor carried by the graphics processing component 204, such as 3Dmark (a widely used benchmark software that can be used to evaluate the performance of the graphics processor, including multiple test scenarios to simulate different types of workloads) stress test mode or NVIDIA-SIM (a command-line utility that provides management and monitoring functions for the graphics processor) stress test instruction. The fault location of the server link 20 refers to the location of the faulty link segment of the server link 20. The fault types of the server link 20 include hardware link failure and software error.

[0019] It can be understood that the acquisition component 101 of the embodiment of the present application includes multiple probe arrays 1011, which are set at different data acquisition points. Each probe array 1011 is composed of a physical layer and a protocol layer, wherein the physical layer is used to collect physical layer data of the corresponding data acquisition point, and the protocol layer is used to collect protocol layer data of the corresponding data acquisition point. The test component 102 includes a control module 1021 and an analysis module 1022. When it is necessary to test the server link 20, the control module 1021 sends a test instruction to the server link 20. The test instruction is a graphics processor stress test instruction, such as a 3DMark stress test mode or an NVIDIA-SMI stress test instruction. After receiving the test instruction, the server link 20 triggers a stress test on the graphics processor. The control module 1021 continuously monitors the server link 20. When a fault is detected in the server link 20, the analysis module 1022 obtains the physical layer data and protocol layer data of each data acquisition point from the acquisition component 101, and determines the specific location and / or fault type of the faulty link segment of the server link 20 based on these data.

[0020] In an embodiment of the present application, the data collection points include multiple collection points of the link between the central processing unit 201 and the server motherboard 202, the collection points of the link between the server motherboard 202 and the switching component 203, the collection points of the link between the switching component 203 and the graphics processing component 204, and the collection points at the power supply port of the graphics processing component 204.

[0021] Among them, the data collection points can be set at the PCIe slot of the server motherboard 202, the output and input ports of the switching component 203, the gold fingers of the graphics processing component 204, and the power supply detection points of the graphics processing component 204, etc. Among them, the gold fingers of the graphics processing component 204 are a series of contact points coated with a layer of conductive material on the edge of the graphics processing component 204 circuit board. These contact points are designed to be inserted into the expansion slot (such as the PCIe slot) on the server motherboard 202, thereby realizing data transmission and power supply between the graphics processing component 204 and other parts of the computer.

[0022] It can be understood that the data collection points of the embodiment of the present application include multiple locations, specifically: the collection point of the link between the central processing unit 201 and the server motherboard 202, the collection point of the link between the server motherboard 202 and the switching component 203, the collection point of the link between the switching component 203 and the graphics processing component 204, and the collection point at the power supply port of the graphics processing component 204, etc. These data collection points can be respectively set at the PCIe slot of the server motherboard 202, the input and output ports of the switching component 203, the gold finger of the graphics processing component 204, and the power supply detection point of the graphics processing component 204.

[0023] In an embodiment of the present application, the physical layer of the probe array 1011 includes a differential probe, an impedance test probe and a power supply noise probe, wherein the high-speed differential probe is connected to the collection point of the link between the central processing unit 201 and the server motherboard 202, the collection point of the link between the server motherboard 202 and the switching component 203, and the collection point of the link between the switching component 203 and the graphics processing component 204 to measure the first bandwidth data of each link segment of the server link 20; the impedance test probe is connected to the collection point of the link between the server motherboard 202 and the switching component 203 and the collection point of the link between the switching component 203 and the graphics processing component 204 through the RF coaxial connector to measure the impedance resolution data of the time domain reflectometer; the power supply noise probe is connected to the collection point at the power supply port of the graphics processing component 204 to measure the second bandwidth data and the noise floor data.

[0024] Among them, the high-speed differential probe is a test device for measuring high-speed signals, which can support a very high frequency range and can accurately capture first bandwidth data through its high bandwidth capability; the impedance test probe is a tool designed to be connected to a circuit or device and can be used to detect its impedance characteristics; the RF coaxial connector is a connection device for transmitting RF signals, has good shielding performance, can effectively reduce the influence of external electromagnetic interference, and is widely used in occasions where high-frequency signals need to be processed. The RF coaxial connector of the embodiment of the present application can use an SMA connector (SubMiniature version A RF coaxial connector); the time domain reflectometer impedance resolution data, namely TDR (Time Domain Reflectometry, time domain reflectometer) impedance resolution, can distinguish the minimum distance or minimum impedance change between two adjacent impedance mutation points; the power supply noise probe is connected to the collection point at the power supply port of the graphics processing component 204 in parallel with the power supply port of the graphics processing component 204.

[0025] It can be understood that the physical layer of the probe array 1011 of the embodiment of the present application includes a high-speed differential probe, an impedance test probe and a power supply noise probe, wherein the high-speed differential probe is connected to the collection point of the link between the central processing unit 201 and the server motherboard 202, the collection point of the link between the server motherboard 202 and the switch component 203, and the collection point of the link between the switch component 203 and the graphics processing component 204, for measuring the first bandwidth data of each link segment of the server link 20; the impedance test probe is connected to the collection point of the link between the server motherboard 202 and the switch component 203 and the collection point of the link between the switch component 203 and the graphics processing component 204 through an RF coaxial connector, for measuring the impedance resolution data of the time domain reflectometer, which represents the minimum distance or minimum impedance change between two adjacent impedance mutation points; the power supply noise probe is connected to the collection point at the power supply port of the graphics processing component 204, and the access method is to be connected in parallel to the power supply port, for measuring the second bandwidth data and noise floor data. The first bandwidth data, the impedance resolution data of the time domain reflectometer, the second bandwidth data and the noise floor data constitute the required physical layer data.

[0026] In an embodiment of the present application, a time synchronization clock is provided on the physical layer and the protocol layer, and the data collection time of the physical layer data and the protocol layer is aligned based on the time synchronization clock.

[0027] Among them, after aligning the data collection time of the physical layer data and the protocol layer based on the time synchronization clock, it is also necessary to check whether the data collection time error of the aligned physical layer data and the protocol layer is less than or equal to the preset time error threshold. If it is greater than the preset time error threshold, time alignment is required again. The time error threshold is set according to actual needs and is not specifically limited here.

[0028] It can be understood that a time synchronization clock is provided on both the physical layer and the protocol layer of the embodiment of the present application, which is used to align the data acquisition time of the physical layer data and the protocol layer. Moreover, after the time alignment is completed based on the time synchronization clock, it is necessary to further check whether the time error between the aligned physical layer data and the protocol layer data is less than or equal to a preset time error threshold (for example, 100ns); if the time error is greater than the time error threshold, the time alignment operation needs to be re-executed to ensure that the physical layer and protocol layer data remain highly synchronized in the time dimension.

[0029] In an embodiment of the present application, the control module 1021 sends a fault instruction to the acquisition component 101; the acquisition component 101 responds to the fault instruction and obtains physical layer data and protocol layer data within a target time period, wherein the target time period is the time period before the fault is sent.

[0030] Among them, the fault instruction is issued when the control module 1021 detects a fault in the server link 20; the interception of the target time period is achieved by presetting the target time period length. The preset target time period length is specifically set according to actual needs and is not specifically limited here. For example, the preset target time period length is 30s. When the acquisition component 101 responds to the fault instruction and selects the physical layer data and the protocol layer data, it starts from the moment the fault is sent and goes back 30s as the target time period.

[0031] It should be noted that the control module 1021 monitors the server link 20. When it detects that the graphics processor in the fixed slot fails, it locates the abnormal bus identifier of the device in the server link 20. The bus identifier is the bus ID (Identifier). When the bus identifier is abnormal, it is judged that the server link 20 is faulty. Among them, the bus identifier abnormality includes: bus identifier jump, upstream node loss, device category code change, etc. The bus identifier jump refers to the sudden change of the originally stable bus identifier, indicating that there may be a hardware failure, link instability or configuration error; upstream node loss means that the device cannot communicate through its directly connected upper node; the device category code change means that the control module 1021 has recognized the different type or function of the device, even if the actual hardware has not changed, usually due to firmware damage, accidental modification of the configuration register or driver error.

[0032] It can be understood that when the control module 1021 of the embodiment of the present application detects an abnormality in the device bus identifier in the server link 20, it determines that the server link 20 has failed and sends a fault instruction to the acquisition component 101; after the acquisition component 101 responds to the fault instruction, it intercepts a time period of the preset target time period length before the time when the fault occurs as the target time period according to the preset target time period length. For example, the preset target time period length is 30 seconds, that is, starting from the time when the fault occurs, it traces back 30 seconds as the target time period, and simultaneously obtains the physical layer data and protocol layer data within the target time period, and uses the physical layer data and protocol layer data within the target time period as the analysis basis of the analysis module 1022.

[0033] In an embodiment of the present application, the test component 102 also includes a display, on which an interactive interface is provided, on which the topology map and eye diagram of the server link 20 are displayed; the analysis module 1022 renders the topology map and eye diagram of the server link 20 based on the physical layer data and protocol layer data of the data acquisition point, and marks the topology map and eye diagram of the server link 20 based on the fault location and fault type.

[0034] Among them, the topology diagram of the server link 20 depicts the relationship between all hardware devices and their connection methods between the central processing unit 201 and the graphics processing component 204 in the server link 20; the eye diagram is obtained through physical layer data, which is formed by superimposing multiple cycles of signal waveforms on a chart to form an "eye" shape pattern, thereby providing information about signal quality; the topology diagram and eye diagram of the server link 20 are marked based on the fault location and fault type, which means that the server link segment at the fault location is marked with different colors on the topology diagram. The different colors refer to colors different from those displayed in the topology diagram when the link is in normal state, thereby distinguishing the faulty link segment.

[0035] It can be understood that the test component 102 of the embodiment of the present application also includes a display, which is provided with an interactive interface for displaying the topology map and eye diagram of the server link 20. The topology map and eye diagram are obtained by analyzing the physical layer data and protocol layer data of the data acquisition point by the analysis module 1022, and the analysis module 1022 will mark the topology map of the server link 20 based on the location and type of the fault, and finally display it intuitively on the display. Specifically, the topology map of the server link 20 describes in detail the relationship between all hardware devices and their connection methods between the central processing unit 201 and the graphics processing component 204; the eye diagram is generated by superimposing multiple cycles of signal waveforms to provide information about signal quality. When a fault is detected, the analysis module 1022 will highlight the server link segment at the fault location on the topology map in a color different from that of the normal link, so that testers can quickly identify the problem link segment.

[0036] According to the embodiments of the present application, the physical layer data and protocol layer data of the data collection point can be collected by the collection component, and the test component can be controlled to perform server link testing. The test component includes a control module and an analysis module. The control module sends a test instruction to the server link. When a server link failure is detected, the analysis module obtains the physical layer data and protocol layer data of each data collection point from the collection component, and determines the fault location and / or fault type of the server link based on the obtained data, thereby realizing automated link quality evaluation, and being able to determine the fault location and fault type, achieving the technical effect of improving troubleshooting efficiency.

[0037] The server link fault location system is further described below through a specific embodiment.

[0038] First, the main hardware link components of the server in this embodiment are described in detail. Figure 2 As shown, the server of this embodiment is an AI server, and its main hardware links include a central processing unit (CPU), a server motherboard, a switch component (SW) board, and a graphics processing component (GPU) board. Specifically: The central processing unit (CPU) is the brain of the server. Its main functions include task scheduling and management, logical processing, data pre-processing and post-processing, load sharing and hybrid computing, and fault-tolerant processing.

[0039] The server motherboard is the core platform that connects all hardware components. Its design directly impacts system performance, scalability, and stability. Compared to standard server motherboards, AI server motherboards must be optimized for HPC (High-Performance Computing), massively parallel processing, and multi-accelerator collaboration. As the core carrier of high-performance computing, AI server motherboards must strike a balance between multi-accelerator support, high-bandwidth interconnection, reliability, and scalability.

[0040] The switch board, or SW board, is a key component for high-speed interconnection and communication. It manages data flow, optimizes communication efficiency, and reduces latency, especially in multi-GPU / multi-accelerator scenarios. The SW board is a critical hub for high-performance computing in AI servers, and its design directly impacts the collaborative efficiency of multiple GPUs / multi-accelerators.

[0041] The graphics processing unit (GPU) is the core hardware module that hosts the graphics processing unit (GPU) or dedicated AI accelerator. It is responsible for executing large-scale parallel computing tasks (such as deep learning training / inference, scientific computing, etc.). The core functions of the GPU board include parallel computing acceleration, multi-GPU collaborative expansion, high-bandwidth memory support, and energy efficiency and thermal management. The GPU board is the core carrier of AI computing power and is designed around high parallel computing, high-bandwidth interconnection, and reliable operation.

[0042] PCIe is a high-speed serial computer expansion bus standard. The CPU's built-in PCIe controller manages device enumeration and address allocation. PCIe also connects various boards for communication services. For example, in this example, the SW board and GPU board are connected using a high-speed PCIe interface. PCIe is the backbone bus connecting computing, storage, and networking in AI servers. Its bandwidth and latency directly impact the efficiency of multi-GPU collaboration. Figure 2 In the IEEE 802.11 specification, PCIe ×16 is a link width configuration of PCIe, indicating that the PCIe link consists of 16 physical channels. PCIe Gen4 is the fourth generation of Peripheral Component Interconnect Express bus (PCI Express Generation 4).

[0043] OCuLink is a high-speed external expansion interface standard based on the PCIe protocol and connected via copper or fiber cables. It is primarily used to address the bandwidth and latency limitations of traditional interfaces such as USB.

[0044] The main components of this embodiment are as follows Figure 3 Shown, including: Control module: Mainly used to trigger GPU stress tests, such as 3DMark stress test mode or NVIDIA-SIM stress test, and monitor the server topology model, namely the bus ID-physical link mapping table.

[0045] Acquisition components: multi-level link probes, including motherboard PCIe slot probes / SW board input and output port probes / GPU board gold finger probes and GPU board power supply detection points, etc. It is mainly a PCIe bus probe array, including physical layer and protocol layer. The physical layer measurement probes mainly include: (1) high-speed differential probes for measuring bandwidth, welded to PCIe Lane test points (2) impedance test probes for measuring TDR resolution, connected through SMA connectors (between motherboard, SW board, GPU) (3) power supply noise probes for measuring bandwidth and noise floor, connected in parallel to the GPU board 12V power supply line. The physical layer measures impedance / eye diagrams, and the protocol layer parses TLP packets, including abnormal types such as link training failure, DLLP check errors, and transaction layer delays. At the same time, the physical layer and protocol layer set time synchronization clocks to ensure that the time alignment error is less than 100ns. The above collected data are summarized as link quality data at all levels.

[0046] Analysis module: In the event of an anomaly, it mainly performs bus ID topology analysis and link quality analysis at all levels. It can also perform machine learning predictions. Machine learning is an artificial intelligence technology that enables computers to learn from data and improve their performance without explicit programming. It can be mainly divided into supervised learning, unsupervised learning, and semi-supervised learning.

[0047] Display: This is a web-based visualization interface. It provides real-time bus topology and dynamic eye diagram rendering. It includes an interactive interface that allows you to configure some CLI (Command-Line Interface) console functions. It also displays faulty links, marks them in red, generates diagnostic reports, and provides maintenance recommendations.

[0048] as follows Figure 4 This is a flowchart of the specific implementation and processing process of this embodiment.

[0049] (1) The control module triggers GPU stress testing (such as dcgmi stress –memtest, dcgmi is a tool set for managing and monitoring GPUs in data centers, dcgmi stress –memtest is a dcgmi control instruction for stress testing; or NVIDIA-SMI stress testing instruction), monitors the complete bus ID topology structure and transmits it to the display and data acquisition components through signals.

[0050] (2) When a fixed-slot GPU failure is detected, the current device tree is generated through lspci -tvnn (lspci is a command-line tool used to display detailed information of all PCI (Peripheral Component Interconnect) devices. The lspci -tvnn command provides a tree view that shows the hierarchical relationship between devices) and compared with the baseline topology map to locate the abnormal bus ID. The collection component automatically traces back the link quality data 30 seconds before the failure (bandwidth information of each link segment, impedance deviation rate, eye diagram, and PCIe TLP packet retransmission rate, etc.).

[0051] (3) The analysis module uses the fault correlation engine to analyze the topology and quality of each link based on the data collected by the acquisition component. This includes comparing the abnormal topology with the baseline topology, analyzing the link quality quantitative model, and predicting the link quality through machine learning methods. The details are as follows: Define a quality score Q for each link:

[0052] in The impedance deviation rate refers to the percentage fluctuation of the measured link impedance compared to the standard impedance. The eye diagram signal jitter refers to the deviation of the eye diagram signal edge from the ideal position, which can be obtained by analyzing the eye diagram with a jitter analyzer. The power supply fluctuation refers to the percentage fluctuation of the measured power supply port voltage compared to the standard voltage. The lower the quality score Q, the worse the link quality. (The value of this embodiment is ).

[0053] The main parameters of link monitoring are shown in Table 1, where Table 1 is a table of the main parameters of link monitoring and the corresponding fault threshold ranges.

[0054] Table 1

[0055] In Table 1, the PCIe lane width negotiation state refers to the process by which PCIe devices negotiate the optimal communication channel width (e.g., x1, x2, x4, x8, x16, where xm indicates the use of m physical channels for communication) between the devices through link training. This ensures that devices at both ends can transmit data at the optimal bandwidth. The differential signal eye diagram opening area can typically be obtained using an oscilloscope using the first bandwidth data. The smaller the differential signal eye diagram opening area compared to the standard eye diagram area, the greater the signal distortion. Conversely, the larger the differential signal eye diagram opening area compared to the standard eye diagram area, the less interference or attenuation the signal is experiencing.

[0056] Each parameter is compared with the corresponding fault threshold range. If any parameter is within the fault threshold range, it indicates that a link fault has occurred. The fault threshold range is set according to actual needs and is not specifically limited here. In Table 1, the fault threshold range is set as follows: if the PCIe channel width negotiation state is in non-x16 mode for more than 5 seconds, the link is considered faulty; if the differential signal eye diagram opening area is less than 60% of the standard eye diagram area, the link is considered faulty; if the absolute value of the impedance deviation rate is greater than 5%, the link is considered faulty; if the absolute value of the power supply fluctuation is greater than 5%, the link is considered faulty.

[0057] (4) Output the diagnostic report and generate a fault probability distribution diagram based on the quality score of each link. For example: CPU-server motherboard link: 12%; server motherboard-SW board link: 83%, then the server motherboard-SW board link is identified as the faulty segment; SW board-GPU board link: 5%; GPU body: 0%, and output the maintenance recommendations.

[0058] For example, in an 8-GPU server, the GPU fails a continuous stress test. Traditional solutions require replacing all GPU boards for troubleshooting. However, with the system of this embodiment, there is no need to replace all GPU boards for troubleshooting. Specifically: (1) Using lspci, it was found that the bus ID of the faulty GPU changed to 04:00.0 (the normal value should be 03:00.0); (2) Trace the upstream bus ID 02:00.1 corresponding to SW board port 4; (3) The impedance test probe measured the link impedance curve under normal test conditions to be 85Ω±2Ω, and under abnormal conditions to be 112Ω (exceeding the standard). At the same time, the delay abnormality of the protocol layer signal of this link further confirmed the abnormal link; (4) It was confirmed that the trace of the 4th channel PCB (Printed Circuit Board) of the SW board was broken.

[0059] Next, the server link fault location method provided by the embodiment of the present application is described with reference to the accompanying drawings. Figure 5 A flow chart of a server link fault location method provided in an embodiment of the present application is provided. The method is applied to the test component of the above-mentioned server link fault location system, such as Figure 5 As shown, the method includes the following steps: In step S301, a test instruction is sent to the server link.

[0060] The test instruction is issued by the control module of the embodiment of the present application; the test instruction is a stress test instruction for the graphics processor.

[0061] It can be understood that the embodiment of the present application first sends a test instruction to the server link through the control module. After the server link receives the instruction, it triggers a stress test on the graphics processor. At the same time, the control module monitors the server link in real time.

[0062] In step S302, when a server link failure is detected, the physical layer data and protocol layer data of each data collection point are acquired from the collection component.

[0063] Among them, monitoring a server link failure means that the control module detects the failure of the graphics processor in the fixed slot of the server link. At this time, the control module locates the abnormal bus identifier of the device in the server link and sends a fault instruction to the acquisition component; obtaining the physical layer data and protocol layer data of each data acquisition point from the acquisition component means that the analysis module of the embodiment of the present application obtains the physical layer data and protocol layer data of each data acquisition point by the acquisition component. The data obtained is the physical layer data and protocol layer data within the target time period obtained after the acquisition component responds to the fault instruction, and according to the preset target time period length, the time period of the preset target time period length before the time when the fault occurs is intercepted as the target time period.

[0064] It can be understood that when the control module of the embodiment of the present application detects that the graphics processor in the fixed slot of the server link fails, it locates the abnormal bus identifier of the device in the server link and sends a fault instruction to the acquisition component. The analysis module obtains the physical layer data and protocol layer data of each data acquisition point within the target time period from the acquisition component. For example, after the acquisition component responds to the fault instruction, the preset target time period length is 30 seconds, that is, starting from the time the fault occurs, it traces back 30 seconds as the target time period, obtains the physical layer data and protocol layer data within the target time period, and uses the physical layer data and protocol layer data within the target time period as the analysis basis of the analysis module.

[0065] In step S303, the fault location and / or fault type of the server link is determined based on the physical layer data and protocol layer data of each data collection point.

[0066] The server link fault location refers to the location of the faulty link segment of the server link. Server link fault types include hardware link faults and software errors.

[0067] It can be understood that the embodiment of the present application can analyze and locate the location of the server link fault link segment through the physical layer data and protocol layer data of each data collection point, and can distinguish the types of server link faults without replacing all graphics processing components for troubleshooting, thereby achieving the technical effect of improving troubleshooting efficiency.

[0068] In an embodiment of the present application, the fault location and / or fault type of the server link is determined based on the physical layer data and protocol layer data of each data collection point, including: calculating the monitoring parameters of each link segment of the server link through the physical layer data; determining the fault link of the server link based on the monitoring parameters of each link segment and the corresponding fault threshold, and locating the position of the fault link according to the identification of the data collection point; and determining the fault type of the server link based on the protocol layer data.

[0069] Among them, the monitoring parameters of each link segment include PCIe channel width negotiation status, differential signal eye closure, impedance deviation rate and power supply fluctuation, which are calculated based on the physical layer data of each data collection point. Specifically, PCIe channel width negotiation status refers to the process of PCIe devices negotiating the optimal communication channel width (for example, x1, x2, x4, x8, x16) between devices through link training, which can ensure that the devices at both ends can transmit data with the optimal bandwidth; the differential signal eye diagram can usually be obtained from the first bandwidth data through an oscilloscope. The wider the "eye" in the differential signal eye diagram, that is, the larger the differential signal eye diagram opening area, the smaller the signal distortion; conversely, the smaller the differential signal eye diagram opening area, the greater the signal interference or attenuation; the impedance deviation rate refers to the percentage fluctuation of the measured link impedance compared to the standard impedance; the power supply fluctuation refers to the percentage fluctuation of the measured power supply port voltage compared to the standard voltage; the fault threshold corresponding to the monitoring parameters of each link segment is set according to actual needs and is not specifically limited here. For example, PCIe The fault threshold for the channel width negotiation status is set to a non-x16 mode duration greater than 5s, which is considered a fault; the fault threshold for the differential signal eye diagram is set to a differential signal eye diagram opening area less than 60% of the standard eye diagram area, which is considered a fault; the fault threshold for the impedance deviation rate is set to an absolute value greater than 5%, which is considered a fault; the fault threshold for power supply fluctuation is set to an absolute value greater than 5%, which is considered a fault; the data collection point is identified by the bus identifier of the data collection point; the server link fault type is determined based on the protocol layer data, and by analyzing the protocol layer data, it is determined that there is a software problem, such as excessive retransmissions, TLP packet loss or format errors.

[0070] It can be understood that the embodiment of the present application can calculate the monitoring parameters of each link segment, such as PCIe channel width, eye diagram quality, bit error rate and power supply fluctuation, by collecting and analyzing the physical layer and protocol layer data of each data collection point in the server link, and compare it with the preset fault threshold to determine whether there is a faulty link. If a link is determined to be faulty, the specific position of the faulty link is located in combination with the bus identifier at the data collection point. At the same time, by analyzing the protocol layer data (such as TLP packet loss, number of retransmissions, etc.), it can be further identified whether the link failure is caused by a software error, thereby distinguishing the fault type of the server link and realizing efficient fault detection and diagnosis of the server link.

[0071] In an embodiment of the present application, the fault location and / or fault type of the server link is determined based on the physical layer data and protocol layer data of each data collection point, and the method also includes: calculating the quality parameters of each section of the server link through the target physical layer data; calculating the quality score of each section of the server link through a quality scoring formula based on the quality parameters of each section of the server link; analyzing the quality of each section of the server link based on the quality score; and generating a fault probability distribution map based on the quality of each section of the server link.

[0072] The formula for calculating the quality parameter of each segment of the server link is:

[0073] in , The value of is set according to actual needs and is not specifically limited here; the lower the quality score Q, the worse the quality of the link segment.

[0074] It can be understood that the embodiment of the present application collects and analyzes the physical layer data of each segment in the server link, extracts key parameters reflecting the communication quality, such as impedance deviation rate, eye diagram quality, power supply fluctuation, etc., and uses the set quality scoring formula to evaluate each link segment to obtain its link quality status. On this basis, a fault probability distribution map of each link segment is further drawn to visually display which link segments have a higher failure risk, helping to achieve refined monitoring and fault warning of server links.

[0075] According to the server link fault locating method provided in the embodiment of the present application, a test instruction is sent to the server link through the control module. When a server link fault is detected, the analysis module obtains the physical layer data and protocol layer data of each data collection point from the collection component, and determines the fault location and / or fault type of the server link based on the obtained data, thereby realizing automated link quality evaluation, and being able to determine the fault location and fault type, achieving the technical effect of improving troubleshooting efficiency.

[0076] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0077] For the description of the features in the embodiment corresponding to the server link fault locating method, please refer to the relevant description of the embodiment corresponding to the server link fault locating system, which will not be repeated here.

[0078] The embodiment of the present application also provides an electronic device, such as Figure 6As shown, it includes a memory 401 and a processor 402. The memory 401 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in the above-mentioned server link fault location method embodiment.

[0079] An embodiment of the present application further provides a non-volatile computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in the above-mentioned server link fault location method embodiment when running.

[0080] In an exemplary embodiment, the non-volatile computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0081] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0082] The above is a detailed introduction to a server link fault location system, method, electronic device and storage medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A server link fault location system, characterized in that: The server link includes a central processing unit, a server motherboard, a switching component, and a graphics processing component. Multiple data collection points are set on the server link, wherein the system includes: A collection component, wherein the collection component includes a plurality of probe arrays, the probe arrays are arranged at the data collection points, the probe arrays include a physical layer and a protocol layer, the physical layer collects physical layer data of the data collection points, and the protocol layer collects protocol layer data of the data collection points; A test component, wherein the test component includes a control module and an analysis module, the control module sends a test instruction to the server link, and when a server link failure is detected, the analysis module obtains the physical layer data and protocol layer data of each data collection point from the collection component, and determines the fault location and / or fault type of the server link based on the physical layer data and protocol layer data of each data collection point.

2. The server link fault location system according to claim 1, characterized in that: The data collection points include multiple collection points of the link between the central processing unit and the server motherboard, the collection points of the link between the server motherboard and the switching component, the collection points of the link between the switching component and the graphics processing component, and the collection points at the power supply port of the graphics processing component.

3. The server link fault location system according to claim 2, characterized in that: The physical layer of the probe array includes a high-speed differential probe, an impedance test probe and a power supply noise probe, wherein: The high-speed differential probe is connected to a collection point of the link between the central processing unit and the server mainboard, a collection point of the link between the server mainboard and the switch component, and a collection point of the link between the switch component and the graphics processing component to measure first bandwidth data of each link segment of the server link; The impedance test probe is connected to the collection point of the link between the server motherboard and the switching component and the collection point of the link between the switching component and the graphics processing component through a radio frequency coaxial connector to measure the impedance resolution data of the time domain reflectometer; The power supply noise probe is connected to a collection point at a power supply port of a graphics processing component to measure second bandwidth data and noise floor data.

4. The server link fault location system according to claim 1, characterized in that: The physical layer and the protocol layer are provided with a time synchronization clock, and the data collection time of the physical layer data and the data collection time of the protocol layer are aligned based on the time synchronization clock.

5. The server link fault location system according to claim 1, characterized in that: The control module sends a fault instruction to the acquisition component; The acquisition component responds to the fault instruction to obtain the physical layer data and the protocol layer data within a target time period, wherein the target time period is a time period before the fault is sent.

6. The server link fault location system according to claim 1, characterized in that: The test component also includes a display, which is provided with an interactive interface, and the interactive interface displays the topology map and eye diagram of the server link; the analysis module renders the topology map and eye diagram of the server link based on the physical layer data and protocol layer data of the data acquisition point, and marks the topology map and eye diagram of the server link based on the fault location and fault type.

7. A server link fault location method, characterized in that: The method is applied to the test component of the server link fault location system according to any one of claims 1 to 6, wherein the method comprises: Send test instructions to the server link; When a server link failure is detected, obtaining physical layer data and protocol layer data of each data collection point from the collection component; The fault location and / or fault type of the server link is determined based on the physical layer data and the protocol layer data of each data collection point.

8. The server link fault location method according to claim 7, characterized in that: The determining the fault location and / or fault type of the server link according to the physical layer data and the protocol layer data of each data collection point includes: Calculating monitoring parameters of each link segment of the server link using the physical layer data; Determining a faulty link of the server link based on the monitoring parameters of each link segment and the corresponding fault threshold, and locating the position of the faulty link based on the identifier of the data collection point; A fault type of the server link is determined based on the protocol layer data.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server link fault locating method according to any one of claims 7 to 8 when executing the computer program.

10. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server link fault locating method according to any one of claims 7 to 8 are implemented.

Citation Information

Patent Citations

  • Multiple path servers fault locating system

    CN106815108A

  • Fault processing method and device, electronic equipment and storage medium

    CN113918375A

  • Fault positioning method and device based on hardware link, equipment and medium

    CN118631641A

  • Judgment method and device for link between switch and server

    CN119052128A

  • Link test method, electronic device, storage medium, product and computing device

    CN119473744A

Cited By

  • Dynamic partition scheduling method, system and equipment for heterogeneous computing resources of server

    CN120950268A

  • Server heterogeneous computing resource dynamic partition scheduling method, system and device

    CN120950268B

  • Graphics processor risk processing method and device, equipment and medium

    CN121256232A

  • Data interface communication protocol dynamic adaptation method and system

    CN122093486A

  • Method for handling link failure, monitor, and server

    CN122450728A