Method and apparatus for determining faulty device, and non-volatile readable storage medium and electronic device

By finding the target device identifier in the preset N+M link information and determining its location, the problem of inaccurate fault positioning of PCIE devices with Switch boards in the prior art is solved, and fast and accurate fault positioning and repair are achieved.

WO2025130240A1PCT designated stage expired Publication Date: 2025-06-26INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
PCT/CN2024/122097
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-09-29
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately locate PCIE equipment failures for AI models with Switch boards. Usually, they can only locate silk screens on the motherboard, but cannot directly locate silk screens on the Switch boards, resulting in difficulty in operation and maintenance.

Method used

By obtaining the target device identification of the target fault device, the link information including the identification is found in the preset N+M link information, whether the target fault device is located between the processor and the first switch, and whether the target fault device is a device on the switch board is determined based on the preset record item information and the target equipment identification.

Benefits of technology

It realizes accurate positioning of PCIE equipment failures with AI models with Switch boards, reduces troubleshooting time, and improves operation and maintenance efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024122097_26062025_PF_FP_ABST
    Figure CN2024122097_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a method and apparatus for determining a faulty device, and a non-volatile readable storage medium and an electronic device. The method for determining a faulty device comprises: acquiring a target device identifier of a target faulty device; searching N+M pieces of preset link information for link information which comprises the target device identifier; when jth link information which comprises the target device identifier is found from the N+M pieces of preset link information and jth indication information in the jth link information indicates that there is a first switch on a jth device link which is formed between a processor and a jth connection device, determining whether the target faulty device on the jth device link is located between the processor and the first switch; and when the target faulty device is not located between the processor and the first switch, on the basis of preset recorded-item information and the target device identifier, determining whether the target faulty device is a device on a switch board.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for determining faulty equipment, non-volatile readable storage medium, and electronic device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311747507.9, entitled “Method and device for determining faulty equipment, storage medium and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The embodiments of the present application relate to the field of computers, and more specifically, to a method and apparatus for determining a faulty device, a non-volatile readable storage medium, and an electronic device. Background Art

[0004] The Basic Input and Output System (BIOS) is a set of programs embedded in a ROM (Read-Only Memory) chip on a computer's motherboard. It stores the computer's most important power-on self-test, hardware initialization routines, and low-level system services. PCI Express (PCI Express) is a high-speed serial bus technology used to connect computer motherboards and other devices, such as graphics cards, network drives, NVME (Non-Volatile Memory Express) drives, and GPUs (Graphics Processing Units). PCI devices typically offer higher data transfer speeds and bandwidth, providing better performance and scalability than traditional PCI buses. In the PCI depth-first algorithm, the SecBus (Secondary Bus) is often referred to as the "secondary bus" and is used to identify the number or ID of switches within the PCI bus. Each switch has a unique SecBus number. The SubBus (Subordinate Bus) is often referred to as the "slave bus" and is used to identify the number or ID of devices within the PCI bus. When multiple devices are connected to a switch, each device has a unique SubBus number. In summary, SecBus is used to identify the switch, while SubBus is used to identify the device. These two concepts are used in the PCIE depth-first algorithm to determine the hierarchical relationship between various devices and switches on the PICe bus to achieve data transmission and management. BMC is the abbreviation of Baseboard Management Controller, which is an independent chip or integrated circuit located on the computer motherboard, which is configured to monitor and manage the hardware and software of the computer system.

[0005] Related technologies can accurately locate faults on standard models and standard PCIE devices, but they are less accurate for AI models with switch boards. Often, the fault can only be located on the silkscreen on the motherboard, or not at all. Direct locating the fault is impossible. For Smart NICs with multiple virtual network ports, due to the limited storage and processing capabilities of the BMC, when a Bus Device Function (BDF) on a Smart NIC fails, it's often difficult to pinpoint the device causing the fault. This inability to accurately locate the fault makes it difficult for maintenance personnel to quickly perform repairs in some scenarios.

[0006] Regarding the related technologies, the related technologies can accurately locate the faults of ordinary PCIE devices, but the technical problem of inaccurately locating the faults of AI models with Switch boards has not yet been proposed. An effective solution has not yet been proposed.

[0007] Summary of the Invention

[0008] The embodiments of the present application provide a method and apparatus for determining a faulty device, a non-volatile readable storage medium, and an electronic device, to at least solve the problem in the related art that the related art can accurately locate the faults of ordinary PCIE devices, but is not accurate enough in locating the faults of AI models with switch boards.

[0009] According to one embodiment of the present application, a method for determining a faulty device is provided, comprising: obtaining a target device identifier of a target faulty device; searching for link information including the target device identifier in preset N+M link information, wherein the N+M link information has a one-to-one correspondence with N+M connection devices, the N+M connection devices include N connection devices on a mainboard and M connection devices on a switch board, each of the M connection devices is connected to one of the P switches on the switch board, the i-th link information in the N+M link information includes the device identifiers of multiple devices on the i-th device link formed from the processor on the mainboard to the i-th connection device, N, M, and P are all positive integers, i and j are less than or equal to N +M positive integer; when the j-th link information including the target device identifier is found in the N+M link information, and the j-th indication information in the j-th link information indicates that a first switch exists on the j-th device link formed from the processor to the j-th connection device, determining whether the target faulty device is located between the processor and the first switch on the j-th device link, wherein the P switches include the first switch; when the target faulty device is not located between the processor and the first switch, determining whether the target faulty device is a device on the switch board according to preset record item information and the target device identifier, wherein the record item information includes the device identifier of each switch in the P switches.

[0010] In an exemplary embodiment, after searching for link information including the target device identifier in the preset N+M link information, the method further includes: when the j-th link information including the target device identifier is found in the N+M link information, and the j-th indication information indicates that the first switch is not included in the j-th device link, determining that the target faulty device is a device on the mainboard.

[0011] In an exemplary embodiment, after determining whether the target faulty device on the jth device link is located between the processor and the first switch, the method further includes: when the target faulty device is located between the processor and the first switch, determining that the target faulty device is a device on the mainboard.

[0012] In an exemplary embodiment, determining whether the target faulty device on the j-th device link is located between the processor and the first switch includes: when the target device identifier in the j-th link information is between the device identifier of the processor and the device identifier of the first switch, determining that the target faulty device on the j-th device link is located between the processor and the first switch.

[0013] In an exemplary embodiment, determining whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier includes: searching the record item information for a record item including the target device identifier, wherein the record item information includes P record items, and the kth record item among the P record items includes the device identifier of the kth switch among the P switches, where k is a positive integer less than or equal to P; and when the pth record item including the target device identifier is found in the record item information, determining that the target faulty device is a device on the switch board, and the target faulty device is one of the P switches, wherein p is a positive integer less than or equal to P, and the device identifier of the pth switch among the P switches included in the pth record item is equal to the target device identifier.

[0014] In an exemplary embodiment, determining whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier includes: determining P bus number ranges based on P secondary bus numbers and P slave bus numbers corresponding to the P switches included in the record item information, wherein the record item information includes P record items, a kth record item among the P record items includes the device identifier of a kth switch among the P switches and the secondary bus number and slave bus number of the kth switch, k is a positive integer less than or equal to P, a minimum value of the kth bus number range among the P bus number ranges is the secondary bus number of the kth switch, and a maximum value of the kth bus number range is the slave bus number of the kth switch; determining whether the target device identifier is within the P bus number ranges; and if it is determined that the target device identifier is within one of the P bus number ranges, determining that the target faulty device is a connected device among the M connected devices on the switch board.

[0015] In an exemplary embodiment, after determining whether the target device identifier is located in the P bus number ranges, the method includes: if it is determined that the target device identifier is not located in each bus number range in the P bus number ranges, determining N+M bus number ranges based on the N+M secondary bus numbers and N+M slave bus numbers corresponding to the N+M connection devices included in the N+M link information, wherein the i-th link information in the N+M link information also includes the secondary bus number and the slave bus number of the root port where the i-th connection device is located, the minimum value of the i-th bus number range in the N+M bus number ranges is the secondary bus number of the root port where the i-th connection device is located, and the maximum value of the i-th bus number range is the slave bus number of the root port where the i-th connection device is located; determining whether the target device identifier is located in the N+M bus number ranges; and if it is determined that the target device identifier is located in one of the N+M bus number ranges, determining that the target faulty device is a device on the mainboard.

[0016] In an exemplary embodiment, after determining whether the target device identifier is located in the N+M bus number ranges, the method further includes: in a case where it is determined that the target device identifier is not located in each bus number range in the N+M bus number ranges, displaying a first prompt message, wherein the first prompt message is used for failing to determine the location of the target faulty device.

[0017] In an exemplary embodiment, after searching for link information including the target device identifier in the preset N+M link information, the method further includes: when link information including the target device identifier is not found in the N+M link information, determining whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier.

[0018] In an exemplary embodiment, determining whether the target faulty device on the j-th device link is located between the processor and the first switch includes: determining that the target faulty device on the j-th device link is not located between the processor and the first switch when the target device identifier in the j-th link information is not located between the device identifier of the processor and the device identifier of the first switch.

[0019] In an exemplary embodiment, when it is determined that the target fault device is a device on the switch board, the method further includes: obtaining an identification of the switch board and displaying a second prompt information, wherein the second prompt information includes the identification of the switch board, and the second prompt information is used to indicate that the target fault device is a device on the switch board; or obtaining the identification of the switch board, and when the target fault device is one of the M connected devices, displaying a third prompt information, wherein the third prompt information includes the identification of the switch board, and the third prompt information is used to indicate that the target fault device is one of the M connected devices on the switch board; or obtaining the identification of the switch board, and when the target fault device is one of the P switches, displaying a fourth prompt information, wherein the fourth prompt information includes the identification of the switch board, and the fourth prompt information is used to indicate that the target fault device is one of the P switches on the switch board.

[0020] In an exemplary embodiment, obtaining the identification of the switch board includes: obtaining the identification of the switch board from the record item information, wherein the record item information includes the identification of the switch board and P record items, and the kth record item among the P record items includes the device identification of the kth switch among the P switches.

[0021] In an exemplary embodiment, when it is determined that the target faulty device is a device on the mainboard, the method further includes: obtaining an identification of the mainboard and displaying a fifth prompt message, wherein the fifth prompt message includes the identification of the mainboard, and wherein the fifth prompt message is used to indicate that the target faulty device is a device on the mainboard; or obtaining an identification of the mainboard, and when the target faulty device is one of the N connected devices, displaying a sixth prompt message, wherein the sixth prompt message includes the identification of the mainboard, and wherein the sixth prompt message is used to indicate that the target faulty device is one of the N connected devices on the mainboard; or obtaining an identification of the mainboard, and when the target faulty device is a device other than the N connected devices on the mainboard, displaying a seventh prompt message, wherein the seventh prompt message includes the identification of the mainboard, and wherein the seventh prompt message is used to indicate that the target faulty device is a device other than the N connected devices on the mainboard.

[0022] In an exemplary embodiment, obtaining the identification of the mainboard includes: obtaining the identification of the mainboard from predetermined connection device description information, wherein the connection device description information includes the identification of the mainboard and the N+M link information.

[0023] In an exemplary embodiment, before searching for link information including the target device identifier in the preset N+M link information, the method further includes: obtaining device identifiers of multiple devices on each device link of the N+M device links, wherein the N+M device links include a device link formed from the processor to each device of the N+M connected devices, and the device identifiers of multiple devices on the i-th device link of the N+M device links include the device identifier of the i-th connected device and the device identifier of the root port where the i-th connected device is located; obtaining the identifier of the mainboard; and obtaining the secondary bus number and the slave bus number of the root port where each connected device of the N+M connected devices is located.

[0024] In an exemplary embodiment, before searching for link information including the target device identifier in the preset N+M link information, the method further includes: determining whether one of the P switches exists on each of the N+M device links, thereby obtaining N+M indication information, wherein the i-th indication information among the N+M indication information is used to indicate whether one of the P switches exists on the i-th device link.

[0025] In an exemplary embodiment, obtaining the device identifications of multiple devices on each device link in the N+M device links includes: when none of the N+M connected devices are virtual network port devices, obtaining the device identifications of multiple devices on each device link in the N+M device links sent by the N+M connected devices.

[0026] In an exemplary embodiment, before determining whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier, the method further includes: obtaining a device identifier of each of the P switches; obtaining an identifier of the switch board; and obtaining a secondary bus number and a slave bus number of each of the P switches.

[0027] In an exemplary embodiment, before determining whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier, the method further includes: recording the device identifier of each switch in the P switches and the secondary bus number and the slave bus number of each switch in the P switches in P record items in the record item information, wherein the kth record item in the P record items includes the device identifier of the kth switch in the P switches and the secondary bus number and the slave bus number of the kth switch, and k is a positive integer less than or equal to P.

[0028] According to another embodiment of the present application, a device for determining a faulty device is provided, characterized in that it includes: an acquisition module, configured to obtain a target device identifier of a target faulty device; a search module, configured to search for link information including the target device identifier in preset N+M link information, wherein the N+M link information has a one-to-one correspondence with N+M connection devices, the N+M connection devices include N connection devices on the main board and M connection devices on the switch board, each of the M connection devices is connected to one of the P switches on the switch board, the i-th link information in the N+M link information includes the device identifiers of multiple devices on the i-th device link formed from the processor on the main board to the i-th connection device, N, M and P are all positive integers, i and j are less than or equal to a positive integer N+M; a first determination module, configured to determine whether the target faulty device is located between the processor and the first switch on the jth device link, if the jth link information including the target device identifier is found in the N+M link information, and the jth indication information in the jth link information indicates that a first switch exists on the jth device link formed from the processor to the jth connection device, wherein the P switches include the first switch; a second determination module, configured to determine whether the target faulty device is a device on the switch board according to preset record item information and the target device identifier, if the target faulty device is not located between the processor and the first switch, wherein the record item information includes the device identifier of each switch in the P switches.

[0029] According to another embodiment of the present application, a computer non-volatile readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0030] According to another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.

[0031] Through the present application, link information containing the target device identifier of the target fault device is searched in the preset N+M link information, wherein each link information corresponds to a connection device, the connection devices include N connection devices on the mainboard and M connection devices on the switch board, each connection device in the M connection devices is connected to one of the P switches on the switch board, and the link information includes the device identifiers of multiple devices on the device link formed from the processor on the mainboard to the connection device; if the jth link information containing the target device identifier is found, and the indication information contained in the link information indicates that there is a first switch on the device link, it is determined that the target fault device is on the device Whether the target fault device is located between the processor and the first switch in the link. If the target fault device is not located between the processor and the first switch, whether the target fault device is a device on the switch board is determined according to the preset record item information and the target device identifier, wherein the record item information includes the device identifier of each switch; adopting the above scheme, the position of the faulty device can be accurately located, so that the faulty device can be repaired as soon as possible, the time spent on troubleshooting the fault location is reduced, and the user experience is improved, thereby solving the technical problem in the related technology that the related technology can accurately locate the device fault of ordinary PCIE, but the device fault location of AI models with Switch boards is not accurate enough. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIG1 is a block diagram of the hardware structure of a BIOS according to a method for determining a faulty device according to an embodiment of the present application;

[0033] FIG2 is a flow chart of a method for determining a faulty device according to an embodiment of the present application;

[0034] FIG3 is a structural block diagram of an optional PCIE device fault detection system according to an embodiment of the present application;

[0035] FIG4 is a flow chart of an optional PCIE device fault detection method according to an embodiment of the present application;

[0036] FIG5 is a schematic diagram of the system architecture of an optional system for determining a faulty device according to an embodiment of the present application;

[0037] FIG6 is a structural block diagram of a device for determining a faulty device according to an embodiment of the present application. DETAILED DESCRIPTION

[0038] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0039] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0040] The following is an explanation of the professional terms used in this application:

[0041] BIOS: computer basic input and output system;

[0042] BMC: Baseboard Management Controller;

[0043] OS: operating system;

[0044] RpSecBus: Secondary Bus Number, CPU Root Port secondary bus;

[0045] RpSubBus: Subordinate Bus, CPU Root Port slave bus;

[0046] SwitchSecBus: Secondary Bus Number, Switch Port secondary bus;

[0047] SwitchSubBus: Subordinate Bus, Switch Port slave bus;

[0048] Ep: End Point Device, pointing to devices such as GPU, network card, NVME disk, etc.

[0049] It should be noted that Ep refers to all PCIE devices in this application;

[0050] AI model PCIE link: CPU Rp→Bridge1→Bridge2→Bridge3→Switch Bridge→Ep;

[0051] Directly connected Smart NIC link: CPU Rp → Smart NIC;

[0052] Switch board: I / O expansion board, containing 4 Switch bridge device chips, used to expand the interface of PCIE devices;

[0053] SlotId: slot number.

[0054] The method embodiments provided in the embodiments of the present application can be executed in a BIOS or a similar computing device. Taking running on a BIOS as an example, FIG1 is a hardware structure block diagram of a BIOS of a method for determining a faulty device in an embodiment of the present application. As shown in FIG1 , the BIOS may include one or more (only one is shown in FIG1 ) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU (Microcontroller Unit) or a programmable logic device FPGA (Field-Programmable Gate Array)) and a memory 104 configured to store data, wherein the above-mentioned BIOS may also include a transmission device 106 and an input / output device 108 configured to have a communication function. It will be understood by those skilled in the art that the structure shown in FIG1 is merely illustrative and does not limit the structure of the above-mentioned BIOS. For example, the BIOS may also include more or fewer components than those shown in FIG1 , or have a configuration different from that shown in FIG1 .

[0055] The memory 104 can be configured to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining a faulty device in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the BIOS via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0056] The transmission device 106 is configured to receive or transmit data via a network. Examples of such a network may include a wireless network provided by the BIOS's communication provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module configured to communicate with the Internet wirelessly.

[0057] In this embodiment, a method for determining a faulty device is provided. FIG2 is a flow chart of the method for determining a faulty device according to an embodiment of the present application. As shown in FIG2 , the flow chart includes the following steps:

[0058] Step S202, obtaining the target device identifier of the target faulty device;

[0059] Step S204: Searching for link information including a target device identifier in preset N+M link information, wherein the N+M link information has a one-to-one correspondence with the N+M connection devices, the N+M connection devices include N connection devices on the mainboard and M connection devices on the switch board, each of the M connection devices is connected to one of the P switches on the switch board, and the i-th link information in the N+M link information includes the device identifiers of multiple devices on the i-th device link formed from the processor on the mainboard to the i-th connection device, N, M, and P are all positive integers, and i and j are positive integers less than or equal to N+M.

[0060] Step S206: If the j-th link information including the target device identifier is found in the N+M link information, and the j-th indication information in the j-th link information indicates that the first switch exists on the j-th device link formed from the processor to the j-th connected device, determining whether the target faulty device is located between the processor and the first switch on the j-th device link, where the P switches include the first switch;

[0061] Step S208: If the target faulty device is not located between the processor and the first switch, determine whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier, wherein the record item information includes the device identifier of each switch in the P switches.

[0062] Through the above steps, link information containing the target device identifier of the target fault device is searched in the preset N+M link information, wherein each link information corresponds to a connection device, the connection devices include N connection devices on the mainboard and M connection devices on the switch board, each connection device in the M connection devices is connected to one of the P switches on the switch board, and the link information includes the device identifiers of multiple devices on the device link formed from the processor on the mainboard to the connection device; if the jth link information containing the target device identifier is found, and the indication information contained in the link information indicates that there is a first switch on the device link, it is determined that the target fault device is on the device link. Whether the target fault device is located between the processor and the first switch in the link. If the target fault device is not located between the processor and the first switch, whether the target fault device is a device on the switch board is determined according to the preset record item information and the target device identifier, wherein the record item information includes the device identifier of each switch; adopting the above scheme, the position of the faulty device can be accurately located, so that the faulty device can be repaired as soon as possible, the time spent on troubleshooting the fault location is reduced, and the user experience is improved, thereby solving the technical problem in the related technology that the related technology can accurately locate the device fault of ordinary PCIE, but the device fault location of AI models with Switch boards is not accurate enough.

[0063] The execution subject of the above steps may be BIOS, computer terminal, etc., but is not limited thereto.

[0064] The execution order of step S202 and step S204 can be interchanged, that is, step S204 can be executed first, and then step S202.

[0065] In an exemplary embodiment, the above-mentioned step S204 is executed: after searching for link information including the target device identifier in the preset N+M link information, the method further includes: when the j-th link information including the target device identifier is found in the N+M link information, and the j-th indication information indicates that the first switch is not included in the j-th device link, determining that the target faulty device is a device on the mainboard.

[0066] In some embodiments, if link information containing the target device identifier is matched among the preset N+M link information, and it is determined based on the indication information contained in the link information that there is no Switch bridge device (i.e., the first switch) in the device link corresponding to the link information, then the target faulty device is determined to be the device on the mainboard.

[0067] Through this embodiment, after matching the link where the target device identifier is located, it is further confirmed whether a switch bridge device exists on the link, so as to avoid misjudgment and improve the accuracy of locating the faulty device.

[0068] Based on the above steps, after executing step S206: determining whether the target faulty device is located between the processor and the first switch on the jth device link, the method further includes: when the target faulty device is located between the processor and the first switch, determining that the target faulty device is a device on the mainboard.

[0069] If it is determined that the target faulty device is between the processor and the first switch, it is determined that the faulty device is on the uplink of the Switch bridge device, and the faulty device is determined to be a device on the motherboard, and the matching process exits.

[0070] In some embodiments, the above step of determining whether the target faulty device is located between the processor and the first switch on the j-th device link may be implemented by the following scheme, including: determining that the target faulty device is located between the processor and the first switch on the j-th device link when the target device identifier in the j-th link information is between the device identifier of the processor and the device identifier of the first switch.

[0071] Each link information not only contains the device IDs of multiple devices passed from the processor to the connected device, but also the storage order between the device IDs is used to represent the order of different devices in the link. Therefore, it is possible to determine whether the target faulty device is located in the uplink of the switch bridge device by determining whether the target device ID is between the device ID of the processor and the device ID of the first switch.

[0072] Through this embodiment, after determining that there is a Switch bridge device in the link, it is further determined whether the faulty device is located on the uplink of the Switch bridge device to determine whether the faulty device is a device on the mainboard; thereby accurately locating the position of the faulty device.

[0073] In some embodiments, the above-mentioned step S208, determining whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier, can be implemented in the following manner, including: searching the record item information for a record item including the target device identifier, wherein the record item information includes P record items, the kth record item among the P record items includes the device identifier of the kth switch among the P switches, and k is a positive integer less than or equal to P; when the pth record item including the target device identifier is found in the record item information, determining that the target faulty device is a device on the switch board, and the target faulty device is one of the P switches, wherein p is a positive integer less than or equal to P, and the device identifier of the pth switch among the P switches included in the pth record item is equal to the target device identifier.

[0074] The process of determining whether the target faulty device is a device on the switch board includes: searching for a record item including the target device identifier in the record item information (the PCIE Switch asset information table corresponding to the Switch linked list), each record item in the record item information including the device identifier of the switch (Switch bridge device); if a record item including the target device identifier is found, it is determined that the target faulty device is a device on the switch board, and the category of the target faulty device is a Switch bridge device.

[0075] Through this embodiment, on the basis of matching the faulty link, the BDF (device identifier) ​​of the Switch bridge device is further matched in the PCIE Switch asset information table, thereby accurately locating the position of the target faulty device.

[0076] In some embodiments, the above-mentioned step S208, determining whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier, can also be implemented in the following manner, including: determining P bus number ranges based on P secondary bus numbers and P slave bus numbers corresponding to the P switches included in the record item information, wherein the record item information includes P record items, the kth record item among the P record items includes the device identifier of the kth switch among the P switches and the secondary bus number and slave bus number of the kth switch, k is a positive integer less than or equal to P, the minimum value of the kth bus number range among the P bus number ranges is the secondary bus number of the kth switch, and the maximum value of the kth bus number range is the slave bus number of the kth switch; determining whether the target device identifier is located in the P bus number ranges; and if it is determined that the target device identifier is located in one of the P bus number ranges, determining that the target faulty device is a connection device among the M connection devices on the switch board.

[0077] Each record item in the record item information also includes the secondary bus number (SwitchSecBus) and the subordinate bus number (SwitchSubBus) corresponding to the switch, thereby determining the P bus number ranges corresponding to the P switches, and the bus number range is [SwitchSecBus, SwitchSubBus]. Determine whether the target device identifier is located in any bus number range in the P bus ranges, that is, (SwitchSecBus<=BDF<=SwitchSubBus). If a match is successful, it is determined that the target faulty device is one of the M connected devices on the switch board.

[0078] Based on the above steps, after determining whether the target device identifier is located in the P bus number ranges, the method includes: if it is determined that the target device identifier is not located in each bus number range in the P bus number ranges, determining N+M bus number ranges based on the N+M secondary bus numbers and N+M slave bus numbers corresponding to the N+M connection devices included in the N+M link information, wherein the i-th link information in the N+M link information also includes the secondary bus number and the slave bus number of the root port where the i-th connection device is located, the minimum value of the i-th bus number range in the N+M bus number ranges is the secondary bus number of the root port where the i-th connection device is located, and the maximum value of the i-th bus number range is the slave bus number of the root port where the i-th connection device is located; determining whether the target device identifier is located in the N+M bus number ranges; and if it is determined that the target device identifier is located in one of the N+M bus number ranges, determining that the target faulty device is a device on the mainboard.

[0079] If the target device identifier is not located in any of the P bus number ranges, it is necessary to continue traversing the PCIE asset information table, which includes the N+M link information; the link information also includes the secondary bus number (i.e., RpSecBus) and the subordinate bus number (i.e., RpSubBus) of the root port (Rootport) where the connected device is located, and determine the N+M bus number ranges ([RpSecBus, RpSubBus]), and continue to determine whether the target device identifier is located in any of the N+M bus number ranges; if the match is successful (i.e., RpSecBus<=BDF<=RpSubBus), it is determined that the target faulty device is a device on the mainboard.

[0080] Based on the above steps, after determining whether the target device identifier is located in the N+M bus number ranges, the method also includes: when it is determined that the target device identifier is not located in each bus number range in the N+M bus number ranges, displaying a first prompt message, wherein the first prompt message is used to indicate that the location of the target faulty device cannot be determined.

[0081] After the above matching process, if it is determined that the target device identifier is not located in any of the N+M bus number ranges, it means that the matching has failed. At this time, a first prompt message will be displayed to the user. The first prompt message is used to inform the user that the location of the target faulty device cannot be determined and to exit the matching process.

[0082] Based on the above steps, after searching for link information including the target device identifier in the preset N+M link information, the method also includes: when the link information including the target device identifier is not found in the N+M link information, determining whether the target faulty device is a device on the switch board based on the preset record item information and the target device identifier.

[0083] If the BDF (target device identifier) ​​to be queried does not match the link information containing the target device identifier in the PCIE asset information table, that is, the BDF to be queried does not match the BDF of Ep, the BDF of the four upper-level devices of Ep, and the RootPort BDF where Ep is located, then the PCIE Switch asset information (that is, the above-mentioned preset record item information) must be matched. The PCIE asset information table is set to store the above-mentioned preset N+M link information.

[0084] In some embodiments, determining whether the target faulty device is located between the processor and the first switch on the j-th device link includes: determining that the target faulty device is not located between the processor and the first switch on the j-th device link if the target device identifier in the j-th link information is not between the device identifier of the processor and the device identifier of the first switch.

[0085] If the target device identifier is not located between the processor's device identifier and the first switch's device identifier in the jth link information, then the target faulty device is determined not to be located between the processor and the first switch, that is, the faulty device is not located on the uplink of the Switch bridge device.

[0086] In some embodiments, when it is determined that the target faulty device is a device on a switch board, the method further includes: obtaining an identifier of the switch board and displaying a second prompt message, wherein the second prompt message includes the identifier of the switch board, and the second prompt message is used to indicate that the target faulty device is a device on the switch board; or obtaining an identifier of the switch board, and when the target faulty device is one of M connected devices, displaying a third prompt message, wherein the third prompt message includes the identifier of the switch board, and the third prompt message is used to indicate that the target faulty device is one of the M connected devices on the switch board; or obtaining an identifier of the switch board, and when the target faulty device is one of P switches, displaying a fourth prompt message, wherein the fourth prompt message includes the identifier of the switch board, and the fourth prompt message is used to indicate that the target faulty device is one of the P switches on the switch board.

[0087] After determining that the target faulty device is a device on the switch board, different prompt information is displayed to the user according to the different categories of the target faulty device, where the categories of the target faulty device on the switch board include: one connection device (and PCIE device) among the M connection devices on the switch board, one switch among P switches (i.e., Switch bridge device), and other devices on the link; if it is determined that the category of the target faulty device is other devices on the link, the switch board identifier (i.e., the silkscreen information of the Switch board) is displayed to the user, and the user is informed that the target faulty device is other devices on the link; if it is determined that the category of the target faulty device is one connection device among the M connection devices on the switch board, the switch board identifier is displayed to the user, and the user is informed that the target faulty device is one connection device among the M connection devices; if it is determined that the category of the target faulty device is one switch among P switches, the switch board identifier is displayed to the user, and the user is informed that the category of the target faulty device is one switch among P switch boards.

[0088] In an optional embodiment, after locating the target faulty device as a device on the switch board and displaying the silkscreen information of the switch board and the category information of the target faulty device to the user, the user can also be shown which device the target faulty device is, that is, the BDF information of the target faulty device is also directly displayed to the user, and / or the device ID information of the target faulty device is matched according to the BDF information and displayed to the user, as shown in Figure 5, to help the user confirm which device on the Switch board the target faulty device is, such as Switch Bridge2, EP3, etc.

[0089] Through this embodiment, the silkscreen information of the switch board is determined and the category of the faulty device is clarified, so that the user can accurately determine the repair strategy and locate the faulty device based on this information, thereby improving the user experience.

[0090] In some embodiments, obtaining the identification of the switch board includes: obtaining the identification of the switch board in the record item information, wherein the record item information includes the identification of the switch board and P record items, and the kth record item among the P record items includes the device identification of the kth switch among the P switches.

[0091] The switch identification (i.e., the silkscreen information of the Switch board) can be queried in the record information. The record information is the PCIE Switch asset information generated based on the Switch linked list, which stores the BDF of each Switch bridge device and the silkscreen information of the Switch board.

[0092] In some embodiments, when it is determined that the target faulty device is a device on the mainboard, the method further includes: obtaining an identification of the mainboard and displaying a fifth prompt message, wherein the fifth prompt message includes the identification of the mainboard, and wherein the fifth prompt message is used to indicate that the target faulty device is a device on the mainboard; or obtaining an identification of the mainboard, and when the target faulty device is one of N connected devices, displaying a sixth prompt message, wherein the sixth prompt message includes the identification of the mainboard, and the sixth prompt message is used to indicate that the target faulty device is one of the N connected devices on the mainboard; or obtaining an identification of the mainboard, and when the target faulty device is a device other than the N connected devices on the mainboard, displaying a seventh prompt message, wherein the seventh prompt message includes the identification of the mainboard, and the seventh prompt message is used to indicate that the target faulty device is a device other than the N connected devices on the mainboard.

[0093] If it is determined that the target faulty device is a device on the mainboard, it is necessary to obtain the silkscreen information of the corresponding mainboard (i.e., the identification of the above-mentioned mainboard), and display different prompt information to the user according to the different categories of the target faulty device, wherein the categories of the target faulty device located on the mainboard include: one of the N connected devices on the mainboard, the device on the mainboard, and the device other than the N connected devices on the mainboard; if the category of the target faulty device is a device on the mainboard, the silkscreen information of the mainboard is displayed to the user, and the user is informed that the target faulty device is a device on the mainboard; or, in the case of determining that the target faulty device is one of the N connected devices, the silkscreen information of the mainboard is displayed to the user, and the user is informed that the target faulty device is one of the N connected devices; or, if it is determined that the target faulty device is a device other than the N connected devices on the mainboard, the private message information of the mainboard is displayed to the user, and the user is informed that the target faulty device is a device other than the N connected devices on the mainboard.

[0094] In an optional embodiment, after locating the target faulty device as a device on the mainboard and displaying the silkscreen information of the mainboard and the category information of the target faulty device to the user, the user can also be shown which device the target faulty device is, that is, the BDF information of the target faulty device is also directly displayed to the user, and / or the device ID information of the target faulty device is matched according to the BDF information and displayed to the user, as shown in Figure 5, to help the user confirm which device on the mainboard the target faulty device is, such as EP1, Bridge3, etc., according to the BDF information.

[0095] Through this embodiment, based on determining that the target faulty device is a device on the mainboard, the category of the target faulty device is further determined, thereby helping the user to more accurately repair the fault according to the category and location of the faulty device.

[0096] In some embodiments, obtaining the mainboard identifier includes: obtaining the mainboard identifier from predetermined connection device description information, wherein the connection device description information includes the mainboard identifier and N+M link information.

[0097] The silk screen information of the motherboard can be found from the predetermined connection device description information. The connection device description information is the PCIE asset information table, which stores the motherboard identification (ie, motherboard silk screen information) and N+M link information.

[0098] In an exemplary embodiment, before searching for link information including a target device identifier in preset N+M link information, the method further includes: obtaining device identifiers of multiple devices on each device link of the N+M device links, wherein the N+M device links include a device link formed from a processor to each device of the N+M connected devices, and the device identifiers of multiple devices on the i-th device link of the N+M device links include the device identifier of the i-th connected device and the device identifier of the root port where the i-th connected device is located; obtaining an identifier of the mainboard; and obtaining a secondary bus number and a slave bus number of the root port where each of the N+M connected devices is located.

[0099] Before starting fault detection, you need to build an Ep linked list (i.e., the aforementioned connection device description information). The Ep linked list contains: the device identifiers of multiple devices on each of the N+M device links. Each device link contains multiple devices from the processor to the connection device, including the connection device and the four-level devices above the connection device. In addition, you need to obtain the BDF of the RootPort where the connection device is located and the silkscreen information of the motherboard, and obtain the RpSecBus (secondary bus number) and RpSubBus (slave bus number) of the root port where each connection device is located.

[0100] In some embodiments, before searching for link information including a target device identifier in preset N+M link information, the method further includes: determining whether one of the P switches exists on each of the N+M device links, and obtaining N+M indication information, wherein the i-th indication information in the N+M indication information is used to indicate whether one of the P switches exists on the i-th device link.

[0101] During the process of traversing and generating the Ep linked list, the link from Rp (root port) to the Ep device is also scanned to see if there is a Switch Bridge device (i.e., a switch), and the scan result (indication information) is also added to the Ep linked list. The indication information is used to indicate whether there is a switch on the device link.

[0102] In some embodiments, obtaining device identifications of multiple devices on each device link in N+M device links includes: when none of the N+M connected devices are virtual network port devices, obtaining device identifications of multiple devices on each device link in the N+M device links sent by the N+M connected devices.

[0103] During the traversal process, the device's DeviceId and VendorId will be used to determine whether the connected device is connected to a smart network card. If so, the virtual network port device report inside the network card will be filtered out. That is, only when the connected device is not a virtual network port device will the device identifications of multiple devices in the corresponding device link be obtained.

[0104] Through this embodiment, by filtering the reports of the virtual network card devices, the storage resources of the BMC can be saved and the BMC's ability to accurately locate faults of the smart network card can be improved.

[0105] In some embodiments, before determining whether the target faulty device is a device on a switch board based on preset record item information and a target device identifier, the method further includes: obtaining a device identifier of each of the P switches; obtaining an identifier of the switch board; and obtaining a secondary bus number and a subordinate bus number of each of the P switches.

[0106] The module will continue to traverse all bridge devices (Bridge), identify the Switch bridge device on the Switch board, and parse the information of each valid Switch bridge device in turn to obtain the corresponding Switch BDF (device identification of the switch), SwitchSecBus (secondary bus number), SwitchSubBus (slave bus number) and Switch board silkscreen information (switch board identification).

[0107] In some embodiments, before determining whether the target faulty device is a device on a switch board based on preset record item information and the target device identifier, the method further includes: recording the device identifier of each switch among the P switches and the secondary bus number and the slave bus number of each switch among the P switches in P record items in the record item information, wherein the kth record item among the P record items includes the device identifier of the kth switch among the P switches and the secondary bus number and the slave bus number of the kth switch, and k is a positive integer less than or equal to P.

[0108] After obtaining the above information, a switch linked list, namely the above record information, may be generated based on the above information. The record information includes: a device identifier of each switch in the P switches, a secondary bus number, and a slave bus number of each switch.

[0109] This application, through collaborative code development between the BIOS and BMC, covers PCIE device fault detection for AI models with switch boards and fault detection for multiple virtual network ports on smart network cards. This reduces the space required to store PCIE asset information tables on the BMC, improves search efficiency, and enhances the operational efficiency of AI data centers. It also avoids errors caused by manually collecting relevant error information when automatic location is not possible, greatly facilitating operational maintenance and achieving significant results in application scenarios involving precise fault diagnosis and location. The code is also highly scalable and adaptable to different AI server platforms, making it easy to improve the technology and highly valuable for promotion.

[0110] In some embodiments, the above-mentioned method for determining a faulty device can be applied to a PCIE device fault detection system proposed in this application. As shown in Figure 3, the system is composed of a BIOS PCIE asset information reporting module 32, a BMC PCIE asset information storage module 34, a BMC PCIE fault location module 36 and a BMC log alarm module 38.

[0111] The PCIE asset information reporting module 32 will traverse the PICE device (Ep) during the BIOS Post (power-on) process, initialize and store the Ep linked list, identify the PCIE devices directly connected to the motherboard and under the Switch board, and perform information analysis for each valid PCIE device in turn, obtain the BDF of the Ep, the BDF of the four-level device above the Ep, and the RootPort BDF where the Ep is located, obtain the SlotId corresponding to the RootPort, and obtain the corresponding motherboard silk screen information, RpSecBus and RpSubBus according to the SlotId and put them into the Ep linked list. During the traversal process, it also scans whether there is a Switch Bridge device in the link from Rp to the Ep device, and adds the scan result information (i.e. the above-mentioned indication information) to the linked list. At the same time, it determines whether the smart network card is connected and filters out the virtual network card device report inside the network card based on the DeviceId and VendorId of the Ep, so as to save the BMC's storage resources and improve the BMC's ability to accurately locate faults of the smart network card. The module continues to traverse all bridge devices, initializes and stores the Switch linked list, identifies the Switch bridge devices on the switch board, and analyzes the information of each valid switch device in turn, obtaining the corresponding Switch BDF, SwitchSecBus, SwitchSubBus, and Switch board silkscreen information, and adds it to the Switch linked list. Finally, it stores the Ep linked list and Switch linked list in shared memory and notifies the BMC PCIE asset information storage module.

[0112] The asset information storage module 34 parses the linked list to obtain the PCIE asset information table and the PCIE Switch asset information table. When a fault is actually sent, the system triggers an SMI interrupt to send the BDF and error register information of the error-reporting device to the BMC PCIE fault location module 36.

[0113] The PCIE fault location module 36 matches and searches the PCIE asset information table and the PCIE switch asset information table according to specific rules based on the faulty device's BDF. If the search is successful, the corresponding silkscreen is displayed; if not, "Not Found" is displayed and the BMC log alarm module 38 is called to generate an alarm log.

[0114] In an optional embodiment, the present application further provides an optional PCIE device fault detection method, the implementation process of which is shown in FIG4 and includes the following steps:

[0115] Step S401: Power on the machine and start the BIOS;

[0116] Step S402: The PCIE asset information reporting module begins to traverse PCIE devices (Ep), initializes and stores the Ep linked list, identifies PCIE devices directly connected to the motherboard and under the switch board, and analyzes information for each valid PCIE device in turn to obtain the Ep's BDF, the BDFs of the Ep's four upper-level devices, and the RootPort BDF where the Ep is located. The module then obtains the SlotId corresponding to the RootPort and, based on the SlotId, obtains the corresponding motherboard silkscreen information, RpSecBus, and RpSubBus, and stores them in the Ep linked list.

[0117] Step S403: During the traversal process, the link from Rp to Ep is also scanned to see if there is a Switch Bridge device. The scan result information is also added to the linked list. If there is, the BDF is marked as a Switch Bridge and the link is updated to the Ep linked list. If not, the link is marked as having no Switch Bridge and the Ep linked list is updated.

[0118] Step S404: During the traversal process, determine whether a smart network card is connected based on the DeviceId and VendorId of the Ep, so as to filter out the virtual network port device reports inside the network card;

[0119] Step S405: The module continues to traverse all bridge devices, initializes and stores the Switch linked list, identifies the Switch devices on the Switch board, and parses the information of each valid Switch device in turn to obtain the corresponding Switch BDF, SwitchSecBus, SwitchSubBus and Switch board silkscreen information, and adds them to the Switch linked list;

[0120] Step S406: Put the Ep list and the Switch list into the shared memory and notify the BMC PCIE asset information storage module;

[0121] Step S407: The asset information storage module parses the linked list to obtain the PCIE asset information table and the PCIE Switch asset information table, and stores them in JSON format.

[0122] Step S408: When a fault actually occurs, the system triggers an SMI interrupt to send the BDF and error register information of the fault-reporting device to the BMC PCIE fault location module.

[0123] Step S409: traverse the PCIE asset information table;

[0124] Step S410: Determine whether the BDF to be queried matches the BDF of Ep. The BDF of Ep includes: BDF / RootPort BDF / LastRoot BDF / SecondRoot BDF / ThirdRoot BDF / FourRoot BDF. If the BDF to be queried matches the BDF of Ep, execute step S411. If the BDF to be queried does not match the BDF of Ep, execute step S413.

[0125] Step S411: Check whether the matched link has a Switch Bridge device. If there is no Switch device, the matching is successful, the faulty device is the motherboard silkscreen on the matching link, and the matching process exits. If there is, execute step S412.

[0126] Step S412: Confirm whether the faulty device BDF is in the uplink or downlink of the Switch device. If it is in the uplink, the matching is successful. The faulty device is the motherboard silk screen on the matching link, and the matching process exits; if it is in the downlink, execute step S413;

[0127] Step S413: traverse the PCIE Switch asset information table;

[0128] Step S414: Determine whether the BDF to be queried matches the BDF of the Switch, and whether the BDF to be queried falls between SwitchSecBus and SwitchSubBus. If the BDF to be queried matches the BDF of the Switch, or falls between SwitchSecBus and SwitchSubBus (SwitchSecBus<=BDF<=SwitchSubBus), the match is successful, the faulty device is the Switch board silkscreen on the matching link, and the matching process exits. If the BDF to be queried does not match, execute step S415.

[0129] Step S415: traverse the PCIE asset information table;

[0130] Step S416: Determine whether the BDF to be queried falls between RpSecBus and RpSubBus. If the BDF to be queried falls between RpSecBus and RpSubBus (RpSecBus<=BDF<=RpSubBus), it indicates that the match is successful, the faulty device is the motherboard silk screen on the matching link, and the matching process exits; if the BDF to be queried has not been matched, it indicates NOT FOUND, and the matching process exits.

[0131] After executing the above steps, the BMC log alarm module generates an alarm log based on the error register information and the matching result (NOT FOUND / motherboard silkscreen x / switch board silkscreen x). The log is displayed on the web interface and the BMC log alarm module determines whether to send an email to the operation and maintenance personnel for repair based on the error level (correctable error / uncorrectable error) analyzed from the error register information. This allows the module to accurately locate the faulty device and the error level, enabling faster fault repair.

[0132] In an optional embodiment, the present application provides an optional system architecture, as shown in Figure 5. Figure 5 describes a connection method between various devices such as the mainboard, switch board, bridge, switch bridge device, etc. in the present application, so as to facilitate a better understanding of the present application.

[0133] Among them, PCIE devices include Ep1, Ep2 and Ep3, among which Ep1 and Ep2 are PCIE devices directly connected to the motherboard, and Ep1 is connected to the smart network card, and Ep3 is a PCIE device connected to the switch board; for Ep1, since it is connected to the smart network card, when generating the Ep chain table, the BDF information reported by the virtual network port device on the device link will be filtered out. Therefore, the Ep chain table corresponding to Ep1 stores the BDF of the root port RP1 corresponding to Ep1 and the BDF of Ep1; and for the Ep chain table corresponding to Ep2, the BDF of RP2, the BDF of Bridge1-Bridge3 (the figure is only used as an example, there may be fewer or more Bridges in actual applications), and the BDF of Ep2 are stored; and for the Ep chain table corresponding to Ep3, it contains the BDF of RP3, the BDF of Bridge1, the BDF of Switch Bridge2, and the BDF of Ep3.

[0134] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which is stored in a non-volatile readable storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0135] This embodiment also provides a device for determining a faulty device. The device is configured to implement the above-described embodiments and optional implementations, and details already described are omitted. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0136] FIG6 is a structural block diagram of a device for determining a faulty device according to an embodiment of the present application. As shown in FIG6 , the device includes:

[0137] An acquisition module 62 is configured to acquire a target device identifier of a target faulty device;

[0138] a search module 64 configured to search for link information including a target device identifier in preset N+M link information, wherein the N+M link information have a one-to-one correspondence with the N+M connection devices, the N+M connection devices include N connection devices on the mainboard and M connection devices on the switch board, each of the M connection devices is connected to one of the P switches on the switch board, and the i-th link information in the N+M link information includes the device identifiers of multiple devices on the i-th device link formed from the processor on the mainboard to the i-th connection device, N, M, and P are all positive integers, and i and j are positive integers less than or equal to N+M;

[0139] The first determining module 66 is configured to determine whether the target faulty device is located between the processor and the first switch on the jth device link, if the jth link information including the target device identifier is found in the N+M link information and the jth indication information in the jth link information indicates that the first switch exists on the jth device link formed from the processor to the jth connected device, wherein the P switches include the first switch;

[0140] The second determination module 68 is configured to determine whether the target faulty device is a device on the switch board based on preset record item information and the target device identifier when the target faulty device is not located between the processor and the first switch, wherein the record item information includes the device identifier of each switch in the P switches.

[0141] By means of the above-mentioned device, link information containing the target device identifier of the target fault device is searched for in the preset N+M link information, wherein each link information corresponds to a connection device, the connection devices include N connection devices on the mainboard and M connection devices on the switch board, each connection device of the M connection devices is connected to one of the P switches on the switch board, and the link information includes the device identifiers of multiple devices on the device link formed from the processor on the mainboard to the connection device; if the jth link information containing the target device identifier is found, and the indication information contained in the link information indicates that there is a first switch on the device link, it is determined that the target fault device is on the device link. Whether the target fault device is located between the processor and the first switch in the link. If the target fault device is not located between the processor and the first switch, whether the target fault device is a device on the switch board is determined according to the preset record item information and the target device identifier, wherein the record item information includes the device identifier of each switch; adopting the above scheme, the position of the faulty device can be accurately located, so that the faulty device can be repaired as soon as possible, the time spent on troubleshooting the fault location is reduced, and the user experience is improved, thereby solving the technical problem in the related technology that the related technology can accurately locate the device fault of ordinary PCIE, but the device fault location of AI models with Switch boards is not accurate enough.

[0142] In some embodiments, the above-mentioned search module 64 is also configured to determine that the target faulty device is a device on the mainboard when the j-th link information including the target device identifier is found in the N+M link information and the j-th indication information indicates that the first switch is not included on the j-th device link.

[0143] In some embodiments, if link information containing the target device identifier is matched among the preset N+M link information, and it is determined based on the indication information contained in the link information that there is no Switch bridge device (i.e., the first switch) in the device link corresponding to the link information, then the target faulty device is determined to be the device on the mainboard.

[0144] Through this embodiment, after matching the link where the target device identifier is located, it is further confirmed whether a switch bridge device exists on the link, so as to avoid misjudgment and improve the accuracy of locating the faulty device.

[0145] In some embodiments, the first determining module 66 is further configured to determine that the target faulty device is a device on the mainboard when the target faulty device is located between the processor and the first switch.

[0146] If it is determined that the target faulty device is between the processor and the first switch, it is determined that the faulty device is on the uplink of the Switch bridge device, and the faulty device is determined to be a device on the motherboard, and the matching process exits.

[0147] In some embodiments, the first determining module 66 is further configured to determine that the target faulty device is located between the processor and the first switch on the jth device link when the target device identifier in the jth link information is between the device identifier of the processor and the device identifier of the first switch.

[0148] Each link information not only contains the device IDs of multiple devices passed from the processor to the connected device, but also the storage order between the device IDs is used to represent the order of different devices in the link. Therefore, it is possible to determine whether the target faulty device is located in the uplink of the switch bridge device by determining whether the target device ID is between the device ID of the processor and the device ID of the first switch.

[0149] Through this embodiment, after determining that there is a Switch bridge device in the link, it is further determined whether the faulty device is located on the uplink of the Switch bridge device to determine whether the faulty device is a device on the mainboard; thereby accurately locating the position of the faulty device.

[0150] In some embodiments, the second determination module 68 is further configured to search for a record item including a target device identifier in the record item information, wherein the record item information includes P record items, and the kth record item among the P record items includes the device identifier of the kth switch among the P switches, and k is a positive integer less than or equal to P; when the pth record item including the target device identifier is found in the record item information, it is determined that the target faulty device is a device on the switch board, and the target faulty device is one of the P switches, wherein p is a positive integer less than or equal to P, and the device identifier of the pth switch among the P switches included in the pth record item is equal to the target device identifier.

[0151] The process of determining whether the target faulty device is a device on the switch board includes: searching for a record item including the target device identifier in the record item information (the PCIE Switch asset information table corresponding to the Switch linked list), each record item in the record item information including the device identifier of the switch (Switch bridge device); if a record item including the target device identifier is found, it is determined that the target faulty device is a device on the switch board, and the category of the target faulty device is a Switch bridge device.

[0152] Through this embodiment, on the basis of matching the faulty link, the BDF (device identifier) ​​of the Switch bridge device is further matched in the PCIE Switch asset information table, thereby accurately locating the position of the target faulty device.

[0153] In some embodiments, the second determination module 68 is further configured to determine P bus number ranges based on the P secondary bus numbers and P subordinate bus numbers corresponding to the P switches included in the record item information, wherein the record item information includes P record items, the kth record item among the P record items includes the device identifier of the kth switch among the P switches and the secondary bus number and subordinate bus number of the kth switch, k is a positive integer less than or equal to P, the minimum value of the kth bus number range among the P bus number ranges is the secondary bus number of the kth switch, and the maximum value of the kth bus number range is the subordinate bus number of the kth switch; determine whether the target device identifier is located in the P bus number ranges; and when it is determined that the target device identifier is located in one of the P bus number ranges, determine that the target faulty device is a connection device among the M connection devices on the switch board.

[0154] Each record item in the record item information also includes the secondary bus number (SwitchSecBus) and the subordinate bus number (SwitchSubBus) corresponding to the switch, thereby determining the P bus number ranges corresponding to the P switches, and the bus number range is [SwitchSecBus, SwitchSubBus]. Determine whether the target device identifier is located in any bus number range in the P bus ranges, that is, (SwitchSecBus<=BDF<=SwitchSubBus). If a match is successful, it is determined that the target faulty device is one of the M connected devices on the switch board.

[0155] In some embodiments, the above-mentioned second determination module 68 is further configured to determine N+M bus number ranges based on the N+M secondary bus numbers and N+M slave bus numbers corresponding to the N+M connection devices included in the N+M link information when it is determined that the target device identifier is not located in each bus number range in the P bus number ranges, wherein the i-th link information in the N+M link information also includes the secondary bus number and the slave bus number of the root port where the i-th connection device is located, the minimum value of the i-th bus number range in the N+M bus number ranges is the secondary bus number of the root port where the i-th connection device is located, and the maximum value of the i-th bus number range is the slave bus number of the root port where the i-th connection device is located; determine whether the target device identifier is located in the N+M bus number ranges; and when it is determined that the target device identifier is located in one of the N+M bus number ranges, determine that the target faulty device is a device on the mainboard.

[0156] If the target device identifier is not located in any of the P bus number ranges, it is necessary to continue traversing the PCIE asset information table, which includes the N+M link information; the link information also includes the secondary bus number (i.e., RpSecBus) and the subordinate bus number (i.e., RpSubBus) of the root port (Rootport) where the connected device is located, and determine the N+M bus number ranges ([RpSecBus, RpSubBus]), and continue to determine whether the target device identifier is located in any of the N+M bus number ranges; if the match is successful (i.e., RpSecBus<=BDF<=RpSubBus), it is determined that the target faulty device is a device on the mainboard.

[0157] In some embodiments, the above-mentioned second determination module 68 is also configured to display a first prompt message when it is determined that the target device identifier is not located in each bus number range in the N+M bus number ranges, wherein the first prompt message is used to indicate that the location of the target faulty device cannot be determined.

[0158] After the above matching process, if it is determined that the target device identifier is not located in any of the N+M bus number ranges, it means that the matching has failed. At this time, a first prompt message will be displayed to the user. The first prompt message is used to inform the user that the location of the target faulty device cannot be determined and to exit the matching process.

[0159] In some embodiments, the search module 64 is also used to determine whether the target faulty device is a device on the switch board based on preset record item information and target device identification when no link information including the target device identification is found in the N+M link information.

[0160] If the BDF (target device identifier) ​​to be queried does not match the link information containing the target device identifier in the PCIE asset information table, that is, the BDF to be queried does not match the BDF of Ep, the BDF of the four upper-level devices of Ep, and the RootPort BDF where Ep is located, then the PCIE Switch asset information (that is, the above-mentioned preset record item information) must be matched. The PCIE asset information table is set to store the above-mentioned preset N+M link information.

[0161] In some embodiments, the first determination module 66 is further configured to determine that the target faulty device is not located between the processor and the first switch on the jth device link when the target device identifier in the jth link information is not located between the device identifier of the processor and the device identifier of the first switch.

[0162] If the target device identifier is not located between the processor's device identifier and the first switch's device identifier in the jth link information, then the target faulty device is determined not to be located between the processor and the first switch, that is, the faulty device is not located on the uplink of the Switch bridge device.

[0163] In some embodiments, the above-mentioned second determination module 68 is also configured to obtain the identification of the switch board and display a second prompt message, wherein the second prompt message includes the identification of the switch board, and the second prompt message is used to indicate that the target faulty device is a device on the switch board; or obtain the identification of the switch board, and when the target faulty device is one of the M connection devices, display a third prompt message, wherein the third prompt message includes the identification of the switch board, and the third prompt message is used to indicate that the target faulty device is one of the M connection devices on the switch board; or obtain the identification of the switch board, and when the target faulty device is one of the P switches, display a fourth prompt message, wherein the fourth prompt message includes the identification of the switch board, and the fourth prompt message is used to indicate that the target faulty device is one of the P switches on the switch board.

[0164] After determining that the target faulty device is a device on the switch board, different prompt information is displayed to the user according to the different categories of the target faulty device, where the categories of the target faulty device on the switch board include: one connection device (and PCIE device) among the M connection devices on the switch board, one switch among P switches (i.e., Switch bridge device), and other devices on the link; if it is determined that the category of the target faulty device is other devices on the link, the switch board identifier (i.e., the silkscreen information of the Switch board) is displayed to the user, and the user is informed that the target faulty device is other devices on the link; if it is determined that the category of the target faulty device is one connection device among the M connection devices on the switch board, the switch board identifier is displayed to the user, and the user is informed that the target faulty device is one connection device among the M connection devices; if it is determined that the category of the target faulty device is one switch among P switches, the switch board identifier is displayed to the user, and the user is informed that the category of the target faulty device is one switch among P switch boards.

[0165] In an optional embodiment, after locating the target faulty device as a device on the switch board and displaying the silkscreen information of the switch board and the category information of the target faulty device to the user, the user can also be shown which device the target faulty device is, that is, the BDF information of the target faulty device is also directly displayed to the user, and / or the device ID information of the target faulty device is matched according to the BDF information and displayed to the user, as shown in Figure 5, to help the user confirm which device on the Switch board the target faulty device is, such as Switch Bridge2, EP3, etc.

[0166] Through this embodiment, the silkscreen information of the switch board is determined and the category of the faulty device is clarified, so that the user can accurately determine the repair strategy and locate the faulty device based on this information, thereby improving the user experience.

[0167] In some embodiments, the acquisition module 62 is further configured to obtain the identifier of the switch board from the record item information, wherein the record item information includes the identifier of the switch board and P record items, and the kth record item among the P record items includes the device identifier of the kth switch among the P switches.

[0168] The switch identification (i.e., the silkscreen information of the Switch board) can be queried in the record information. The record information is the PCIE Switch asset information generated based on the Switch linked list, which stores the BDF of each Switch bridge device and the silkscreen information of the Switch board.

[0169] In some embodiments, the above-mentioned second determination module 68 is also configured to obtain the identification of the mainboard and display a fifth prompt message, wherein the fifth prompt message includes the identification of the mainboard, wherein the fifth prompt message is used to indicate that the target faulty device is a device on the mainboard; or obtain the identification of the mainboard, and when the target faulty device is one of the N connected devices, display a sixth prompt message, wherein the sixth prompt message includes the identification of the mainboard, and the sixth prompt message is used to indicate that the target faulty device is one of the N connected devices on the mainboard; or obtain the identification of the mainboard, and when the target faulty device is a device other than the N connected devices on the mainboard, display a seventh prompt message, wherein the seventh prompt message includes the identification of the mainboard, and the seventh prompt message is used to indicate that the target faulty device is a device other than the N connected devices on the mainboard.

[0170] If it is determined that the target faulty device is a device on the mainboard, it is necessary to obtain the silkscreen information of the corresponding mainboard (i.e., the identification of the above-mentioned mainboard), and display different prompt information to the user according to the different categories of the target faulty device, wherein the categories of the target faulty device located on the mainboard include: one of the N connected devices on the mainboard, the device on the mainboard, and the device other than the N connected devices on the mainboard; if the category of the target faulty device is a device on the mainboard, the silkscreen information of the mainboard is displayed to the user, and the user is informed that the target faulty device is a device on the mainboard; or, in the case of determining that the target faulty device is one of the N connected devices, the silkscreen information of the mainboard is displayed to the user, and the user is informed that the target faulty device is one of the N connected devices; or, if it is determined that the target faulty device is a device other than the N connected devices on the mainboard, the private message information of the mainboard is displayed to the user, and the user is informed that the target faulty device is a device other than the N connected devices on the mainboard.

[0171] In an optional embodiment, after locating the target faulty device as a device on the mainboard and displaying the silkscreen information of the mainboard and the category information of the target faulty device to the user, the user can also be shown which device the target faulty device is, that is, the BDF information of the target faulty device is also directly displayed to the user, and / or the device ID information of the target faulty device is matched according to the BDF information and displayed to the user, as shown in Figure 5, to help the user confirm which device on the mainboard the target faulty device is, such as EP1, Bridge3, etc., according to the BDF information.

[0172] Through this embodiment, based on determining that the target faulty device is a device on the mainboard, the category of the target faulty device is further determined, thereby helping the user to more accurately repair the fault according to the category and location of the faulty device.

[0173] In some embodiments, the acquisition module 62 is further configured to acquire the motherboard identifier from predetermined connection device description information, wherein the connection device description information includes the motherboard identifier and N+M link information.

[0174] The silk screen information of the motherboard can be found from the predetermined connection device description information. The connection device description information is a PCIE asset information table that stores the motherboard identifier (ie, motherboard silk screen information) and N+M link information.

[0175] In some embodiments, the search module 64 is further configured to obtain device identifications of multiple devices on each device link among N+M device links, wherein the N+M device links include a device link formed from a processor to each device among N+M connected devices, and the device identifications of multiple devices on the i-th device link among the N+M device links include the device identification of the i-th connected device and the device identification of the root port where the i-th connected device is located; obtain the identification of the mainboard; obtain the secondary bus number and the slave bus number of the root port where each connected device among the N+M connected devices is located.

[0176] Before starting fault detection, you need to build an Ep linked list (i.e., the aforementioned connection device description information). The Ep linked list contains: the device identifiers of multiple devices on each of the N+M device links. Each device link contains multiple devices from the processor to the connection device, including the connection device and the four-level devices above the connection device. In addition, you need to obtain the BDF of the RootPort where the connection device is located and the silkscreen information of the motherboard, and obtain the RpSecBus (secondary bus number) and RpSubBus (slave bus number) of the root port where each connection device is located.

[0177] In some embodiments, the search module 64 is further configured to determine whether one of the P switches exists on each of the N+M device links, and obtain N+M indication information, wherein the i-th indication information among the N+M indication information is used to indicate whether one of the P switches exists on the i-th device link.

[0178] During the process of traversing and generating the Ep linked list, the link from Rp (root port) to the Ep device is also scanned to see if there is a Switch Bridge device (i.e., a switch), and the scan result (indication information) is also added to the Ep linked list. The indication information is used to indicate whether there is a switch on the device link.

[0179] In some embodiments, the acquisition module 62 is further configured to acquire device identifiers of multiple devices on each device link of the N+M device links sent by the N+M connection devices when none of the N+M connection devices are virtual network port devices.

[0180] During the traversal process, the device's DeviceId and VendorId will be used to determine whether the connected device is connected to a smart network card. If so, the virtual network port device report inside the network card will be filtered out. That is, only when the connected device is not a virtual network port device will the device identifications of multiple devices in the corresponding device link be obtained.

[0181] Through this embodiment, by filtering the reports of the virtual network card devices, the storage resources of the BMC can be saved and the BMC's ability to accurately locate faults of the smart network card can be improved.

[0182] In some embodiments, the second determining module 68 is further configured to obtain a device identification of each of the P switches; obtain an identification of a switch board; and obtain a secondary bus number and a slave bus number of each of the P switches.

[0183] The module will continue to traverse all bridge devices (Bridge), identify the Switch bridge device on the Switch board, and parse the information of each valid Switch bridge device in turn to obtain the corresponding Switch BDF (device identification of the switch), SwitchSecBus (secondary bus number), SwitchSubBus (slave bus number) and Switch board silkscreen information (switch board identification).

[0184] In some embodiments, the second determining module 68 is further configured to record the device identification of each of the P switches and the secondary bus number and the slave bus number of each of the P switches in P record items in the record item information, wherein the kth record item in the P record items includes the device identification of the kth switch in the P switches and the secondary bus number and the slave bus number of the kth switch, where k is a positive integer less than or equal to P.

[0185] After obtaining the above information, a switch linked list, namely the above record information, may be generated based on the above information. The record information includes: a device identifier of each switch in the P switches, a secondary bus number, and a slave bus number of each switch.

[0186] This application, through collaborative code development between the BIOS and BMC, covers PCIE device fault detection for AI models with switch boards and fault detection for multiple virtual network ports on smart network cards. This reduces the space required to store PCIE asset information tables on the BMC, improves search efficiency, and enhances the operational efficiency of AI data centers. It also avoids errors caused by manually collecting relevant error information when automatic location is not possible, greatly facilitating operational maintenance and achieving significant results in application scenarios involving precise fault diagnosis and location. The code is also highly scalable and adaptable to different AI server platforms, making it easy to improve the technology and highly valuable for promotion.

[0187] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0188] An embodiment of the present application further provides a computer non-volatile readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0189] In an exemplary embodiment, the above-mentioned computer non-volatile readable storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other non-volatile readable storage media that can store computer programs.

[0190] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0191] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0192] The examples in this embodiment can refer to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.

[0193] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0194] The above are merely examples of the present application and are not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for determining a faulty device, characterized in that: include: Obtain the target device ID of the target faulty device; Searching for link information including the target device identifier in preset N+M link information, wherein the N+M link information has a one-to-one correspondence with the N+M connection devices, the N+M connection devices include N connection devices on the mainboard and M connection devices on the switch board, each of the M connection devices is connected to one of the P switches on the switch board, the i-th link information in the N+M link information includes device identifiers of multiple devices on the i-th device link formed from the processor on the mainboard to the i-th connection device, N, M and P are all positive integers, and i and j are positive integers less than or equal to N+M; When the j-th link information including the target device identifier is found in the N+M link information, and the j-th indication information in the j-th link information indicates that a first switch exists on the j-th device link formed from the processor to the j-th connection device, determining whether the target faulty device is located between the processor and the first switch on the j-th device link, wherein the P switches include the first switch; In the case that the target faulty device is not located between the processor and the first switch, it is determined whether the target faulty device is a device on the switch board according to preset record item information and the target device identifier, wherein the record item information includes a device identifier of each switch in the P switches.

2. The method according to claim 1, characterized in that After searching the preset N+M link information for link information including the target device identifier, the method further includes: When the j-th link information including the target device identifier is found in the N+M link information, and the j-th indication information indicates that the j-th device link does not include the first switch, it is determined that the target faulty device is a device on the mainboard.

3. The method according to claim 1, characterized in that After determining whether the target faulty device on the jth device link is located between the processor and the first switch, the method further includes: In a case where the target faulty device is located between the processor and the first switch, it is determined that the target faulty device is a device on the mainboard.

4. The method according to claim 1, characterized in that: The determining whether the target faulty device on the jth device link is located between the processor and the first switch comprises: When the target device identifier in the j-th link information is located between the device identifier of the processor and the device identifier of the first switch, it is determined that the target faulty device is located between the processor and the first switch on the j-th device link.

5. The method according to claim 1, characterized in that: The step of determining whether the target faulty device is a device on the switch board according to the preset record item information and the target device identifier includes: Searching for a record item including the target device identifier in the record item information, wherein the record item information includes P record items, a k-th record item among the P record items includes a device identifier of a k-th switch among the P switches, and k is a positive integer less than or equal to P; When a p-th record item including the target device identifier is found in the record item information, it is determined that the target faulty device is a device on the switch board, and the target faulty device is one of the P switches, wherein p is a positive integer less than or equal to P, and the device identifier of the p-th switch among the P switches included in the p-th record item is equal to the target device identifier.

6. The method according to claim 1, characterized in that The step of determining whether the target faulty device is a device on the switch board according to the preset record item information and the target device identifier includes: Determine P bus number ranges according to the P secondary bus numbers and P subordinate bus numbers corresponding to the P switches included in the record item information, wherein the record item information includes P record items, a k-th record item among the P record items includes a device identifier of a k-th switch among the P switches and a secondary bus number and a subordinate bus number of the k-th switch, k is a positive integer less than or equal to P, a minimum value of the k-th bus number range among the P bus number ranges is the secondary bus number of the k-th switch, and a maximum value of the k-th bus number range is the subordinate bus number of the k-th switch; Determining whether the target device identifier is located within the P bus number ranges; When it is determined that the target device identifier is located in one bus number range among the P bus number ranges, it is determined that the target faulty device is a connection device among the M connection devices on the switch board.

7. The method according to claim 6, characterized in that After determining whether the target device identifier is located in the P bus number ranges, the method further includes: In a case where it is determined that the target device identifier is not located in each bus number range in the P bus number ranges, determining N+M bus number ranges according to the N+M secondary bus numbers and the N+M subordinate bus numbers corresponding to the N+M connection devices included in the N+M link information, wherein the i-th link information in the N+M link information also includes the secondary bus number and the subordinate bus number of the root port where the i-th connection device is located, the minimum value of the i-th bus number range in the N+M bus number ranges is the secondary bus number of the root port where the i-th connection device is located, and the maximum value of the i-th bus number range is the subordinate bus number of the root port where the i-th connection device is located; Determining whether the target device identifier is located in the N+M bus number range; When it is determined that the target device identifier is located in one bus number range among the N+M bus number ranges, it is determined that the target faulty device is a device on the mainboard.

8. The method according to claim 7, characterized in that After determining whether the target device identification is located in the N+M bus number ranges, the method further includes: When it is determined that the target device identifier is not located in each bus number range of the N+M bus number ranges, first prompt information is displayed, wherein the first prompt information is used to indicate that the location of the target faulty device cannot be determined.

9. The method according to claim 1, characterized in that: After searching the preset N+M link information for link information including the target device identifier, the method further includes: When no link information including the target device identifier is found in the N+M link information, it is determined whether the target faulty device is a device on the switch board according to preset record item information and the target device identifier.

10. The method according to any one of claims 1 to 9, characterized in that The determining whether the target faulty device on the jth device link is located between the processor and the first switch comprises: When the target device identifier in the jth link information is not located between the device identifier of the processor and the device identifier of the first switch, it is determined that the target faulty device is not located between the processor and the first switch on the jth device link.

11. The method according to any one of claims 1 to 9, characterized in that In the case where it is determined that the target faulty device is a device on the switch board, the method further includes: Obtaining the identifier of the switch board and displaying second prompt information, wherein the second prompt information includes the identifier of the switch board, and the second prompt information is used to indicate that the target faulty device is a device on the switch board; or obtaining an identifier of the switch board, and displaying third prompt information when the target faulty device is one of the M connection devices, wherein the third prompt information includes the identifier of the switch board, and the third prompt information is used to indicate that the target faulty device is one of the M connection devices on the switch board; or Obtain an identifier of the switch board, and display a fourth prompt message when the target faulty device is one of the P switches, wherein the fourth prompt message includes the identifier of the switch board, and the fourth prompt message is used to indicate that the target faulty device is one of the P switches on the switch board.

12. The method according to claim 11, characterized in that The obtaining the identifier of the switch board includes: The identifier of the switch board is obtained from the record item information, wherein the record item information includes the identifier of the switch board and P record items, and the kth record item among the P record items includes the device identifier of the kth switch among the P switches.

13. The method according to any one of claims 2, 3, 6 and 7, characterized in that In the case where it is determined that the target faulty device is a device on the mainboard, the method further includes: Obtaining the identification of the mainboard, and displaying fifth prompt information, wherein the fifth prompt information includes the identification of the mainboard, wherein the fifth prompt information is used to indicate that the target faulty device is a device on the mainboard; or obtaining an identification of the mainboard, and displaying sixth prompt information when the target faulty device is one of the N connected devices, wherein the sixth prompt information includes the identification of the mainboard, and the sixth prompt information is used to indicate that the target faulty device is one of the N connected devices on the mainboard; or Obtain the identification of the mainboard, and when the target faulty device is a device other than the N connected devices on the mainboard, display a seventh prompt message, wherein the seventh prompt message includes the identification of the mainboard, and the seventh prompt message is used to indicate that the target faulty device is a device other than the N connected devices on the mainboard.

14. The method according to claim 13, characterized in that The obtaining the identification of the mainboard includes: The identification of the mainboard is obtained from predetermined connection device description information, wherein the connection device description information includes the identification of the mainboard and the N+M link information.

15. The method according to any one of claims 1 to 9, characterized in that Before searching for link information including the target device identifier in the preset N+M link information, the method further includes: Acquire device identifiers of multiple devices on each device link among N+M device links, wherein the N+M device links include a device link formed from the processor to each device among the N+M connected devices, and the device identifiers of multiple devices on the i-th device link among the N+M device links include the device identifier of the i-th connected device and the device identifier of a root port where the i-th connected device is located; Obtaining an identification of the mainboard; The secondary bus number and the slave bus number of the root port where each connection device in the N+M connection devices is located are obtained.

16. The method according to claim 15, characterized in that Before searching for link information including the target device identifier in the preset N+M link information, the method further includes: Determine whether one of the P switches exists on each of the N+M device links, and obtain N+M indication information, wherein an i-th indication information among the N+M indication information is used to indicate whether one of the P switches exists on the i-th device link.

17. The method according to claim 15, characterized in that The obtaining of device identifiers of multiple devices on each device link of the N+M device links includes: In a case where none of the N+M connection devices is a virtual network port device, device identifiers of multiple devices on each of the N+M device links sent by the N+M connection devices are obtained.

18. The method according to any one of claims 1 to 9, characterized in that Before determining whether the target faulty device is a device on the switch board according to the preset record item information and the target device identifier, the method further includes: Obtaining a device identification of each switch in the P switches; Obtaining an identification of the switch board; The secondary bus number and the slave bus number of each switch in the P switches are obtained.

19. The method according to claim 18, characterized in that Before determining whether the target faulty device is a device on the switch board according to the preset record item information and the target device identifier, the method further includes: The device identification of each switch in the P switches and the secondary bus number and the slave bus number of each switch in the P switches are recorded in P record items in the record item information, wherein the kth record item in the P record items includes the device identification of the kth switch in the P switches and the secondary bus number and the slave bus number of the kth switch, and k is a positive integer less than or equal to P.

20. A device for determining a faulty device, characterized in that: include: An acquisition module is configured to acquire a target device identifier of a target faulty device; A search module is configured to search for link information including the target device identifier in preset N+M link information, wherein the N+M link information has a one-to-one correspondence with N+M connection devices, the N+M connection devices include N connection devices on the mainboard and M connection devices on the switch board, each of the M connection devices is connected to one of the P switches on the switch board, the i-th link information in the N+M link information includes device identifiers of multiple devices on the i-th device link formed from the processor on the mainboard to the i-th connection device, N, M and P are all positive integers, and i and j are positive integers less than or equal to N+M; a first determining module, configured to determine whether the target faulty device is located between the processor and the first switch on the jth device link, when the jth link information including the target device identifier is found in the N+M link information and the jth indication information in the jth link information indicates that there is a first switch on the jth device link formed from the processor to the jth connection device, wherein the P switches include the first switch; The second determination module is configured to determine whether the target faulty device is a device on the switch board according to preset record item information and the target device identifier when the target faulty device is not located between the processor and the first switch, wherein the record item information includes a device identifier of each switch in the P switches.

21. A computer non-volatile readable storage medium, characterized in that: The computer non-volatile readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 19 when executed by a processor.

22. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 19 are implemented.

Citation Information

Patent Citations

  • PCIe equipment management method and device

    CN103763129A

  • Fault processing method and device and server

    CN111414268A

  • Method and system for positioning fault of intelligent network card

    CN113645056A

  • Server, mainboard and external equipment fault positioning method of server

    CN116340068A

  • Fault equipment determination method and device, storage medium and electronic equipment

    CN117499214A

Cited By

  • Basic input and output system diagnosis method and device, storage medium and server

    CN120872793A