Server system, faulty device positioning method, computer system, program product, and storage medium
By setting specific structures and connection methods in the server system, and utilizing the baseboard management controller and switching unit to establish dynamic and fixed interconnection tables, the problem of difficulty in locating faulty graphics processor devices in heterogeneous resource pooling is solved, enabling rapid and accurate location and reducing maintenance costs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2026-04-02
AI Technical Summary
In heterogeneous resource pooling, it is difficult to accurately locate faulty graphics processors, leading to increased maintenance time and costs.
By setting specific structures and connection methods in the server system, including the central processing unit, graphics processing unit, switching devices and network switches, and by using the baseboard management controller and switching units, dynamic and fixed interconnection tables are established to achieve physical location of faulty devices.
It enables rapid and accurate location of faulty graphics processor devices, reducing maintenance time and costs.
Smart Images

Figure CN2025087672_02042026_PF_FP_ABST
Abstract
Description
Server system and fault device positioning method, computer system, program product and storage medium
[0001] Cross-reference to Related Applications
[0002] This application claims priority to the Chinese patent application No. 202411370081.4, filed on September 29, 2024, entitled "Server system and fault device positioning method, computer system, program product and storage medium", the entire content of which is incorporated herein by reference. TECHNICAL FIELD
[0003] The present application relates to the technical field of computing, in particular to a server system and a fault device positioning method, system, computer program product and readable storage medium thereof. BACKGROUND
[0004] In order to meet the resource allocation needs of artificial intelligence, machine learning, intelligent computing and the like, data centers increasingly configure pooled heterogeneous resources to improve the utilization rate of server hardware resources and meet the demand for heterogeneous computing.
[0005] Due to the pooling of heterogeneous resources, the allocation relationship between the heterogeneous resources as graphics processing units (GPUs) and computing resources is uncertain. For example, at time A, the 0th computing resource pool is allocated with the 0th GPU and the 1st GPU in the 0th GPU resource pool. After a period of time, the 0th GPU and the 1st GPU are removed from the 0th computing resource, and at the subsequent time B, the 0th computing resource pool is allocated with the 2nd GPU and the 3rd GPU in the 1st GPU resource pool. For the 0th computing resource pool, at times A and B when there is a demand for GPU resources, two GPU resources are obtained. However, the 0th computing resource and the GPU do not have a fixed allocation relationship. Therefore, the basic input / output system firmware in the 0th computing resource saves the link identification information of the allocated GPU, but cannot save the silk print information of the allocated GPU. This will result in the inability to locate the fault GPU when the allocated GPU fails. If the fault GPU is to be accurately located, the maintenance time and cost will be increased. SUMMARY
[0006] The present application provides the following technical solutions:
[0007] In a first aspect, a server system is provided, comprising: a central processing unit, a switching device, a graphics processing unit, and a network switch, wherein the central processing unit is arranged in a first chassis, the graphics processing unit is arranged in a second chassis, and the switching device is arranged in a third chassis, and the first chassis, the second chassis, and the third chassis are arranged in a same server cabinet.
[0008] The first chassis, the second chassis, and the third chassis are respectively connected with the network switch through a first cable; and
[0009] The first chassis and the second chassis are respectively connected with the third chassis through a second cable.
[0010] Further, the first chassis has: a first network port and a first communication port;
[0011] The second chassis has: a second network port, a second communication port;
[0012] The third chassis has: a third network port, a third communication port, and a fourth communication port; and
[0013] The third communication port is connected with the first communication port through the first cable, the fourth communication port is connected with the second communication port through the first cable, and the first network port, the second network port, and the third network port are respectively connected with the network switch through the second cable.
[0014] Further, the first chassis comprises: a first baseboard management controller, the central processing unit, and a first connector;
[0015] The first baseboard management controller is connected with the central processing unit, and the central processing unit is connected with the first connector;
[0016] The first connector serves as the first communication port; and
[0017] The first baseboard management controller has a first baseboard management controller network port, and the first baseboard management controller network port serves as the first network port.
[0018] Further, the first connector is a high-speed pluggable input / output connector.
[0019] Further, the second chassis comprises: a second baseboard management controller, the graphics processing unit, and a second connector;
[0020] The second baseboard management controller is connected with the graphics processing unit, and the graphics processing unit is connected with the second connector;
[0021] The second connector serves as the second communication port; and
[0022] The second baseboard management controller has a second baseboard management controller network port, which serves as a second network port.
[0023] Further, the second connector is a high-speed pluggable input / output connector.
[0024] Further, the third chassis comprises a third baseboard management controller, a transmission expansion device, a switching unit, a third connector and a fourth connector.
[0025] The switching unit is connected to the third baseboard management controller through the transmission expansion device.
[0026] The third baseboard management controller has a third baseboard management controller network port, which serves as a third network port.
[0027] The switching unit has a computing communication port and a graphics card communication port.
[0028] The computing communication port is connected to the third connector, which serves as a third communication port; and
[0029] The graphics card communication port is connected to the fourth connector, which serves as a fourth communication port.
[0030] Further, the switching unit comprises a plurality of switching devices, and the plurality of switching devices are fully connected.
[0031] Further, each switching device has:
[0032] a host port, which serves as a computing communication port;
[0033] a device port, which serves as a graphics card communication port; and
[0034] an internal interconnection port, which is used to connect the plurality of switching devices into full connection.
[0035] Further, the third connector and the fourth connector are high-speed pluggable input / output connectors.
[0036] Further, the second cable is a CDFP (400G Form Factor Pluggable, 400G pluggable module) high-speed cable.
[0037] In a second aspect, a server system fault device positioning method is provided. The server system includes a central processor, a graphics processor, a plurality of switching devices, and a network switch. The central processor is disposed in a first chassis, the graphics processor is disposed in a second chassis, and the plurality of switching devices are disposed in a third chassis. The first chassis, the second chassis, and the third chassis are disposed in the same server cabinet. The first chassis, the second chassis, and the third chassis are connected to the network switch through cables. The first chassis and the second chassis are connected to the third chassis through cables. The third chassis includes a third baseboard management controller. The method includes:
[0038] In response to the central processor detecting a server link fault, triggering a system management interrupt;
[0039] Analyzing a high-level error report register of the fault device to determine fault information of the fault device;
[0040] According to the fault information of the fault device, determining device asset information of the fault device; and
[0041] In response to the device asset information of the fault device not containing slot silk screen information of the fault device, determining a physical location of the fault device according to the fault information of the fault device and an allocation relationship table through the third baseboard management controller, wherein the allocation relationship table at least includes a dynamic mapping relationship and a fixed interconnection relationship. The dynamic mapping relationship includes a connection relationship between the plurality of switching devices, and the fixed interconnection relationship includes a corresponding relationship between the slot silk screen information and the second communication port.
[0042] Further, the first chassis has a first network port and a first communication port; the second chassis has a second network port and a second communication port; the third chassis has a third network port, a third communication port, and a fourth communication port; the third communication port is connected to the first communication port through a cable, and the fourth communication port is connected to the second communication port through a cable; the first network port, the second network port, and the third network port are connected to the network switch through cables; and the fault information of the fault device includes link identification information of the fault device.
[0043] According to the fault information of the fault device and the allocation relationship table, determining the physical location of the fault device includes:
[0044] According to the link identification information of the fault device, determining a port number of the plurality of switching devices;
[0045] Based on the port number of the plurality of switching devices and the dynamic mapping relationship, determining the fourth communication port; and
[0046] Based on the fourth communication port and the fixed interconnection relationship, the second communication port and the fault device slot silk print information connected with the corresponding second communication port are determined, and then the fault device slot is determined.
[0047] Further, the first chassis comprises a first baseboard management controller, a central processing unit and a first connector; the first baseboard management controller is connected with the central processing unit, and the central processing unit is connected with the first connector; the first connector serves as the first communication port; the first baseboard management controller has a first baseboard management controller network port, which serves as the first network port; the second chassis comprises a second baseboard management controller, a graphics processing unit and a second connector; the second baseboard management controller is connected with the graphics processing unit, and the graphics processing unit is connected with the second connector; the second connector serves as the second communication port; the second baseboard management controller has a second baseboard management controller network port, which serves as the second network port; the third chassis further comprises a transmission expansion device, a switching unit, a third connector and a fourth connector; the switching unit is connected with the third baseboard management controller through the transmission expansion device; the third baseboard management controller has a third baseboard management controller network port, which serves as the third network port; the switching unit has a computing communication port and a graphics card communication port; the computing communication port is connected with the third connector, and the third connector serves as the third communication port; the graphics card communication port is connected with the fourth connector, and the fourth connector serves as the fourth communication port.
[0048] Before triggering the system management interrupt in response to the detection of the server link failure by the central processing unit, the method further comprises:
[0049] Powering on the server system, and establishing a network connection among the first baseboard management controller, the second baseboard management controller and the third baseboard management controller;
[0050] Obtaining the fixed interconnection relationship, and configuring the fixed interconnection relationship in a configuration file of the third baseboard management controller;
[0051] Obtaining the connection relationship of the switching device in the third chassis as a dynamic configuration relationship; and
[0052] Determining the distribution relationship table according to the fixed interconnection relationship and the dynamic configuration relationship.
[0053] Further, the configuration of the fixed interconnection relationship in the configuration file of the third baseboard management controller comprises:
[0054] Numbering the hardware devices in the first chassis, the second chassis and the third chassis; wherein, the hardware devices comprise the central processing unit, the graphics processing unit, the plurality of switching devices, the first connector, the second connector, the third connector and the fourth connector.
[0055] numbering ports of the plurality of switching devices; and
[0056] associating the ports of the plurality of switching devices with the numbering of the third connector and the fourth connector, associating the numbering of the third connector with the numbering of the first connector, associating the numbering of the fourth connector with the numbering of the second connector, and associating the numbering of the second connector with the slot silk screen information, obtaining the fixed interconnection relationship.
[0057] Further, obtaining the connection relationship of the switching device in the third chassis as the dynamic configuration relationship, comprising:
[0058] obtaining the internal interconnection relationship between the plurality of switching devices;
[0059] obtaining the mapping relationship between the port numbering of the plurality of switching devices and the device link identifier;
[0060] obtaining the mapping relationship between the asset information of the graphics processor and the second connector; and
[0061] taking the internal interconnection relationship, the mapping relationship between the port numbering of the plurality of switching devices and the device link identifier, and the mapping relationship between the asset information of the graphics processor and the third connector as the dynamic mapping relationship.
[0062] Further, the method further comprises:
[0063] parsing the fault information of the faulty device, obtaining the fault type and the fault level of the faulty device, wherein the fault type at least includes: non-recoverable fault, and the fault level at least includes: non-fatal fault and fatal fault; and
[0064] in response to the fault type of the faulty device being non-recoverable fault, and the fault level of the faulty device being non-fatal fault, resetting the faulty device.
[0065] Further, in response to the fault type of the faulty device being non-recoverable fault, and the fault level of the faulty device being fatal fault, issuing a device fault alarm.
[0066] In a third aspect, a computer system is provided, comprising one or more processors; and a memory associated with the one or more processors, the memory being configured to store a server system faulty device positioning program, the server system faulty device positioning program being configured to be executed by the one or more processors to implement the server system faulty device positioning method of the second aspect.
[0067] In a fourth aspect, a computer program product is provided, comprising a computer program, the computer program being configured to be executed by one or more processors to implement the server system faulty device positioning method of the second aspect.
[0068] In a fifth aspect, a non-transitory computer readable storage medium is provided, having stored thereon a server system fault device locating program, which, when executed by one or more processors, implements the server system fault device locating method according to the second aspect. BRIEF DESCRIPTION OF DRAWINGS
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0070] Fig. 1 is a structural schematic diagram of a server system according to an embodiment of the present application;
[0071] Fig. 2 is a schematic diagram of full connection between switching devices according to an embodiment of the present application;
[0072] Fig. 3 is a flow schematic diagram of a server system fault device locating method according to an embodiment of the present application;
[0073] Fig. 4 is a structural schematic diagram of a server system fault device locating apparatus according to an embodiment of the present application;
[0074] Fig. 5 is a structural schematic diagram of a computer system according to an embodiment of the present application;
[0075] Fig. 6 is a flow schematic diagram of a server system fault device locating method according to an embodiment of the present application;
[0076] Fig. 7 is a flow schematic diagram of another server system fault device locating method according to an embodiment of the present application;
[0077] Fig. 8 is a structural schematic diagram of a computer program product according to an embodiment of the present application;
[0078] Fig. 9 is a structural schematic diagram of a non-transitory computer readable storage medium according to an embodiment of the present application. DETAILED DESCRIPTION
[0079] In order to make the purpose, technical solutions and advantages of the present application more clear, the technical solutions in the embodiments of the present application will be described clearly and completely with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0080] Unless otherwise defined, technical terms or scientific terms used in the present disclosure shall have the ordinary meaning as understood by a person of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are used to distinguish different components. Similarly, the terms "one", "a", or "the" and similar terms do not denote a quantity of particular mentioned items, but indicate the existence of at least one of the items. The numbering in the drawings of the specification only indicates the distinction of respective functional components or modules, and does not indicate the logical relationship between the components or modules. The terms "comprise", "comprising", and similar terms mean that the elements or objects before the term encompass the elements or objects listed after the term and equivalents thereof, and do not exclude other elements or objects. The terms "connected" or "coupled" and similar terms do not limit to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like only indicate relative positional relationships, which can change when the absolute positions of the described objects change.
[0081] In the following, various embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that in the drawings, the same reference signs are assigned to components having substantially the same or similar structure and function, and repeated descriptions thereof will be omitted.
[0082] In view of the problem that it is difficult to physically locate the faulty device when the pooled heterogeneous resources fail, the present application provides the following technical solutions.
[0083] In some embodiments, as shown in FIG. 1, the server system includes at least one central processor 120, at least one graphics processor 320, at least one switching device 230, and a network switch 4. The at least one central processor 120 is arranged in at least one first chassis 10, the at least one graphics processor 320 is arranged in at least one second chassis 30, and the at least one switching device 230 is arranged in a third chassis 2. The at least one first chassis 10, the at least one second chassis 30, and the third chassis 2 are all arranged in the same server cabinet (not shown). Any first chassis 10 in the at least one first chassis 10, any second chassis 30 in the at least one second chassis 30, and the third chassis 2 are respectively connected to the network switch 4 through a first cable. Any first chassis 10 in the at least one first chassis 10 and any second chassis 30 in the at least one second chassis 30 are respectively connected to the third chassis 2 through a second cable.
[0084] The at least one first chassis 10 constitutes a first chassis 1, which is a computing resource pool. The third chassis 2 constitutes a third chassis 2, which is an input / output resource pool. The at least one second chassis 30 constitutes a second chassis 3, which is a graphics processor resource pool.
[0085] The first chassis 1 has a first chassis network port 1 L and at least one first input / output port 1 IO . The first chassis network port 1 L is communicatively connected with a network switch 4. The third chassis 2 has a third chassis network port 2 L , at least one second input / output port 2 IO1 and at least one third input / output port 2 IO2 . The third chassis network port 2 L is communicatively connected with the network switch 4. The number of the at least one second input / output port 2 IO1 corresponds to the number of the at least one first input / output port 1 IO . Any one of the at least one second input / output port 2 IO1 correspondingly connects one of the at least one first input / output port 1 IO , forming a communicative connection between the first chassis 1 and the third chassis 2. The second chassis 3 has a second chassis network port 3 L and at least one fourth input / output port 3 IO . The second chassis network port 3 L is communicatively connected with the network switch 4. The number of the at least one fourth input / output port 3 IO corresponds to the number of the at least one third input / output port 2 IO2 . Any one of the at least one fourth input / output port 3 IO correspondingly connects one of the at least one third input / output port 2 IO2 , forming a communicative connection between the second chassis 3 and the third chassis 2.
[0086] Generally, the first chassis, the third chassis and the second chassis are separately arranged in different chassis. The first chassis 1, the third chassis 2 and the second chassis 3, which are respectively connected with the network switch 4, constitute a local area network.
[0087] In some embodiments, the at least one first input / output port 1 IO , the at least one second input / output port 2 IO1 , the at least one third input / output port 2 IO2 , and the at least one fourth input / output port 3 IOThis is a high-speed connection port designed to achieve extremely high data transmission speeds, capable of reaching a data rate of 25Gbps per channel across 16 channels, resulting in a total data transmission speed of 400Gbps. In some embodiments, this high-speed connection port is a CDFP (400G Form Factor Pluggable) port. These interconnected ports are connected via CDFP cables.
[0088] The first, third, and second chassis are housed in separate enclosures, but they can all be housed within the same rack. Each resource pool contains a baseboard management controller responsible for managing its respective pool. The firmware of the baseboard management controllers uses OpenBMC (Open Baseboard Management Controller). The baseboard management controllers in the first, third, and second chassis are all connected to a network switch, forming a local area network. The pooled management controllers can communicate with the baseboard management controllers in each resource pool to achieve overall system management and control.
[0089] In some embodiments, the first chassis 1 includes at least one first chassis 10. Each of the at least one first chassis 10 has: a first network port 10. L and the first communication port 10 IO The first network port 10 of any first chassis 10 L After interconnection, it serves as the first chassis network port 1. L The first communication port 10 of any first chassis 10 IO As at least one first input / output port 1 IO One of the first input / output ports 1 IO The first chassis 10 includes a first baseboard management controller 110, a central processing unit 120, and a first connector 130. The first baseboard management controller 110 is connected to the central processing unit 120, and the central processing unit 120 is connected to the first connector 130, which serves as the first chassis communication port 10. IO The first baseboard management controller 110 has a first baseboard management controller network port 110. L As the first chassis network port 10 L .
[0090] In some embodiments, the third chassis 2 includes: a third baseboard management controller 21, a transmission expansion device 22, and a switching unit 23. The third baseboard management controller 21 has: a third baseboard management controller network port 21. LThe switching unit 23 is connected with the third baseboard management controller 21 through the transmission expansion device 22. The third baseboard management controller network port 21 L As the third chassis network port 2 L .
[0091] The switching unit 23 also has: at least one computing communication port 23 IO1 , and at least one heterogeneous communication port 23 IO2 . The number of at least one computing communication port 23 IO1 corresponds to the number of at least one second input / output port 2 IO1 . The number of at least one heterogeneous communication port 23 IO2 corresponds to the number of at least one third input / output port 2 IO2 . The switching unit 23 includes at least one switching device 230. Full connection is formed between multiple switching devices 230, as shown in FIG. 2. FIG. 2 shows the form of full connection between 8 switching devices.
[0092] Any switching device 230 in the at least one switching device has: a host port H, as a computing communication port 23 IO1 . A device port D, as a heterogeneous communication port 23 IO2 . An internal interconnection port F, for connecting multiple switching devices into full connection.
[0093] The switching unit contains several layers of PCIe (Peripheral Component Interconnect Express, high-speed serial computer expansion bus standard) switching boards, each layer of PCIe switching board has two switching devices, and the switching device is an input / output device for expanding a group of PCIe signals into multiple groups of PCIe signals, for expanding, networking, and distributing PCIe resources in the first chassis. Each switching device can lead out 5 high-speed signal ports (such as CDFP ports) for connection with the first chassis and the second chassis to form a fixed topology. In some embodiments, CDFP cables are used to connect the high-speed signal ports with the first and second chassis.
[0094] The pool management controller is connected with the switching device through an I 2 C (Inter Integrated Circuit, integrated circuit) link. The firmware of the switching device provides basic read / write register functions. The pool management controller sends commands to the switching device through the I 2 C link to read and write the internal registers of the switching device. Unified management and configuration of the switching device are realized. Device temperature, port flow information, port type, and device routing information are monitored.
[0095] In some embodiments, the second chassis 3 includes at least one second chassis 30. Each of the at least one second chassis 30 has a second network port 30. L and at least one second communication port 30 IO The second network port 30 of any second chassis 30 L After interconnection, it serves as the second chassis network port 3. L Communication port 30 of any second chassis 30 IO As one of at least four input / output ports, port 3 is a fourth input / output port. IO .
[0096] Further, the second chassis 30 includes: a second baseboard management controller 310, at least one graphics processor 320, and at least one third connector 330. The second baseboard management controller 310 is connected to any one of the at least one graphics processors 320, and any one of the at least one graphics processors 320 is connected to one of the at least one third connector 330. Each of the at least one third connector 330 serves as one of the at least one fourth input / output ports 3. IO .
[0097] The first baseboard management controller 110, the third baseboard management controller 21, and the second baseboard management controller 310 are all connected to the network switch 4 to form a local area network. Data transmission between the first baseboard management controller 110, the third baseboard management controller 21, and the second baseboard management controller 310 can be performed through the network switch 4.
[0098] In other embodiments, as shown in FIG3, the server system fault location method, applied to the server system described in the first aspect, includes:
[0099] S100: In response to the detection of a link failure, a system management interrupt is triggered;
[0100] S200: Parse the advanced error reporting register of the faulty device to determine the fault information of the faulty device;
[0101] S300: Determine the equipment asset information of the faulty equipment based on the fault information of the faulty equipment;
[0102] S400: In response to the fact that the equipment asset information of the faulty equipment does not contain the silkscreen information of the faulty equipment, the physical location of the faulty equipment is determined according to the fault information of the faulty equipment and the allocation relationship table, wherein the allocation relationship table includes at least: dynamic mapping relationship and fixed interconnection relationship.
[0103] The failure information of the failed device at least includes link identification information of the failed device. The link identification information includes bus number, device number and function number of the device in the link. In a PCIe link, the link identification information refers to PCIe link identification information of the device. Generally, the link identification information of the device in the PCIe link is represented by a set of BDF values. Illustratively, the bus number of the graphics processor device is 0000, the device number is 02, and the function number is 00.0. The link identification information, i.e., BDF, of the graphics processor device in the PCIe link can be represented as 0000:02:00.0. Generally, the link identification information of the device can be obtained by the lspci command.
[0104] However, since the graphics processor resources and the central processor resources are arranged in different chassis respectively, the basic input and output system cannot record the graphics processor slot positions not in the same chassis in the device asset information. Therefore, the link identification information and the allocation relationship table are needed to determine the physical position of the failed graphics processor.
[0105] Generally, when the whole system is powered on, the basic input and output system enumerates and processes PCIe devices, and allocates a set of link identification information (BDF) for each PCIe device. Then, the basic input and output system assembles the identified asset information into a Json data format and writes it into the shared memory, which includes the PCIe device link identification information in the first chassis and the slot silk screen information corresponding to the device, and the link identification information of the allocated graphics processor device. The first baseboard management controller obtains the asset information reported by the basic input and output system from the shared memory and stores it in the file system of the baseboard management controller, illustratively, as an assetInfo.json file.
[0106] According to the failure information of the failed device, the position of the failed device is determined, including:
[0107] S410: determining the port number of the switching device according to the link identification information of the failed device;
[0108] S420: determining the third communication port based on the port number of the switching device and the dynamic mapping relationship;
[0109] S430: determining the fourth communication port and the slot position of the failed device connected to the corresponding fourth communication port based on the third communication port and the fixed interconnection relationship, and further determining the physical position of the failed device.
[0110] Before triggering the system management interrupt in response to detecting the link failure, further including:
[0111] S010: powering on the server system and establishing a network connection between the first baseboard management controller, the second baseboard management controller and the third baseboard management controller;
[0112] S020: Obtain the fixed interconnection relationship and configure the fixed interconnection relationship as a configuration file in the third baseboard management controller.
[0113] S030: Obtain the dynamic configuration relationship between the third chassis and the second chassis.
[0114] S040: Determine the allocation relationship table according to the fixed interconnection relationship and the dynamic configuration relationship.
[0115] The allocation relationship table can be pre-configured, saved in the configuration file of the third baseboard management controller, and saved in the firmware of the switching device. It can also be configured when the system is powered on, saved in the configuration file of the third baseboard management controller, and saved in the firmware of the switching device.
[0116] The fixed interconnection relationship configured in the configuration file of the third baseboard management controller includes:
[0117] S021: Number the devices in the first chassis, the third chassis, and the second chassis.
[0118] S022: Associate the switching device port number in the third chassis with the switching device number.
[0119] S023: Associate the switching device port number in the third chassis with the first chassis number in the first chassis, the second subsystem number in the second chassis, and the communication port number of any second subsystem.
[0120] S024: Associate the communication port number of any second subsystem with the device slot.
[0121] S025: The above association relationship is the fixed interconnection relationship.
[0122] Obtaining the dynamic configuration relationship between the third chassis and the second chassis includes:
[0123] S031: Obtain the internal interconnection relationship between the switching device port numbers maintained by the firmware of the switching device.
[0124] S032: Obtain the mapping relationship between the switching device port number maintained by the firmware of the switching device and the device link identifier.
[0125] S033: Obtain the mapping relationship between the asset information of any graphics processor in the at least one graphics processor maintained by the third baseboard management controller and the at least one third connector.
[0126] S034: Take the mapping relationship between the internal interconnection relationship, the switch device port number and the device link identifier, and the mapping relationship between the asset information of any one of the at least one graphics processor and the at least one third connector as a dynamic mapping relationship.
[0127] The server system fault device positioning method further comprises:
[0128] S500: Analyze the fault information of the fault device to obtain the fault type and the fault level of the fault device, wherein the fault type at least includes an unrecoverable fault, and the fault level at least includes a non-fatal fault and a fatal fault;
[0129] S600: In response to the fault type of the fault device being an unrecoverable fault and the fault level of the fault device being a non-fatal fault, replace the fault device with an idle device of the same type in the resource pool where the fault device is located.
[0130] S700: Update the allocation relationship table.
[0131] Further, S700': in response to the fault type of the fault device being an unrecoverable fault and the fault level of the fault device being a fatal fault, perform device fault alarm.
[0132] The system management is established on the basis of the above-mentioned server system structure. The first baseboard management controller, the second baseboard management controller and the third baseboard management controller have LLDP (Link Layer Discovery Protocol) function. The pooling management controller can discover the IP (Internet Protocol) information of the corresponding baseboard management controller according to the broadcast of each baseboard management controller.
[0133] In order to facilitate the maintenance of the allocation relationship of the graphics processor of the system, the hardware devices need to be numbered. Illustratively, the first chassis is numbered with HostBox_ID, which is numbered as 0 to 7 respectively. The GPU resource pool is numbered with GPUBox_ID, which is numbered as 0 and 1 respectively. The CDFP port of the switch device is numbered with Switch_CDFP_ID, as follows:
[0134] First layer: CDFP0-0~CDFP0-4; CDFP1-0~CDFP1-4;
[0135] Second layer: CDFP2-0~CDFP2-4; CDFP3-0~CDFP3-4;
[0136] Third layer: CDFP4-0~CDFP4-4; CDFP5-0~CDFP5-4;
[0137] Fourth layer: CDFP6-0 ~ CDFP6-4; CDFP7-0 ~ CDFP7-4.
[0138] The Switch_Port number defined in the firmware of the switching device is set, such as 0x200060 corresponding to a set of internal registers of the switching device. The Box_CDFP_ID number is set for the third connector, such as CDFP0-CDFP15.
[0139] Through the above number setting, the mapping relationship between Switch_CDFP_ID and Switch_Port, the mapping relationship between Switch_CDFP_ID and HostBox_ID, GPUBox_ID, Box_CDFP_ID, and the mapping relationship between Box_CDFP_ID and the silk screen of the graphics processor slot are fixed interconnection relationships, which are configured in the configuration file of the third baseboard management controller.
[0140] The internal interconnection relationship between the switching device port numbers maintained by the firmware of the switching device, the mapping relationship between the switching device port numbers maintained by the firmware of the switching device and the device link identifier, and the mapping relationship between the asset information of any graphics processor in the at least one graphics processor maintained by the third baseboard management controller and the at least one third connector are dynamic configuration relationships dynamically maintained by the firmware of the switching chip.
[0141] On the basis of having the latest number information mentioned above, when a certain graphics processor fails, first, the switching device port (Switch_Port) number is determined according to the link identifier information, the switching device internal interconnection relationship is determined according to the switching device port number, the CDFP port number of the switching device is determined according to the switching device internal interconnection relationship, and then the third network port connected thereto is determined, so as to determine the slot of the graphics processor.
[0142] The pooling management controller, i.e., the third baseboard management controller, summarizes the fixed interconnection relationship and the dynamic configuration relationship to form an allocation relationship table. Illustratively, taking a Json format data structure as an example, the allocation relationship between the graphics processor device and the computing resource is shown as follows:
[0143] The pooling management controller provides a display interface and a Redfish interface for adjusting the graphics processor resource allocation and adjustment. For example, when the No. 0 graphics processor is removed from the computing resource under HostBox_ID 0 and allocated to the computing resource under HostBox_ID 1, the I 2The C command interacts with the firmware of the corresponding switching device to set the corresponding switching device register. After the setting is successful, the pooling management controller updates the allocation relationship table to ensure the correctness of the allocation relationship.
[0144] It should be understood that although the steps in the flowchart of FIG. 3 are shown in a sequence indicated by arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other sequences. Moreover, at least part of the steps in FIG. 3 can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of the sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.
[0145] In some embodiments, a server system fault device positioning method is applied to a server system, the server system comprising: at least one central processor 120, at least one graphics processor 320, at least one switching device 230, and a network switch 4; the at least one central processor 120 is arranged in at least one first chassis 10, the at least one graphics processor 320 is arranged in at least one second chassis 30, and the at least one switching device 230 is arranged in a third chassis 2. The at least one first chassis 10, the at least one second chassis 30, and the third chassis 2 are arranged in the same server cabinet. Any first chassis 10 in the at least one first chassis 10, any second chassis 30 in the at least one second chassis 30, and the third chassis 2 are respectively connected to the network switch 4 through a first cable. Any first chassis 10 in the at least one first chassis 10 and any second chassis 30 in the at least one second chassis 30 are respectively connected to the third chassis 2 through a second cable, and the third chassis 2 comprises a third baseboard management controller 21.
[0146] The server system fault device positioning method comprises: triggering a system management interrupt in response to the at least one central processor detecting a server link fault. Parsing a high-level error report register of the fault device to determine fault information of the fault device. According to the fault information of the fault device, determining device asset information of the fault device. In response to the slot silk screen information of the fault device not being included in the device asset information of the fault device, determining the physical location of the fault device through the third baseboard management controller according to the fault information of the fault device and an allocation relationship table, wherein the allocation relationship table at least comprises: a dynamic mapping relationship and a fixed interconnection relationship, the dynamic mapping relationship comprises a connection relationship between a plurality of switching devices, and the fixed interconnection relationship comprises a corresponding relationship between the slot silk screen information and the at least one second communication port.
[0147] In addition, any first chassis 10 of the at least one first chassis 10 has a first network port 10 L and a first communication port 10 IO . Any second chassis 30 of the at least one second chassis 30 has a second network port 30 L and at least one second communication port 30 IO . The third chassis 2 has a third network port (a third baseboard management controller network port 21 L ), at least one third communication port (a computing communication port 23 IO1 ), and at least one fourth communication port (a heterogeneous communication port 23 IO2 ). Any third communication port of the at least one third communication port is connected to one first communication port 10 IO through a corresponding connection by a second cable. Any fourth communication port of the at least one fourth communication port is connected to one second communication port 30 IO of the at least one second communication port 30 IO through a corresponding connection by a cable. The first network port 10 L , the second network port 30 L , and the third network port are respectively connected to the network switch 4 through a first cable. The fault information of the faulty device includes link identification information of the faulty device.
[0148] At this time, according to the fault information of the faulty device and the allocation relationship table, the physical location of the faulty device is determined, including: determining the port number of the switching device according to the link identification information of the faulty device. Based on the port number of the switching device and the dynamic mapping relationship, the fourth communication port is determined. Based on the fourth communication port and the fixed interconnection relationship, the second communication port and the slot silk screen information of the faulty device slot connected to the corresponding second communication port are determined, and then the faulty device slot is determined.
[0149] Further, any of the at least one first chassis 10 comprises a first baseboard management controller 110, a central processing unit 120 and a first connector 130. The first baseboard management controller 110 is connected with the central processing unit 120, and the central processing unit 120 is connected with the first connector 130. The first connector 130 serves as a first communication port. The first baseboard management controller 110 has a first baseboard management controller network port, which serves as a first network port. Any of the at least one second chassis 30 comprises a second baseboard management controller 310, at least one graphics processing unit 320 and at least one second connector 330. The second baseboard management controller 310 is connected with the at least one graphics processing unit 320, and any of the at least one graphics processing unit 320 is connected with one of the at least one second connector 330. The at least one second connector 330 serves as at least one second communication port. The second baseboard management controller 310 has a second baseboard management controller network port, which serves as a second network port. The third chassis 2 further comprises a transmission expansion device 22, a switching unit 23, at least one third connector and at least one fourth connector. The switching unit 23 is connected with the third baseboard management controller 21 through the transmission expansion device. The third baseboard management controller 21 has a third baseboard management controller network port, which serves as a third network port. The switching unit 23 has at least one computing communication port and at least one graphics card communication port. Any of the at least one computing communication port is connected with one of the at least one third connector, which serves as at least one third communication port. Any of the at least one graphics card communication port is connected with one of the at least one fourth connector, which serves as at least one fourth communication port.
[0150] At this time, in response to the central processing unit detecting the server link failure, before triggering the system management interrupt, further comprising: powering on the server system, establishing network connections between the first baseboard management controller, the second baseboard management controller and the third baseboard management controller. Obtaining the fixed interconnection relationship, and configuring the fixed interconnection relationship in the configuration file of the third baseboard management controller. Obtaining the connection relationship of the at least one switching device in the third chassis as a dynamic configuration relationship. According to the fixed interconnection relationship and the dynamic configuration relationship, determining the allocation relationship table.
[0151] Further, the fixed interconnection relationship is configured in a configuration file of the third baseboard management controller, including: numbering the hardware devices in the at least one first chassis, the at least one second chassis and the third chassis. The hardware devices include: a central processing unit, a graphics processing unit, a switching device, a first connector, a second connector, a third connector and a fourth connector. The ports of the switching device are numbered. The ports of the switching device are associated with the numbers of the third connector and the fourth connector. The number of the third connector is associated with the number of the first connector, and the number of the fourth connector is associated with the number of the second connector. The number of the second connector is associated with the slot silk screen information. The above association relationship is taken as the fixed interconnection relationship.
[0152] Further, the connection relationship of the at least one switching device in the third chassis is obtained as a dynamic configuration relationship, including: obtaining the internal interconnection relationship between the switching devices. The mapping relationship between the port numbers of the switching devices and the device link identifiers is obtained. The mapping relationship between the asset information of any one of the at least one graphics processing unit and the at least one second connector is obtained. The internal interconnection relationship, the mapping relationship between the port numbers of the switching devices and the device link identifiers, and the mapping relationship between the asset information of any one of the at least one graphics processing unit and the at least one third connector are taken as the dynamic mapping relationship.
[0153] Further, the fault positioning method further comprises: analyzing the fault information of the fault device, obtaining the fault type and the fault level of the fault device, wherein the fault type at least includes: non-recoverable fault, and the fault level at least includes: non-fatal fault and fatal fault. In response to the fault type of the fault device being non-recoverable fault, and the fault level of the fault device being non-fatal fault, the fault device is reset.
[0154] Further, in response to the fault type of the fault device being non-recoverable fault, and the fault level of the fault device being fatal fault, a device fault alarm is issued. The device fault alarm contains a device fault log. The fault log contains the chassis where the fault PCIe device is located, and the slot in the chassis. According to the device fault alarm, the operation and maintenance personnel can determine the physical location of the fault device in time and replace it.
[0155] In other embodiments, as shown in Figure 4, the server system fault device location device includes: an interrupt triggering module, used to trigger a system management interrupt in response to at least one central processing unit detecting a server link failure; a fault determination module, used to parse the advanced error report register of the faulty device to determine the fault information of the faulty device; an asset identification module, used to determine the device asset information of the faulty device based on the fault information of the faulty device; and a physical location module, used to determine the physical location of the faulty device through a third baseboard management controller based on the fault information of the faulty device and an allocation relationship table if the device asset information of the faulty device does not contain the slot silkscreen information of the faulty device. The allocation relationship table includes at least: dynamic mapping relationships and fixed interconnection relationships. The dynamic mapping relationships include the connection relationships between multiple switching devices, and the fixed interconnection relationships include the correspondence between slot silkscreen information and at least one second communication port.
[0156] The limitations of the server system fault location device described above can be found in the limitations of the server system fault location method described above, and will not be repeated here. Each module in the aforementioned server system fault location device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0157] In other embodiments, as shown in FIG5, the computer system includes one or more processors and a memory associated with the one or more processors. The memory stores a server system fault location program, which, when executed by the one or more processors, implements the server system fault location method described above. The server system fault location method described above will not be repeated here.
[0158] In other embodiments, as shown in FIG8, this application also provides a computer program product, including a computer program that, when executed by one or more processors, implements the server system fault device location method described above. The server system fault device location method described above will not be repeated here.
[0159] In other embodiments, Figure 9 shows a schematic diagram of the structure of a non-volatile computer-readable storage medium provided in an embodiment of this application. A non-volatile computer-readable storage medium stores a server system fault device location program. When the server system fault device location program is executed by one or more processors, it implements the server system fault device location method described above. The server system fault device location method described above will not be repeated here.
[0160] The technical scheme provided by the embodiments of the present application has the beneficial effects that: by implementing the server system and the fault device positioning method, device, system, computer program product and readable storage medium disclosed in the embodiments of the present application, flexible modification and rapid upgrade of key devices such as CPUs and GPUs can be facilitated, and the utilization rate of server hardware resources can be greatly improved. The device that causes the PCIe fault in the whole cabinet can be quickly located. For GPU devices, if the fault is non-recoverable and non-fatal, hot removal, hot addition operation can be performed on the GPU device, the device is reset, the normal operation of the business is recovered in time, and the maintenance pressure of the operation and maintenance personnel is reduced.
[0161] All the above technical schemes can be combined to form the embodiments of the present application, which will not be described one by one here.
[0162] FIG. 6 shows a flowchart of a server system fault device positioning method.
[0163] When the system is powered on, the basic input / output system of the first chassis enumerates and processes PCIe devices, and allocates a set of link identification information for each PCIe device. Then the basic input / output system assembles the identified asset information into a Json data format and writes it to the shared memory, which includes the PCIe device link identification information in the first chassis and the slot silk screen information corresponding to the device, and the link identification information of the allocated graphics processor device. The first baseboard management controller stores the asset information reported by the basic input / output system in the file system of the baseboard management controller from the shared memory, assuming as an assetInfo.json file.
[0164] When the central processing unit in the first chassis detects that a certain graphics processor device has an advanced fault report, it triggers a system management interrupt, which is processed by the basic input / output system of the first chassis.
[0165] The basic input / output system of the first chassis parses the advanced error report register of the faulty graphics processor to determine the error information of the faulty device and record the link identification information of the faulty device.
[0166] The basic input / output system of the first chassis reports the link identification information and error information of the faulty graphics processor device to the first baseboard management controller through the IPMI interface.
[0167] The first baseboard management controller queries the device asset information sent by the basic input / output system during boot-up according to the link identification information. If the corresponding silk screen information can be found, it is determined that the device is a PCIe device in the first chassis, then the related alarm log information is recorded, and the log information is reported to the pooled resource manager.
[0168] In response to the query not finding corresponding silk screen information, the device reports fault information to the pooled resource manager through an IPMI (Intelligent Platform Management Interface) or Redfish interface for the allocated graphics processor device, the pooled resource manager queries the allocation relationship table of the graphics processor device and the first chassis according to the HostBox_ID number and link identification information to determine the Box number, silk screen information, Box CDFP number and graphics processor asset information, records an alarm log by the pooled management controller, and synchronizes the log information to the first baseboard management controller.
[0169] The pooled management controller parses the error information of the faulty graphics processor device, and in response to determining that the error information is an unrecoverable and non-fatal fault, the pooled management controller performs a hot removal and hot addition operation on the faulty graphics processor device by an I 2 C command interacts with the firmware of the switching device to perform a hot removal, hot addition operation and device reset on the faulty graphics processor device.
[0170] If it is a fatal fault, the Box number and corresponding slot of the graphics processor device that fails can be determined by the operation and maintenance personnel according to the log for repair. The operation and maintenance personnel can query the log information through the first baseboard management controller and the pooled management controller.
[0171] When allocating graphics processor resources to the entire system, the allocation relationship table of the graphics processor device and the first chassis needs to be updated at the same time. For example, if the connection between the graphics processor with number 0 and the first chassis with number 0 is removed and allocated to the first chassis with number 1, the I 2 C command interacts with the firmware of the switching device to set the related registers of the corresponding switching firmware. After successful setting, the pooled management controller updates the allocation relationship table of the graphics processor device and the first chassis to ensure the accuracy of the allocation relationship and enable accurate positioning when the GPU fails.
[0172] FIG. 7 shows a flowchart of another server system fault device positioning method.
[0173] When the whole system is powered on and the host is started, the basic input / output system of the first chassis enumerates and processes the PCIe devices, and assigns a set of link identification information for each PCIe device. Then the basic input / output system assembles the identified asset information into a Json data format and writes it to the shared memory, which includes the PCIe device link identification information in the first chassis and the slot silk screen information corresponding to the device, and the link identification information of the assigned graphics processor device. The first baseboard management controller obtains the asset information reported by the basic input / output system from the shared memory and stores it in the file system of the baseboard management controller, assuming it is stored as an assetInfo.json file.
[0174] The pooled management controller establishes a distribution relationship table of the graphics processor device and the first chassis, taking the first chassis as the dimension, and sends the distribution relationship to the first baseboard management controller through the redfish interface.
[0175] The first baseboard management controller sets the silk screen information in the distribution relationship to the corresponding device asset information according to the obtained distribution relationship (which includes link identification information) and the device asset information assetInfo.json file according to the link identification information.
[0176] When the central processor in the first chassis detects that a certain graphics processor device has generated an advanced fault report, it triggers a system management interrupt, which is processed by the basic input / output system of the first chassis.
[0177] The basic input / output system of the first chassis parses the advanced error report register of the faulty graphics processor to determine the error information of the faulty device and records the link identification information of the faulty device.
[0178] The basic input / output system of the first chassis reports the link identification information and error information of the faulty graphics processor device to the first baseboard management controller through the IPMI interface.
[0179] The first baseboard management controller queries the device asset information sent by the basic input / output system at startup according to the link identification information, and in response to determining that the corresponding silk screen information can be found, it is determined that the device is a PCIe device in the first chassis, then records the related alarm log information, and reports the log information to the pooled resource manager.
[0180] If the corresponding silk screen information cannot be queried, the device is an allocated graphics processor device, and fault information is reported to the pooled resource manager through an IPMI or Redfish interface. The pooled resource manager queries a distribution relationship table of graphics processor devices and the first chassis according to a HostBox_ID number and link identification information, determines a Box number, silk screen information, a Box CDFP number, and graphics processor asset information, records an alarm log by the pooled management controller, and synchronizes the log information to the first baseboard management controller.
[0181] The pooled management controller parses error information of the faulty graphics processor device, and in response to determining that the error information is an unrecoverable and non-fatal fault, the pooled management controller performs a hot removal, hot addition, or device reset operation on the faulty graphics processor device by interacting with firmware of the switching device through an I 2 C command interacts with firmware of the switching device, performs a hot removal, hot addition, or device reset operation on the faulty graphics processor device.
[0182] If it is a fatal fault, an operation and maintenance personnel can determine the Box number and corresponding slot of the graphics processor device that fails according to the log, and perform repair. The operation and maintenance personnel can query the log information through the first baseboard management controller and the pooled management controller.
[0183] When allocating graphics processor resources in the entire system, the distribution relationship table of the graphics processor device and the first chassis needs to be updated at the same time. For example, when the graphics processor numbered 0 is removed from the first chassis numbered 0 and is allocated to the first chassis numbered 1, the pooled management controller needs to call an I 2 C command interacts with firmware of the switching device, sets a corresponding switching firmware related register, and after the setting is successful, the pooled management controller updates the distribution relationship table of the graphics processor device and the first chassis, to ensure the accuracy of the distribution relationship and to accurately locate the GPU when it fails.
[0184] In particular, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, as shown in FIG. 8, the embodiments of the present application include a computer program product including a computer program loaded on a computer readable medium, the computer program containing program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from a memory, or installed from a ROM. When the computer program is executed by an external processor, the above-mentioned functions defined in the method of the embodiments of the present application are performed.
[0185] It should be noted that the computer readable medium in the embodiments of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the embodiments of the present application, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, an optical fiber, an RF (Radio Frequency) or the like, or any suitable combination of the above.
[0186] The computer readable medium described above can be contained in the server described above; or can exist separately and not be assembled into the server. The computer readable medium described above carries one or more programs, when the one or more programs are executed by the server, the server: in response to detecting that the peripheral mode of the terminal is not activated, acquires the frame rate of the application on the terminal; when the frame rate meets the screen-off condition, judges whether the user is acquiring the screen information of the terminal; in response to the judgment result that the user is not acquiring the screen information of the terminal, controls the screen to enter the immediate dim mode.
[0187] Computer program code for carrying out operations of embodiments of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0188] The various embodiments in the specification are described in progressive manner, and the same or similar parts among the various embodiments can be mutually referred to. Each embodiment focuses on the difference from other embodiments. In particular, the system or system embodiments are described in a relatively simple manner, because they are substantially similar to the method embodiments. The relevant parts can be referred to the description of the method embodiments. The system and system embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0189] The above describes the technical solutions provided by the present application in detail, and the principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only to help understand the method and core idea of the present application; meanwhile, those skilled in the art can make changes in the specific implementation manners and application ranges according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.
[0190] The above only describes the preferred embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A server system, characterized by The application relates to a server system, comprising: a central processing unit, a graphics processing unit, a switching device and a network switch; wherein, the central processing unit is arranged in a first chassis, the graphics processing unit is arranged in a second chassis, and the switching device is arranged in a third chassis, and the first chassis, the second chassis and the third chassis are arranged in the same server cabinet; the first chassis, the second chassis and the third chassis are connected with the network switch through first cables; and the first chassis and the second chassis are connected with the third chassis through second cables.
2. The server system of claim 1, wherein, The first chassis has a first network port and a first communication port; the second chassis has a second network port and a second communication port; the third chassis has a third network port, a third communication port and a fourth communication port; wherein, the third communication port is connected with the first communication port through a first cable, the fourth communication port is connected with the second communication port through a first cable, and the first network port, the second network port and the third network port are connected with the network switch through second cables.
3. The server system of claim 2, wherein, The first chassis comprises a first baseboard management controller, the central processing unit and a first connector; the first baseboard management controller is connected with the central processing unit, and the central processing unit is connected with the first connector; the first connector serves as the first communication port; and the first baseboard management controller has a first baseboard management controller network port, which serves as the first network port.
4. The server system of claim 3, wherein, The first connector is a high-speed pluggable input / output connector.
5. The server system of claim 2, wherein, The second chassis comprises a second baseboard management controller, the graphics processing unit and a second connector; the second baseboard management controller is connected with the graphics processing unit, and the graphics processing unit is connected with the second connector; the second connector serves as the second communication port; and the second baseboard management controller has a second baseboard management controller network port, which serves as the second network port.
6. The server system of claim 5, wherein, The second connector is a high-speed pluggable input / output connector.
7. The server system of claim 2, wherein, The third chassis comprises a third baseboard management controller, a transmission expansion device, a switching unit, a third connector and a fourth connector; the switching unit is connected with the third baseboard management controller through the transmission expansion device; the third baseboard management controller has a third baseboard management controller network port, which serves as the third network port; the switching unit has a computing communication port and a graphics card communication port; the computing communication port is connected with the third connector, and the third connector serves as the third communication port; and the graphics card communication port is connected with the fourth connector, and the fourth connector serves as the fourth communication port.
8. The server system of claim 7, wherein, The switching unit comprises a plurality of switching devices; the plurality of switching devices are fully connected.
9. The server system of claim 8, wherein, Each switching device has: a host port, which serves as the computing communication port; a graphics card port, which serves as the graphics card communication port; and a network port, which serves as the network port. The device port is the communication port of the graphic card. The internal interconnection port is used to connect the plurality of switching devices into full connection.
10. The server system according to any of claims 7-9, characterized by The third connector and the fourth connector are high-speed pluggable input / output connectors.
11. The server system of claim 1, wherein, The second cable is a CDFP (400G Form Factor Pluggable) high-speed cable.
12. A server system fault device locating method, characterized by, The server system comprises a central processing unit, a graphic processing unit, a plurality of switching devices, and a network switch; the central processing unit is arranged in a first chassis, the graphic processing unit is arranged in a second chassis, and the plurality of switching devices are arranged in a third chassis; the first chassis, the second chassis, and the third chassis are arranged in the same server cabinet. The first chassis, the second chassis, and the third chassis are connected with the network switch through cables respectively. The first chassis and the second chassis are connected with the third chassis through cables respectively, and the third chassis comprises a third baseboard management controller. In response to the central processing unit detecting a server link fault, a system management interrupt is triggered. The advanced error report register of the fault device is parsed to determine the fault information of the fault device. According to the fault information of the fault device, the device asset information of the fault device is determined. In response to the slot silk screen information of the fault device not being included in the device asset information of the fault device, the physical location of the fault device is determined through the third baseboard management controller according to the fault information of the fault device and the allocation relationship table, wherein the allocation relationship table at least comprises a dynamic mapping relationship and a fixed interconnection relationship, the dynamic mapping relationship comprises the connection relationship between the plurality of switching devices, and the fixed interconnection relationship comprises the corresponding relationship between the slot silk screen information and the second communication port.
13. The server system fault device locating method of claim 12, wherein, The first chassis has a first network port and a first communication port; the second chassis has a second network port and a second communication port; the third chassis has a third network port, a third communication port, and a fourth communication port; the third communication port is connected with the first communication port through a cable, the fourth communication port is connected with the second communication port through a cable, and the first network port, the second network port, and the third network port are connected with the network switch through cables respectively. The fault information of the fault device comprises link identification information of the fault device. According to the fault information of the fault device and the allocation relationship table, the physical location of the fault device is determined, which comprises: According to the link identification information of the fault device, the port number of the plurality of switching devices is determined; Based on the port number of the plurality of switching devices and the dynamic mapping relationship, the fourth communication port is determined; and Based on the fourth communication port and the fixed interconnection relationship, the second communication port and the slot silk screen information of the fault device connected with the corresponding second communication port are determined, and then the slot of the fault device is determined.
14. The server system fault device locating method of claim 13, wherein, The first chassis comprises a first baseboard management controller, the central processor and a first connector; the first baseboard management controller is connected with the central processor, and the central processor is connected with the first connector; the first connector serves as the first communication port; the first baseboard management controller has a first baseboard management controller network port, which serves as the first network port; the second chassis comprises a second baseboard management controller, the graphic processor and a second connector; the second baseboard management controller is connected with the graphic processor, and the graphic processor is connected with the second connector; the second connector serves as the second communication port; the second baseboard management controller has a second baseboard management controller network port, which serves as the second network port; the third chassis further comprises a transmission expansion device, a switching unit, a third connector and a fourth connector; the switching unit is connected with the third baseboard management controller through the transmission expansion device; the third baseboard management controller has a third baseboard management controller network port, which serves as the third network port; the switching unit has a computing communication port and a graphic card communication port; the computing communication port is connected with the third connector, and the third connector serves as the third communication port; the graphic card communication port is connected with the fourth connector, and the fourth connector serves as the fourth communication port. Before triggering the system management interrupt in response to the detection of the server link failure by the central processor, the method further comprises: powering on the server system, and establishing a network connection among the first baseboard management controller, the second baseboard management controller and the third baseboard management controller; obtaining the fixed interconnection relationship and configuring the fixed interconnection relationship in a configuration file of the third baseboard management controller; obtaining the connection relationship of the plurality of switching devices in the third chassis as a dynamic configuration relationship; and determining an allocation relationship table according to the fixed interconnection relationship and the dynamic configuration relationship.
15. The server system fault device locating method of claim 14, wherein, The configuration of the fixed interconnection relationship in the configuration file of the third baseboard management controller comprises: numbering hardware devices in the first chassis, the second chassis and the third chassis; wherein the hardware devices comprise the central processor, the graphic processor, the plurality of switching devices, the first connector, the second connector, the third connector and the fourth connector; numbering ports of the plurality of switching devices; and associating the numbering of the ports of the plurality of switching devices with the numbering of the third connector and the fourth connector, associating the numbering of the third connector with the numbering of the first connector, associating the numbering of the fourth connector with the numbering of the second connector, and associating the numbering of the second connector with slot silk screen information, to obtain the fixed interconnection relationship.
16. The server system fault device locating method of claim 14, wherein, The obtaining of the connection relationship of the plurality of switching devices in the third chassis as a dynamic configuration relationship comprises: obtaining an internal interconnection relationship among the plurality of switching devices; obtaining a mapping relationship between port numbers of the plurality of switching devices and device link identifiers; obtaining a mapping relationship between asset information of the graphic processor and the second connector; and taking the internal interconnection relationship, the mapping relationship between port numbers of the plurality of switching devices and device link identifiers, and the mapping relationship between asset information of the graphic processor and the third connector as dynamic mapping relationships.
17. The server system fault device locating method of claim 12, wherein, The method further comprises: analyzing fault information of the faulty device to obtain a fault type and a fault level of the faulty device, wherein the fault type at least includes an unrecoverable fault, and the fault level at least includes a non-fatal fault and a fatal fault; and in response to the fault type of the faulty device being the unrecoverable fault and the fault level of the faulty device being the non-fatal fault, resetting the faulty device.
18. The server system fault device locating method according to claim 17, wherein in response to the fault type of the faulty device being the unrecoverable fault and the fault level of the faulty device being the fatal fault, issuing a device fault alarm.
19. A computer system, characterized in that comprise one or more processors; and a memory associated with the one or more processors, the memory being configured to store a server system fault device locating program, the server system fault device locating program being configured to be executed by the one or more processors to implement the server system fault device locating method according to any one of claims 12 to 18. The computer program is configured to be executed by one or more processors to implement the server system fault device locating method according to any one of claims 12 to 18.
20. A computer program product comprising a computer program, characterized in that, The computer program is configured to be executed by one or more processors to implement the server system fault device locating method according to any one of claims 12 to 18.
21. A non-transitory computer readable storage medium, comprising: The computer program is configured to be executed by one or more processors to implement the server system fault device locating method according to any one of claims 12 to 18.
Citation Information
Patent Citations
Data communication system and method
CN108768899A
Graphic processor system
CN109408451A
Extended memory error processing method and system, electronic equipment and storage medium
CN115686896A
Server system, fault device positioning method, computer system, program product and storage medium
CN118885324A
Big data server
CN205721583U