Server and server monitoring method

By adding a connector and a management board to the graphics processor board, the problem of untimely and incomplete information monitoring caused by the addition of devices on the GPU board is solved, and comprehensive coverage and rapid response to multiple computing devices are achieved, improving the stability and security of the graphics processor board.

CN120336128AInactive Publication Date: 2025-07-18INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510836142.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The increase in devices on GPU boards in traditional servers leads to problems such as untimely and incomplete information monitoring, which cannot meet the management needs of complex GPU boards.

Method used

A first connector is added to the graphics processor board and an independent management board is equipped with a first connector, which is connected to the graphics processor board to achieve comprehensive monitoring and rapid response to multiple computing devices.

Benefits of technology

It realizes comprehensive coverage and rapid response to information monitoring of multiple computing devices on the graphics processor board, improving the stability and security of the graphics processor board.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336128A_ABST
    Figure CN120336128A_ABST
Patent Text Reader

Abstract

The invention discloses a server and a server monitoring method, and relates to the technical field of servers, the server comprises a graphics processor board and a management board, a first connector is additionally arranged on the graphics processor board, the management board independent of the graphics processor board is additionally arranged, and the management board is connected with the graphics processor board through the first connector. The information monitoring device is used for monitoring the running state of the graphics processor board, so that on the basis of not changing the structure of the graphics processor board, comprehensive coverage and quick response of information monitoring of a plurality of computing devices on the graphics processor board are realized, and the problems of untimely and incomplete information monitoring caused by increase of the devices are solved; and the stability and the safety of the graphics processor board are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and in particular, to a server and a server monitoring method. Background Art

[0002] With the development of big data and artificial intelligence, Internet customers' demand for GPU (Graphics Processing Unit) computing resources has been increasing day by day. The hardware system architecture of traditional servers with GPUs has been widely used, and GPU boards used to carry GPUs and interconnect with computing nodes have emerged as the times require. However, as the functions of GPU boards become more and more complex and the components on the GPU boards become more numerous, the GPU board management function is insufficient and cannot meet the management requirements of complex GPU boards. Summary of the Invention

[0003] This application provides a server and a server monitoring method to at least solve the problem in the related art that the increase in components on the GPU board leads to untimely and incomplete information monitoring.

[0004] This application provides a server, including: a graphics processor board integrated with a first connector for performing a specified task; and a management board located on the graphics processor board and connected to the graphics processor board through the first connector for monitoring the operating state of the graphics processor board.

[0005] This application also provides a server monitoring method, including: communicating with a plurality of computing components on the graphics processor board through a first processor on the graphics processor board to obtain first state data of the plurality of computing components, and feeding back the first state data of the plurality of computing components to the management board through a first connector on the graphics processor board; the plurality of computing components are used for performing a specified task; the first state data of the plurality of computing components includes timing control data and communication link switching data of the plurality of computing components; receiving, by a first management controller on the management board, the first state data of the plurality of computing components sent by the first processor, communicating with the plurality of computing components to obtain second state data of the plurality of computing components, and monitoring the plurality of computing components based on the first state data and the second state data of the plurality of computing components; the second state data of the plurality of computing components includes the operating state of the plurality of computing components.

[0006] Through this application, a first connector is added to the graphics processing unit (GPU) board, and a management board independent of the GPU board is added. The management board is connected to the GPU board through the first connector and is used to monitor the operating status of the GPU board. Thus, without changing the structure of the GPU board, comprehensive coverage and rapid response of information monitoring for multiple computing devices on the GPU board are achieved, solving the problems of untimely and incomplete information monitoring caused by the increase in devices, and improving the stability and security of the GPU board. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0008] Figure 1 FIG. [X] is a topology structure diagram of an optional server provided by an embodiment of this application.

[0009] Figure 2 FIG. [X] is a topology structure diagram of an optional GPU board provided by an embodiment of this application.

[0010] Figure 3 FIG. [X] is a block diagram of an optional GPU board management and monitoring provided by an embodiment of this application.

[0011] Figure 4 FIG. [X] is a schematic flow diagram of an optional health management and anomaly detection provided by an embodiment of this application.

[0012] Figure 5 FIG. [X] is a layout diagram of an optional graphics manager board provided by an embodiment of this application.

[0013] Figure 6 FIG. [X] is a schematic diagram of the PCIe link interconnection relationship between optional OAM modules and the PCIe interconnection between OAM and PHY Retimer provided by an embodiment of this application.

[0014] Figure 7 FIG. [X] is a schematic diagram of the PCIe link connection relationship between an optional OAM module, a PCIe Retimer, and a fourth connector provided by an embodiment of this application.

[0015] Figure 8 FIG. [X] is a layout diagram of an optional management board provided by an embodiment of this application.

[0016] Figure 9 FIG. [X] is a schematic flow diagram of a server monitoring method provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0018] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0019] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] The following explains the professional terms involved in the present application:

[0021] GPU: Graphics Processing Unit, graphics processing unit.

[0022] CPLD: Complex Programmable Logic Device, complex programmable logic device.

[0023] BMC: Baseboard Management Controller, baseboard management controller.

[0024] CPU: Central Processing Unit, central processing unit.

[0025] SENSOR: sensor.

[0026] OAM: Operation and Administration Module, modular hardware unit, which is a hardware unit used to implement operation and management functions in modular design.

[0027] PCIe Retimer: Peripheral Component Interconnect Express Retimer, expansion bus retimer, which is a device used for signal regeneration and clock recovery on the PCIe link.

[0028] EXP Card: Expansion Card, an expansion card is a daughter board used on the GPU board to expand functions and connectivity.

[0029] PHY Retimer: Physical Layer Retimer, a key chip integrated on the expansion card, responsible for signal regeneration and timing adjustment. Especially during high-speed data transmission, it can restore signal integrity and ensure accurate data transmission at the physical layer.

[0030] QSFP: Quad Small Form Factor Pluggable, a four-channel small form factor pluggable interface.

[0031] PAM4: 4-Level Pulse Amplitude Modulation, four-level pulse amplitude modulation.

[0032] VR: Voltage Regulator, an electronic device used to convert the input power supply into a stable output voltage for use in an electronic system or a specific circuit.

[0033] FRU: Field Replaceable Unit, refers to those components or modules that can be directly replaced by the user at the system site without interrupting system operation or requiring professional technicians. These units usually include easily damaged or regularly maintained hardware parts such as power supplies, fan modules, hard disk drives, etc.

[0034] UART: Universal Asynchronous Receiver / Transmitter, a commonly used serial communication protocol for data transmission between computers or electronic devices.

[0035] JTAG: Joint Test Action Group, a standard interface protocol mainly used for testing and debugging integrated circuit chips, especially microprocessors, microcontrollers, and complex digital systems.

[0036] SPI: Serial Peripheral Interface, serial peripheral interface.

[0037] I2C: Inter-Integrated Circuit, an integrated circuit interconnection, a two-wire serial bus communication protocol for short-distance communication between microcontrollers, sensors, memories, and other integrated circuits (ICs).

[0038] GPIO: General Purpose Input / Output, a general-purpose input / output.

[0039] MDC: Management Data Clock, a management data clock.

[0040] MDIO: Management Data Input / Output, a management data input / output.

[0041] Redfish API: An application programming interface.

[0042] According to one aspect of the embodiments of the present application, a server is provided. Optionally, as Figure 1 shown, in this embodiment, the above server includes a graphics processor board 102 and a management board 104. Among them, a first connector 106 is integrated on the graphics processor board 102. The graphics processor board 102 is used to process complex artificial intelligence computing tasks and feedback the status data of the graphics processor board 102 to the management board 104 through the first connector 106. The management board 104 is located on the graphics processor board 102 and is used to monitor the operating status of the graphics processor board 102.

[0043] Graphics Processing Unit Board (GPU board), refers to a high-performance computing board designed for servers. Its core function is to carry and manage multiple computing devices to support complex artificial intelligence computing tasks. In the embodiments of this application, the graphics processing unit board is integrated with a first processor, multiple computing devices, and a first connector; the first processor is electrically connected to the multiple computing devices. Electrical connection refers to the physical connection established between two or more points in a circuit through conductive media (such as wires, cables, connectors, terminals, PCB traces, etc.). The first processor is used to communicate with the multiple computing devices to obtain the first status data of the multiple computing devices; the first status data of the multiple computing devices includes the timing control data and communication link switching data of the multiple computing devices. Among them, the first processor refers to the processor integrated on the GPU board, such as a complex programmable logic device, whose main function is to monitor the multiple computing devices and obtain the first status data of the multiple computing devices. Computing devices generally refer to various hardware components on the GPU board used to execute specific computing tasks, such as modular hardware units, modular hardware units, expansion cards, voltage regulators, etc.; the multiple computing devices are used to execute specified tasks. The first status data of the multiple computing devices refers to the information reflecting the operating conditions of the multiple computing devices on the GPU board, including but not limited to timing control data and communication link switching data. Among them, the timing control data refers to the information used to manage the power-on time points and operating sequences of the computing devices on the GPU board. By processing the timing control data by the first processor, it can ensure that each component is powered on and operates in the preset order at the correct moment, avoiding equipment damage or system instability caused by improper power-on order. The communication link switching data is the dynamic switching information of the communication path of the first processor between different devices. The first connector is used to connect to the management board and transmit the first status data of the multiple computing devices to the management board. The first connector refers to a dedicated interface for connecting the GPU board and the management board, establishing the physical and electrical connections between the key computing devices on the GPU board and the management board, so that the management board can collect and monitor the status data of these computing devices. For example, the first connector can adopt a GEN-Z 4C connector. GEN-Z is a new high-speed interconnection protocol designed to improve the data transmission speed and efficiency inside servers and data centers. The GEN-Z 4C connector is used to support the physical layer connection of the Gen-Z protocol.

[0044] Management board, integrated with a first management controller; the first management controller is electrically connected to a first processor and multiple computing devices, and is configured to receive first status data of the multiple computing devices sent by the first processor, communicate with the multiple computing devices to obtain second status data of the multiple computing devices, and monitor the multiple computing devices based on the first status data and the second status data of the multiple computing devices; the second status data of the multiple computing devices includes the operating status of the multiple computing devices. Among them, the management board is vertically inserted into the graphics processor board through a first connector, so that the management board is located on the graphics processor board, thereby realizing the function of monitoring the operating status of the graphics processor board without increasing the volume of the server. The management board refers to a dedicated board card integrated with a first management controller, and its main task is to be responsible for the monitoring and management of all computing devices on the GPU board. The management board communicates with the computing devices through its rich interfaces, collects and processes the first status data reported by the first processor, and the second status data directly obtained from the computing devices, ensuring the stability and efficiency of the GPU board operation. The first management controller is the core component of the management board. For example, the first management controller can be a BMC, which is responsible for communicating with the computing devices on the GPU board through rich interfaces. The first management controller not only receives the first status data of the computing devices reported by the first processor, but also directly obtains the second status data from the multiple computing devices, including operation logs and status information, for real-time monitoring and maintaining the healthy operation of the GPU board. The second status data refers to the operating status information directly obtained from the multiple computing devices through the first management controller. The second status data is crucial for evaluating the performance, health status, and fault diagnosis of the computing devices. The first management controller can provide comprehensive monitoring and management capabilities for the operating status of the GPU board by parsing the first status data and the second status data.

[0045] In this embodiment, the first processor on the graphics processor board and the first management controller on the management board are used as the core monitoring devices for implementing the information monitoring function. By dividing the functions of the first processor and the first management controller, the advantages of the large number of pins and stable operation of the first processor, and the rich interface functions of the first management controller are utilized, thereby fully enhancing the monitoring and management capabilities of the operating status of the graphics processor board; the first management controller not only collects the first status data reported by the first processor on the graphics processor board, but also directly queries the operating status information from the computing devices to form the second status data. The first management controller comprehensively analyzes and processes these two types of status data, realizing the full coverage and rapid response of information monitoring for multiple computing devices on the graphics processor board, solving the problems of untimely and incomplete information monitoring caused by the increase of devices, and improving the stability and security of the graphics processor board.

[0046] Optionally, when the first processor on the graphics processing unit (GPU) board starts up, it first completes its own initialization and prepares to receive and process signals from various computing devices. The first processor controls the power-on sequence of various computing devices by monitoring the voltage status of various computing devices, dynamically switches communication links, and simultaneously collects status information such as abnormal power-off monitoring and temperature reading to form first status data, and transmits the first status data to the first management controller through the first connector. The first management controller is integrated on the management board, receives the first status data about the computing devices sent by the first processor, and directly communicates with multiple computing devices using rich interfaces to obtain their operation logs and status information as second status data. The first management controller comprehensively analyzes the first status data and the second status data to achieve real-time monitoring of the computing devices.

[0047] Through the embodiments of the present application, a first connector is added to the GPU board, and a management board independent of the GPU board is added. The management board is connected to the GPU board through the first connector and is used to monitor the operating status of the GPU board. Thus, without changing the structure of the GPU board, comprehensive coverage and rapid response of information monitoring for multiple computing devices on the GPU board are achieved, solving the problems of untimely and incomplete information monitoring caused by the increase of devices, and improving the stability and security of the GPU board.

[0048] In an exemplary embodiment, Figure 2 is a topological structure diagram of an optional GPU board provided by the embodiments of the present application. As Figure 2 shown, multiple computing devices include multiple modular hardware units, multiple extended bus retimers, and multiple expansion cards.

[0049] Figure 3 is a management and monitoring block diagram of an optional GPU board provided by the embodiments of the present application. As Figure 3 shown, the first status data of multiple computing devices further includes GPU board timing control data, EXP card timing control data, UART path switching data, JTAG path switching data, and GPU board temperature reading data. The second status data of multiple computing devices includes the operating status of multiple modular hardware units, the operating status of multiple extended bus retimers, and the operating status of multiple expansion cards. The first processor (CPLD_1) on the GPU board can also implement the UART switching and JTAG switching functions, that is, switch the UART and JTAG of the first management controller (BMC_1) to eight OAMs respectively to enable the first management controller (BMC_1) to obtain the operating status of each OAM. After the first processor (CPLD_1) on the GPU board completes the acquisition of the board status, it finally uploads the information to the first management controller (BMC_1) through I2C after internal logic processing.

[0050] Among them, the GPU board timing control data refers to the data generated by the first processor (CPLD_1), which records the power-on sequence and timing control information of each computing device during the power-on and initialization of the GPU board. The EXP card timing control data is the data generated by the first processor (CPLD_1), which records the timing control information of key devices during the power-on and initialization of the EXP card. The UART path switching data is the data generated by the first processor (CPLD_1), which records the UART communication path switching status between the first management controller (BMC_1) and multiple OAM modules on the GPU board. When the first management controller (BMC_1) needs to read or control a specific OAM module, the first processor (CPLD_1) switches the path of the UART signal by controlling the GPIO interface, enabling the first management controller (BMC_1) to establish communication with the target OAM. The UART path switching data includes the path switching instruction and the execution result. The JTAG path switching data is the data generated by the first processor (CPLD_1), which records the JTAG access path switching status of each chip on the GPU board. When the first management controller (BMC_1) needs to access a specific chip through the JTAG interface for fault detection or firmware update, the first processor (CPLD_1) switches the JTAG signal path to that chip. The JTAG path switching data includes the switching command and the chip access status. As Figure 2 shown, sensors are usually integrated on the GPU board to monitor the operating temperature of the board and prevent system failures caused by overheating. The GPU board temperature reading data is the data obtained by the first processor (CPLD_1) through interaction with the sensor via the I2C master module, which reflects the real-time temperature status of the GPU board.

[0051] In this embodiment, a management board is added to the GPU board, and the first management controller on the management board is used to achieve efficient management of key devices (OAM, CPLD, VR, PCIe Retimer, EXP card) on the GPU board. In addition, through the cooperation of the first processor and the first management controller, functions such as timing control, voltage monitoring, abnormal power-off monitoring, temperature acquisition, OAM status monitoring, and Retimer chip status monitoring on the GPU board are realized.

[0052] In this embodiment, the modular hardware unit (OAM) refers to the modular hardware unit integrated on the GPU board, which communicates with the first processor and is used to execute specific computing tasks and management tasks. Each OAM can work independently, communicate with other components of the GPU board through interfaces such as PCIe (Peripheral Component Interconnect Express, expansion bus), and support horizontal expansion (scale out) to enhance the computing power of the server.

[0053] In this embodiment, the number of multiple expansion bus retimers is equal to the number of multiple modular hardware units. As Figure 2 shown, the GPU board integrates eight OAM modules. Therefore, there are eight PCIe Retimers (denoted as PCIeRetimer0~8 respectively), and their status information is obtained through the first management controller (such as BMC_1) to monitor the signal transmission efficiency and stability of the GPU board.

[0054] The expansion card (EXP card) is a daughter board on the GPU board used to expand functions and connectivity. It communicates with multiple modular hardware units to realize the interconnection of multiple modular hardware units with the modular hardware units of other node servers, supporting the communication requirements of the server cluster. The first management controller (such as BMC_1) monitors the status of the EXP card to ensure the normal operation of the GPU board and its expansion functions.

[0055] In this embodiment, the number of multiple modular hardware units is twice the number of multiple expansion cards; every two modular hardware units among the multiple modular hardware units are connected to one expansion card among the multiple expansion cards. Such a design can enable each expansion card to carry twice the communication volume of the modular hardware units. Compared with each hardware unit being independently connected to the expansion card, it can significantly improve the signal transmission bandwidth and efficiency, and centralize multiple hardware units to communicate on fewer expansion cards, which can reduce the wiring complexity inside the server, shorten the signal transmission path, help improve the signal quality, reduce signal delay and mutual electromagnetic interference, and ensure the stable transmission of high-speed signals.

[0056] Through this embodiment, the direct communication between the OAM module and the expansion card reduces the signal transmission levels and distances, and speeds up the data exchange speed; the interaction between the OAM module and the expansion bus retimer helps improve the signal quality.

[0057] In an exemplary embodiment, as Figure 2 shown, the expansion card among the multiple expansion cards includes a physical layer retimer and multiple pluggable optical modules; every two modular hardware units among the multiple modular hardware units are connected to the physical layer retimer in the expansion card among the multiple expansion cards; the physical layer retimer in the expansion card among the multiple expansion cards is electrically connected to the multiple pluggable optical modules in the expansion card among the multiple expansion cards.

[0058] As Figure 2 shown, through the data input / output interface (such as the MDC / MDIO interface) on the management board, the first management controller (BMC_1) can obtain the operating status of the PHY Retimer, including information such as signal quality and error rate.

[0059] The pluggable optical module is a component of the GPU board expansion card (EXP card) and is used to achieve high-speed data optical signal transmission. It is not only responsible for communication between GPU boards, but also for information exchange with external node servers or network devices. It is a key device for building server clusters and network connections. The pluggable optical module cooperates with the PHY Retimer chip on the EXP card through pulse amplitude modulation to ensure the stability and high bandwidth of data transmission. For example, Figure 2 As shown, the pluggable optical module can be QSFP. In this embodiment, four EXP cards (respectively recorded as EXP card 0, EXP card 1, EXP card 2, and EXP card 3) are installed on the GPU board. Each EXP card has a PHY Retimer chip and four QSFPs (respectively recorded as QSFP0, QSFP1, QSFP2, and QSFP3). The EXP card is used to realize the interconnection between the OAM in the server and the OAM of other node servers.

[0060] In this embodiment, the physical layer retimer in the expansion card among the multiple expansion cards transmits information with the multiple pluggable optical modules in the expansion card among the multiple expansion cards through pulse amplitude modulation; the multiple pluggable optical modules in the expansion card among the multiple expansion cards transmit information with each other through pulse amplitude modulation; the physical layer retimer in the expansion card among the multiple expansion cards transmits information with any two modular hardware units among the multiple modular hardware units through pulse amplitude modulation. Among them, the pulse amplitude modulation method is a key technology for internal signal transmission of the expansion card (EXP card). For example, the pulse amplitude modulation method can adopt PAM4, which transmits two bits of information by encoding four different levels in the signal, thereby increasing the data transmission rate and reducing electromagnetic interference.

[0061] like Figure 2 As shown, in this embodiment, eight OAMs (respectively denoted as OAM0-OAM7) are arranged on the GPU board, and the scale out (horizontal expansion refers to expanding the capacity of the system by adding more nodes to form a distributed cluster system) of each OAM is 8 channels, and the scale out of two OAMs is connected to one PHY Retimer, and the PHY Retimer is downstream connected to the QSFP chip, and information is transmitted between the PHY Retimer chip, the QSFP chip, and the OAM through PAM4; the first management controller (BMC_1) obtains the operating status data of the PHY Retimer chip through the data input and output interface (MDC / MDIO).

[0062] like Figure 2 As shown, each EXP card also integrates a serial peripheral interface flash memory and a field replaceable unit.

[0063] likeFigure 2 As shown, the expansion card among the multiple expansion cards further includes a voltage regulator (VR); the voltage regulator (VR) is used to provide an operating voltage for the expansion card among the multiple expansion cards; the voltage regulator (VR) is a key component on the expansion card and is responsible for converting the input raw voltage into a stable voltage suitable for use by each device on the expansion card.

[0064] Through this embodiment, the communication between the PHY Retimer and the modular hardware unit (OAM) ensures the stability and efficiency of high-speed data transmission; deploying multiple pluggable optical modules on each EXP card enables dynamic adjustment of the network connection ability of the GPU board according to actual needs. This modular design allows the GPU board to easily access different types of network devices or form a cluster with other GPU boards, enhancing the flexibility and adaptability of the GPU board.

[0065] In an exemplary embodiment, as Figure 2 shown, the management board is further integrated with a second processor, and the second processor is electrically connected to the first management controller and the first processor respectively. The second processor refers to the processor integrated on the management board, such as a complex programmable logic device (CPLD). Its main function is to receive and process the log data from the first processor, and then transmit the processed information to the first management controller. In this way, the second processor assists in realizing the centralized management and monitoring of the operating states of each modular hardware unit on the GPU board, enhancing the control ability over the complex hardware architecture, and ensuring the stable operation and efficient management of the server.

[0066] In this embodiment, when obtaining the operating states of multiple modular hardware units, as Figure 2As shown in the figure, the first management controller (BMC_1) is connected to the second processor (CPLD_2) using the UART interface or the I2C interface. The second processor (CPLD_2) sends specific instructions to the first processor (such as GPLD_1) on the GPU board through the UART interface, requesting to obtain the running data of multiple modular hardware units (such as OAM modules), such as data instructions for querying the temperature, power consumption, core frequency, video memory usage, etc. of the OAM. After receiving the instructions, the first processor (such as GPLD_1) on the GPU board uses its integrated UART interface or JTAG interface to query the status information of multiple modular hardware units (such as OAM modules), read the debug information of multiple modular hardware units (such as OAM modules), or perform a firmware health check, so as to obtain comprehensive log data. After the first processor processes the collected log data, it transmits the log data of multiple modular hardware units to the first management controller through the JTAG interface, or uploads the log data of multiple modular hardware units to the second processor on the management board through the UART interface. The second processor (CPLD_2) receives the log data from the first processor, further integrates and processes these data to ensure the accuracy and integrity of the information. The processed log data is transmitted to the first management controller (BMC_1) through the I2C interface or UART interface on the first management controller (BMC_1) to achieve the final aggregation of data. The first management controller (BMC_1) analyzes the comprehensive log data of multiple modular hardware units received to obtain the running status of each modular hardware unit.

[0067] Through this embodiment, a second processor is integrated on the management board, forming a dual-core management architecture with the first processor on the GPU board. Through division of labor and cooperation, the first processor is responsible for real-time monitoring and preliminary data processing, while the second processor undertakes in-depth analysis and reporting of the data, achieving efficient and comprehensive management of the running status of the GPU board.

[0068] In an exemplary embodiment, the multiple computing devices further include a voltage regulator; the voltage regulator is used to provide a working voltage for the multiple computing devices; the first status data of the multiple computing devices further includes abnormal power-off monitoring data of the graphics processor board.

[0069] Among them, the voltage regulator (VR) is a power management component in the multiple computing devices, and the voltage regulator is responsible for converting the input voltage into a stable voltage suitable for the operation of each computing device on the GPU board.

[0070] The abnormal power-off monitoring data is generated by the first processor (CPLD on the GPU board) and is used to record and report information on whether an abnormal power-off (i.e., abnormal power failure caused by a power supply fault) occurs during the operation of the voltage regulator. Such as Figure 2As shown, the first processor (CPLD_1) controls the enable signal of the voltage regulator on the GPU board and reads the power good signal through the GPIO interface. When the enable signal is high and the power good signal becomes low, its abnormal state is recorded (abnormal power-off monitoring). The abnormal power-off monitoring data can be used to evaluate the operating state of the computing devices on the GPU board, prevent system failures, and quickly locate the cause after a failure. The generation and reporting mechanism of the abnormal power-off monitoring data ensures that the GPU board can provide timely feedback in case of abnormal power supply, facilitating system management and maintenance.

[0071] Optionally, as Figure 2 shown, the first processor (CPLD_1) first initializes its GPIO interface, communicates with the voltage regulator through the timing control module to ensure the correct interaction of the power good signal (POWER GOOD) and the enable signal (ENABLE). The first processor (CPLD_1) reads the output voltage of the voltage regulator through its GPIO. The voltage information is transmitted through the GPIO cable, and the analog-to-digital conversion module (ADC) inside the first processor (CPLD_1) converts the analog voltage value into a digital signal for processing. The first processor (CPLD_1) compares the read output voltage with a preset voltage threshold. If the voltage fluctuation of the output voltage exceeds the threshold range, the first processor (CPLD_1) will record the abnormal state and generate voltage monitoring data. The first processor (CPLD_1) integrates the comparison result, the output voltage value, and any abnormal information to generate voltage monitoring data. Finally, the first processor (CPLD_1) uploads the generated voltage monitoring data to the first management controller (BMC_1) through the I2C slave module, completing the preliminary processing and reporting of the data, providing a basis for subsequent comprehensive status analysis.

[0072] Through this embodiment, the direct interaction between the first processor and the voltage regulator realizes the real-time monitoring of the voltage on the GPU board. When the output voltage of the voltage regulator deviates from the preset voltage threshold, the first processor can quickly identify and record the abnormal state, generating detailed voltage monitoring data. This immediate response mechanism significantly improves the problem of untimely monitoring, ensuring the server's quick response in case of abnormal situations such as voltage fluctuations.

[0073] In an exemplary embodiment, in a related anomaly detection scheme, methods of fixed threshold and unimodal analysis are often adopted. The fixed threshold strategy means that the system preset an invariant anomaly boundary in advance, and all monitoring data are compared with this threshold. However, this method ignores the impact of changes in the operating environment on the device performance, such as external factors like temperature and humidity, as well as the aging of the device itself over time and usage. Therefore, the fixed threshold often fails to accurately reflect the normal operating range of the device under different conditions, easily leading to false alarms or missed anomaly reports. At the same time, unimodal analysis only focuses on one type of data or feature (such as only monitoring temperature or only monitoring power consumption), ignoring the multi-dimensional complexity of the device operating state and unable to capture the interactions and dependencies between these parameters, thus limiting the accuracy and comprehensiveness of anomaly detection.

[0074] Therefore, in this embodiment, the first management controller is further configured to generate baseline thresholds corresponding to different operating parameters of multiple computing devices according to the second state data of the multiple computing devices through a pre-constructed threshold generation model; generate threshold compensation factors corresponding to different operating parameters of the multiple computing devices according to the environmental data of the multiple computing devices; determine dynamic parameter thresholds corresponding to different operating parameters of the multiple computing devices according to the baseline thresholds and threshold compensation factors corresponding to different operating parameters of the multiple computing devices; compare the operating parameter values corresponding to different operating parameters in the second state data of the multiple computing devices with the dynamic parameter thresholds corresponding to different operating parameters; determine the computing devices that meet the preset conditions among the multiple computing devices as abnormal computing devices, and feedback anomaly information to the first processor to instruct the first processor to take anomaly decisions corresponding to the abnormal computing devices; where the preset condition refers to the condition that the operating parameter value is greater than the corresponding dynamic parameter threshold.

[0075] Among them, the threshold generation model is constructed based on deep learning and is a model for dynamically generating parameter thresholds according to the working state of the current computing device. Figure 4 is a schematic flowchart of an optional health management and anomaly detection provided by an embodiment of the present application, as Figure 4As shown, the threshold generation model includes a Long Short-Term Memory network (LSTM) and a Gaussian Mixture Model (GMM). The LSTM is deployed in the first management controller and is used to capture and understand the trends and patterns of the operating status data of various computing devices (such as GPUs, CPLDs, VRs, etc.) on the GPU board over time. In this embodiment, the LSTM model extracts meaningful long-term dependence features (such as the power consumption change and temperature fluctuation of the GPU, etc.) from the second status data obtained from the GPU board and outputs the hidden state to the GMM model. Here, the hidden state refers to the state vector that is dynamically updated over time series progress inside the network when the LSTM model analyzes the operating status data of the computing devices (such as GPUs, CPLDs, VRs, etc.) on the GPU board. This state vector synthesizes all the information from the start of the sequence to the current time point, including the historical trends of operating parameters such as GPU load changes, power consumption, and temperature, as well as the performance of these parameters in different working modes (such as training, inference, and idle). The GMM model generates a threshold distribution curve of the operating parameters of multiple computing devices that matches the current working mode based on the hidden state of the LSTM model. A series of points on the curve represent the expected normal parameter range under a given mode, that is, the baseline thresholds corresponding to different operating parameters of multiple computing devices.

[0076] The environmental data contains external condition information such as the temperature, humidity, and air flow rate of the environment where the GPU board is located. The threshold compensation factor is a dynamically adjusted value calculated based on the environmental data and is used to compensate for the impact of environmental changes on the baseline threshold to ensure the accuracy of the dynamic parameter threshold. For example, as Figure 4 shown, by fusing the external sensor data such as environmental temperature and humidity and rack position through a Kalman filter, the threshold compensation factor is obtained, including the temperature change amount ΔT_env and the air flow rate Airflow_rate. The dynamic parameter threshold is the actual threshold adaptively generated based on the current working mode and environmental conditions of the computing device by combining the baseline threshold and the threshold compensation factor. For example, the dynamic parameter threshold Threshold_adj can be expressed as: Threshold_adj = f(Threshold_base, ΔT_env, Airflow_rate), where the dynamic parameter threshold Threshold_adj can be calculated by using a weighted summation method of the threshold compensation factor and the baseline threshold.

[0077] An abnormal computing device refers to a computing device whose operating parameter value in the second state data exceeds the dynamic parameter threshold. The abnormal information is the alarm information generated by the first management controller when detecting an abnormal computing device, which includes the identifier of the abnormal computing device, the abnormal parameter, and the exceeded dynamic parameter threshold, and is used to assist the first processor in making abnormal decisions. The abnormal decision refers to the processing actions taken by the first processor after receiving the abnormal information according to the preset maintenance strategy, such as resetting the abnormal device, reducing the load, or isolating the faulty component, etc., to ensure the stable operation of the server system.

[0078] Optionally, as Figure 4 shown, the first management controller collects the second state data from multiple computing devices in real time, including operating parameters such as temperature, power consumption, and frequency, as well as environmental data such as temperature, humidity, and air flow rate; uses the pre-trained LSTM-GMM model to input the second state data of each computing device within a preset time period to generate the baseline threshold reflecting the normal operation range of each computing device; fuses the environmental data through the Kalman filter, calculates the threshold compensation factor according to the environmental change, and adjusts the baseline threshold using the threshold compensation factor to adapt to the current environment (such as in the way of weighted summation) to obtain the dynamic parameter thresholds corresponding to different operating parameters in multiple computing devices; compares the operating parameter values corresponding to different operating parameters in the second state data of multiple computing devices with the corresponding dynamic parameter thresholds; if the operating parameter value exceeds the corresponding dynamic parameter threshold, marks the corresponding computing device as an abnormal computing device, and feeds back the abnormal information to the first processor. The first processor takes the abnormal decision corresponding to the abnormal computing device.

[0079] For example, taking the OAM module as an example, after the first management controller receives the second operating states of multiple modular hardware units, it identifies the data formats of the second operating states of the multiple modular hardware units. According to the data formats of the second operating states of the multiple modular hardware units, with reference to the communication protocol SOP provided by the OAM module manufacturer, it extracts the operating parameter values representing different operating parameters from the received data. For example, it extracts information such as temperature values and power consumption values according to the positions and lengths specified in the communication protocol. The first management controller performs format conversion and processing on the extracted operating parameter values of different operating parameters to obtain operating state values with practical significance. For example, it converts hexadecimal data into decimal numerical values to obtain parameter values such as temperature and frequency. Since there are some missing data or data with large deviations in the operating parameter values of different operating parameters being processed, the first management controller uses the "filling method" to handle the missing values. For a small number of missing values, the mean, median, or mode can be used for filling. Taking temperature data as an example, the mean value of temperature will be used to fill the missing values. After that, based on the machine learning method, the first management controller uses the Isolation Forest algorithm to effectively identify the outliers in the data and perform data elimination. After the above preprocessing, it compares the operating parameter values corresponding to different operating parameters in the operating states of the processed multiple modular hardware units with the dynamic parameter thresholds corresponding to different operating parameters, so as to judge whether the OAM is operating normally. For example, if the temperature of the OAM exceeds the set maximum temperature threshold, then it can be judged that the OAM is in an overheated state and corresponding measures may need to be taken, such as adjusting the fan speed to enhance heat dissipation. The first management controller determines the modular hardware units with operating parameter values greater than the corresponding dynamic parameter thresholds among the multiple modular hardware units as abnormal modular hardware units. When the first management controller obtains that the OAM module is operating abnormally by parsing the OAM module data, at this time, the first management controller can inform the first processor (such as CPLD) of the GPU board through I2C in the form of writing registers which OAM module is currently abnormal and what type of abnormality exists. At this time, the first processor (such as CPLD) restarts the abnormal OAM module by resetting or powering off and then powering on the current abnormal OAM. After the restart is completed, if the first management controller determines that the current abnormal state has been lifted, it can continue to operate normally; if the first management controller still judges that the current OAM module is abnormal, the first management controller informs the first processor (such as CPLD) through I2C to directly power off this OAM module and isolate it to avoid affecting other OAM modules.

[0080] Through this embodiment, the first management controller generates baseline thresholds for different operating parameters according to the second state data of multiple computing devices on the GPU board through a pre-constructed threshold generation model. At the same time, a threshold compensation factor is also generated based on environmental data (such as temperature and humidity, air flow rate, etc.). By combining the baseline threshold and the compensation factor, the dynamic parameter threshold suitable for the current environmental conditions is determined dynamically, solving the problem in the related art that the fixed threshold method cannot adapt to environmental changes. The adaptive generation of the dynamic threshold in this embodiment ensures that the threshold can be adjusted in real time according to the actual environmental conditions, so as to more accurately judge the abnormal state and improve the accuracy and efficiency of abnormal detection.

[0081] In an exemplary embodiment, the first management controller is further configured to determine the fault source according to the abnormal data of the abnormal computing device through a pre-constructed fault propagation graph network; predict the causal characteristics of the abnormal data of the abnormal computing device through a pre-constructed causal discovery model; and parse the causal characteristics of the abnormal data of the abnormal computing device through a pre-constructed counterfactual interpreter to obtain an interpretability report of the abnormal computing device.

[0082] Among them, the fault propagation graph network (FPGN) is a graph structure based on the physical layout and power supply topology among components, and uses a graph attention network (GAT) to model the fault propagation relationship among various components on the GPU board. When only one component is abnormal, the abnormal component is the fault source; when multiple components are abnormal at the same time, FPGN traces and determines the most likely fault source through a random walk algorithm, providing technical support for the rapid location and solution of faults.

[0083] The causal discovery model is a model that uses the PC (Peter-Clark) algorithm to extract key causal features from historical maintenance logs. It can identify the causal relationships between different component anomalies on the GPU board, providing a theoretical basis for the in-depth analysis of abnormal data and fault prediction. The training process of the causal discovery model is as follows: obtain training samples; each sample contains a set of feature data and labels, and a set of feature data includes environmental parameters (such as temperature, humidity), hardware operating status (such as power consumption, voltage), software activities (such as task load, firmware version), etc.; the labels are used to annotate the causal relationships contained in the sample data. For example, a sample may be labeled as "high load causes the GPU temperature to rise" or "power failure triggers abnormal power-off of the server". In each training iteration, a set of feature data in the training sample is input into the causal discovery model. The causal discovery model outputs the predicted causal relationship according to the current parameters, uses a loss function (such as cross-entropy loss) to measure the difference between the causal relationship output by the model and the sample label, and adjusts the model parameters according to the gradient of the loss function through the backpropagation algorithm (such as gradient descent). The goal is to minimize the loss function, that is, to make the causal relationship predicted by the model as consistent as possible with the true label in the sample. Repeat the forward propagation and backpropagation processes until the model meets the preset stop conditions, such as reaching the maximum number of iterations, the loss function converges to a very small value, or the performance on the validation set no longer improves, to obtain the trained causal discovery model.

[0084] The counterfactual interpreter is a special module that, based on the causal discovery model, conducts counterfactual analysis on the causal characteristics of abnormal data, that is, explores questions such as "if condition X did not occur, would anomaly Y still appear?", generates an interpretability report on the operating state of the abnormal computing device, and helps maintenance personnel understand the potential causes and influencing factors of the fault. For example, the counterfactual interpreter can generate an interpretability report of "if the voltage decreased by 5% at that time, the anomaly probability would decrease by X%" for each abnormal event. The training process of the counterfactual interpreter is as follows: obtain training samples, each sample contains multiple features and labels of the device state, the multiple features include temperature, power consumption, voltage, system load, fan speed, etc., and the operating state of the device (normal or abnormal), and the label is used to indicate which features (or feature combinations) may bring about a state recovery when the device is in an abnormal state. For example, if the overheating of the GPU board causes an anomaly, the label may indicate that reducing the temperature (by increasing the fan speed or reducing the load) is a possible counterfactual scenario. Input the samples containing the abnormal state into the counterfactual interpreter, the counterfactual interpreter generates one or more counterfactual explanations (i.e., the changed feature values), calculates the gap between the generated counterfactual explanations and the true labels, and updates the model parameters through backpropagation according to the gradient of the loss function. The goal is to minimize the loss, that is, the generated counterfactual explanations should be as close as possible to the actual effective counterfactual conditions until the preset stop condition is met, and the trained counterfactual interpreter is obtained.

[0085] Optionally, as Figure 4 shown, the first management controller continuously collects the operating states and interaction data of each device on the GPU board, constructs potential fault propagation paths between GPU components using the Graph Attention Network (GAT) based on the PCB layout and power supply topology, trains the GAT model using historical fault data to identify and quantify the possibility of fault propagation. When the first management controller detects an abnormal computing device, it applies the random walk technique in the fault propagation graph network to trace the most likely path of abnormal data propagation and locate the fault source; inputs the abnormal data of the fault source into the pre-constructed causal discovery model. The causal discovery model outputs the causal characteristics of the abnormal data of the fault source. The first management controller uses the counterfactual interpreter to conduct counterfactual reasoning based on the causal feature prediction results, explores the possible changes of the abnormal data under different conditions, integrates the counterfactual analysis results, generates an interpretability report describing the operating state of the abnormal computing device and its possible causes, and sends the report to the second management controller on the server motherboard through the Redfish API or a similar interface for further processing or display in the management interface.

[0086] Through this embodiment, by introducing a fault propagation graph network, it is possible to quickly locate the real source of the anomaly based on the component relationship graph and the fault propagation relationship model, reducing the time and labor costs of fault troubleshooting and improving the maintenance efficiency; by using a causal discovery model, it is not limited to the analysis of surface fault phenomena, but delves into the causal relationships behind the abnormal data, helping to understand the root causes of faults, prevent the recurrence of similar faults, and enhance the overall system stability; by using a counterfactual interpreter, it provides valuable hypothesis analysis and decision-making basis for the maintenance team.

[0087] In an exemplary embodiment, the first management controller is further configured to determine the computing devices that do not meet the preset conditions among the multiple computing devices as normal computing devices; through a pre-constructed health index model, perform fusion processing on the first state data and the second state data of the normal computing devices to obtain the cross-modal features of the normal computing devices, and based on the cross-modal features of the normal computing devices, predict the health index and the remaining useful life of the normal computing devices; determine the optimal maintenance strategy according to the health index and the remaining useful life of the normal computing devices, and use the optimal maintenance strategy to maintain the normal computing devices.

[0088] Among them, the health index model is a multi-modal fusion model that comprehensively analyzes the first state data (such as temperature, power consumption) and the second state data (such as log records, topological relationships) of the computing devices on the GPU board, predicts the health status and the remaining useful life (RUL) of the computing devices, and provides data support for maintenance decisions. Preferably, the large health index model can be compressed through knowledge distillation technology, and the compressed health index model can be deployed in the first management controller, or the health index model can be dynamically allocated to multiple cloud computing nodes through model slicing technology.

[0089] Cross-modal features refer to comprehensive features that integrate temporal sensing data, maintenance log texts, and topological information in the first-state data and the second-state data. A hierarchical feature fusion network is designed in the health index model. The underlying layer of this network uses a one-dimensional convolutional neural network (1D-CNN) to process temporal sensing data, the middle layer uses a Transformer to process maintenance log texts, and the upper layer integrates topological information through a graph convolutional network. Through an attention gating mechanism, the weight allocation of each modal feature is dynamically adjusted to obtain cross-modal features. The cross-modal features can be expressed as: HI = Σ(α_i·HI_sensor + β_j·HI_log + γ_k·HI_topology), where HI_sensor represents the output of the underlying layer; HI_log represents the output of the middle layer; HI_topology represents the output of the upper layer; α_i, β_j, and γ_k represent weights respectively.

[0090] The optimal maintenance strategy refers to the most effective and economical maintenance plan obtained through algorithm optimization by considering factors such as maintenance costs and downtime losses based on the health index and remaining useful life (RUL) of the computing device.

[0091] In some embodiments, in traditional maintenance decision-making schemes, rule-based scheduling is the most common practice. This method relies on preset rules and thresholds, and once the state of the computing device exceeds the normal range, corresponding maintenance actions will be triggered. However, this method has obvious limitations: First, it is too rigid and cannot adapt to the dynamic changes of the computing environment because the rules are usually static and cannot be automatically adjusted to cope with new operating conditions. Second, rule-based scheduling often only focuses on one or a few goals, such as simply reducing costs or minimizing downtime, while ignoring the long-term stability of system performance. Finally, it is difficult to verify the effectiveness of the maintenance strategy with this method. Once the rules are set improperly, it may lead to over-maintenance or under-maintenance, which will instead damage the overall health of the system.

[0092] Therefore, to solve the above problems, in this embodiment, multiple objectives of the optimization engine are defined, such as minimizing maintenance costs, minimizing downtime, maximizing system performance, etc.; for each normal computing device, according to its health index and remaining service life, an initial population including a variety of maintenance strategies corresponding to the health index and remaining service life of the normal computing device is created, such as replacing spare parts, upgrading firmware, adjusting the load, etc., and each strategy is encoded into a binary or real number sequence for the optimization engine to process; an evaluation function is designed to evaluate the costs, downtime, system performance, etc. of each maintenance strategy; the Non-dominated Sorting Genetic Algorithm II (NSGA-II) algorithm is used for iterative optimization. Through genetic operations such as selection, crossover, and mutation, a new generation of population is generated, and at the same time, non-dominated sorting and crowding degree calculation are performed on the population to screen out the Pareto optimal solution set, that is, there is no strategy that can simultaneously outperform the strategies in the current set among different maintenance objectives. Finally, the NSGA-II optimization engine will output a set of Pareto optimal maintenance plans, and the operation and maintenance personnel can select the most suitable maintenance strategy (i.e., the best maintenance strategy) from the set of Pareto optimal maintenance plans according to the specific business requirements and cost budget to execute, so as to achieve the best balance among cost, time, and performance.

[0093] Among them, a pre-established maintenance cost model can be used to evaluate the cost of each maintenance strategy. The maintenance cost model is constructed with a multi-dimensional loss function, which takes the holding cost of spare part inventory, the cost loss of downtime, and the cost of manual maintenance, etc. as input variables, and the output is the total maintenance cost. Using historical maintenance data and cost records, the maintenance cost model is trained and calibrated to ensure that it can accurately predict the cost changes under different maintenance decisions.

[0094] Among them, in the process of finding the optimal maintenance plan, each maintenance strategy in the population is a strategy verified by a virtual maintenance simulator. Among them, the virtual maintenance simulator is a simulation tool based on digital twin technology. By establishing a life prediction model driven by the Arrhenius equation (a relational expression describing the change of the chemical reaction rate constant with temperature), it simulates the performance changes of components under different environmental conditions and workloads, so as to simulate and predict the long-term operation effect and possible failure modes of the server GPU board under different maintenance strategies.

[0095] This embodiment adopts a method of multi-objective optimization combined with a virtual maintenance simulator. First, through the NSGA-II optimization engine, based on the health index prediction and cost model, a set of Pareto-optimal maintenance plan sets is output. This method goes beyond the limitations of traditional rules and can find the best balance among multiple objectives such as cost, downtime, and system performance, ensuring that maintenance actions are both economical and efficient. More importantly, this embodiment also introduces a virtual maintenance simulator. Based on physics-based degradation modeling and digital twin technology, it can simulate the long-term effects of different maintenance strategies in a virtual environment, thereby verifying the feasibility and effectiveness of the plan. This approach greatly reduces the cost and risk of actual trial and error, provides scientific and flexible decision-making support for the maintenance team, and ensures the accuracy of maintenance decisions and the continuous optimization of system operation. In summary, through intelligent multi-objective optimization and virtual verification, this embodiment significantly improves the accuracy and efficiency of maintenance decisions, contributing to the construction of a more stable and reliable server GPU board operation and maintenance system.

[0096] Optionally, as Figure 4 shown, the first management controller determines the computing devices that do not meet the preset conditions among the multiple computing devices as normal computing devices; fuses the first state data and the second state data of the normal computing devices to generate cross-modal features including various operating states and maintenance record information; uses the pre-trained health index model to analyze the cross-modal features of the normal computing devices, predicts the health index and remaining useful life (RUL) of each normal computing device; based on the health index and RUL prediction results of each normal computing device, constructs an initial population containing various maintenance strategies corresponding to the health index and remaining useful life of the normal computing devices, and adopts a multi-objective optimization scheduling algorithm to comprehensively consider maintenance costs, downtime losses, etc., to generate a Pareto-optimal maintenance strategy set; selects the best maintenance strategy from the Pareto-optimal maintenance strategy set according to the requirements, and performs maintenance on the normal computing devices according to the determined best maintenance strategy, such as adjusting the cooling system, performing preventive replacement, etc., to ensure the continuous and stable operation of the GPU board.

[0097] Through this embodiment, by using the health index model to analyze the cross-modal features of normal computing devices, it is possible to comprehensively consider multiple factors, predict the health status and remaining useful life of the equipment, provide a scientific basis for maintenance personnel, formulate more reasonable maintenance strategies, avoid blind maintenance or over-maintenance, and save resources; the maintenance strategy based on health prediction can plan the maintenance or replacement of the equipment in advance, avoiding the risks of emergency shutdown and data loss caused by sudden equipment failures, thereby reducing economic losses and business interruptions.

[0098] In an exemplary embodiment, the first management controller is further configured to generate a health index of a normal computing device based on the cross-modal features of the normal computing device through a pre-constructed neural network; and perform time series extrapolation on the health index of the normal computing device through a pre-constructed exponential decay prediction model to obtain the remaining useful life of the normal computing device.

[0099] The neural network is a machine learning algorithm that mimics the structure of human brain neurons and can learn the complex relationship between input data and output results through training. For example, the neural network can be a Bayesian neural network with Monte Carlo dropout integrated in its output layer to provide a confidence interval for the health index. In this embodiment, the first management controller uses the neural network to analyze the relationship between cross-modal features and the health index to generate a more accurate health assessment result.

[0100] The exponential decay prediction model is a prediction model that predicts the future health state and remaining useful life of a computing device based on the decay law of the health index of the computing device over time. The training process of the exponential decay prediction model is as follows: Obtain training samples, each of which contains a series of time series data points recording the performance metrics of specific components (such as capacitors, radiators) at different time points, as well as labels that indicate the final failure time point of the corresponding components of each sample or are labeled in the form of remaining useful life. Input the time series data of the samples into the model, and the model predicts the future state or remaining useful life of specific components. Use a loss function (such as mean squared error MSE, mean absolute error MAE) to measure the deviation between the predicted value of the model and the sample label, and adjust the model parameters according to the gradient of the loss function through a backpropagation algorithm (such as stochastic gradient descent SGD). The goal is to minimize the prediction error until a preset training stop condition is reached, such as the maximum number of iterations, the model loss no longer decreases significantly, etc., to obtain a trained exponential decay prediction model.

[0101] In relevant health assessment schemes, the health status of equipment is usually represented by a single health index. Although this approach is intuitive, a single health index is difficult to capture the complexity and variability of the equipment's operating status; secondly, this simplified assessment method cannot provide confidence or uncertainty in the health status of the equipment, which means that users cannot know the accuracy of the health index or the possible error range, which is not conducive to making accurate maintenance decisions. In contrast, this embodiment can provide a confidence interval for the health index of each computing device by introducing advanced technologies such as Bayesian neural networks, which means that users can understand the credibility and uncertainty of health evaluations and make more cautious and accurate maintenance decisions. On the other hand, an exponential decay prediction model is also developed to predict the RUL probability distribution of key components, which not only provides maintenance predictability, but also allows the maintenance team to optimize maintenance plans and resource allocation based on the actual situation and future expectations of the equipment, avoiding unnecessary downtime and cost waste. Overall, this embodiment significantly improves the operating efficiency and reliability of the server GPU board through refined health assessments and forward-looking maintenance planning, while reducing maintenance costs and improving the overall operation and maintenance management level.

[0102] Alternatively, if Figure 4 As shown, the first management controller inputs the constructed cross-modal features into a pre-trained neural network model (such as a Bayesian neural network), which outputs the health index of the computing device by learning the relationship between different features and the health status; the first management controller inputs the health index of the computing device into an exponential decay prediction model, which obtains the trend of the natural decay of the health index over time based on the historical health index of the computing device and the current health index of the computing device, and predicts the future health status through time series analysis based on the trend, and uses an extrapolation algorithm to calculate the remaining useful life (RUL) of the computing device according to the predicted health index decay trend.

[0103] Through this embodiment, the neural network analyzes cross-modal features and can comprehensively consider various data of computing devices, including but not limited to hardware sensor readings, operation logs, topology information, etc., so as to generate a more accurate and comprehensive health index that reflects the actual operating status of the device; using the exponential decay prediction model, the changing trend of the device health index is calculated based on time series analysis, which can predict its future health status and remaining service life, providing the maintenance team with a basis for forward-looking maintenance planning, and effectively preventing service interruptions caused by sudden failures.

[0104] In an exemplary embodiment, the server further includes: a server mainboard, a switch board, a mid-back board, a power board, and a fan board.

[0105] The server motherboard is a platform that houses the core computing resources (such as CPUs) and management functions of the server. It is the foundation of the hardware system and the core of the entire server. It provides multiple hardware interfaces to connect various components, integrates chip management resources to achieve data transmission and processing, and has the ability to be expanded and upgraded. At the same time, it emphasizes stability and reliability. The server motherboard integrates a second management controller; the second management controller is used to interact with the first management controller for information; the first management controller is also used to report the first status data and the second status data of multiple computing devices to the second management controller. Among them, the second management controller (such as BMC_2) is located on the server motherboard and is the central management unit of the server, responsible for monitoring the overall health status of the server and coordinating the management activities of each component. The second management controller (BMC_2) communicates with the first management controller (BMC_1) on the management board through the Redfish API, collects key information of the GPU board, including the status data and abnormal monitoring data of the computing device, to achieve real-time monitoring and efficient management.

[0106] The server motherboard is connected to the switch board through connectors and cables to achieve signal transmission. The switch board is connected to the server motherboard upstream and to the graphics processor board downstream. The switch board integrates multiple switching components, and the switching components among the multiple switching components are used to expand the peripheral resources of the server motherboard to connect multiple peripheral devices. The midplane, as an intermediate board connecting the switch board, power supply board, fan board, and GPU board, realizes the interconnection of signals in each of the above boards. The midplane is used to connect the switch board and the graphics processor board through high-density connectors. The power supply board is used to provide the working voltage for the server; the power supply board is connected to the midplane. The fan board is used to dissipate heat from the server; the fan board is connected to the midplane through a connector. The fan board integrates a fan and a CPLD for controlling the fan, which is used to realize the power supply function of the fan and the direct control of the fan speed, thus providing guarantee for the heat dissipation of the entire server.

[0107] Through this embodiment, the server motherboard provides multiple interfaces to support the access of various external devices, so that the functions of the server can be flexibly expanded according to different business requirements; the switch board realizes the expansion and sharing of peripheral resources through its built-in switching components, enabling the server to connect more peripheral devices; the midplane, as a key bridge connecting the motherboard, switch board, power supply board, fan board, and GPU board, ensures the efficient transmission and routing of signals between each board, simplifies the physical wiring, and reduces the risk of signal interference; the power supply board is designed with redundant power supplies and an intelligent power management system, which can effectively avoid single-point failures and dynamically adjust the power supply according to the server load to achieve the energy-saving goal; the fan board is responsible for the heat dissipation of the server. By precisely controlling the fan speed, it ensures that the server can still maintain an appropriate temperature under long-term high load, avoiding performance degradation or hardware damage caused by overheating.

[0108] In an exemplary embodiment, multiple modular hardware units are symmetrically arranged and mounted on a graphics processor board through a second connector, so that the impedance of the signal path is matched when the signals of the multiple modular hardware units are interconnected. Herein, the second connector refers to a physical interface specifically used to connect the modular hardware unit and the GPU board. For example, the second connector can be a socket. Figure 5 is a layout diagram of an optional graphics manager board provided by an embodiment of the present application. As Figure 5 shown, assuming that there are eight modular hardware units (OAM modules) on the GPU board, the eight OAM modules are connected to the GPU board through the second connector and fixed to the GPU board through bolts via the OAM module fixing holes. The eight OAM modules are symmetrically arranged at the middle position of the GPU board to ensure the impedance matching and layout symmetry of the signal path when the eight OAM signals are interconnected, so as to reduce signal reflection and interference and ensure the stable transmission of signals.

[0109] In an exemplary embodiment, in the length direction, multiple expansion cards are symmetrically arranged and mounted on the graphics processor board through a third connector. Herein, in the length direction, the installation positions of the multiple expansion cards are below the installation position of the first processor. Herein, on the graphics processor board, the smallest area including the installation positions of the multiple expansion cards and the first processor is the first area, and the smallest area including the installation positions of the multiple modular hardware units is the second area. The first area and the second area do not overlap in the length direction. Herein, the second connector refers to a physical interface connecting the expansion card to the GPU board. For example, the second connector can be a high-speed serial connector, a dedicated interface connector, etc. In this embodiment, the expansion card is used to realize signal interconnection between different servers. Since the connection cables between different servers are routed at the front end of the server, for the convenience of cable management and to ensure short circuit lengths, the first processor is installed in the first area, and the multiple expansion cards are installed in the first area through the third connector and are below the first processor. Thus, the multiple OAM modules are installed below the multiple expansion cards. Figure 6 is a schematic diagram of the PCIe link interconnection relationship between OAM modules and the PCIe interconnection between OAM and PHY Retimer provided by an embodiment of the present application. As Figure 6 shown, there is a PCIe link between the same ports of the eight OAM modules (denoted as OAM0 - OAM7) to realize the communication of multiple OAMs. Every two of the eight OAM modules are connected to the PHY Retimer of the same expansion card. Therefore, the two OAM modules connected to the same expansion card and the connected expansion card are mounted on the same straight line along the length direction of the GPU board. As Figure 5As shown in the figure, it is assumed that there are four expansion cards, and the PHY Retimers in the four expansion cards are respectively denoted as PHY Retimer0 - PHY Retimer3. The four expansion cards are connected to the GPU board through a third connector and fixed to the GPU board through bolts via the expansion card fixing holes. Eight OAM modules are installed below the four expansion cards and are symmetrically arranged at the middle position of the GPU board.

[0110] In an exemplary embodiment, multiple expansion bus retimers are hardware devices for improving the signal quality of PCIe (Peripheral Component Interconnect Express, expansion bus). Its main functions include signal regeneration, protocol transparency, and link equalization to ensure the stability and reliability of the PCIe high - speed link signal transmission from the upstream main board and switch board to the OAM module. Therefore, in order to reduce the trace length, in this embodiment, multiple expansion bus retimers are symmetrically arranged and installed on the graphics processor board in the length direction of the GPU board. Among them, in the length direction, the smallest area containing the installation positions of multiple expansion bus retimers is the third area, and the first area, the second area, and the third area are adjacent in sequence and do not overlap. Figure 7 It is a schematic diagram of the PCIe link connection relationship between an optional OAM module, a PCIe Retimer, and a fourth connector provided by an embodiment of the present application. As Figure 7 shown, the eight OAM modules are denoted as OAM0 - OAM7. The OAM module not only needs to be connected to the expansion card but also needs to be connected to the expansion bus retimer (PCIe Retimer). There is also a PCIe link between the PCIe Retimer and the fourth connector (such as the board - end ExaMAX connector). Therefore, the OAM module needs to be set between the expansion card and the PCIe Retimer. To facilitate cable management and ensure short lines, the PCIe Retimer needs to be set adjacent to the OAM module and in the position below the OAM module. As Figure 5 shown, it is assumed that there are eight expansion bus retimers (such as PCIe Retimer chips), and the eight PCIe Retimer chips are installed below the OAM module.

[0111] In an exemplary embodiment, in the length direction, a plurality of fourth connectors arranged symmetrically are further installed on the graphics processing unit (GPU) board. The number of the plurality of fourth connectors is the same as the number of the plurality of extended bus retimers. The fourth connectors are used to connect to the midplane. Wherein, on the GPU board, the smallest area including the installation positions of the fourth connectors is the fourth area. In the length direction, the first area, the second area, the third area, and the fourth area are adjacent to each other in sequence and do not overlap. The fourth connector refers to a board-to-board connector for connecting the GPU board to the midplane or other forms of backplane. For example, the fourth connector can be a board-end ExaMAX connector. As Figure 7 shown, each of the eight PCIe Retimer chips corresponds to an OAM module and a fourth connector respectively, that is, there is a PCIe link between one PCIe Retimer chip and an OAM module and a fourth connector respectively. Therefore, the number of the plurality of fourth connectors is the same as the number of the plurality of extended bus retimers, and the installation position of each fourth connector is on the same straight line as the installation position of each PCIe Retimer chip. Since the fourth connector is used to connect the signals of the GPU board to the signals of the upstream main board and the switching board, in this embodiment, as Figure 5 shown, the fourth connector is installed at the rear of the GPU board, so that the GPU board needs to have the function of being separately pulled out from the chassis.

[0112] In an exemplary embodiment, the first connector is used to connect the GPU board and the management board. Considering the installation space limitation on the GPU board (no external plug-in cards other than the EXP card can be installed at the front end. Since the EXP card requires external cables, if there are other external plug-in cards, it will interfere with the cable routing), in this embodiment, in the length direction, the installation position of the first connector does not overlap with the second area and overlaps with the third area, and in the width direction, the installation position of the first connector does not overlap with the fourth area. As Figure 5 shown, the first connector is installed at the right rear position of the GPU board. In this way, the volume of the device can be reduced, and the interference of the first connector on the cable routing of the EXP card can be avoided.

[0113] In an exemplary embodiment, as Figure 5 shown, the GPU board further includes a GPU board grab handle, which is convenient for grabbing the GPU board.

[0114] In this embodiment, the OAM modules are arranged symmetrically and connected by the second connector, ensuring impedance matching during signal interconnection, reducing signal reflection and interference, and thus ensuring stable signal transmission. The first processor is placed at the front lower part, and multiple expansion cards follow closely below it through the third connector, while the fourth connector is located at the rear of the GPU board. This not only enables the GPU board to have an independent pulling function, facilitating maintenance and upgrade, but also maximizes the use of space with this design of lower in the front and higher in the rear and arranged step by step, reducing the volume of the device. At the same time, it avoids cable interference between components and is convenient for cable management operations. The expansion bus retimer is placed between the OAM module and the fourth connector, being both close to the source end (the fourth connector) of the PCIe data and adjacent to the destination end (the OAM module). This layout shortens the signal transmission path, reduces signal delay, and improves the quality and reliability of the PCIe signal. The first connector is located on the right side of the rear end of the GPU board, overlapping with both the second area (the OAM module area) and the expansion bus retimer part, and separated from the third area (the fourth connector area). This design not only ensures the effective connection of the GPU management board but also avoids conflicts with the cables of the EXP card, maintaining good signal integrity and the convenience of device operation.

[0115] In an exemplary embodiment, the minimum area on the management board for installing the second processor is the fifth area, and the minimum area for installing the first management controller is the sixth area. The fifth area and the sixth area are adjacent in sequence and do not overlap in the length direction of the management board. As can be seen from the above embodiment, the main function of the second processor is to receive and process the log data from the first processor, and then transmit the processed information to the first management controller, as well as to implement the function of switching the UART serial port of the management board. To shorten the wiring between the second processor and the first management controller, in this embodiment, Figure 8 is a layout diagram of an optional management board provided by this embodiment. As Figure 8 shown, in the length direction of the management board, the second processor and the first management controller are installed adjacent to each other in sequence. In this embodiment, the management board is used to implement the function of obtaining the running information and status monitoring of key devices on the GPU board. At the same time, it can complete the information interaction function with the second management controller on the server motherboard through the Redfish API.

[0116] In an exemplary embodiment, the management board further includes a main firmware and a standby firmware. The main firmware and the standby firmware are electrically connected to the first management controller respectively and are used to store the firmware files of the first management controller. The main firmware and the standby firmware are arranged and installed on the management board in sequence along the width direction of the management board. In the length direction, the minimum area including the main firmware and the standby firmware is the seventh area. The fifth area, the sixth area, and the seventh area are adjacent in sequence and do not overlap in the length direction of the management board. In this embodiment, since the main firmware and the standby firmware serve the first management controller, therefore, asFigure 8 As shown, the main firmware and the standby firmware are set close to the first management controller, thereby reducing the transmission delay and shortening the wiring length.

[0117] In an exemplary embodiment, the management board further includes a gold finger component electrically connected to the first connector; on the management board, the minimum area including the installation positions of the second processor, the first management controller, the main firmware, and the standby firmware is the eighth area, and the minimum area including the installation position of the gold finger component is the ninth area. The eighth area and the ninth area are adjacent in sequence and do not overlap in the width direction of the management board. In this embodiment, the management board and the graphics processor board are connected through a connector. Among them, the first connector is integrated on the graphics processor board. Therefore, a gold finger component for connecting to the first connector is also provided on the corresponding management board. As Figure 8 shown, to reduce the volume of the management board, the length of the management board is within a preset length range. The minimum value of this preset length range is the length of the gold finger component, and the maximum value is the sum of the length of the gold finger component and a preset error value, which can be set according to requirements.

[0118] Through this embodiment, the second processor is closely adjacent to the first management controller, reducing the signal wiring length, effectively reducing the data transmission delay, and improving the real-time performance and communication efficiency of information processing; the main firmware and the standby firmware are installed close to the first management controller, which not only simplifies the reading path of the firmware file but also facilitates a quick switch to the standby firmware when the first management controller fails, ensuring the continuous operation of the system and the smooth transition of firmware updates, and improving the redundancy and fault recovery ability of the system; the gold finger component of the management board is closely aligned with the first connector on the GPU board, which not only ensures the stability of the electrical connection but also minimizes the volume of the device to the greatest extent by controlling the length of the management board within the preset range, improving the deployment density and space utilization rate of the server, enabling the management board to achieve efficient connection with the GPU board in a limited space.

[0119] An embodiment of the present application provides a server monitoring method. Figure 9 It is a schematic flowchart of an optional server monitoring method provided by an embodiment of the present application. As Figure 9 shown, the server monitoring method includes:

[0120] Step S902, communicate with multiple computing devices on the graphics processor board through the first processor on the graphics processor board to obtain the first state data of the multiple computing devices, and feedback the first state data of the multiple computing devices to the management board through the first connector on the graphics processor board; the multiple computing devices are used to execute specified tasks; the first state data of the multiple computing devices includes the timing control data and communication link switching data of the multiple computing devices.

[0121] Step S904: Receive, via a first management controller on a management board, first status data of multiple computing devices sent by a first processor, communicate with the multiple computing devices to obtain second status data of the multiple computing devices, and monitor the multiple computing devices based on the first status data and the second status data of the multiple computing devices; the second status data of the multiple computing devices includes the operating status of the multiple computing devices.

[0122] For the descriptions of the features in the corresponding embodiments of the server monitoring method, reference may be made to the relevant descriptions of the corresponding embodiments of the above server, which will not be elaborated herein one by one.

[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation.

[0124] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0125] The above has introduced in detail a server provided by this application. Specific examples are used herein to elaborate on the principle and implementation of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A server, characterized in that, Comprising: A graphics processing unit board integrated with a first connector for performing specified tasks; A management board located on the graphics processing unit board and connected to the graphics processing unit board through the first connector for monitoring the operating state of the graphics processing unit board.

2. The server according to claim 1, wherein: The graphics processing unit board is further integrated with a first processor and a plurality of computing devices; the first processor is electrically connected to the plurality of computing devices; The management board is integrated with a first management controller; The first management controller is electrically connected to the first processor and the plurality of computing devices.

3. The server according to claim 2, wherein The plurality of computing devices includes: A plurality of modular hardware units that communicate with the first processor for performing specific computing tasks and management tasks; A plurality of expansion cards that communicate with the plurality of modular hardware units to enable the interconnection of the modular hardware units of the plurality of modular hardware units with the modular hardware units of other node servers.

4. The server according to claim 3, wherein The number of the plurality of modular hardware units is twice the number of the plurality of expansion cards; every two of the plurality of modular hardware units are connected to one of the plurality of expansion cards.

5. The server according to claim 4, characterized in that The expansion card among the plurality of expansion cards includes a physical layer retimer; Every two of the plurality of modular hardware units are connected to the physical layer retimer in the expansion card among the plurality of expansion cards.

6. The server according to claim 5, wherein The expansion card among the plurality of expansion cards further includes a plurality of pluggable optical modules; The physical layer retimer of the expansion card among the plurality of expansion cards is electrically connected to the plurality of pluggable optical modules in the expansion card among the plurality of expansion cards.

7. The server according to claim 3, characterized in that The plurality of computing devices further includes: A plurality of expansion bus retimers that communicate with the plurality of modular hardware units for restoring the clocks of the plurality of modular hardware units; the number of the plurality of expansion bus retimers is equal to the number of the plurality of modular hardware units.

8. The server according to claim 7, wherein The management board is further integrated with a second processor; the second processor is electrically connected to the first management controller and the first processor respectively.

9. The server according to claim 8, wherein The server further includes: A server main board for data transmission and data processing; A switching board electrically connected to the server main board upstream and electrically connected to the graphics processing unit board downstream for expanding the peripheral resources of the server main board to connect a plurality of peripheral devices.

10. The server according to claim 9, characterized in that, The server further includes: A midplane for connecting the switching board and the graphics processing unit board for realizing signal interconnection among the server main board, the switching board and the graphics processing unit board.

11. The server according to claim 10, wherein The server further includes: A power supply board electrically connected to the midplane for providing operating voltage for the server.

12. The server according to claim 11, wherein The server further includes: A fan board electrically connected to the midplane for dissipating heat from the server.

13. The server according to claim 10, wherein In the length direction, the plurality of modular hardware units are symmetrically arranged and installed on the graphics processing unit board through a second connector so that the impedance of the signal path matches when the signals of the plurality of modular hardware units are interconnected.

14. The server according to claim 13, wherein In the length direction, the multiple expansion cards are symmetrically arranged and mounted on the graphics processor board through a third connector. Among them, in the length direction, the mounting positions of the multiple expansion cards are located below the mounting position of the first processor. Among them, on the graphics processor board, the smallest area including the mounting positions of the multiple expansion cards and the first processor is the first area, and the smallest area including the mounting positions of the multiple modular hardware units is the second area. The first area and the second area do not overlap in the length direction.

15. The server according to claim 14, wherein In the length direction, the multiple expansion bus retimers are symmetrically arranged and mounted on the graphics processor board. Among them, in the length direction, the smallest area including the mounting positions of the multiple expansion bus retimers is the third area. The first area, the second area, and the third area are adjacent in sequence and do not overlap.

16. The server according to claim 15, characterized in that In the length direction, the graphics processor board is further mounted with multiple fourth connectors arranged symmetrically. The number of the multiple fourth connectors is the same as the number of the multiple expansion bus retimers. The fourth connectors are used to connect the mid-backplane. Among them, on the graphics processor board, the smallest area including the mounting positions of the fourth connectors is the fourth area. In the length direction, the first area, the second area, the third area, and the fourth area are adjacent in sequence and do not overlap.

17. The server according to claim 16, wherein In the length direction, the mounting position of the first connector does not overlap with the second area and overlaps with the third area. And, in the width direction, the mounting position of the first connector does not overlap with the fourth area.

18. The server according to claim 8, characterized in that, The smallest area for mounting the second processor on the management board is the fifth area, and the smallest area for mounting the first management controller is the sixth area. The fifth area and the sixth area are adjacent in sequence and do not overlap in the length direction of the management board.

19. The server according to claim 18, wherein The management board further includes a main firmware and a standby firmware. The main firmware and the standby firmware are electrically connected to the first management controller respectively and are used to store the firmware files of the first management controller. The main firmware and the standby firmware are arranged and mounted on the management board in sequence along the width direction of the management board. In the length direction, the smallest area including the main firmware and the standby firmware is the seventh area. The fifth area, the sixth area, and the seventh area are adjacent in sequence and do not overlap in the length direction of the management board.

20. The server according to claim 19, wherein The management board further includes a gold finger component, which is electrically connected to the first connector. On the management board, the smallest area including the mounting positions of the second processor, the first management controller, the main firmware, and the standby firmware is the eighth area, and the smallest area including the mounting position of the gold finger component is the ninth area. The eighth area and the ninth area are adjacent in sequence and do not overlap in the width direction of the management board.

21. A server monitoring method, characterized in that, Including: Communicate with a plurality of computing devices on the graphics processing unit board through a first processor on the graphics processing unit board to obtain first status data of the plurality of computing devices, and feedback the first status data of the plurality of computing devices to a management board through a first connector on the graphics processing unit board; the plurality of computing devices are used to execute specified tasks; the first status data of the plurality of computing devices includes timing control data and communication link switching data of the plurality of computing devices; Receive the first status data of the plurality of computing devices sent by the first processor through a first management controller on the management board, communicate with the plurality of computing devices to obtain second status data of the plurality of computing devices, and monitor the plurality of computing devices based on the first status data and the second status data of the plurality of computing devices; The second status data of the plurality of computing devices includes the operating status of the plurality of computing devices.

Citation Information

Patent Citations

  • Graphic processor board card

    CN109408445A

  • Monitoring management system and monitoring management method for Phytium server

    CN111881002A

  • Graphics processor board card and graphics processor management method

    CN112000545A

  • Monitoring management system of Feiteng server

    CN212411186U

Cited By

  • Information processing method and electronic equipment

    CN120821634A

  • Information processing method and electronic device

    CN120821634B