Management board, universal substrate and monitoring method
Through the collaborative design of the processing unit and logic unit of the management board and the hardware acceleration engine, the problems of poor hardware reusability and insufficient protocol compatibility of the management board on the UBB board are solved, and rapid iteration and efficient real-time monitoring are achieved, which reduces R&D costs and delays, and improves the system's real-time monitoring capabilities and resource utilization efficiency.
Patent Information
- Application Number
- CN202510908383.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-02
AI Technical Summary
The management boards used on UBB boards have problems such as poor hardware reusability, insufficient protocol compatibility, weak real-time monitoring capabilities, and inefficient resource utilization, resulting in long R&D cycle, high cost, lagging technology iteration, weak real-time monitoring capabilities and inefficient resource utilization.
It provides a management board, which adopts a hybrid architecture designed in collaboration with processing units and logic units, and uses the target firmware to configure the logic unit to adapt to different types of general substrates, combines the hardware acceleration engine to realize high-speed data real-time parsing, and realizes hardware multiplexing and protocol adaptation through a hierarchical dynamic software system.
It supports more than 80% of general-purpose substrate types without redesigning hardware, shortening the R&D cycle by 70%, reducing the hardware development cost by 50%, real-time analysis of 200Gbps data, meeting the real-time monitoring needs of high-speed buses, and improving system reliability and resource utilization efficiency.
Smart Images

Figure CN120407490A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of servers, and particularly to a management board, a general baseboard, and a monitoring method. Background Art
[0002] In the fields of artificial intelligence and high-performance computing, as the core carrier of the open accelerator infrastructure, the UBB board (Universal Baseboard) needs to achieve coordinated management of multiple OAM modules (Open Compute Project Accelerator Module) through the head motherboard. However, in related technologies, the head motherboard and the UBB board are designed independently, resulting in problems such as poor hardware reusability, insufficient protocol compatibility, weak real-time monitoring ability, and low resource utilization on the management board used on the UBB board. Summary of the Invention
[0003] In view of the above problems, this application provides a management board, a general baseboard, and a monitoring method.
[0004] According to the first aspect of this application, a management board is provided, including: a processing unit, configured to select a target firmware from multiple firmwares based on the type of the general baseboard, and configure a logic unit through the target firmware, so that the management board adapts to the general baseboard of the target type; a bus interface unit, configured to implement an electrical connection between the management board and an acceleration computing module disposed on the general baseboard; and the logic unit, electrically connected to the bus interface unit, configured to process data from the acceleration computing module to implement monitoring of the acceleration computing module.
[0005] According to an embodiment of this application, the management board further includes: a first memory, configured to store logic unit firmware; a second memory, configured to store data from the acceleration computing module and monitoring results of the acceleration computing module; a debugging interface, connecting the processing unit and the logic unit; and a reconfigurable hardware, including a preset number of logic units.
[0006] According to an embodiment of this application, the management board further includes: a power management unit, connected to the logic unit, configured to supply power to the acceleration computing module; and a clock generator, connected to the logic unit, configured to provide a clock signal to the acceleration computing module.
[0007] According to an embodiment of this application, the bus interface unit includes: a data bus, configured to electrically connect a serializer / deserializer on the management board and a serializer / deserializer on the acceleration computing module to implement data interaction between the management board and the acceleration computing module; and the serializer / deserializer, configured to convert serial data and parallel data.
[0008] According to an embodiment of the present application, the above management board further includes: a monitoring unit, configured to collect eye diagram parameters from the above acceleration computing module, and trigger a fault response when the above eye diagram parameters do not meet the preset conditions; parse the data transmission protocol between the above management board and the above acceleration computing module to monitor protocol layer faults.
[0009] According to an embodiment of the present application, the above management board further includes: a diagnostic engine, configured to locate the fault location of the above acceleration computing module and perform recovery processing on the faulty acceleration computing module.
[0010] According to an embodiment of the present application, the above management board further includes a hierarchical dynamic software system, and the above hierarchical dynamic software system includes a hardware layer, a management layer, and an application interface layer.
[0011] According to an embodiment of the present application, the above hardware layer is used to provide a unified interface for the above management board to adapt the hardware of the above management board to multiple types of general substrates; implement device enumeration and hot plug detection for the above acceleration computing module to obtain detection results; generate an acceleration computing module list based on the above detection results, where the above acceleration computing module list includes at least one of the connection status, hot plug status, load, and data transmission rate of the above acceleration computing module.
[0012] According to an embodiment of the present application, the above management layer further includes: a power and clock management module, configured to adjust the current output of the above power management unit according to the load of the above acceleration computing module in the above acceleration computing module list; adjust the frequency of the clock signal output by the above clock generator according to the data transmission rate in the above acceleration computing module list.
[0013] According to an embodiment of the present application, the above management layer further includes: a protocol parsing module, configured to control the above monitoring unit to parse the data transmission protocol according to a preset priority; control the above logic unit to process the data from the above acceleration computing module to generate a monitoring result for the above acceleration computing module.
[0014] According to an embodiment of the present application, the above management board further includes a fault detection module, configured to: control the above diagnostic engine to perform fault detection on the above acceleration computing module based on a preset fault knowledge base; perform a fault response when a fault is detected.
[0015] According to an embodiment of the present application, the above application interface layer includes: an on-board interaction interface, configured to display the status of the above acceleration computing module in real time; a remote management interface, configured to obtain the status of the above management board and issue configuration instructions for the above management board.
[0016] The second aspect of the present application provides a general substrate, and the above general substrate is integrated with the above management board.
[0017] According to an embodiment of the present application, the above general substrate includes an acceleration computing module disposed on one side of the general substrate close to the bus interface unit of the above management board.
[0018] The third aspect of the present application provides a monitoring method, which applies the above management board. The method is characterized in that the method includes: in response to detecting that the general substrate is powered on, using the processing unit to select a target firmware from multiple firmware based on the type of the general substrate, and configuring the logic unit through the above target firmware, so that the above management board adapts to the above general substrate of the target type; using the bus interface unit to realize the electrical connection between the above management board and the acceleration computing module disposed on the above general substrate; using the above logic unit to process the data from the above acceleration computing module to realize the monitoring of the above acceleration computing module.
[0019] According to the management board, general substrate and monitoring method of the present application, the processing unit is adapted to the logic unit, that is, the processing unit and the logic unit work together. The processing unit can select the target firmware according to the type of the general substrate integrated into the management board, and configure the logic unit with the target firmware, so that the management board can adapt to the general substrate integrated into it. Thus, based on different types of general substrates integrated into it, the processing unit can configure the logic unit with different firmware to adapt to different types of general substrates. On this basis, the management board of the present application can support more than 80% of the general substrate types, without re-designing the hardware, the R & D cycle is shortened by 70%, the hardware development cost is reduced by 50%, and it adapts to the rapid iteration of the acceleration computing module. And, through the hardware acceleration engine of the logic unit, real-time parsing of 200Gbps-level data is realized, meeting the real-time monitoring requirements of the high-speed bus. Description of the Drawings
[0020] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other objects, features and advantages of the present application will be clearer. In the drawings:
[0021] Figure 1 A schematic diagram of a management board according to an embodiment of the present application is shown;
[0022] Figure 2 A schematic diagram of the storage and configuration unit on the management board according to an embodiment of the present application is shown;
[0023] Figure 3 A schematic diagram of the power supply and clock subsystem on the management board according to an embodiment of the present application is shown;
[0024] Figure 4Shows a schematic diagram of the connection between a power supply and clock subsystem and a logic unit according to an embodiment of the present application;
[0025] Figure 5 Shows a schematic diagram of a hierarchical dynamic software system on a management board according to an embodiment of the present application;
[0026] Figure 6 Shows a schematic diagram of an application interface layer on a management board according to an embodiment of the present application;
[0027] Figure 7 Shows a schematic diagram of a management board according to another embodiment of the present application;
[0028] Figure 8 Shows a flowchart of a monitoring method according to an embodiment of the present application;
[0029] Figure 9 Shows a flowchart of an application on a management board according to an embodiment of the present application. Detailed implementation manners
[0030] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present application.
[0031] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0032] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0033] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0034] In the process of implementing this application, it is found that for the management board used on the UBB board, the solutions in the related technologies include the following: stand-alone head management board, decentralized control architecture, and external management unit.
[0035] For the stand-alone head management board, a dedicated hardware design is adopted, which is connected to the UBB board through a low-speed bus, but only supports power monitoring of specific OAM modules and cannot be compatible with the CXL protocol (Compute Express Link Protocol) module.
[0036] For the decentralized control architecture, power management, clock generation, and protocol parsing are dispersed to multiple chips and interconnected through PCB (Printed Circuit Board) traces. For example, the power management chip, clock generator, and ASIC (Application-Specific Integrated Circuit) parsing chip are independently deployed, but this will cause the data interaction delay between chips to exceed 10 μs, and fault diagnosis requires cross-module coordination.
[0037] For the external management unit, a management controller is externally connected through an RJ45 (Registered Jack 45) or USB (Universal Serial Bus) interface, but this relies on network transmission to achieve remote monitoring. For example, in the traditional server management solution, the real-time performance is poor (the error reporting delay exceeds 50 μs), and the underlying signals of the high-speed bus cannot be directly accessed.
[0038] Therefore, in the related technologies, the management board used on the UBB board has the following problems: hardware customization requires independent design of management boards for different models of UBB boards, which results in a long R&D cycle (single design cycle exceeds 6 months) and high cost (hardware repeated development cost increases by 40%), and is unable to adapt to the rapid iteration of market demand, resulting in high reuse costs; the software parsing engine is rigid and only supports fixed bus protocols. When new protocols are added, the hardware logic needs to be redeveloped and designed or the firmware needs to be upgraded. For example, supporting CXL3.0 requires an additional investment of 2000+ man-hours, and the cycle is as long as 9 months, which lags behind in technology iteration, resulting in single protocol support and insufficient protocol compatibility; high-speed bus data parsing relies on software to be gradually updated. Packet processing results in long delays, making it impossible to meet the microsecond-level synchronization requirements between OAM modules and to capture complete signals in real time. This reduces system throughput by more than 15%, reduces fault location efficiency, and leads to weak real-time monitoring capabilities. The independent head is separated from the UBB board, and hardware resources such as power filtering and clock buffering cannot be shared, resulting in an increase of more than 25% in system power consumption. In addition, the redundant board layout occupies more than 30% of the board area in high-density computing scenarios, resulting in inefficient resource utilization. Independent components increase board complexity, the power conversion efficiency is only 85%, and the clock synchronization error exceeds 100ps, affecting the stability of high-speed data transmission and resulting in insufficient integration.
[0039] To this end, an embodiment of the present application provides a management board to solve the problems of poor hardware reusability, insufficient protocol compatibility, weak real-time monitoring capabilities, and inefficient resource utilization existing in the management board in related technologies.
[0040] Figure 1 A schematic diagram of a management board according to an embodiment of the present application is shown.
[0041] like Figure 1 As shown, the management board 110 is arranged on the universal substrate 100 , the universal substrate 100 is provided with an OAM (accelerated computing module), and the management board 110 is provided with a processing unit 111 , a logic unit 112 and a bus interface unit 113 .
[0042] In one embodiment, the universal substrate 100 may have eight accelerated computing modules, such as Figure 1 The accelerated computing module 1, accelerated computing module 2, ..., accelerated computing module 8 shown in FIG.
[0043] In one embodiment, the processing unit 111 may be an AST2600 (AST2600 Baseboard Management Controller) processing unit.
[0044] Specifically, the processing unit 111, which is connected to the logic unit 112, can be used to select a target firmware from multiple firmwares based on the type of the general substrate 100, and configure the logic unit 112 with the target firmware, so that the management board 110 adapts to the general substrate 100 of the target type.
[0045] According to an embodiment of the present application, when integrating the management board into the general substrate, the target firmware corresponding to the general substrate can be selected based on the type of the integrated general substrate. Configure the logic unit with the target firmware so that the management board adapts to the integrated general substrate.
[0046] Among them, the target firmware can be used to represent the firmware related to the logic unit and configure the logic unit.
[0047] Since different types of general substrates have different requirements, and the management board needs to adapt to the general substrate and the requirements of the general substrate, therefore, the target firmware can be selected according to the type of the general substrate integrated into the management board.
[0048] In one embodiment, the processing unit 111 can be responsible for system resource scheduling, remote communication, and user interaction logic, support dynamic loading of the accelerated computing module driver, and achieve software flexibility.
[0049] According to an embodiment of the present application, the bus interface unit 113 can be used to implement the electrical connection between the management board 110 and the accelerated computing module arranged on the general substrate 100 to ensure the stability of data transmission between the management board and the accelerated computing module.
[0050] In one embodiment, the logic unit 112 can be an FPGA (Field-Programmable Gate Array).
[0051] Specifically, the logic unit 112 can be electrically connected to the bus interface unit 113 and is used to process the data from the accelerated computing module to realize the monitoring of the accelerated computing module.
[0052] In one embodiment, the logic unit 112 can implement a high-speed protocol preprocessing engine with a single-channel data processing rate greater than or equal to 200 Gbps, and support parallel processing of data from 8 accelerated computing modules to achieve hardware acceleration parsing.
[0053] Based on the above, the management board 110 adopts the collaborative design of the processing unit 111 and the logic unit 112 to form a "software control + hardware acceleration" hybrid architecture, which is also a heterogeneous processing architecture, realizing the combination of software flexibility and hardware acceleration, and supporting protocol dynamic loading and high-speed data parallel processing.
[0054] According to an embodiment of the present application, the processing unit is adapted to the logic unit, that is, the processing unit and the logic unit work together. The processing unit can select a target firmware according to the type of the general substrate integrated into the management board, and configure the logic unit with the target firmware, so that the management board can be adapted to the general substrate integrated into it. Thus, based on different types of general substrates integrated into it, the processing unit can configure the logic unit with different firmwares to adapt to different types of general substrates. On this basis, the management board of the present application can support more than 80% of the general substrate types, without re-designing the hardware, shortening the R & D cycle by 70%, reducing the hardware development cost by 50%, and adapting to the rapid iteration of the acceleration computing module. Moreover, through the hardware acceleration engine of the logic unit, real-time parsing of 200Gbps-level data is achieved, meeting the real-time monitoring requirements of the high-speed bus.
[0055] Figure 2 FIG. shows a schematic diagram of a storage and configuration unit on a management board according to an embodiment of the present application.
[0056] As Figure 2 shown, a storage and configuration unit 200 is also arranged on the management board. The storage and configuration unit 200 may include a first memory 210, a second memory 220, a debugging interface 230, and a reconfigurable hardware 240.
[0057] In one embodiment, the first memory 210 may be a Flash (Flash Memory).
[0058] Specifically, the first memory 210 can be used to store the logic unit firmware.
[0059] Among them, the logic unit firmware can represent the firmware related to the logic unit.
[0060] According to an embodiment of the present application, since the logic unit firmware is stored in the first memory 210, the processing unit 111 can select a target firmware corresponding to the type of the general substrate 100 from the logic unit firmware stored in the first memory 210, so that the management board 110 is adapted to the general substrate 100.
[0061] According to an embodiment of the present application, the first memory 210 can also be used to store the management board firmware, the protocol parsing rule library, the acceleration computing module firmware library, etc., support remote online upgrade, so that the firmware update time is short, facilitating technology iteration and maintenance.
[0062] In one embodiment, the second memory 220 may be 4GB LPDDR4 (Low-Power Double Data Rate 4, 4GB low-power double data rate 4th generation memory).
[0063] Specifically, the second memory 220 can be used to store data from the acceleration computing module and the monitoring results of the acceleration computing module.
[0064] According to an embodiment of the present application, the second memory 220 can be used as a high-speed data buffer to store the data collected in real time from the acceleration computing module and the monitoring results of the acceleration computing module. Thus, the second memory 220 can support dual-channel reading and writing to meet the requirements of high-speed data caching.
[0065] In one embodiment, the debugging interface 230 can be a JTAG (Joint Test Action Group) debugging interface.
[0066] Specifically, the debugging interface 230 can connect the processing unit 111 and the logic unit 112, that is, the processing unit 111 and the logic unit 112 can be connected through the debugging interface 230.
[0067] According to an embodiment of the present application, the debugging interface 230 supports the processing unit 111 to perform on-line debugging of hardware logic and configuration loading of the logic unit 112. By cooperating with the on-board DIP switch, the working mode of the management board can be manually switched, which is convenient for debugging and configuration.
[0068] In one embodiment, the reconfigurable hardware 240 refers to the logic unit.
[0069] Specifically, the reconfigurable hardware 240 includes a preset number of logic units.
[0070] Among them, the preset number can be set according to needs.
[0071] For example, the preset number of logic units can be 20% of all the logic units, that is, 20% of the logic units are reserved.
[0072] Specifically, the reconfigurable hardware 240 supports dynamically loading a new protocol parsing IP (Intellectual Property) core through software instructions, that is, configuring the logic unit through the target firmware, so that the hardware reuse rate can be increased to more than 80% without hardware modification.
[0073] Thus, the processing unit 111 can configure the reserved preset number of logic units through the target firmware.
[0074] According to an embodiment of the present application, online upgrade is supported through the first memory, the firmware update time is short; the high-speed data caching requirements are met through the second memory; online debugging and logic unit configuration loading are supported through the debugging interface; through the reconfigurable hardware, the hardware reuse rate can be increased to more than 80% without hardware modification.
[0075] Figure 3 Shows a schematic diagram of the power supply and clock subsystem on the management board according to an embodiment of the present application.
[0076] As Figure 3 shown, a power supply and clock subsystem 300 is also arranged on the management board, and the power supply and clock subsystem 300 includes a power management unit 310 and a clock generator 320.
[0077] According to an embodiment of the present application, the power management unit 310 can be connected to the logic unit 112 for powering the acceleration calculation module.
[0078] Specifically, the power management unit 310 supports a main power input of 54V / 48V and an auxiliary power supply of 12V. The power management unit 310 can supply power to 8 independent acceleration calculation modules through a multi-phase power controller, and can perform dynamic power adjustment on each acceleration calculation module, and can also perform overcurrent protection on each acceleration calculation module and respond in a timely manner.
[0079] In one embodiment, the power supply and clock subsystem 300 may further include an integrated power status sensor for real-time monitoring of the voltage, current, and temperature of the power supply on the management board 110 to achieve intelligent power distribution and protection of the power supply.
[0080] According to an embodiment of the present application, the clock generator 320 can also be connected to the logic unit 112 for providing a clock signal to the acceleration calculation module.
[0081] Specifically, the clock generator 320 can use a high-precision differential crystal oscillator combined with a clock recovery circuit to generate a reference clock, and through a low-jitter clock buffer to reduce the jitter in the reference clock and distribute the clock signal with reduced jitter to each acceleration calculation module, so as to ensure a low clock phase error between each acceleration calculation module. On this basis, the clock generator 320 also supports dynamic clock frequency switching to ensure high-speed data acquisition synchronization for each acceleration calculation module.
[0082] Figure 4 Shows a schematic diagram of the connection between the power supply and clock subsystem and the logic unit according to an embodiment of the present application.
[0083] As Figure 4 shown, both the power management unit 310 and the clock generator 320 in the power supply and clock subsystem 300 arranged on the management board are connected to the logic unit 112.
[0084] Specifically, the power management unit 310 can be connected to the logic unit 112 through a low-impedance power line, and the clock generator 310 can be connected to the logic unit 112 through a low-impedance clock line. Thus, by using low-impedance power lines and clock lines for connection, noise interference can be reduced.
[0085] According to an embodiment of the present application, through the power management unit, dynamic power regulation and over-current protection response are supported in a timely manner; through the clock generator, dynamic clock frequency switching is supported to ensure high-speed data acquisition synchronization.
[0086] According to an embodiment of the present application, the bus interface unit includes: a data bus for electrically connecting the serializer / deserializer on the management board to the serializer / deserializer on the acceleration computing module to achieve data interaction between the management board and the acceleration computing module; and a serializer / deserializer for converting serial data and parallel data.
[0087] According to an embodiment of the present application, the bus interface unit 113 disposed on the management board 110 may include a data bus and a serializer / deserializer.
[0088] Since in high-speed data transmission, using serial transmission can significantly improve transmission efficiency and reduce the number of physical cables, the communication between the management board and the acceleration computing module is carried out through the data bus and in a serial data format.
[0089] Specifically, a first serializer and a first deserializer may be disposed on the management board 110, and a second serializer and a second deserializer may be provided on the acceleration computing module.
[0090] In one embodiment, the data bus may be used to connect the first serializer on the management board 110 to the second deserializer on the acceleration computing module, and may also be used to connect the first deserializer on the management board 110 to the second serializer on the acceleration computing module.
[0091] Specifically, during the process of the management board 110 transmitting data to the acceleration computing module, the first serializer on the management board 110 can be used to convert the parallel data of the management board into serial data and send it to the data bus, so that the data bus sends the serial data to the acceleration computing module. The second deserializer on the acceleration computing module can convert the received serial data into parallel data for the acceleration computing module to process. During the process of the acceleration computing module transmitting data to the management board 110, the data conversion process is similar and will not be elaborated here.
[0092] In one embodiment, the serializer / deserializer may be a Serializer / Deserializer; the data bus may be a high-speed data bus for high-speed data interaction between the management board and the acceleration computing module to achieve fast transmission of fault logs and configuration parameters.
[0093] In another embodiment, the bus interface unit may further include a low-speed control bus, which may be multiple groups of I2C (Inter-Integrated Circuit) buses, supporting a communication rate of 100 kHz - 400 kHz. The low-speed control bus can connect the management board to the sensors on the general substrate and the configuration registers on the acceleration computing module respectively. Thus, the low-speed control bus supports simultaneous access to multiple devices, enabling a relatively high data acquisition frequency and realizing real-time interaction of low-speed data.
[0094] According to an embodiment of the present application, when 8-way acceleration computing modules are arranged on the general substrate 100, a high-speed data path is arranged on the management board 110, and the high-speed data path can integrate 8 groups of SerDes (Serializer / Deserializer) hard cores.
[0095] Specifically, the high-speed data path supports high-speed interfaces and adapts to the signal characteristics of different acceleration computing modules by dynamically configuring PHY layer (Physical Layer) parameters. Among them, the impedance matching of the high-speed data path is controlled within 50Ω ± 5% to ensure the stability of high-speed signal transmission.
[0096] According to an embodiment of the present application, data transmission between the management board and the acceleration computing module through a serializer / deserializer and a data bus can achieve fast data transmission while ensuring the stability of high-speed signal transmission.
[0097] According to an embodiment of the present application, the management board further includes: a monitoring unit, which is used to collect eye diagram parameters from the acceleration computing module and trigger a fault response when the eye diagram parameters do not meet the preset conditions; and analyze the data transmission protocol between the management board and the acceleration computing module to monitor protocol layer faults.
[0098] According to an embodiment of the present application, a monitoring unit is also arranged on the management board 110, and the monitoring unit can be used for signal integrity monitoring and protocol layer error detection.
[0099] Specifically, for signal integrity monitoring, the monitoring unit can collect eye diagram parameters from the acceleration computing module in real time, that is, it can collect SerDes signal eye diagram parameters in real time.
[0100] In one embodiment, the signal transmitted on the SerDes is analyzed in real time to obtain eye diagram parameters to evaluate the quality and transmission performance of the transmitted signal.
[0101] Among them, the preset conditions can indicate that the collected eye diagram parameters do not exceed the preset range, and the preset range is set according to requirements.
[0102] According to an embodiment of the present application, when the collected eye diagram parameters exceed the preset range and the collected eye diagram parameters do not meet the preset conditions, a fault response can be triggered.
[0103] Specifically, the monitoring unit supports parallel monitoring of 8 acceleration calculation modules and real-time feedback on the quality of transmitted data.
[0104] According to an embodiment of the present application, for protocol layer error detection, the monitoring unit can parse the data transmission protocol between the management board and the acceleration calculation module and monitor protocol layer faults based on the parsed errors.
[0105] In one embodiment, the data transmission protocol may include PCI-ETLP packets (PCI Express Transaction Layer Packets), CXL instructions (Compute Express Link), and Gen-Z messages (Generation Z).
[0106] According to an embodiment of the present application, parsing the data transmission protocol mainly involves parsing the request / response latency of the CXL protocol, the ECRC (End-to-End Cyclic Redundancy Check) error of PCI-E, and sequence errors, and counting the error types through hardware counters and updating the statistical results per second to achieve real-time detection of protocol layer faults.
[0107] According to an embodiment of the present application, the monitoring unit supports real-time monitoring of the acceleration calculation module, real-time feedback on signal quality, and real-time monitoring of protocol layer faults.
[0108] According to an embodiment of the present application, the management board further includes: a diagnostic engine for locating the fault location of the acceleration calculation module and performing recovery processing on the faulty acceleration calculation module.
[0109] According to an embodiment of the present application, a diagnostic engine is also arranged on the management board 110. The diagnostic engine can be used to locate the fault location of the acceleration calculation module and can also execute a self-healing mechanism.
[0110] Specifically, the diagnostic engine can track the bus state transition based on a finite state machine, combine with the Bayesian network algorithm, locate the fault level, and quickly locate the fault location of the fault point in the acceleration calculation module.
[0111] Among them, the finite state machine tracking the bus state transition means managing and controlling different states of the bus in the system in the way of a finite state machine and tracking the transition process of the bus state.
[0112] Specifically, the diagnostic engine can also execute a self-healing mechanism when detecting communication anomalies in the acceleration computing module to recover the faulty acceleration computing module.
[0113] In one embodiment, the self-healing mechanism can include automatically restarting the power supply corresponding to the acceleration computing module, resetting the link, etc.
[0114] According to the embodiments of the present application, the diagnostic engine also supports the isolation of faulty acceleration computing modules and the system to operate at a reduced level, which can improve system reliability.
[0115] According to the embodiments of the present application, through the common grounding design of the hardware circuit and the intelligent self-healing mechanism, the anti-interference ability is improved, making it suitable for the harsh industrial environment. At the same time, based on the fault location algorithm of the finite state machine and the Bayesian network, microsecond-level anomaly response and system recovery can be achieved, improving system reliability.
[0116] Figure 5 Shows a schematic diagram of the hierarchical dynamic software system on the management board according to the embodiments of the present application.
[0117] As Figure 5 shown, a hierarchical dynamic software system 500 is also arranged on the management board. The hierarchical dynamic software system 500 can include a hardware layer 510, a management layer 520, and an application interface layer 530.
[0118] Based on the above, the processing unit 111, the bus interface unit 113, the logic unit 112, the storage and configuration unit 200, the power supply and clock subsystem 300, the data bus, the serializer / deserializer, the monitoring unit, and the diagnostic engine arranged on the management board 110 belong to the hardware architecture on the management board 110.
[0119] According to the embodiments of the present application, adopting the architecture of "heterogeneous hardware integration + hierarchical software collaboration", hardware reuse and protocol adaptation between the nose and the general substrate can be realized through standardized interfaces.
[0120] According to the embodiments of the present application, the hardware layer is used to provide a unified interface for the management board to adapt the hardware of the management board to multiple types of general substrates; device enumeration and hot plug detection are implemented for the acceleration computing module to obtain detection results; and an acceleration computing module list is generated based on the detection results, where the acceleration computing module list includes at least one of the connection status, hot plug status, load, and data transfer rate of the acceleration computing module.
[0121] According to an embodiment of the present application, the hardware layer 510 is responsible for encapsulating hardware interface drivers, specifically including a power management unit driver, a clock generator driver, and a SerDes PHY (Serializer / Deserializer Physical Layer) control interface.
[0122] Specifically, the hardware layer 510 also provides a unified API (Application Programming Interface) for the management board 110 to call.
[0123] Therefore, the management board can be compatible with the hardware of various types of universal baseboards, shielding the hardware differences between different types of universal baseboards and achieving hardware independence.
[0124] In addition, the hardware layer 510 is also used to implement device enumeration and hot plug detection for the accelerated computing module, and obtain the detection results, thereby generating an accelerated computing module list based on the detection results.
[0125] Device enumeration refers to identifying and listing all currently connected hardware devices, that is, listing all currently connected accelerated computing modules.
[0126] Specifically, the hardware layer 510 detects whether a new accelerated computing module has been connected to obtain the connection status of the accelerated computing module. When an accelerated computing module is inserted or removed, the hardware layer 510 senses the change in hot-plug status and triggers an interrupt or event. Hot-plug detection ensures that the accelerated computing module can be automatically identified when inserted and that related resources are promptly cleaned up when removed. The detection results may include the detected connection status of the accelerated computing module and the hot-plug status of the accelerated computing module.
[0127] Thus, a list of accelerated computing modules can be generated based on the detection results.
[0128] After the hardware layer identifies a new accelerated computing module and generates a list of accelerated computing modules, it automatically loads the protocol parsing plug-in associated with the new module, enabling plug-and-play. Furthermore, the list of accelerated computing modules is updated whenever a new module is added or the status of an existing module changes.
[0129] According to the embodiments of this application, the hardware layer, through a unified interface, can shield hardware differences between different universal baseboards, achieving hardware independence and automatically loading corresponding protocol parsing plug-ins, supporting plug-and-play. Furthermore, the hardware layer can automatically adapt to new devices when hardware changes occur and generate or update the accelerated computing module list without intervention, greatly improving system usability and user experience.
[0130] According to an embodiment of the present application, the management layer further includes: a power supply and clock management module, configured to adjust the current output of the power management unit according to the load of the acceleration computing modules in the acceleration computing module list; and adjust the frequency of the clock signal output by the clock generator according to the data transmission rate in the acceleration computing module list.
[0131] According to an embodiment of the present application, the management layer 520 in the hierarchical dynamic software system 500 further includes a power supply and clock management module. The power supply and clock management module can be used to implement dynamic power distribution and clock frequency adaptation.
[0132] According to an embodiment of the present application, for dynamic power distribution, the power supply and clock management module can be used to adjust the current output of the power management unit 310 according to the load of the acceleration computing modules in the acceleration computing module list, so as to achieve dynamic power distribution.
[0133] Specifically, the power supply and clock management module can adjust the output current of the power management unit according to the load of the acceleration computing module through hardware PWM (Pulse Width Modulation).
[0134] For example, if the power management unit supports a main power supply of 54V, the power supply and clock management module can adjust the output current of the 54V main power supply through hardware PWM.
[0135] According to an embodiment of the present application, the power supply and clock management module supports an energy-saving mode and a full-load mode, and can be used to integrate thermal power consumption limit protection, prevent overheating and frequency reduction, and improve the energy efficiency ratio.
[0136] According to an embodiment of the present application, for clock frequency adaptation, the power supply and clock management module can adjust the frequency of the clock signal output by the clock generator 320 according to the data transmission rate.
[0137] In one embodiment, the power supply and clock management module can achieve glitch-free switching through a hardware phase-locked loop to ensure optimal utilization of clock resources.
[0138] According to an embodiment of the present application, through the power supply and clock management module, dynamic power adjustment and clock frequency adaptation can be performed, so that the system energy efficiency ratio is improved, the full-load power consumption is reduced, the power consumption in the idle state is decreased, and the heat dissipation and energy-saving requirements of a high-density data center are adapted.
[0139] According to an embodiment of the present application, the management layer further includes: a protocol parsing module, configured to control the monitoring unit to parse the data transmission protocol according to a preset priority; and a control logic unit to process the data from the acceleration computing module to generate a monitoring result for the acceleration computing module.
[0140] According to an embodiment of the present application, the management layer 520 may further include a protocol parsing module, which may be used to parse data transmission protocols in sequence according to priorities.
[0141] Specifically, when the data transmission protocols include PCI-ETLP packets, CXL instructions, and Gen-Z messages, the order of priorities may be CXL memory access, PCI-E configuration access, and normal data transmission in sequence. Thus, by parsing CXL memory access, PCI-E configuration access, and normal data transmission in sequence, the parsing efficiency can be improved.
[0142] In one embodiment, the protocol parsing module may schedule parsing tasks through a software queue and parse data transmission protocols in sequence according to priorities.
[0143] According to an embodiment of the present application, the protocol parsing module is further used to control the logic unit 112 to process data from the acceleration computing module to generate a monitoring result for the acceleration computing module.
[0144] Specifically, the monitoring result may include a real-time throughput curve of the acceleration computing module, an error rate histogram, a bus utilization report, etc.
[0145] In one embodiment, the monitoring result for the acceleration computing module may be stored in the second memory 220 and uploaded to a remote management platform through an IPMI (Intelligent Platform Management Interface) interface, facilitating real-time monitoring and data analysis.
[0146] According to an embodiment of the present application, through the protocol parsing module, the hardware acceleration engine can achieve real-time parsing of 200Gbps-level data, with low parsing latency, short error detection response time, meeting the real-time monitoring requirements of high-speed buses, and shortening the fault location time from minutes to seconds. Moreover, it can achieve hierarchical parsing from the physical layer to the transaction layer, support hybrid processing and priority scheduling of protocols such as CXL and PCI-E, and improve the parsing efficiency.
[0147] According to an embodiment of the present application, the management layer further includes: a fault detection module, which is used to control a diagnostic engine to perform fault detection on the acceleration computing module based on a preset fault knowledge base; and perform a fault response in case of detecting a fault.
[0148] According to an embodiment of the present application, the management layer 520 may further include a fault detection module, and the fault detection module may be used for fault detection and fault response.
[0149] In one embodiment, a preset fault knowledge base can be established in advance, and the preset fault knowledge base can include more than 200 bus error codes. The fault detection module determines the faults of the acceleration calculation module from the preset fault knowledge base to a certain extent.
[0150] According to an embodiment of the present application, the fault detection module can control the diagnostic engine to perform fault detection on the acceleration calculation module based on the preset fault knowledge base.
[0151] Specifically, based on the preset fault knowledge base, the fault detection module supports rule-based rapid diagnosis and machine learning-based anomaly prediction to improve the fault handling efficiency.
[0152] According to an embodiment of the present application, in the case where the fault detection module detects a fault in the acceleration calculation module, a fault response can be performed.
[0153] In one embodiment, the fault detection module supports a three-level fault response. Specifically, the three-level fault response can include early warning, degradation, and shutdown.
[0154] Among them, the fault response strategy can be customized through the remote management interface to meet the requirements of different scenarios.
[0155] According to an embodiment of the present application, the fault detection module performs fault detection on the acceleration calculation module based on the preset fault knowledge base, improves the fault handling efficiency, and supports different fault response strategies to meet the requirements of different scenarios.
[0156] Figure 6 Shows a schematic diagram of the application interface layer on the management board according to an embodiment of the present application.
[0157] As Figure 6 shown, the application interface layer 530 can include an on-board interaction interface 610 and a remote management interface 620.
[0158] In one embodiment, the on-board interaction interface 610 can be used to display the status of the acceleration calculation module in real time.
[0159] Specifically, the on-board interaction interface 610 can display the status of the acceleration calculation module in real time through an OLED (Organic Light Emitting Diode) display screen.
[0160] Among them, the status of the acceleration calculation module can include the connection status of the acceleration calculation module, monitoring results, etc.
[0161] According to an embodiment of the present application, the on-board interaction interface 610 also supports button input to configure parameters, so that the interaction response time is short, which is convenient for local operation and monitoring.
[0162] In one embodiment, the remote management interface 620 can be used to obtain the status of the management board and issue configuration instructions for the management board.
[0163] Specifically, the remote management interface 620 supports the IPMI2.0 (Intelligent Platform Management Interface 2.0) and SNMPv3 (Simple Network Management Protocol version 3) protocols, and can remotely obtain the status of the management board, issue configuration instructions for the management board, and upgrade the firmware. At the same time, the remote management interface 620 is also compatible with mainstream data center management platforms, supports batch device monitoring and configuration, and realizes remote operation and maintenance.
[0164] According to an embodiment of the present application, the on-board interaction interface can facilitate local operation and monitoring; the remote management interface can support batch device monitoring and configuration, and realize remote operation and maintenance.
[0165] Figure 7 FIG. shows a schematic diagram of a management board according to another embodiment of the present application.
[0166] As Figure 7 shown, the management board can be specifically divided into a hierarchical dynamic software system 500 and a hardware architecture 700.
[0167] According to an embodiment of the present application, the hardware architecture 700 can include a monitoring and diagnosis module 710, a storage and configuration unit 200, a power supply and clock subsystem 300, a bus interface unit 113, and a control module 720.
[0168] Among them, the monitoring and diagnosis module 710 can include a monitoring unit and a diagnosis engine; the control module 720 can include a processing unit 111, a logic unit 112, and a bus interface unit 113.
[0169] According to an embodiment of the present application, Figure 7 the layout of each layer and each unit in the management board 110 shown in
[0170] The present application also provides a general substrate, which is integrated with the above-mentioned management board.
[0171] In Figure 1 the management board 110 is integrated onto the general substrate 100.
[0172] According to an embodiment of the present application, the management board is deeply integrated with the general substrate, the volume is reduced compared with the traditional solution, the board area occupation is reduced, and high integration and reliability are achieved.
[0173] According to an embodiment of the present application, a general substrate includes an acceleration computing module disposed on one side of the general substrate close to the bus interface unit of the management board.
[0174] In Figure 1 the acceleration computing module disposed on the general substrate 100 is close to the bus interface unit 113 on the management board 110.
[0175] According to an embodiment of the present application, the bus interface unit on the management board is close to the acceleration computing module on the general substrate to reduce the signal trace length.
[0176] Taking Figure 1 the layout of the general substrate shown as an example, the SerDes interface (bus interface unit 113) can be directly soldered beside the acceleration computing module of the general substrate 100. The SerDes interface can also be connected to the logic unit 112 through 20 - mil differential traces, and the impedance matching of the differential traces is controlled within 50Ω ± 5%; the power pins are connected with thick copper wires to ensure the stability of large - current transmission.
[0177] Based on the above, the management board provided by the present application can achieve the following goals: deep hardware reuse, compatible with multiple generations of UBB boards and OAM modules through programmable logic FPGA and standardized interfaces, supporting plug - and - play, reducing the R & D cost by more than 50%; multi - protocol adaptation: hardware acceleration and software collaboration, real - time parsing of protocols such as CXL3.0 and PCI - E5.0, with the parsing delay less than or equal to 50ns and the error detection response time less than 1μs; intelligent resource collaboration: dynamically allocate power and clock resources, support power consumption adaptive adjustment of the OAM module, improve the energy efficiency ratio by 20%, and the power consumption in the idle state is less than 150mW; high integration and low latency: deeply integrated with the UBB board, reducing the volume by 40%, realizing hardware resource sharing, and improving the system reliability and density.
[0178] Figure 8 The flowchart of the monitoring method according to an embodiment of the present application is shown.
[0179] As Figure 8 shown, the monitoring method 800 includes operations S810 to S830.
[0180] According to an embodiment of the present application, the monitoring method 800 can be implemented by using the above - mentioned management board.
[0181] In operation S810, in response to detecting the power - on of the general substrate, the processing unit selects a target firmware from multiple firmwares based on the type of the general substrate, and configures the logic unit through the target firmware so that the management board adapts to the general substrate of the target type.
[0182] In operation S820, an electrical connection between the management board and the acceleration computing module disposed on the general substrate is realized by using the bus interface unit.
[0183] In operation S830, the data from the acceleration computing module is processed by using the logic unit to monitor the acceleration computing module.
[0184] According to an embodiment of the present application, the processing unit is adapted to the logic unit, that is, the processing unit and the logic unit work together. The processing unit can select the target firmware according to the type of the general substrate integrated into the management board, and configure the logic unit with the target firmware, so that the management board can be adapted to the general substrate integrated therein. Thus, based on different types of general substrates integrated, the processing unit can configure the logic unit with different firmwares to adapt to different types of general substrates. On this basis, the management board of the present application can support more than 80% of the general substrate types, without re-designing the hardware, shortening the R & D cycle by 70%, reducing the hardware development cost by 50%, and adapting to the rapid iteration of the acceleration computing module. Moreover, real-time parsing of 200Gbps-level data is realized through the hardware acceleration engine of the logic unit, meeting the real-time monitoring requirements of the high-speed bus.
[0185] Figure 9 A flowchart of the application of the management board according to an embodiment of the present application is shown.
[0186] As Figure 9 shown, this embodiment includes operation S910 to operation S922.
[0187] In operation S910, in response to detecting that the general substrate is powered on, the target firmware is loaded by using the processing unit to configure the logic unit.
[0188] In operation S911, the clock generator outputs a clock signal to synchronize the clocks of the acceleration computing module and the management board.
[0189] According to an embodiment of the present application, after operations S910 and S911 are executed, the system initialization is completed.
[0190] In operation S912, the data of the acceleration computing module collected is preprocessed.
[0191] According to an embodiment of the present application, preprocessing the data may include oversampling by an ADC (Analog-to-Digital Converter) and hardware equalization processing, restoring the original data stream, calibrating the signal DC offset and cable attenuation, and improving the signal quality.
[0192] In operation S913, the monitoring unit analyzes the data transmission protocol and monitors protocol layer faults.
[0193] According to an embodiment of the present application, the parsing engine identifies CXL instructions or PCI-E TLPs, extracts key information, and the real-time monitoring engine counts the error rate.
[0194] In operation S914, it is determined whether the error rate exceeds a threshold.
[0195] According to an embodiment of the present application, in the case where the error rate exceeds the threshold, operation S915 is executed; in the case where the error rate does not exceed the threshold, operation S917 is executed.
[0196] In operation S915, a hardware interrupt occurs.
[0197] In operation S916, the processing unit records the exception and reports it.
[0198] In operation S917, the protocol parsing module controls the logic unit to process the data from the acceleration computing module to generate a monitoring result for the acceleration computing module.
[0199] In operation S918, the monitoring result is uploaded to the remote management platform.
[0200] In operation S919, it is determined whether the user operates to trigger.
[0201] According to an embodiment of the present application, in the case where the user operation triggers, operation 920 is executed; in the case where the user operation does not trigger, operation S921 is executed.
[0202] In operation S920, configuration parameters are executed or the self-healing mechanism is triggered.
[0203] According to an embodiment of the present application, the user can configure parameters or trigger the self-healing mechanism through the interface to achieve fast fault handling.
[0204] In operation S921, the current state is maintained.
[0205] In operation S922, it ends.
[0206] Those skilled in the art can understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0207] The embodiments of the present application have been described above. However, these embodiments are merely for illustrative purposes and not for limiting the scope of the present application. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.
Claims
1. A management board, characterized in that, The management board includes: A processing unit, configured to select a target firmware from multiple firmwares based on the type of the general substrate, and configure a logic unit through the target firmware, so that the management board adapts to the general substrate of the target type; A bus interface unit, configured to implement an electrical connection between the management board and an acceleration computing module disposed on the general substrate; The logic unit, electrically connected to the bus interface unit, configured to process data from the acceleration computing module to implement monitoring of the acceleration computing module.
2. The management board according to claim 1, wherein The management board further includes: A first memory, configured to store logic unit firmware; A second memory, configured to store data from the acceleration computing module and monitoring results of the acceleration computing module; A debugging interface, connecting the processing unit and the logic unit; Reconfigurable hardware, including a preset number of logic units.
3. The management board according to claim 1, characterized in that, The management board further includes: A power management unit, connected to the logic unit, configured to supply power to the acceleration computing module; A clock generator, connected to the logic unit, configured to provide a clock signal to the acceleration computing module.
4. The management board according to claim 1, characterized in that, The bus interface unit includes: A data bus, configured to electrically connect a serializer / deserializer on the management board and a serializer / deserializer on the acceleration computing module to implement data interaction between the management board and the acceleration computing module; The serializer / deserializer, configured to convert serial data and parallel data.
5. The management board according to claim 3, characterized in that The management board further includes: A monitoring unit, configured to Collect eye diagram parameters from the acceleration computing module, and trigger a fault response when the eye diagram parameters do not meet preset conditions; Analyze a data transmission protocol between the management board and the acceleration computing module to monitor protocol layer faults.
6. The management board according to claim 5, characterized in that, The management board further includes: A diagnostic engine, configured to locate a fault location of the acceleration computing module and perform a recovery process on the faulty acceleration computing module.
7. The management board according to claim 6, wherein, The management board further includes a hierarchical dynamic software system, and the hierarchical dynamic software system includes a hardware layer, a management layer, and an application interface layer.
8. The management board according to claim 7, wherein The hardware layer is configured to Provide a unified interface for the management board to adapt the hardware of the management board to multiple types of general substrates; Implement device enumeration and hot plug detection for the acceleration computing module to obtain a detection result; Generate an acceleration computing module list based on the detection result, wherein the acceleration computing module list includes at least one of a connection state, a hot plug state, a load, and a data transmission rate of the acceleration computing module.
9. The management board according to claim 8, characterized in that The management layer further includes: A power and clock management module, configured to Adjust the current output of the power management unit according to the load of the acceleration computing module in the acceleration computing module list; Adjust the frequency of the clock signal output by the clock generator according to the data transmission rate in the acceleration computing module list.
10. The management board according to claim 9, characterized in that, The management layer further includes: A protocol analysis module, configured to Control the monitoring unit to analyze the data transmission protocol according to a preset priority; Control the logic unit to process the data from the acceleration computing module to generate a monitoring result for the acceleration computing module.
11. The management board according to claim 9, wherein The management layer further includes a fault detection module for: Controlling the diagnostic engine to perform fault detection on the acceleration computing module based on a preset fault knowledge base; Performing a fault response in the case of detecting a fault.
12. The management board according to claim 7, characterized in that, The application interface layer includes: An on-board interaction interface for real-time display of the status of the acceleration computing module; A remote management interface for obtaining the status of the management board and issuing configuration instructions for the management board.
13. A general substrate, characterized in that, The general substrate is integrated with the management board according to any one of claims 1 to 12.
14. The general substrate according to claim 13, wherein The general substrate includes an acceleration computing module disposed on one side of the general substrate close to the bus interface unit of the management board.
15. A monitoring method, applied to the management board according to any one of claims 1 to 12, characterized in that, The method includes: In response to detecting that the general substrate is powered on, using the processing unit to select a target firmware from multiple firmware based on the type of the general substrate, and configuring the logic unit through the target firmware so that the management board adapts to the general substrate of the target type; Using the bus interface unit to implement an electrical connection between the management board and the acceleration computing module disposed on the general substrate; Using the logic unit to process the data from the acceleration computing module to implement monitoring of the acceleration computing module.
Citation Information
Patent Citations
Test fixture, test system and test method for OAM type accelerator card
CN116560920A
Baseboard management control system, controller, server and supervision method
CN116841820A
Baseboard management system and server
CN118093474A
Firmware upgrading system and method, server device, program product and storage medium
CN118409775A
Storage and calculation integrated calculation module based on OAM form
CN119862151A