PC Farm array server management and control device

By combining the embedded BMC control module, the multi-protocol hardware channel module, and the intelligent monitoring and diagnostic module, hardware-level full-domain control of the PC Farm array server is achieved, solving the problems of manual dependence, management loopholes, and monitoring blind spots in existing technologies, and improving operation and maintenance efficiency and the integrity of cluster management.

CN121523992APending Publication Date: 2026-02-13启朔(深圳)科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511382462.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing PC Farm array server management technologies suffer from problems such as heavy reliance on manual operation and maintenance, management architecture vulnerabilities, bottlenecks in large-scale deployment, and monitoring blind spots, making it difficult to achieve efficient management and control.

Method used

Design a PC Farm array server management and control device, which includes an embedded BMC management and control module, a multi-protocol hardware channel module, an out-of-band management module, and an intelligent monitoring and diagnostic module to achieve hardware-level full-domain management and control.

Benefits of technology

Improve operational efficiency and continuity, fill gaps in the management architecture, overcome bottlenecks in large-scale deployment, enhance monitoring and fault handling capabilities, and reduce costs and risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121523992A_ABST
    Figure CN121523992A_ABST
Patent Text Reader

Abstract

The invention provides a PC Farm array server management and control device, belongs to the technical field of computer hardware and server management, and aims to solve the problems of tedious manual operation and maintenance, lack of management architecture, large-scale deployment bottleneck and monitoring blind areas in existing PC Farm management and control. The device takes an embedded BMC management and control module as a core, and is in communication connection with a multi-protocol hardware channel module, an out-of-band management module and an intelligent monitoring and diagnosis module; the BMC module independently controls the node cluster, the out-of-band management module realizes remote batch power supply control, KVM operation, system deployment and firmware updating, and the intelligent monitoring and diagnosis module acquires hardware data and disposes abnormity. The scheme gets rid of dependence of a node operating system, improves operation and maintenance efficiency and cluster stability, and adapts to large-scale management and control requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer hardware and server management technology, specifically to a PC Farm array server management and control device. Background Technology

[0002] PC Farm array servers typically consist of clusters of multiple independent blade nodes to meet large-scale computing needs. However, existing technologies have many problems in practical applications and are difficult to adapt to the requirements of efficient management and control.

[0003] 1. Heavy reliance on manual operation and maintenance - Whether it is system configuration, BIOS update or application deployment, it must be done manually on a node-by-node basis; troubleshooting also requires on-site visits, which is not only cumbersome, but also prone to interruption of the process due to human error, affecting the normal operation of the cluster;

[0004] 2. The management architecture has obvious flaws. Consumer-grade hardware (such as home motherboards) does not have industrial-grade out-of-band management interfaces and cannot use standard management protocols such as IPMI and Redfish. External management tools such as Ansible rely on the network function of the node's operating system to work. Once the system crashes, they become completely useless and cannot even control hardware such as power supplies and fans.

[0005] 3. Bottleneck in large-scale deployment: Different node hardware (such as different brands of GPUs and different generations of CPUs) leads to numerous driver compatibility issues, requiring manual adaptation and installation for each node. The more nodes there are, the longer the deployment time becomes.

[0006] 4. The solutions themselves have flaws—semi-automated management platforms require additional DHCP servers and network configurations, which increases hardware costs and amplifies the risk of network attacks; software-level monitoring tools like Zabbix can only collect system data while the nodes are running, and cannot monitor hardware status (such as voltage and temperature) or offline nodes, resulting in monitoring blind spots.

[0007] To address this, a PC Farm array server management and control device is proposed. Summary of the Invention

[0008] The present invention aims to solve the problems mentioned in the background art by providing a PC Farm array server management and control device.

[0009] The specific technical solution is as follows:

[0010] A PC Farm array server management device includes a chassis, a node cluster consisting of multiple blade nodes housed within the chassis, an embedded BMC management module, and a multi-protocol hardware channel module, an out-of-band management module, and an intelligent monitoring and diagnostic module, all of which establish communication connections with the embedded BMC management module. The embedded BMC management module interacts with each blade node via the multi-protocol hardware channel module to independently manage the entire node cluster. The out-of-band management module includes a power management submodule, a KVM-over-IP submodule, a batch deployment submodule, and a firmware update submodule. The power management submodule controls the power status of the node cluster, the KVM-over-IP submodule enables remote visualization operation of the node cluster, the batch deployment submodule performs batch installation of the system and drivers for the node cluster, and the firmware update submodule performs parallel firmware updates for the node cluster. The intelligent monitoring and diagnostic module collects hardware status data from the node cluster, performs threshold judgments on the data, and executes log recording, abnormal alarms, and subsequent processing operations based on the judgment results.

[0011] The aforementioned PC Farm array server management device includes an embedded BMC management module designed based on the Arm Cortex-A series core chip, with the core chip being the RK3568, and equipped with 8GB of memory and 64GB of storage. The storage unit is used to store the logical ID and physical port mapping table of the blade nodes, the fault code library, and the automated management scripts.

[0012] The aforementioned PC Farm array server management device includes a multi-protocol hardware channel module comprising four UART interfaces and four I / O ports. 2 It features a USB-C interface, four USB Host interfaces, one USB OTG interface, and multiple SPI interfaces; the UART interface is used for blade node debugging, BIOS-level fault diagnosis, and log capture. 2 The C interface is used to collect real-time data on the blade node's temperature, voltage, fan speed, and power consumption. The USB Host interface is used to mount a virtual optical drive for system deployment and to simulate mouse and keyboard operations. The USB OTG interface is used for data exchange with external devices. The SPI interface is used for BIOS and EC firmware updates of the blade node.

[0013] In the aforementioned PC Farm array server management device, the power management submodule of the out-of-band management module is connected to three power modules inside the chassis, namely power module 1, power module 2, and power module 3. The power management submodule can realize the power-on, power-off, restart, and forced power-off operations of a single blade node or a batch of blade nodes. It can also control the on / off status of the 12VStandBy power supply and monitor the output voltage stability of each power module in real time.

[0014] The aforementioned PC Farm array server management device includes an out-of-band management module whose KVM-over-IP submodule is connected to an HDMI-to-USB chip and an external HDMIOut interface. The HDMI-to-USB chip converts the HDMI signal output from the blade node into a USB signal and transmits it to the embedded BMC management module. The external HDMIOut interface is used to connect an external display device to view the screen of any blade node in real time. The KVM-over-IP submodule also supports remote keyboard and mouse operation of the blade node via an external USB input device.

[0015] The aforementioned PC Farm array server management device includes an out-of-band management module with a batch deployment submodule comprising a USB switching module and a virtual media mounting unit. The USB switching module is used to switch the USB communication link between the embedded BMC management module and each blade node. The virtual media mounting unit can load an ISO image file and mount it as a virtual optical drive to multiple blade nodes simultaneously. The batch deployment submodule can also push preset automated deployment scripts to each blade node to achieve automatic disk partitioning and driver installation of the blade nodes.

[0016] The aforementioned PC Farm array server management device includes an out-of-band management module whose firmware update submodule supports parallel flashing of the BIOS and graphics card vBIOS of each blade node and has a resume function. The firmware update submodule establishes communication with the motherboard EC and graphics card power supply module of the blade node and monitors the voltage stability of the motherboard 24-pin power supply and CPU 8-pin power supply before flashing the firmware. The graphics card power supply module has three reserved 8-pin interfaces, and the firmware update submodule controls the output voltage of the graphics card power supply module to remain stable during the flashing of the graphics card vBIOS.

[0017] The aforementioned PC Farm array server management device, wherein the intelligent monitoring and diagnostic module includes a sensor data acquisition unit, a threshold judgment unit, a log recording unit, an alarm processing unit, and an action execution unit; the sensor data acquisition unit communicates with I... 2 The C interface collects data on the CPU temperature, GPU temperature, power supply voltage, and fan speed of the blade node. The fans are two PWM-controlled 12000RPM fans, one on each side of the GPU PCIE Gen4x16 adapter board. The threshold judgment unit compares the collected data with preset thresholds. If the data meets the preset threshold range, the log recording unit records the current hardware status log. If the data exceeds the preset threshold range, the alarm processing unit triggers an alarm, and the action execution unit performs automatic frequency reduction or shutdown of the blade node.

[0018] The aforementioned PC Farm array server management device includes an intelligent monitoring and diagnostic module that further comprises a fault code library. The fault code library contains various fault modes such as heat dissipation failure mode and power instability mode. The intelligent monitoring and diagnostic module can match the collected temperature curve change trend and voltage fluctuation data with the fault modes in the fault code library to determine the fault type, and generate corresponding maintenance suggestions based on the fault type.

[0019] In the aforementioned PC Farm array server management device, the embedded BMC management module assigns a unique logical ID to each blade node during the management initialization phase. The logical ID ranges from 1 to 8, and a mapping table between the logical ID and the physical port of the blade node is established. After the out-of-band management module completes system deployment, the batch deployment submodule obtains the installation interface image of each blade node through the KVM-over-IP submodule, uses OCR recognition technology to identify keywords in the installation interface to verify the system deployment results, and automatically triggers a retry deployment operation for blade nodes that fail to deploy.

[0020] The present invention has the following beneficial effects:

[0021] 1. Improve operational efficiency and continuity – BMC can be independently managed, and out-of-band management modules can be operated remotely in batches, eliminating the need for manual node-by-node processing. Faults can also be checked remotely, reducing the need for manual intervention. Problems can be handled without going to the site, and cluster operation is less prone to interruption.

[0022] 2. Filling the gaps in the management architecture - It integrates BMC modules and multi-protocol channels, and can directly manage hardware such as power supplies and fans. It can operate regardless of whether the node system is normal or not, which solves the problem that consumer-grade hardware does not have external interfaces and external tools depend on the system. It forms a complete management link from the hardware layer to the control layer.

[0023] 3. Break through the bottleneck of large-scale deployment - the batch deployment sub-module can automatically push scripts and adapt to the drivers of heterogeneous nodes, without the need for manual adjustment one by one. No matter how many nodes there are, the deployment will not take more time and can adapt to the needs of cluster expansion.

[0024] 4. Enhanced monitoring and fault handling capabilities: The intelligent monitoring module can monitor hardware status and offline nodes, with no blind spots. It can also automatically handle abnormalities to avoid hardware damage. The fault code library can directly match fault types and provide maintenance suggestions, eliminating the need for blind troubleshooting based on experience, making repairs faster and more accurate.

[0025] 5. Reduced costs and risks – No additional DHCP server is required, reducing hardware costs and avoiding the attack risks associated with additional network configurations. This enhances management capabilities while controlling costs and ensuring security. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the architecture of a PC Farm array server management device provided in an embodiment of the present invention. Detailed Implementation

[0027] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0028] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this application. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0029] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present application. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0030] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0031] Example

[0032] The PC Farm array server management device provided in this embodiment, such as Figure 1As shown, the system includes a chassis, a node cluster consisting of multiple blade nodes housed within the chassis, an embedded BMC control module, and a multi-protocol hardware channel module, an out-of-band management module, and an intelligent monitoring and diagnostic module, all of which establish communication connections with the embedded BMC control module. The embedded BMC control module interacts with each blade node via the multi-protocol hardware channel module to independently manage the entire node cluster. The out-of-band management module includes a power management submodule, a KVM-over-IP submodule, a batch deployment submodule, and a firmware update submodule. The power management submodule controls the power status of the node cluster, the KVM-over-IP submodule enables remote visualization operation of the node cluster, the batch deployment submodule performs batch installation of the system and drivers for the node cluster, and the firmware update submodule performs parallel firmware updates for the node cluster. The intelligent monitoring and diagnostic module collects hardware status data from the node cluster, performs threshold checks on the data, and executes log recording, abnormal alarms, and subsequent processing operations based on the results.

[0033] This solution achieves hardware-level full-domain control of PC Farm array servers by constructing an architecture with an embedded BMC control module at its core, complemented by multi-protocol hardware channel modules, out-of-band management modules, and intelligent monitoring and diagnostic modules. The embedded BMC control module directly interacts with each blade node, eliminating dependence on the node's operating system and resolving the issue of traditional external tools (such as Ansible) becoming completely ineffective after a system crash. The out-of-band management module's four sub-modules cover power control, remote operation, batch deployment, and firmware updates, automating the entire cluster control process and avoiding the tedious manual operation of each node. The intelligent monitoring and diagnostic module provides real-time monitoring and anomaly handling of hardware status, filling the gap in hardware-level data coverage that traditional software-level monitoring tools (such as Zabbix) cannot provide. The overall architecture forms a closed loop from the control center to the functional modules, comprehensively addressing the pain points of existing solutions' lack of management architecture and reliance on manual maintenance, achieving independence and completeness in cluster control.

[0034] Specifically, in this embodiment, the embedded BMC control module is designed based on the Arm Cortex-A series core chip, which is the RK3568 and is equipped with 8GB of memory and 64GB of storage. The storage unit is used to store the logical ID and physical port mapping table of the blade node, the fault code library and the automated control script.

[0035] This solution explicitly designes the embedded BMC control module based on Arm Cortex-A series core chips (such as the RK3568) and configures it with fixed memory and storage. The storage content includes a logical ID mapping table, a fault code library, and automation scripts. This design enables the BMC module to operate stably and independently, carrying the core data required for cluster control without relying on external storage devices, thus avoiding the risk of control interruption due to external storage failure. The selection of the RK3568 chip and the matching of memory and storage configurations ensure that the module can efficiently handle tasks such as multi-node data interaction and batch command issuance without experiencing lag due to insufficient computing power or storage. The pre-set logical ID mapping table and fault code library provide basic support for subsequent initialization configuration and fault diagnosis, reducing temporary data processing steps in the control process and improving overall control efficiency.

[0036] Specifically, in this embodiment, the multi-protocol hardware channel module includes 4 UART interfaces and 4 I / O ports. 2 It features a USB-C interface, four USB Host interfaces, one USB OTG interface, and multiple SPI interfaces; the UART interface is used for blade node debugging, BIOS-level fault diagnosis, and log capture. 2 The C interface is used to collect real-time data on the blade node's temperature, voltage, fan speed, and power consumption. The USB Host interface is used to mount a virtual optical drive for system deployment and to simulate mouse and keyboard operations. The USB OTG interface is used for data exchange with external devices. The SPI interface is used for BIOS and EC firmware updates of the blade node.

[0037] This solution clarifies the interface types and functions of the multi-protocol hardware channel module (UART, I...). 2 (C, USB, SPI) constructs a full-scenario hardware interaction link covering "debugging and diagnostics - data acquisition - deployment and operation - firmware update". Among them, the UART interface directly connects to node debugging and fault log capture, obtaining BIOS-level data without going through the system, solving the problem of traditional solutions requiring on-site operation for fault diagnosis; I 2 The Type-C interface collects real-time hardware data such as temperature and voltage, providing a precise data source for intelligent monitoring and diagnostics, avoiding the limitations of software-level monitoring that can only acquire system data. The USB interface (including Host and OTG) supports virtual optical drive deployment and keyboard / mouse emulation, providing hardware-level support for system installation and remote operation without relying on the node's own USB interface. The SPI interface is dedicated to BIOS and EC firmware updates, ensuring the stability and specificity of firmware flashing, avoiding the dependence on the system environment of traditional update methods. These multiple interfaces, each with its own function and working collaboratively, achieve seamless hardware-level interaction, providing a reliable channel for the implementation of subsequent management and control functions.

[0038] Specifically, in this embodiment, the power management submodule of the out-of-band management module is connected to three power modules inside the chassis, namely power module 1, power module 2, and power module 3. The power management submodule can realize the power-on, power-off, restart, and forced power-off operations of a single blade node or a batch of blade nodes. It can also control the on / off status of the 12V StandBy power supply and monitor the output voltage stability of each power module in real time.

[0039] This solution connects the power management submodule to three power modules within the chassis, providing them with individual / batch power control, 12V StandBy power management, and voltage monitoring capabilities, thus achieving refined and safe management of the cluster power supply. Individual / batch power control allows for on-demand operation of node power status, eliminating the need for manual plugging and unplugging of power supplies at each node, solving the problem of low efficiency in traditional manual maintenance. 12V StandBy power management ensures that nodes maintain basic power supply even when offline, providing a guarantee for remote wake-up and offline data acquisition. Voltage stability monitoring can detect abnormal power output in real time, preventing node hardware damage or data loss due to power instability. The overall design covers both the flexibility of power control and the safety of hardware operation, filling the gap in traditional solutions that cannot achieve hardware-level power management.

[0040] Specifically, in this embodiment, the KVM-over-IP submodule of the out-of-band management module is connected to an HDMI-to-USB chip and an external HDMIOut interface; the HDMI-to-USB chip is used to convert the HDMI signal output by the blade node into a USB signal and transmit it to the embedded BMC control module; the external HDMIOut interface is used to connect an external display device to view the screen of any blade node in real time; the KVM-over-IP submodule also supports remote keyboard and mouse operation of the blade node through an external USB input device.

[0041] This solution enables remote, visual management of blade nodes through a combination of an HDMI-to-USB chip, an external HDMIOut interface, and an external USB input device. The HDMI-to-USB chip converts the node's HDMI signal into a USB signal recognizable by the BMC module, ensuring real-time remote access to the node's screen without the need for a local monitor. The external HDMIOut interface supports external display devices, allowing maintenance personnel to view the remote node's status locally, balancing remote and local operation needs. Support for external USB input devices ensures seamless remote keyboard and mouse operation, avoiding the latency or functional limitations of traditional remote tools. This design completely eliminates reliance on on-site equipment for node management, achieving "visual operation anytime, anywhere" and significantly improving the convenience of remote maintenance.

[0042] Specifically, in this embodiment, the batch deployment submodule of the out-of-band management module includes a USB switching module and a virtual media mounting unit. The USB switching module is used to switch the USB communication link between the embedded BMC control module and each blade node. The virtual media mounting unit can load ISO image files and mount them as virtual optical drives to multiple blade nodes simultaneously. The batch deployment submodule can also push preset automated deployment scripts to each blade node to realize automatic disk partitioning and driver installation of the blade nodes.

[0043] This solution achieves efficient batch deployment of multi-node systems and drivers through a USB switching module, a virtual media mounting unit, and automated script push functionality. The USB switching module flexibly switches the USB communication links between the BMC module and each node, ensuring that multiple nodes receive deployment commands simultaneously, avoiding the time-consuming problem of establishing links node by node in traditional deployments. The virtual media mounting unit mounts the ISO image to multiple nodes simultaneously, eliminating the need for a separate physical optical drive or USB flash drive for each node, reducing deployment hardware costs and operational steps. The automated script push automatically executes disk partitioning and driver installation, eliminating the need for manual intervention to address compatibility issues with heterogeneous nodes (such as different CPUs and GPUs), solving the bottleneck of increased time consumption in traditional large-scale deployments. The overall functionality forms a closed-loop deployment process of "link switching - image mounting - automatic installation," significantly improving multi-node deployment efficiency and adapting to the needs of large-scale cluster expansion.

[0044] Specifically, in this embodiment, the firmware update submodule of the out-of-band management module supports parallel flashing of the BIOS and graphics card vBIOS of each blade node, and has the function of resuming interrupted downloads; the firmware update submodule establishes communication with the motherboard EC and graphics card power supply module of the blade node, and monitors the voltage stability of the motherboard 24-pin power supply and CPU 8-pin power supply before flashing the firmware; the graphics card power supply module has three reserved 8-pin interfaces, and the firmware update submodule controls the output voltage of the graphics card power supply module to remain stable during the flashing of the graphics card vBIOS.

[0045] This solution achieves both high efficiency and security in firmware updates through parallel refresh, breakpoint resume, power supply monitoring, and graphics card power supply control. The parallel refresh function supports simultaneous updates of multiple BIOS nodes and the graphics card vBIOS, avoiding the time-consuming issues of traditional node-by-node updates. The breakpoint resume function ensures that updates do not need to be restarted after an interruption, reducing repetitive operations caused by temporary network or hardware failures. Pre-flash motherboard and CPU power supply monitoring can preemptively eliminate the impact of unstable power supply on firmware updates, preventing hardware failure due to update failure. The voltage control and 8-pin connector reservation of the graphics card power supply module not only adapt to the current graphics card power supply requirements but also provide compatibility for future graphics card upgrades, solving the problems of poor compatibility and high failure rates in traditional firmware updates. The overall design balances update efficiency and hardware security, ensuring a stable and reliable firmware update process.

[0046] Specifically, in this embodiment, the intelligent monitoring and diagnostic module includes a sensor data acquisition unit, a threshold judgment unit, a log recording unit, an alarm processing unit, and an action execution unit; the sensor data acquisition unit uses I... 2 The C interface collects data on the CPU temperature, GPU temperature, power supply voltage, and fan speed of the blade node. The fans are two PWM-controlled 12000RPM fans, one on each side of the GPU PCIE Gen4x16 adapter board. The threshold judgment unit compares the collected data with preset thresholds. If the data meets the preset threshold range, the log recording unit records the current hardware status log. If the data exceeds the preset threshold range, the alarm processing unit triggers an alarm, and the action execution unit performs automatic frequency reduction or shutdown of the blade node.

[0047] This solution, through the collaboration of sensor data acquisition, threshold judgment, log recording, alarm processing, and action execution units, combined with the configuration of fans of specific specifications, achieves real-time monitoring of hardware status and automatic handling of anomalies. The sensor data acquisition unit relies on I... 2 The C interface acquires core hardware data such as CPU and graphics card temperatures, ensuring monitoring coverage of key hardware indicators. The threshold judgment unit filters abnormal data using preset thresholds, avoiding the tediousness of manual real-time monitoring. The log recording unit automatically stores normal state data, providing a basis for subsequent maintenance and traceability. The alarm handling and action execution unit promptly alerts and triggers frequency reduction / shutdown in case of anomalies, preventing hardware damage due to continuous abnormalities. The PWM-controlled 12000RPM fan, paired with the graphics card adapter board, ensures stable graphics card cooling, reducing hardware anomalies caused by insufficient heat dissipation. The overall functionality forms a closed-loop monitoring system of "collection-judgment-handling," solving the problems of traditional software monitoring's inability to cover hardware and lack of automatic anomaly handling, ensuring long-term stable hardware operation.

[0048] Specifically, in this embodiment, the intelligent monitoring and diagnostic module also includes a fault code library, which contains various fault modes such as heat dissipation failure mode and power instability mode. The intelligent monitoring and diagnostic module can match the collected temperature curve change trend and voltage fluctuation data with the fault modes in the fault code library to determine the fault type, and generate corresponding maintenance suggestions based on the fault type.

[0049] This solution achieves precise fault location and efficient handling by pre-setting a fault code library (including modes such as heat dissipation failure and power instability) in the intelligent monitoring and diagnostic module. The fault code library provides pre-defined judgment criteria for hardware anomalies. When abnormal temperature curves or voltage fluctuations are detected, the corresponding fault mode can be directly matched, eliminating the need for maintenance personnel to check hardware one by one, thus solving the problems of directionless and time-consuming traditional fault diagnosis. Maintenance suggestions generated based on fault modes provide clear operational guidance for maintenance personnel, avoiding maintenance errors or delays caused by insufficient experience. This design transforms fault diagnosis from "experience-dependent" to "data-matching," significantly improving the accuracy and efficiency of fault handling and shortening the fault recovery cycle.

[0050] Specifically, in this embodiment, the embedded BMC control module assigns a unique logical ID to each blade node during the control initialization phase. The logical ID ranges from 1 to 8, and a mapping table between the logical ID and the physical port of the blade node is established. After the out-of-band management module completes the system deployment, the batch deployment submodule obtains the installation interface image of each blade node through the KVM-over-IP submodule, uses OCR recognition technology to identify keywords in the installation interface to verify the system deployment results, and automatically triggers a retry deployment operation for blade nodes that fail to deploy.

[0051] This solution achieves orderly cluster management and reliable deployment results by initializing and allocating logical IDs, establishing mapping tables, and implementing post-deployment OCR verification and automatic retry for failed nodes. The logical ID allocation and mapping table establishment during the initialization phase ensure that each blade node is uniquely identified during management, avoiding mis-issuance of management commands due to node confusion. Post-deployment OCR verification eliminates the need for manual inspection of each node's installation interface, automatically determining deployment success and reducing the tediousness and errors of manual verification. The automatic retry function for failed nodes eliminates the need for manual re-triggering of the deployment process, preventing deployment omissions due to temporary failures. The overall design covers both "orderly configuration" in the early stages of management and "result assurance" in the later stages, ensuring the accuracy and efficiency of the entire cluster process from initialization to deployment, and minimizing manual intervention.

[0052] In summary, the working principle of the PC Farm array server management device provided in this embodiment is as follows:

[0053] The core of this control system is the embedded BMC control module. Through its cooperation with multi-protocol hardware channel modules, out-of-band management modules, and intelligent monitoring and diagnostic modules, it achieves hardware-level control of the PC Farm cluster. The overall logic consists of four steps:

[0054] 1. Core Management and Control – The BMC module is based on Arm Cortex-A series chips. It can store the logical ID and physical port mapping table of blade nodes, fault code library and automation scripts. It can independently manage the entire cluster without relying on the operating system of any node. It can complete everything from data storage to command issuance on its own.

[0055] 2. Data Interaction – The multi-protocol hardware channel module acts as the “connector” between the BMC and the blade node: the UART interface is used to debug the node and capture BIOS-level fault logs; 2 The C interface collects real-time hardware data such as node temperature, voltage, and fan speed; the USB interface (including Host and OTG) can connect a virtual optical drive to install the system and simulate keyboard and mouse operations; the SPI interface is dedicated to updating the node's BIOS and EC firmware. Each interface performs its own function to ensure that the data is transmitted accurately and is usable.

[0056] 3. Out-of-band operation – The out-of-band management module is divided into four sub-modules: The power management sub-module connects to the three power supplies in the chassis, enabling individual or batch control of node power-on, power-off, and restart, as well as monitoring power supply voltage stability; The KVM-over-IP sub-module uses an HDMI-to-USB chip to convert the node's HDMI signal into a USB signal that the BMC can recognize, allowing remote viewing and operation of the node screen with an external monitor and keyboard / mouse; The batch deployment sub-module uses a USB switching module to connect different nodes, loads an ISO image as a virtual optical drive for multiple nodes simultaneously, and pushes automated scripts to allow the nodes to partition and install drivers themselves; The firmware update sub-module can flash the BIOS and graphics card vBIOS of multiple nodes simultaneously, and can resume from the breakpoint if interrupted. Before flashing, it also checks the stability of the motherboard and CPU power supply, and can also control the graphics card power supply.

[0057] 4. Intelligent Diagnosis – The intelligent monitoring and diagnosis module relies first on I... 2 The C interface collects hardware data and then uses a threshold to compare the unit with a preset standard: if the data is normal, it is logged for easy later review; if the data exceeds the limit, an alarm is triggered, and the node will be automatically reduced in frequency or shut down to prevent hardware failure; the module also has a fault code library that stores common fault patterns such as heat dissipation failure and unstable power supply, and can match the fault type based on the collected temperature changes and voltage fluctuations to provide maintenance suggestions.

[0058] How to use

[0059] The use of this device follows the process of "initialization - deployment - daily maintenance - troubleshooting", and the specific operations are as follows:

[0060] 1. Initialization Configuration - When the cluster starts up, the BMC module will assign a unique logical ID to each blade node, and then create a mapping table between the ID and the physical port of the node. At the same time, it will adjust the communication links between the BMC and each module and each node to prepare for subsequent management and control.

[0061] 2. Batch System Deployment – ​​When installing a system, first load the ISO image in the batch deployment submodule and mount it to the node to be installed; then BMC pushes the preset automated script to the node, and the node completes disk partitioning and driver installation by itself; after installation, the KVM-over-IP submodule will take a picture of the node installation interface and use OCR to recognize the keywords in the interface to determine whether the installation is successful. Nodes that have not been successfully installed will automatically retry.

[0062] 3. Routine remote operation and maintenance - To manage nodes, you can remotely view the screen and operate them by connecting an external monitor and keyboard and mouse through the KVM-over-IP submodule; to control power, you can use the power management submodule to perform batch or individual operations; to update firmware, you can start the firmware update submodule, select the nodes to flash simultaneously, and you don't need to restart if it is interrupted; to reinstall drivers, just push the corresponding script again.

[0063] 4. Troubleshooting – The intelligent monitoring and diagnostic module continuously monitors the hardware status. Once an alarm is triggered, first check the fault type and repair suggestions provided by the module. For example, if it indicates a heat dissipation failure, check if the fan next to the graphics card is broken; if it indicates an unstable power supply, check the power supply module. After repairing and restarting the node, the module will continue monitoring until it is confirmed that normal operation has been restored.

[0064] The above are merely preferred embodiments of the present invention and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should recognize that any equivalent substitutions and obvious changes made based on the description and illustrations of the present invention should be included within the protection scope of the present invention.

Claims

1. A PC Farm array server management and control device, characterized in that, The system includes a chassis, a node cluster consisting of multiple blade nodes housed within the chassis, an embedded BMC control module, a multi-protocol hardware channel module communicating with the embedded BMC control module, an out-of-band management module, and an intelligent monitoring and diagnostic module. The embedded BMC control module interacts with each blade node via the multi-protocol hardware channel module, enabling independent management of the entire node cluster. The out-of-band management module includes a power management submodule, a KVM-over-IP submodule, a batch deployment submodule, and a firmware update submodule. The power management submodule controls the power status of the node cluster, the KVM-over-IP submodule enables remote visualization operation of the node cluster, the batch deployment submodule handles batch installation of the system and drivers for the node cluster, and the firmware update submodule performs parallel firmware updates for the node cluster. The intelligent monitoring and diagnostic module collects hardware status data from the node cluster, performs threshold checks on the data, and executes log recording, anomaly alarms, and subsequent processing operations based on the results.

2. The PC Farm array server management and control device according to claim 1, characterized in that, The embedded BMC control module is designed based on the Arm Cortex-A series core chip, which uses the RK3568 core chip and is equipped with 8GB of memory and 64GB of storage. The storage unit is used to store the logical ID and physical port mapping table of the blade node, the fault code library and the automated control script.

3. The PC Farm array server management and control device according to claim 1, characterized in that, The multi-protocol hardware channel module includes 4 UART interfaces and 4 I / O ports. 2 It features a USB-C interface, four USB Host interfaces, one USB OTG interface, and multiple SPI interfaces; the UART interface is used for blade node debugging, BIOS-level fault diagnosis, and log capture. 2 The C interface is used to collect real-time data on the blade node's temperature, voltage, fan speed, and power consumption. The USB Host interface is used to mount a virtual optical drive for system deployment and to simulate mouse and keyboard operations. The USB OTG interface is used for data exchange with external devices. The SPI interface is used for BIOS and EC firmware updates of the blade node.

4. The PC Farm array server management and control device according to claim 1, characterized in that, The power management submodule of the out-of-band management module is connected to three power modules inside the chassis, namely power module 1, power module 2, and power module 3. The power management submodule can perform power-on, power-off, restart, and forced power-off operations on a single blade node or a batch of blade nodes. It can also control the on / off status of the 12V StandBy power supply and monitor the output voltage stability of each power module in real time.

5. The PC Farm array server management and control device according to claim 1, characterized in that, The KVM-over-IP submodule of the out-of-band management module is connected to an HDMI-to-USB chip and an external HDMIOut interface. The HDMI-to-USB chip is used to convert the HDMI signal output by the blade node into a USB signal and transmit it to the embedded BMC control module. The external HDMIOut interface is used to connect an external display device to view the screen of any blade node in real time. The KVM-over-IP submodule also supports remote keyboard and mouse operation of blade nodes via an external USB input device.

6. The PC Farm array server management and control device according to claim 1, characterized in that, The out-of-band management module's batch deployment submodule includes one USB switching module and one virtual media mounting unit; The USB switching module is used to switch the USB communication link between the embedded BMC control module and each blade node. The virtual media mounting unit can load ISO image files and mount them as virtual optical drives to multiple blade nodes at the same time. The batch deployment submodule can also push preset automated deployment scripts to each blade node to achieve automatic disk partitioning and driver installation for the blade nodes.

7. The PC Farm array server management and control device according to claim 1, characterized in that, The firmware update submodule of the out-of-band management module supports parallel flashing of the BIOS and graphics card vBIOS of each blade node, and has the function of resuming interrupted downloads; the firmware update submodule establishes communication with the motherboard EC and graphics card power supply module of the blade node, and monitors the voltage stability of the motherboard 24-pin power supply and CPU 8-pin power supply before flashing the firmware; the graphics card power supply module has three reserved 8-pin interfaces, and the firmware update submodule controls the output voltage of the graphics card power supply module to remain stable during the flashing of the graphics card vBIOS.

8. The PC Farm array server management and control device according to claim 1, characterized in that, The intelligent monitoring and diagnostic module includes a sensor data acquisition unit, a threshold judgment unit, a log recording unit, an alarm processing unit, and an action execution unit; the sensor data acquisition unit connects to I... 2 The C interface collects data on the CPU temperature, GPU temperature, power supply voltage, and fan speed of the blade node. The fans are two PWM-controlled 12000RPM fans, one on each side of the GPU PCIE Gen4x16 adapter board. The threshold judgment unit compares the collected data with preset thresholds. If the data meets the preset threshold range, the log recording unit records the current hardware status log. If the data exceeds the preset threshold range, the alarm processing unit triggers an alarm, and the action execution unit performs automatic frequency reduction or shutdown of the blade node.

9. The PC Farm array server management and control device according to claim 1, characterized in that, The intelligent monitoring and diagnostic module also includes a fault code library, which contains various fault modes such as heat dissipation failure mode and power instability mode. The intelligent monitoring and diagnostic module can match the collected temperature curve change trend and voltage fluctuation data with the fault modes in the fault code library to determine the fault type, and generate corresponding maintenance suggestions based on the fault type.

10. The PC Farm array server management and control device according to any one of claims 1-9, characterized in that, During the initialization phase of the management and control module, the embedded BMC control module assigns a unique logical ID to each blade node, with the logical ID ranging from 1 to 8. At the same time, it establishes a mapping table between the logical ID and the physical port of the blade node. After the out-of-band management module completes the system deployment, the batch deployment submodule obtains the installation interface image of each blade node through the KVM-over-IP submodule, uses OCR recognition technology to identify keywords in the installation interface to verify the system deployment results, and automatically triggers a retry deployment operation for blade nodes that fail to deploy.