Whole cabinet server management system, method, device and medium
The whole-cabinet server management system, which operates collaboratively through a wireless backplane bus and modules, solves the problems of low startup efficiency, cumbersome driver installation, and difficult power supply line fault detection in traditional server management. It achieves fast startup, automated driver installation, and power supply anomaly warning, improving operation and maintenance efficiency and reliability.
Patent Information
- Application Number
- CN202510999433.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Traditional server management methods have problems such as low startup efficiency, cumbersome hardware driver installation, difficult power supply line fault detection, and energy waste, resulting in low overall operation and maintenance efficiency.
It uses the coordinated operation of a wireless backplane bus, a programmable node midplane module, and a power management module to achieve rapid node startup and status monitoring through the wireless backplane bus, automatically read hardware device identification information and match drivers, detect power supply line impedance in real time and trigger alarms, and adopt a preset power consumption power supply strategy.
It achieves rapid startup and real-time status monitoring of server nodes, reduces time-consuming manual inspections, automates driver installation, detects power supply line anomalies in advance, reduces downtime risks, optimizes energy consumption management, and improves operation and maintenance efficiency and reliability.
Smart Images

Figure CN120508326B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a whole-cabinet server management system, method, device and medium. Background Art
[0002] With the development of cloud computing, artificial intelligence, and big data, rack-mount servers are widely used in data centers due to their high density and modularity. However, traditional server management methods have several issues. For example, server node startup relies on physical wired connections, which is cumbersome and prone to line connection failures, resulting in inefficient startup. Hardware device driver installation often requires manual matching, which is time-consuming and labor-intensive. Potential faults such as abnormal line impedance cannot be detected promptly, often leading to server downtime due to power supply issues, resulting in wasted energy or insufficient power. These issues lead to low overall server operation and maintenance efficiency, making it difficult to meet the current demand for efficient and stable server management. Summary of the Invention
[0003] The present invention provides a whole-cabinet server management system, method, device and medium, which at least solve the problem of low overall operation and maintenance efficiency of servers in the related art.
[0004] The present invention provides a whole cabinet server management system, comprising:
[0005] The management hub module is used to send a startup signal to the programmable node midboard module via the wireless backplane bus; it is also used to receive an interrupt signal and obtain the hardware device identification information transmitted by the programmable node midboard, and then filter out the corresponding driver according to the hardware device identification information to complete the driver installation;
[0006] The programmable node midboard module is used to indicate its own status by responding to a status code after receiving the startup signal; it is also used to monitor the hot plug signal, trigger the interrupt signal after detecting the hot plug action, and read the hardware device identification information stored in the node memory at the same time;
[0007] The power management module is used to detect the power supply line of each node when the status code of the board module in the programmable node is a preset value, and trigger an alarm if the line impedance value exceeds the set impedance threshold; it is also used to initialize the power supply strategy for each node according to the preset power consumption after the driver is installed.
[0008] The present invention also provides a whole cabinet server management method, comprising:
[0009] The management hub module sends a start signal to the programmable node midplane module via the wireless backplane bus;
[0010] After receiving the start signal, the programmable node midboard module indicates its own status through a response status code;
[0011] When the status code of the board module in the programmable node is a preset value, the power management module detects the power supply line of each node and triggers an alarm if the line impedance value exceeds the set impedance threshold;
[0012] The programmable node midboard module monitors the hot plug signal, triggers an interrupt signal after detecting the hot plug action, and simultaneously reads the hardware device identification information stored in the node memory;
[0013] After receiving the interrupt signal and obtaining the hardware device identification information transmitted by the programmable node midboard, the management center module filters out the corresponding driver according to the hardware device identification information to complete the driver installation;
[0014] After the driver is installed, the power management module initializes a power supply strategy for each node according to the preset power consumption.
[0015] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned whole-rack server methods when executing the computer program.
[0016] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned whole-cabinet server management methods are implemented.
[0017] In the present invention, the whole cabinet server management system can significantly improve the operation and maintenance efficiency through the coordinated operation of the management center module, the programmable node midplane module and the power management module; among them, through the wireless backplane bus startup and status code feedback mechanism, the node can be quickly started and the status is monitored in real time, reducing the time spent on manual inspections; hot-swap monitoring and interrupt signal triggering can automatically read the hardware device identification information and match the driver to complete the installation, avoiding the tedious process of manually searching for the driver and shortening the fault handling time; the power management module's real-time detection and alarm of the power supply line impedance can detect line abnormalities in advance and reduce the risk of downtime due to power supply failures. At the same time, the preset power consumption power supply strategy is initialized to optimize energy consumption management while ensuring equipment operation. In addition, the modules work together automatically to greatly reduce manual intervention, effectively improving the convenience, reliability and efficiency of server cluster operation and maintenance.
[0018] In addition, the present invention also provides a corresponding whole-cabinet server management method, electronic device and computer-readable storage medium for the whole-cabinet server management system, which has the same or corresponding technical features as the above-mentioned whole-cabinet server management system and has the same effect as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A schematic diagram of the structure of a whole-cabinet server management system provided by an embodiment of the present invention;
[0021] Figure 2 A schematic diagram of a hierarchical framework of a whole-cabinet server management system provided by an embodiment of the present invention;
[0022] Figure 3 This is a flowchart of a method for managing a whole-cabinet server provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0024] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0026] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the whole cabinet server management system depends, the specific application environment architecture or specific hardware architecture is described here.
[0027] An embodiment of the present invention provides a whole-rack server management system. The system is described in detail in conjunction with the structure of the whole-rack server management system. Figure 1 A schematic diagram of the structure of the whole cabinet server management system provided by the embodiment of the present invention is shown as follows: Figure 1As shown, the system includes:
[0028] Management hub module 1 is used to send a startup signal to programmable node board (PNB) module 2 via a wireless backplane bus (WBB). It is also used to receive interrupt signals and obtain hardware device identification information transmitted by the programmable node board, and then filter the corresponding driver based on the hardware device identification information to complete the driver installation. The management hub module 1 here can use an intelligent management center (IMC).
[0029] Programmable node midplane module 2 is used to indicate its own status by responding to a status code after receiving a startup signal. It is also used to monitor hot-plug signals and trigger an interrupt signal after detecting a hot-plug action. It also reads the hardware device identification information stored in the node memory.
[0030] Power management module 3 is used to monitor the power supply lines of each node when the status code of the programmable node midplane module 2 reaches a preset value. If the line impedance exceeds the preset impedance threshold, an alarm is triggered. It is also used to initialize the power supply strategy for each node according to the preset power consumption after the driver is installed. Power management module 3 can be an intelligent power management unit (IPU).
[0031] In the above-mentioned whole-cabinet server management system provided by the embodiment of the present invention, the whole-cabinet server management system can significantly improve the operation and maintenance efficiency through the coordinated operation of the management center module 1, the programmable node midplane module 2 and the power management module 3; among them, through the wireless backplane bus startup and status code feedback mechanism, the node can be quickly started and the status is monitored in real time, reducing the time spent on manual inspections; hot-swap monitoring and interrupt signal triggering can automatically read the hardware device identification information and match the driver to complete the installation, avoiding the tedious process of manually searching for the driver and shortening the fault handling time; the power management module 3 can detect and warn the power supply line impedance in real time, detect line abnormalities in advance, reduce the risk of downtime due to power supply failures, and at the same time initialize the power supply strategy with preset power consumption, optimizing energy consumption management while ensuring equipment operation. In addition, the modules work together automatically, greatly reducing manual intervention, and effectively improving the convenience, reliability and efficiency of server cluster operation and maintenance.
[0032] It should be noted that the whole cabinet server is a server form designed for efficient deployment requirements of data centers. In the present invention, the management hub module 1, programmable node midplane module 2 and power management module 3 are integrated into the whole cabinet server to form a whole cabinet server management system.
[0033] The functions of the management hub module 1 can include global resource scheduling, health monitoring, and disaster recovery switching. It can use a low-power Advanced RISC Machine System on a Chip (ARMSoC) to run lightweight Kubernetes (an open source container orchestration platform), manage containerized control services, and broadcast heartbeat signals through a wireless backplane bus with a delay of <1ms.
[0034] The programmable node midplane module 2 features node abstraction, protocol conversion, and local policy enforcement. It utilizes a field-programmable gate array (FPGA) for hardware logic programmability, supports plug-and-play for heterogeneous nodes including central processing units (CPUs), graphics processing units (GPUs), and data processing units (DPUs), and virtualizes a standard management interface for each node, such as the Ethernet-based Intelligent Platform Management Interface (IPMI-over-Ethernet). The programmable node midplane module 2 connects nodes via a peripheral component interconnect express (PCIe) switch and supports hot-swap detection. It also includes a built-in embedded multi-media card (eMMC) to store node configuration templates. For example, it automatically loads the Compute Unified Device Architecture (CUDA) management firmware when a GPU is detected.
[0035] The functions of Power Management Module 3 can include real-time power consumption monitoring, dynamic voltage regulation, and fault isolation. Power Management Module 3 can use a long short-term memory (LSTM) model to predict load in future time periods (e.g., the next 5 minutes) and dynamically switch between 12V and 48V power supply modes, improving efficiency by 15%. Power Management Module 3 can use a digital power supply, automatically bypassing the power supply in the event of a single-line fault to ensure uninterrupted critical services. Each node can be configured with an independent Power Management Integrated Circuit (PMIC) with a sampling accuracy of ±0.5%. Power Management Module 3 can communicate with Management Hub Module 1 via the Controller Area Network (CAN) bus to report real-time power consumption data.
[0036] A wireless backplane bus is an architecture that replaces the internal backplane bus in servers, which is connected by physical cables, with wireless communication. This bus eliminates the limitations of physical cables, improving system flexibility, scalability, and maintenance efficiency. This invention uses a wireless backplane bus to connect modules, achieving low-latency interconnection. This wireless backplane bus can be a high-speed wireless backplane bus. In this invention, the wireless backplane bus can utilize a hybrid communication protocol. Critical data is transmitted via 60GHz millimeter waves with a latency of less than 50μs, while bulk transmission utilizes visible light communication (LiFi) to achieve 10Gbps bandwidth. A self-healing topology is constructed based on the Ad-hoc On-Demand Distance Vector Routing (AODV) protocol for wireless ad hoc networks, automatically switching paths in the event of a connection loss. Each programmable node midplane deploys at least two transceivers to ensure full-duplex communication, with dynamically adjustable signal strength and power consumption of less than 3W per node. The transceivers can include both radio frequency (RF) and optical transceivers.
[0037] During the system initialization process, a hardware power-on self-test (POST) can be performed first: Management Hub Module 1 broadcasts a startup signal via the wireless backplane bus, and Programmable Node Midplane Module 2 responds with a status code, such as 0xAA, indicating normal operation. Simultaneously, Power Management Module 3 detects the impedance of the node power supply line and generates an alarm. For example, the impedance threshold can be set to 5Ω; when the line impedance exceeds 5Ω, an alarm is triggered. Next, the management network is established: Management Hub Module 1 (the master management unit of Management Hub Module 1 can be used here) assigns a logical Internet Protocol Address (IP) to the management device via DHCPv6 (Dynamic Host Configuration Protocol version 6). Programmable Node Midplane Module 2 loads the Field-Programmable Gate Array (FPGA) bitstream. Finally, resource topology discovery is performed: Management Hub Module 1 uses the Link Layer Discovery Protocol (LLDP) to discover the hardware topology. In this way, wireless broadcast startup and status code response can achieve rapid self-test and initial fault screening, improving startup efficiency; power supply line impedance detection provides early warning of risks and reduces the probability of downtime; DHCPv6 dynamically allocates IP and FPGA configuration loading, simplifying network deployment and device initialization processes; the LLDP protocol automatically discovers hardware topology, allowing operation and maintenance personnel to intuitively understand device connection relationships and reduce manual sorting costs, thereby comprehensively improving the automation, reliability and operation and maintenance efficiency of server management.
[0038] During the node hot-swap management process, when an insertion event occurs, the programmable node midplane module 2 detects changes in the Peripheral Component Interconnect Express (PCIe) hot-swap signal and triggers an interrupt to the management hub module 1. The module then reads vendor information, such as the device IP address, from the Electrically Erasable Programmable Read-Only Memory (EEPROM). Based on this information, the management hub module 1 pulls the corresponding driver from the image repository (cloud). Meanwhile, the power management module 3 initializes the power supply strategy based on a preset thermal design power (TDP), such as 400W for a graphics processor. When an unplug event occurs, the management hub module 1 uses checkpointing to migrate node tasks to adjacent devices. The intelligent power management unit quickly cuts off power within a set time (e.g., 10ms) to ensure safety. In this way, through automated monitoring and response, automatic driver installation and precise matching of power supply and cooling strategies can be achieved when hardware is inserted, greatly reducing manual intervention and improving operation and maintenance efficiency; safe unloading and rapid power off of unplugging events can avoid data loss and hardware damage risks, ensuring business continuity. At the same time, task migration technology ensures system load balancing and optimizes resource utilization, thereby enhancing the reliability and intelligence level of server cluster operation.
[0039] Furthermore, in a specific implementation, the above-mentioned whole-cabinet server management system provided in an embodiment of the present invention may also include: an adaptive cooling module (ACM) module, which is used to collect temperature information of hardware equipment within a first set interval time period and obtain ambient humidity temperature data at the same time; transmit the collected temperature information and the obtained ambient humidity temperature data to a long short-term memory (LSTM) model, and output a corresponding temperature heat map to predict the temperature change trend of the hardware equipment; if the predicted local temperature of the hardware equipment exceeds the first set temperature, increase the speed of the corresponding fan and increase the flow rate of the liquid cooling pump; if the predicted overall temperature of the hardware equipment is lower than the second set temperature, switch to silent mode and reduce the fan speed.
[0040] In practice, the adaptive cooling module can be assigned independent temperature control strategies. Within the dynamic cooling control process, precise temperature control can be achieved by building a closed loop. First, data such as the diode temperature (accuracy ±1°C), ambient temperature and humidity (using SHT30 sensors), and liquid cooling circuit flow rate (0-5L / min) are collected at each node (such as the CPU / GPU) every 10 seconds. An LSTM model is then used to infer 120 seconds of historical data to predict thermal field changes in the future (e.g., the next 30 seconds). Finally, hierarchical control is implemented based on the predicted temperature. When the temperature is below 50°C, the fan enters silent mode (PWM ≤ 30%). Fan speed is linearly adjusted (PWM = 30-70%) between 50°C and 80°C. Above 80°C, the liquid cooling pump is activated at full speed and the load is shifted. For example, the corresponding fan speed is increased to 70%, and the liquid cooling pump flow rate is increased by 15%. The high-frequency collection of such data and the precise prediction of the LSTM model can sense cooling needs in advance and avoid sudden temperature rises; the hierarchical control strategy realizes the refined allocation of cooling resources, achieving a balance between silent energy saving (reducing fan noise and power consumption at low loads) and efficient heat dissipation (ensuring equipment stability at high loads), extending the service life of hardware. At the same time, the load migration mechanism ensures that the equipment can still maintain business continuity under extremely high temperatures, effectively improving the reliability and energy utilization of server cluster operations.
[0041] It should be noted that the adaptive cooling module's functions include temperature monitoring, fan / liquid cooling control, and hotspot prediction. It can also be used to construct a three-dimensional thermal field model. Using a deep action-value network algorithm, trained based on historical temperature, load, and ambient humidity data, it adjusts the pulse-width modulation fan speed and the operating status of the liquid cooling pump for intelligent control of the cooling equipment. Each cooling unit is equipped with multiple pulse-width modulation fans and a micro-liquid cooling pump as the hardware foundation.
[0042] In implementation, the adaptive cooling module combines infrared thermal imaging (FLIR Lepton) with onboard temperature sensors to construct a 3D thermal field model. Leveraging the Deep Q-Network (Q stands for action value) algorithm from reinforcement learning control, it dynamically adjusts the PWM parameters of multiple pulse-width modulation (PWM) fans and a micro liquid cooling pump (Alphacool DC-LT). Using historical temperature, load, and ambient humidity (e.g., a 10s sampling period) as training data, it achieves a 20% noise reduction optimization effect. This approach, which fuses multi-source data to construct a 3D thermal field model, accurately locates device hotspots. Combined with dynamic policy adjustments based on the reinforcement learning algorithm, this approach ensures cooling efficiency while allocating cooling resources on demand, avoiding the noise pollution and energy waste associated with traditional fixed-speed cooling. This not only improves the quietness of server cluster operations, but also extends hardware life through intelligent temperature control, providing technical support for the efficient and green operation of high-density data centers.
[0043] Figure 2 Schematic diagram of the hierarchical framework of the whole cabinet server management system provided by the embodiment of the present invention. Figure 2 As shown, the whole cabinet server management system of the present invention can adopt a modular layered design, including an intelligent management layer, a hardware control layer and a node execution layer, and realize low-latency interconnection through a wireless backplane bus. The intelligent management layer may include multiple management hub modules 1, which are used to perform intelligent management, global resource scheduling, health monitoring and disaster recovery switching, and coordinate the operation of each layer. The hardware control layer may include a programmable node midplane module 2, a power management module 3 and an adaptive heat dissipation module, which are used to connect the intelligent management layer and the node execution layer, receive and assign tasks, and send corresponding instructions to the node execution layer. The node execution layer is used to receive instructions from the hardware control layer and control the corresponding nodes (such as central processing units, graphics processing units, etc.) to perform corresponding tasks.
[0044] Furthermore, in a specific implementation, in the above-mentioned whole-cabinet server management system provided in an embodiment of the present invention, the management hub module 1 can also be used to allocate node resources to the programmable node midplane module 2 after receiving a task. The programmable node midplane module 2 can be used to send a request instruction to the adaptive heat dissipation module to obtain the temperature information fed back by the adaptive heat dissipation module. The management hub module 1 can also be used to send a liquid cooling boost signal to the adaptive heat dissipation module when the temperature information exceeds a set range. The adaptive heat dissipation module can also be used to adjust the flow of the liquid cooling pump after receiving the liquid cooling boost signal.
[0045] During implementation, the coordinated operation of the management hub module 1, the programmable node midplane module 2, and the adaptive heat dissipation module can achieve efficient linkage in task allocation and heat dissipation control: after receiving the task, the management hub module 1 accurately allocates node resources, and at the same time obtains the temperature information of the heat dissipation module through the programmable node midplane module 2. When the temperature exceeds the limit, it immediately sends a liquid cooling boost signal to drive the adaptive heat dissipation module to adjust the liquid cooling pump flow, forming a closed-loop management of task allocation, temperature monitoring, and dynamic heat dissipation. In this way, through the linkage of resource allocation and heat dissipation control, abnormal temperatures of nodes due to excessive load are avoided, and stable operation of hardware is guaranteed; real-time monitoring of temperature information and dynamic adjustment of liquid cooling flow can allocate heat dissipation resources on demand, timely cool down under high load, reduce energy consumption under low load, improve heat dissipation efficiency and reduce noise; modular collaboration realizes automated temperature control, reduces manual intervention, and significantly improves the operation and maintenance efficiency and reliability of server clusters, especially suitable for thermal management needs in high-density computing scenarios.
[0046] Furthermore, in a specific implementation, in the above-mentioned whole-cabinet server management system provided in an embodiment of the present invention, the management center module 1 can adopt a distributed redundant architecture, including a master management unit and at least one backup management unit. The master management unit can be used to broadcast a heartbeat signal via a wireless backplane bus within a second set interval time period; the backup management unit can be used to continuously monitor the heartbeat signal. If the heartbeat signal broadcast by the master management unit is not received within the preset time period, it is determined that the master management unit has failed, triggering an election mechanism to elect a new master management unit; the new master management unit can be used to obtain the latest resource topology information from a set storage location and synchronize the unfinished task queue according to the persistent log.
[0047] In practice, the management hub module 1 can be clustered using multiple (e.g., three) nodes. Master election is implemented using the Replicated Atomic File Transfer (RAFT) protocol, with automatic failover in the event of a node failure. Management hub module 1 supports dynamic loading of management policies (e.g., energy optimization, fault prediction) through the application programming interface (API).
[0048] During the fault detection process, the heartbeat packet interval (i.e., the preset time period) can be set to 100ms (with a timeout threshold of 500ms), and the node state machine of the management hub module can be implemented based on the RAFT algorithm. The RAFT algorithm code is as follows:
[0049] .
[0050] As can be seen from the code, when the node is in FOLLOWER state and the electionTimeout is triggered, an election is initiated. Upon becoming a LEADER, a heartbeat signal is broadcast. During the switchover process, the new master management unit reads the latest topology version from Zookeeper (an open source distributed coordination service framework), synchronizes the unfinished task queue with Kafka (an open source distributed stream processing platform), and sends a master switchover notification to all programmable node midplanes via the wireless backplane bus using a multicast address. In this way, a high-availability cluster is built through the RAFT algorithm and heartbeat detection. The 500ms timeout threshold ensures rapid identification of faulty nodes, and the 100ms heartbeat interval ensures real-time status monitoring, avoiding business interruption caused by single point failure and ensuring the disaster recovery capability of the server; the combination of Zookeeper and Kafka realizes reliable synchronization of topology version and task queue, ensuring that data is not lost and tasks are not interrupted during master-slave switching; the multicast notification mechanism quickly synchronizes the switching status through the wireless backplane bus, reduces node response delay, improves the real-time and consistency of cluster management, and finally realizes the automatic failover of the management center module 1, ensuring the high availability and operation and maintenance efficiency of the server cluster.
[0051] Furthermore, in a specific implementation, in the above-mentioned whole-cabinet server management system provided by an embodiment of the present invention, the new master control management unit can also be used to send a master control switch notification signal to the programmable node midplane module 2 via the wireless backplane bus. Upon receiving the master control switch notification signal, the programmable node midplane module 2 can also be used to establish a connection with the new master control management unit and re-establish the communication link; at the same time, it can resume the task that was interrupted due to a failure of the master control management unit.
[0052] During implementation, the new master control management unit sends a master control switch notification signal to the programmable node midplane module 2 via the wireless backplane bus. After receiving the signal, the node midplane establishes a connection with the new master control, reconstructs the communication link, and resumes the task that was interrupted due to the failure of the original master control. In this way, the high-speed transmission capability of the wireless backplane bus ensures the low-latency delivery of the master control switch notification, reduces the node response time, and guarantees the real-time performance of cluster management instructions; the communication link reconstruction mechanism between the node midplane and the new master control can quickly restore the control channel of the distributed system and avoid interruption of the management plane; the task resumption function is based on the state before the failure and can directly resume execution from the breakpoint, avoiding the waste of resources and time loss caused by starting the task from the beginning, significantly improving the system's fault tolerance, and realizing the overall automatic takeover and seamless task migration in the event of a master control failure, ensuring the business continuity of the server cluster and reducing the cost of manual intervention.
[0053] Furthermore, in the specific implementation of the whole-cabinet server management system provided in the embodiment of the present invention, the management hub module 1 can also be used to scan and identify hardware devices, sort out the connection relationships and layout structures between hardware devices, and construct a hardware device topology map to provide a basis for system resource management and scheduling. The programmable node midplane module 2 can also be used to load pre-set configuration files, use FPGAs to program and configure the hardware device logic circuits, and connect hardware devices via PCIe switches.
[0054] In practice, the management hub module 1 scans and identifies hardware devices, organizing the connections and layout between them to construct a topology map, providing a basis for system resource management and scheduling. The programmable node midplane module 2 loads preset configuration files, programs and configures the hardware device logic circuits using FPGAs, and connects the hardware devices via PCIe switches. This enables intelligent management and efficient utilization of hardware resources, laying the foundation for flexible expansion and refined operations and maintenance in complex computing scenarios.
[0055] Furthermore, in specific implementation, in the above-mentioned whole-cabinet server management system provided in the embodiment of the present invention, the power management module 3 can also be used to predict the load situation based on the historical data of processor utilization and the task queue length using the Autoregressive Integrated Moving Average Model (ARIMA), and switch the power supply mode according to the load prediction result; when the load prediction result is lower than the first set value, the set power consumption power supply mode is started; when the load prediction result is between the first set value and the second set value, the balanced power supply mode is started to maintain the baseline energy efficiency state; when the load prediction result exceeds the second set value, the main power supply combined with liquid cooling and boosting power supply mode is started. The power management module 3 can also be used to report the power consumption data to the management center module 1 so that the management center module 1 can perform global resource scheduling and management.
[0056] In implementation, Power Management Module 3 uses an ARIMA model to predict computing demand for future time periods (e.g., the next 5 minutes). Input features include historical CPU utilization data and task queue length. Voltage regulation implements a tiered power supply strategy based on load levels. When the load is below 30%, a 12V low-power mode is activated, improving energy efficiency by 18%. When the load is between 30% and 70%, a 48V balanced mode is used as the baseline. When the load exceeds 70%, a liquid cooling strategy combining a 48V main power supply with a 12V auxiliary power supply is used, improving energy efficiency by 9%. This precise ARIMA model-based prediction of computing demand allows for early detection of load fluctuations, providing data support for voltage regulation and preventing power supply strategies from lagging behind actual demand. The tiered power supply strategy achieves a dynamic balance between energy efficiency and performance. The 12V mode reduces unnecessary power consumption during low loads, while liquid cooling ensures stable operation during high loads. This significantly improves overall energy efficiency compared to a fixed voltage mode. This integrated load prediction and voltage regulation mechanism not only reduces data center electricity costs, but also reduces hardware heat generation and extends equipment life through on-demand power supply, providing technical support for green computing and efficient operations and maintenance.
[0057] It's important to note that the whole-rack server management system of this invention can be expanded to multiple fields, including large-scale data centers and cloud computing, edge computing and edge nodes, high-performance computing (HPC) and supercomputing centers, telecommunications core networks, and network function virtualization (NFV) infrastructure. This not only improves the intelligent whole-rack server management system and method, but also provides a flexible and reliable solution for various industries.
[0058] An embodiment of the present invention also provides a whole-cabinet server management method. Figure 3 Flowchart of the whole cabinet server management method provided by the embodiment of the present invention. Figure 3 As shown, the method includes:
[0059] S301. The management center module sends a start signal to the programmable node midplane module via the wireless backplane bus.
[0060] S302. After receiving the start signal, the programmable node midboard module indicates its own status through a response status code.
[0061] S303. When the status code of the board module in the programmable node is a preset value, the power management module detects the power supply line of each node, and triggers an alarm if the line impedance value exceeds the set impedance threshold.
[0062] S304: The programmable node midboard module monitors the hot-swap signal, triggers an interrupt signal after detecting the hot-swap action, and simultaneously reads the hardware device identification information stored in the node memory.
[0063] S305 , after receiving the interrupt signal and obtaining the hardware device identification information transmitted by the programmable node midboard, the management central module filters out the corresponding driver according to the hardware device identification information to complete the driver installation.
[0064] S306: After the driver is installed, the power management module initializes a power supply strategy for each node according to the preset power consumption.
[0065] In the above-mentioned whole-cabinet server management method provided in the embodiment of the present invention, the operation and maintenance efficiency can be significantly improved through the coordinated operation of the management center module, the programmable node midplane module and the power management module; among them, through the wireless backplane bus startup and status code feedback mechanism, the node can be quickly started and the status is monitored in real time, reducing the time spent on manual inspections; hot-swap monitoring and interrupt signal triggering can automatically read the hardware device identification information and match the driver to complete the installation, avoiding the tedious process of manually searching for the driver and shortening the fault handling time; the power management module's real-time detection and alarm of the power supply line impedance can detect line abnormalities in advance and reduce the risk of downtime due to power supply failures. At the same time, the power supply strategy with preset power consumption is initialized to optimize energy consumption management while ensuring equipment operation. In addition, the modules automatically work together to greatly reduce manual intervention, effectively improving the convenience, reliability and efficiency of server cluster operation and maintenance.
[0066] Since the embodiments of the whole-rack server management method correspond to the embodiments of the whole-rack server management system, the description of the features of the corresponding embodiments of the whole-rack server management method can be found in the description of the corresponding embodiments of the whole-rack server management system, and will not be repeated here. The embodiments of the whole-rack server management method have the same beneficial effects as the above-mentioned whole-rack server management system.
[0067] Furthermore, in specific implementation, the above-mentioned whole-cabinet server management method provided in the embodiment of the present invention may also include: an adaptive heat dissipation module collects temperature information of the hardware device within a first set interval time period, and obtains ambient humidity temperature data at the same time; transmits the collected temperature information and the obtained ambient humidity temperature data to the long-short-term memory model, and outputs a corresponding temperature heat map to predict the temperature change trend of the hardware device; if the predicted local temperature of the hardware device exceeds the first set temperature, the speed of the corresponding fan is increased, and the flow rate of the liquid cooling pump is increased; if the predicted overall temperature of the hardware device is lower than the second set temperature, it is switched to silent mode and the fan speed is reduced.
[0068] Furthermore, in specific implementation, the above-mentioned whole-cabinet server management method provided in the embodiment of the present invention may also include: after the management central module receives the task, it allocates node resources to the programmable node midplane module; the programmable node midplane module sends a request instruction to the adaptive heat dissipation module to obtain the temperature information fed back by the adaptive heat dissipation module; when the temperature information exceeds the set range, the management central module sends a liquid cooling boost signal to the adaptive heat dissipation module; after the adaptive heat dissipation module receives the liquid cooling boost signal, it adjusts the flow rate of the liquid cooling pump.
[0069] Furthermore, in specific implementation, in the above-mentioned whole-cabinet server management method provided in an embodiment of the present invention, when the management center module adopts a distributed redundant architecture, including a master management unit and at least one backup management unit, it can also include: the master management unit broadcasts a heartbeat signal through the wireless backplane bus within a second set interval time period; the backup management unit continuously monitors the heartbeat signal, and if it does not receive the heartbeat signal broadcast by the master management unit within the preset time period, it determines that the master management unit has failed, triggers the election mechanism, and elects a new master management unit; the new master management unit obtains the latest resource topology information from the set storage location, and synchronizes the unfinished task queue according to the persistent log.
[0070] Furthermore, in specific implementation, the above-mentioned whole-cabinet server management method provided in the embodiment of the present invention may also include: the new master control management unit sends a master control switching notification signal to the programmable node midplane module through the wireless backplane bus; after receiving the master control switching notification signal, the programmable node midplane module establishes a connection with the new master control management unit and re-establishes the communication link; at the same time, for tasks that are interrupted due to a failure of the master control management unit, the task is resumed.
[0071] Furthermore, in specific implementation, the above-mentioned whole-cabinet server management method provided in the embodiment of the present invention may also include: the management center module scans and identifies the hardware devices, sorts out the connection relationship and layout structure between the hardware devices, and constructs a hardware device topology map to provide a basis for system resource management and scheduling; the programmable node midboard module loads a pre-set configuration file, uses a field programmable gate array to program and configure the hardware device logic circuit, and connects the hardware devices through a peripheral component interconnection switch.
[0072] Furthermore, in specific implementation, the above-mentioned whole-cabinet server management method provided in the embodiment of the present invention may also include: the power management module predicts the load situation based on the historical data of processor utilization and the task queue length using the autoregressive integral moving average model, and switches the power supply mode according to the load prediction result; when the load prediction result is lower than the first set value, the set power consumption power supply mode is started; when the load prediction result is between the first set value and the second set value, the balanced power supply mode is started to maintain the benchmark energy efficiency state; when the load prediction result exceeds the second set value, the main power supply combined with liquid cooling and boosting power supply mode is started; the power management module reports the power consumption data to the management center module so that the management center module can perform global resource scheduling and management.
[0073] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0074] An embodiment of the present invention further provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned whole-cabinet server management method embodiments.
[0075] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned whole-cabinet server management method embodiments when running.
[0076] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0077] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned whole-cabinet server management method embodiments are implemented.
[0078] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned whole-cabinet server management method embodiments.
[0079] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0080] The above is a detailed introduction to the whole cabinet server management system, method, device and medium provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A whole cabinet server management system, characterized in that: include: The management hub module is used to send a startup signal to the programmable node midplane module via the wireless backplane bus; It is also used to receive an interrupt signal and obtain the hardware device identification information transmitted by the programmable node midboard, and then filter out the corresponding driver according to the hardware device identification information to complete the driver installation; the management center module adopts a distributed redundant architecture, including a main control management unit and at least one backup management unit; the main control management unit is used to broadcast a heartbeat signal through a wireless backplane bus within a second set interval time period; the backup management unit is used to continuously monitor the heartbeat signal, and if the heartbeat signal broadcast by the main control management unit is not received within the preset time period, it is determined that the main control management unit has failed, triggering an election mechanism to elect a new main control management unit; the new main control management unit is used to obtain the latest resource topology information from a set storage location, and synchronize the unfinished task queue according to the persistent log; The programmable node midboard module is used to indicate its own status by responding to a status code after receiving the startup signal; it is also used to monitor the hot plug signal, trigger the interrupt signal after detecting the hot plug action, and read the hardware device identification information stored in the node memory at the same time; A power management module is configured to detect the power supply line of each node when the status code of the board module in the programmable node is a preset value, and trigger an alarm if the line impedance value exceeds a set impedance threshold; It is also used to initialize the power supply strategy for each node according to the preset power consumption after the driver is installed.
2. The whole cabinet server management system according to claim 1, characterized in that: Also includes: The adaptive heat dissipation module is used to collect temperature information of the hardware device within a first set interval time period and obtain ambient humidity and temperature data at the same time; transmit the collected temperature information and the obtained ambient humidity and temperature data to the long-short-term memory model, and output the corresponding temperature heat map to predict the temperature change trend of the hardware device; if the predicted local temperature of the hardware device exceeds the first set temperature, increase the speed of the corresponding fan and increase the flow rate of the liquid cooling pump; if the predicted overall temperature of the hardware device is lower than the second set temperature, switch to silent mode and reduce the fan speed.
3. The whole cabinet server management system according to claim 2, characterized in that: The management hub module is further configured to allocate node resources to the programmable node midboard module after receiving a task; The programmable node midboard module is used to send a request instruction to the adaptive heat dissipation module to obtain temperature information fed back by the adaptive heat dissipation module; The management center module is further configured to send a liquid cooling boost signal to the adaptive heat dissipation module when the temperature information exceeds a set range; The adaptive heat dissipation module is further configured to adjust the flow rate of the liquid cooling pump after receiving the liquid cooling boost signal.
4. The whole cabinet server management system according to claim 1, characterized in that: The new master control management unit is further configured to send a master control switching notification signal to the programmable node midplane module via a wireless backplane bus; The programmable node midboard module is also used to establish a connection with the new master control management unit and re-establish a communication link after receiving the master control switching notification signal; at the same time, it resumes the task that was interrupted due to a failure of the master control management unit.
5. The whole cabinet server management system according to claim 1, characterized in that: The management center module is also used to scan and identify hardware devices, sort out the connection relationship and layout structure between hardware devices, and build a hardware device topology map to provide a basis for system resource management and scheduling; The programmable node midboard module is also used to load pre-set configuration files, use field programmable gate arrays to program and configure hardware device logic circuits, and connect hardware devices through peripheral component interconnection switches.
6. The whole cabinet server management system according to claim 1, characterized in that: The power management module is further configured to predict the load condition using an autoregressive integral moving average model based on historical processor utilization data and task queue length, and switch the power supply mode according to the load prediction result; when the load prediction result is lower than a first set value, start the set power consumption power supply mode; when the load prediction result is between the first set value and a second set value, start the balanced power supply mode to maintain a baseline energy efficiency state; when the load prediction result exceeds the second set value, start the main power supply with liquid cooling and boosting power supply mode; The power management module is further configured to report power consumption data to the management hub module so that the management hub module can perform global resource scheduling and management.
7. A whole cabinet server management method, characterized in that: include: The management hub module sends a start signal to the programmable node midplane module via the wireless backplane bus; When the management center module adopts a distributed redundant architecture, including a main control management unit and at least one backup management unit, the main control management unit broadcasts a heartbeat signal through the wireless backplane bus within a second set interval time period; the backup management unit continuously monitors the heartbeat signal, and if it does not receive the heartbeat signal broadcast by the main control management unit within the preset time period, it is determined that the main control management unit has failed, triggering an election mechanism to elect a new main control management unit; the new main control management unit obtains the latest resource topology information from a set storage location, and synchronizes the unfinished task queue according to the persistent log; After receiving the start signal, the programmable node midboard module indicates its own status through a response status code; When the status code of the board module in the programmable node is a preset value, the power management module detects the power supply line of each node and triggers an alarm if the line impedance value exceeds the set impedance threshold; The programmable node midboard module monitors the hot plug signal, triggers an interrupt signal after detecting the hot plug action, and simultaneously reads the hardware device identification information stored in the node memory; After receiving the interrupt signal and obtaining the hardware device identification information transmitted by the programmable node midboard, the management center module filters out the corresponding driver according to the hardware device identification information to complete the driver installation; After the driver is installed, the power management module initializes a power supply strategy for each node according to the preset power consumption.
8. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the whole-cabinet server management method as claimed in claim 7 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the whole-cabinet server management method according to claim 7.
Citation Information
Patent Citations
Solid state disk hot plug management system and method
CN117591458A
Server fault positioning system and method
CN120256224A