A server management method, device, equipment and readable storage medium

By using a data pre-synchronization mechanism between master and slave RMCs, the problem of data silos caused by master RMC failure is solved, achieving seamless switching and continuity of asset management, and improving system reliability and operational efficiency.

CN122195770APending Publication Date: 2026-06-12XINHUASAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610150744.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In a multi-PowerShelf architecture, a failure of the primary RMC can lead to data silos, requiring operations and maintenance personnel to manually enter asset information, which increases operational complexity and the time when business systems are exposed to risks.

Method used

By using the data pre-synchronization mechanism between master and slave RMCs, and by automatically taking over asset management after migrating from the RMC to the master power frame, a seamless switchover can be achieved, avoiding data loss and reconstruction processes.

Benefits of technology

Shorten maintenance downtime, ensure continuity of thermal management and asset monitoring, improve system reliability and maintenance efficiency, and reduce reliance on manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195770A_ABST
    Figure CN122195770A_ABST
Patent Text Reader

Abstract

The present specification provides a server management method, device, equipment and readable storage medium, the method comprises: the autonomous RMC device obtains and stores the asset information associated with the main power frame, the asset information associated with the main power frame includes the management information associated with the heat dissipation management device; in response to the physical location migration event configured to the main power frame, call the stored asset information associated with the main power frame, the management information associated with the heat dissipation management device; according to the information obtained by calling, the assets of the main power frame and the heat dissipation management device are managed.Through the technical scheme of the present specification, the data pre-synchronization mechanism between master and slave RMCs, when the master RMC fails, the slave RMC can immediately take over the management of all assets including CDU after migrating to the main power frame, without waiting for spare parts or secondary input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of communication technology, and in particular to a server management method, apparatus, device, and readable storage medium. Background Technology

[0002] Compared to traditional rack servers, rack servers adopt a highly integrated and modular design concept, integrating computing nodes, power supply systems, heat dissipation devices, and network switching modules into a unified rack architecture. This results in significant advantages in computing performance, resource utilization, and energy efficiency, making them particularly suitable for high-load application scenarios such as training and inference of large-scale artificial intelligence models.

[0003] As the nerve center of the entire rack, the Rack Management Controller (RMC) is responsible for the core functions of real-time monitoring, data acquisition, and policy control of various hardware assets within the rack. The assets it manages are extensive, including intelligent components that can be automatically discovered via the bus, such as the serial number and operating status parameters of power supply units, and key operating data like temperature and flow rate of coolant distribution units; it also covers a large amount of related device information that requires manual input for the system to recognize, such as asset information of computing nodes that cannot directly establish a communication connection with the RMC. In earlier or simpler application models, the entire rack typically only required a single power supply frame to meet power needs, and only one RMC was used for centralized management. Under this single RMC architecture, all asset information, whether automatically collected or manually entered, was stored in the local storage chip of this single RMC, and its management logic was centralized and singular.

[0004] To meet the ever-increasing demands for power density and extremely high power reliability, rack-mount solutions employing multiple PowerShells (i.e., multiple power supply frames) are gradually becoming a new trend in high-end computing infrastructure. In this architecture, each PowerShell acts as an independent power supply unit, equipped with its own dedicated RMC (Power Management Control), aiming to achieve more granular power management and load balancing.

[0005] Under existing technical solutions, these Resource Management Controllers (RMCs) belonging to different PowerShell domains often operate in isolation. Each manages and stores asset data only within its own PowerShell domain, lacking effective data communication and synchronization mechanisms. This creates a data silo effect. When the primary RMC, which bears the main management responsibilities (e.g., communicating with critical infrastructure such as the Coolant Distribution Unit (CDU), experiences a hardware failure, not only will its management functions be momentarily interrupted, but more seriously, all asset data stored in its local non-volatile memory will be lost. Due to time delays in the procurement and transportation of spare parts, the replacement of the failed RMC cannot be completed immediately, leaving a significant maintenance gap.

[0006] Even if maintenance personnel arrive on-site with a spare RMC board, the replacement process itself involves significant data reconstruction costs. This is because the newly installed RMC has empty local storage, requiring maintenance personnel to manually re-enter all asset information previously managed by the faulty RMC, especially information on compute nodes that cannot be automatically rediscovered. This process is not only tedious, time-consuming, and error-prone, greatly increasing the complexity and manpower costs of maintenance, but also leaves the system's complete asset management functionality unavailable until the new data is entered, further extending the risk exposure time for the business system. Summary of the Invention

[0007] In view of this, this specification provides a server management method, apparatus, device, and readable storage medium to improve the problem of the above-mentioned main RMC failure severely affecting business operations.

[0008] The specific technical solution is as follows: This specification provides a server management method applied to an RMC device. The RMC device is connected to a peer RMC device, the local RMC device is a slave RMC device, and the peer RMC device is a master RMC device. The master RMC device and the slave RMC device are configured in a main power supply frame and a slave power supply frame, respectively. The master RMC device is connected to a thermal management device. The method includes: the autonomous RMC device acquiring and storing asset information associated with the main power supply frame, the asset information associated with the main power supply frame including management information associated with the thermal management device; responding to a physical location migration event configured to the main power supply frame, retrieving the stored asset information associated with the main power supply frame and the management information associated with the thermal management device; managing the assets and thermal management device of the main power supply frame according to the retrieved information; the physical location migration event is triggered by the physical removal of the master RMC device, including physically migrating the slave RMC device to the main power supply frame to replace the physically removed master RMC device.

[0009] As a technical solution, in response to a physical installation event in which a new RMC device is configured from a power supply frame, asset information associated with the main power supply frame is sent to the new RMC device; the physical installation event includes the physical installation and configuration of the new RMC device to the power supply frame where the RMC device was located before the physical location migration event occurred.

[0010] As a technical solution, the asset information includes the BMC information and PSU information of the associated power supply frame.

[0011] As a technical solution, the heat dissipation management device includes a coolant distribution unit.

[0012] This specification also provides a server management device applied to an RMC device. The RMC device has a communication connection with a peer RMC device. The local RMC device is a slave RMC device, and the peer RMC device is a master RMC device. The master RMC device and the slave RMC device are respectively configured in a main power supply frame and a slave power supply frame. The master RMC device is communicatively connected to a thermal management device. The device includes: a first module for the autonomous RMC device to acquire and store asset information associated with the main power supply frame, the asset information associated with the main power supply frame including management information associated with the thermal management device; a second module for responding to a physical location migration event configured to the main power supply frame, and retrieving the stored asset information associated with the main power supply frame and the management information associated with the thermal management device; and a third module for managing the assets and thermal management device of the main power supply frame according to the retrieved information. The physical location migration event is triggered by the physical removal of the master RMC device, including physically migrating the slave RMC device to the main power supply frame to replace the physically removed master RMC device.

[0013] As a technical solution, a fourth module is also included, which is used to send asset information associated with the main power supply frame to the new RMC device in response to a physical installation event in which a new RMC device is configured from the power supply frame; the physical installation event includes the new RMC device being physically installed and configured to the power supply frame where the RMC device was located before the physical location migration event occurred.

[0014] As a technical solution, the asset information includes the BMC information and PSU information of the associated power supply frame.

[0015] As a technical solution, the heat dissipation management device includes a coolant distribution unit.

[0016] This specification also provides an electronic device, including a processor and a readable storage medium storing machine-executable instructions that can be executed by the processor, which executes the machine-executable instructions to implement the aforementioned server management method.

[0017] This specification also provides a readable storage medium storing machine-executable instructions that, when invoked and executed by a processor, cause the processor to implement the aforementioned server management method.

[0018] The technical solutions provided in this specification offer at least the following beneficial effects: Through the data pre-synchronization mechanism between the master and slave RMCs, when the master RMC fails, the system can immediately take over the management of all assets, including CDUs, after migrating from the RMC to the main power supply frame. There is no need to wait for spare parts or re-entry, which significantly shortens the maintenance interruption time, achieves continuity of heat dissipation management and asset monitoring during the failure period, and improves the reliability and maintenance efficiency of the entire cabinet system. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings used in the description of the embodiments of this specification or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this specification.

[0020] Figure 1 This is a flowchart of a server management method according to one embodiment of this specification; Figure 2 This is a structural diagram of a server management device according to one embodiment of this specification; Figure 3 This is a hardware structure diagram of an electronic device according to one embodiment of this specification.

[0021] Figure 4 This is a schematic diagram of the architecture in one embodiment of this specification; Figure 5 This is a schematic diagram of the architecture in one embodiment of this specification.

[0022] Reference numerals: Module 1 21, Module 22, Module 3 23. Detailed Implementation

[0023] The terminology used in the embodiments described herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this specification. The singular forms “a,” “described,” and “the” as used in this specification and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.

[0024] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."

[0025] This specification provides a server management method, apparatus, device, and readable storage medium to at least improve one of the above-mentioned technical problems.

[0026] The specific technical solution is described below.

[0027] In one embodiment, this specification provides a server management method applied to an RMC device, wherein the RMC device is connected to a peer RMC device, the local RMC device is a slave RMC device, the peer RMC device is a master RMC device, the master RMC device and the slave RMC device are respectively configured in a main power supply frame and a slave power supply frame, and the master RMC device is connected to a thermal management device. The method includes: the autonomous RMC device acquiring and storing asset information associated with the main power supply frame, the asset information associated with the main power supply frame including management information associated with the thermal management device; in response to a physical location migration event configured to the main power supply frame, retrieving the stored asset information associated with the main power supply frame and the management information associated with the thermal management device; managing the assets and thermal management device of the main power supply frame according to the retrieved information; the physical location migration event is triggered by the physical removal of the master RMC device, including physically migrating the slave RMC device to the main power supply frame to replace the physically removed master RMC device.

[0028] By leveraging the communication connection and data redundancy mechanism between the master and slave RMCs in a PowerShell architecture, after the master RMC fails and is removed, the slave RMC is physically migrated to the master power supply chassis, automatically taking over all management responsibilities of the original master RMC. This achieves seamless continuity of asset management and thermal control functions. In practice, the entire rack typically has at least two power supply chassis—one master and one slave, each integrating an independent RMC module. The master RMC in the master power supply chassis not only manages the assets within its chassis (such as PSU power modules, sensors, fans, etc.) but also establishes a communication connection with the rack's thermal management equipment (such as CDU, Coolant Distribution Unit) via dedicated interfaces (such as I²C, UART, or Ethernet) to monitor key parameters such as coolant temperature, flow rate, and pump status, and dynamically adjusts the cooling strategy based on the compute node load. Meanwhile, although the slave RMC in the power supply box plays an auxiliary role under normal operating conditions, it is not idle. Instead, it continuously maintains data synchronization with the master RMC through a preset communication link (such as a point-to-point connection based on CAN bus or RJ45 Ethernet).

[0029] Specifically, such as Figure 1 This includes the following steps, the order of which can be changed depending on the needs of the actual application scenario: Step S11: The autonomous RMC device acquires and stores the asset information associated with the main power frame.

[0030] During system initialization or stable operation, the slave RMC will proactively send data requests to the master RMC, or the master RMC will periodically broadcast its complete asset information package. This asset information includes not only hardware identifiers that can be automatically collected within the main power supply chassis (such as the serial number, firmware version, input / output voltage and current of each PSU), but also related data that cannot be automatically discovered through standard protocols and must be manually entered by maintenance personnel during the deployment phase. For example, a specific power supply slot in the main power supply chassis where an AI acceleration computing node is physically plugged in, or a group of GPU servers that rely on the CDU connected to the main power supply chassis for liquid cooling. This type of information constitutes the core content of the asset information associated with the main power supply chassis, and the configuration parameters, control command sets, alarm thresholds, communication addresses, etc. related to the CDU are management information associated with the thermal management equipment.

[0031] After receiving this data, the slave RMC writes it completely into its onboard non-volatile memory (such as SPIFlash or eMMC) and creates an independent data partition for persistent storage. This process can be configured to trigger incremental synchronization every time the primary RMC asset topology changes, or it can be configured to perform a full backup during the low-load period in the early morning every day, ensuring that the slave RMC always holds the latest asset view that is highly consistent with the primary RMC.

[0032] Step S12: In response to the physical location migration event configured to the main power supply frame, retrieve the stored asset information associated with the main power supply frame and the management information associated with the thermal management device.

[0033] When the main RMC needs to be physically removed due to hardware damage, firmware crashes, or other unrecoverable failures, traditional solutions will cause the main power supply chassis and its managed assets (including CDUs) to become unreachable. The upper-level management platform will be unable to obtain the status of the relevant devices or issue control commands, which may trigger the entire rack to reduce its frequency or even shut down for protection purposes. However, using the method described in this invention, maintenance personnel can immediately perform emergency operations without waiting for new spare parts to arrive. First, the power is safely disconnected and the faulty main RMC module is unplugged. Then, the secondary RMC module is removed from the power supply chassis and inserted into the physical slot where the original main RMC was located.

[0034] When the RMC is inserted into the main power box, its startup self-test program will detect the new physical location identifier and automatically switch its role to the main RMC.

[0035] Step S13: Manage the assets and thermal management devices of the main power supply frame based on the information obtained from the call.

[0036] After the role switch is completed, the new master RMC (formerly the slave RMC) initiates the takeover process. First, it retrieves the asset information associated with the master power supply frame and the management information associated with the thermal management equipment that were previously synchronized and saved from the local non-volatile memory.

[0037] The new primary RMC utilizes the CDU communication addresses and protocol stack configurations to proactively attempt to re-establish a connection with the thermal management equipment. Once the handshake is successful, monitoring and control of the cooling system can be immediately restored: for example, if the GPU cluster in the rack is currently performing a large-scale training task, the new primary RMC can dynamically increase the CDU pump speed according to the pre-stored cooling strategy to maintain the liquid cooling circuit temperature within a safe threshold; if an abnormal output of a PSU is detected, it can also accurately determine the affected compute nodes based on asset binding relationships and send precise alarms to the upper-layer BMC (Baseboard Management Controller) or DCIM (Data Center Infrastructure Management) system. It is worth noting that, to avoid misjudgments due to data timeliness issues, the new primary RMC will mark these historical synchronized data as "non-real-time collection" or "recovery mode" when loading them for the first time, and will clearly prompt maintenance personnel in the management interface or logs that the currently displayed asset status may lag behind the actual physical status, and it is recommended to perform a full asset scan to refresh the data when conditions permit.

[0038] In one implementation, in response to a physical installation event in which a new RMC device is configured in a power supply box, asset information associated with the main power supply box is sent to the new RMC device; the physical installation event includes the new RMC device being physically installed and configured in the power supply box where the RMC device was located prior to the occurrence of a physical location migration event.

[0039] As troubleshooting progresses, once the new RMC spare part arrives, maintenance personnel can install it into the previously vacated slave power supply slot. Upon startup, the newly installed RMC recognizes its slave position and automatically assumes the slave RMC role. The system then triggers a new round of automatic negotiation and data synchronization: the new master RMC (formerly the slave RMC), acting as the current authoritative data source, pushes its latest asset database (which may include temporary records added during the takeover period) to the newly added slave RMC via the communication link; simultaneously, the new slave RMC also uploads any locally collected independent data (such as the PSU status of the slave power supply slot itself) to the master RMC, completing bidirectional fusion. During this process, the previously marked "recovery mode" flag is automatically cleared, and all asset information is restored to a "real-time valid" state. The entire recovery and reconstruction process requires no manual intervention to configure IP addresses, re-enter asset binding relationships, or recalibrate CDU parameters, greatly reducing reliance on the professional skills of maintenance personnel and avoiding secondary risks caused by human error.

[0040] In one implementation, the asset information includes the associated power supply unit's BMC information and PSU information.

[0041] In one embodiment, the heat dissipation management device includes a coolant distribution unit.

[0042] In one implementation, to meet high power density and high availability requirements, the system may deploy two or more power supply frames. Each power supply frame not only provides an independent power supply unit (PSU), but also integrates an RMC device responsible for managing the hardware assets within its power supply frame and its associated domain.

[0043] This asset information is diverse. Some of it can be automatically detected and obtained by the RMC, such as the serial number, firmware version, input / output voltage / current, temperature status and other real-time parameters of each PSU in this power supply frame, as well as some basic information of nearby computing nodes that can be accessed through specific buses (such as I²C, SMBus). Other asset information needs to be manually entered by maintenance personnel. Typically, this includes key asset information of certain computing nodes or dedicated accelerator cards that cannot directly establish a communication channel with the RMC, such as equipment model, serial number, physical location code, and business department to which they belong.

[0044] This manually entered information is crucial for accurate asset tracking and operation and maintenance management. Loss of this information necessitates significant manpower and time for re-collection and entry, and is prone to errors. To achieve a unified view and redundant backup of all rack resources, a dedicated communication link needs to be established between RMC devices in the system. This link can be physically implemented through a specific interface (e.g., a ruggedized RJ45 interface) pre-reserved on the power supply frame, using specialized shielded twisted-pair cable or fiber optic cable.

[0045] Regarding the choice of communication protocol, considering the complexity of the electromagnetic environment within the cabinet, the transmission distance (potentially exceeding one meter), and reliability requirements, the Controller Area Network (CAN) bus protocol is a preferred option. The CAN bus features a multi-master architecture, high anti-interference capabilities, and reliable error detection and handling mechanisms, making it well-suited for this type of industrial control scenario. Of course, in other implementations, Ethernet communication based on the TCP / IP protocol can also be used, depending on the network interface capabilities of the RMC hardware itself and the network topology design within the cabinet.

[0046] After the communication link is established, the RMC devices within the system need to negotiate roles or have their master-slave relationship specified by an external mechanism. A common and reliable method is through hardware configuration, i.e., setting physical DIP switches near the hardware backplane or RMC slots of the main and slave power supply frames. When the RMC device powers on and initializes, it reads the state of the corresponding DIP switch through its general purpose input / output (GPIO) pins. The RMC device configured as "master" automatically assumes the role of master RMC, while the RMC configured as "slave" operates as a slave RMC. The master RMC is usually given higher-level management privileges. For example, in a rack design, critical thermal management devices such as coolant distribution units may establish a direct communication connection only with the master RMC on the main power supply frame (e.g., via Modbus, IPMI, etc.) for the sake of simplified wiring or logical management. Therefore, the master RMC is the only device that can directly monitor and manage the CDU, and can obtain the CDU's operating parameters in real time, such as inlet / outlet coolant temperature, flow rate, pump speed, pressure value, and alarm information. This CDU management information is crucial for ensuring proper server heat dissipation and preventing chip overheating, and is part of core asset information.

[0047] Upon initial deployment or any change in asset information (such as manually entering new compute node information, PSU module replacement leading to serial number updates, CDU operating parameter fluctuations, etc.), the master RMC will proactively or according to a preset strategy (such as timed triggering, event-driven), initiate a data synchronization process. It encapsulates the complete local asset information set. This asset information set is a structured dataset containing at least two main components: the first is the associated asset information of the master power supply chassis itself, covering detailed information of all PSUs within its jurisdiction, directly communicable compute node information, and associated asset information added manually; the second, and crucial, is all management information related to the thermal management devices (CDUs), including their static configuration parameters and dynamic operating data. The encapsulated data packet is sent to the slave RMC via an established CAN bus (or Ethernet) link. Upon receiving the data packet, the slave RMC performs verification (such as CRC check), confirms the data is complete and error-free, and then parses and stores it in its own non-volatile storage medium, such as an onboard Flash chip or eMMC storage device. This ensures that a complete copy of asset information, almost identical to that of the main RMC, is persisted locally from the RMC. This includes critical data that the RMC itself cannot directly detect (such as manually entered compute node information) and data that cannot be directly accessed due to physical connection limitations (such as CDU information).

[0048] Synchronization operations can be either full synchronization, which transmits all data under initial or specific conditions, or incremental synchronization, which transmits only the changed asset information during daily operation to improve efficiency and reduce network bandwidth consumption. Correspondingly, the slave RMC also synchronizes asset information related to the power supply frames it manages (mainly the PSUs and automatically discoverable compute nodes) to the master RMC, giving the master RMC a global perspective. Therefore, under stable conditions, both the master and slave RMCs hold complete data backups covering all assets in both power supply frames, forming redundant data backups.

[0049] When the primary RMC device on the main power supply chassis completely fails due to hardware failure (such as chip damage, power module failure), software crash, or other reasons, maintenance personnel receive an alarm and arrive at the site. Since spare RMC parts may not be immediately available, to minimize management downtime, the maintenance strategy is to perform hot-swapping of the RMC board. The specific procedure is as follows: Maintenance personnel first safely power off the device (no power-off is necessary if hot-swapping is supported) and physically remove the faulty primary RMC device from its slot in the main power supply chassis. Then, they unplug the normally functioning secondary RMC device from its current secondary power supply chassis slot. Finally, they physically insert this original secondary RMC device into the newly vacated RMC slot in the main power supply chassis.

[0050] When the original slave RMC device is inserted into the main power frame and powered on again, its initialization process begins. First, the state of the DIP switch next to the RMC slot in the main power frame is read via GPIO pins. Since it is now in the main power frame position, the read state is "master". The firmware or operating system of the migrated RMC responds to this event and performs a series of critical operations. Migrating the RMC does not require rediscovering assets like a brand new blank board; instead, it immediately retrieves a complete copy of the asset information related to the main power frame, previously acquired and stored from the original master RMC during the data synchronization phase, from its local non-volatile memory. This includes detailed information on all PSUs under the main power frame, manually entered associated compute node assets, and crucial CDU management information.

[0051] After retrieving this information, the migration RMC quickly takes over the management responsibilities of the main power chassis assets and CDUs based on its new "master" role status. It immediately attempts to re-establish connections with the CDUs using stored CDU communication parameters (such as IP address, port number, and protocol type). Since the CDUs are physically connected to the main power chassis, and the migration RMC is now physically located there, the connection is intact. Therefore, the migration RMC can quickly restore communication with the CDUs, begin receiving real-time operational data from the CDUs, and resume monitoring and management functions for the CDUs. Simultaneously, it also begins polling or listening to each PSU under the main power chassis based on stored asset information, updating their status. For upper-level operation and maintenance management software (such as the Data Center Infrastructure Management System (DCIM), this switchover is almost transparent. The management software may only detect a brief connection interruption, but quickly (usually within seconds or minutes of the migration RMC initialization) it can see the complete information and status of all assets (including PSUs and CDUs) within the main power chassis domain again, allowing asset management functions to be quickly restored and avoiding prolonged data blind spots. It's worth noting that after the RMC is migrated and taken over, it may intelligently tag asset data. For example, it might mark the power supply frame asset information it collected when it was the RMC as "historical collection, pending update," because it is no longer located at the power supply frame and cannot obtain the latest data there in real time. Meanwhile, the main power supply frame assets and CDU information it is currently managing will be marked as currently valid. This tagging helps with subsequent data synchronization and allows maintenance personnel to distinguish the timeliness of information.

[0052] Once the new RMC spare part arrives at the data center, maintenance personnel do not need to perform any operations on the already stably running main power supply frame (i.e., the main RMC after the migration RMC takes over). They only need to install the new RMC spare part into the empty RMC slot in the secondary power supply frame. After the new RMC spare part is powered on, it confirms its "slave" role by reading the DIP switch status of the secondary power supply frame. Subsequently, it initiates a communication request to the current main RMC (i.e., the previous migration RMC) through the RMC communication link. After the current main RMC detects the addition of the new slave RMC, it initiates a data synchronization process, synchronizing its complete asset information (including the current main power supply frame information, CDU information, and potentially expired historical information from the original slave power supply frame) to the new slave RMC. The new slave RMC receives and stores this data, quickly integrating into the system. Afterward, the main and slave RMCs will resume regular data synchronization to ensure data consistency. For any "historical collection" markers that may exist on the migration RMC, the software logic can automatically reset or clear them after the system stabilizes and data synchronization is completed.

[0053] In one implementation, in a multi-PowerShelf scenario, two PowerShelf instances are connected via a special cable through the RJ45 port on the PowerShelf. The two RMCs communicate using the CAN protocol (the distance between the two RMCs may exceed one meter, and CAN is more reliable than I2C; of course, network communication is also feasible), enabling data communication between the RMCs in the PowerShelf.

[0054] Upon initial listing, data that cannot be automatically acquired, such as compute nodes, is manually entered, and the RMC (Real Estate Management Center) is allowed to capture more asset information (such as PSU and CDU serial numbers). After entry and capture are completed, asset data can be exchanged and synchronized via communication cables, allowing both RMCs to obtain and persistently store data from the other end.

[0055] When the master RMC malfunctions, the slave RMC can be switched to PowerShell where the master RMC resides, replacing the original master RMC. The master-slave configuration is automatically adjusted during installation; this configuration originates from PowerShell's DIP switches, and the RMC can determine this based on the GPIO read results.

[0056] At this moment, RMC still has all the asset information, but it detects that its master-slave status has changed. It then marks the asset information of the original slave node as "historical collection result, which may not be the latest status", so that users can distinguish it.

[0057] Once a new RMC spare part is available, after installation, multiple RMCs will automatically establish communication. The RMCs will automatically reset their historical data collection markers generated during the master-slave switch. The asset assessment and update process requires no manual intervention; maintenance personnel only need to plug and unplug the devices.

[0058] The installation layout and data communication of the master and slave RMCs, as well as the data storage methods, are illustrated in Figure 4. The diagram shows the normal asset information collection and connection of multiple RMCs, ensuring information synchronization during normal data entry and automatic capture. Both RMC1 and RMC2 have asset information groups (containing all information from PowerShell1-master and PowerShell2-slave). When RMC1 malfunctions, RMC2 can be quickly moved to the RMC location of PowerShell1 for replacement, avoiding operational blockage (as shown in Figure 5, when the master RMC malfunctions, the slave RMC is replaced with the master RMC). When spare parts arrive on-site, they can be directly installed on PowerShell2 without processing the RMC of PowerShell1. If there are no changes to the external configuration, no additional manual asset entry is required, making spare part replacement more convenient.

[0059] From a global perspective, RMC1 and RMC2 can perform mutual data backup. RMC can mainly manage CDU (due to physical wiring rules, only the main RMC supports communication with CDU) and PSU information.

[0060] When the RMC at the main RMC location malfunctions, in order to obtain as much information as possible for management, when the main RMC fails, RMC2 in the RMC can be manually replaced to the RMC in the main RMC to obtain the information in the CDU. There is no need to replace external cables or reconfigure the network, achieving the most convenient operation and maintenance. If the main RMC malfunctions, it can be replaced from the RMC to the main RMC. A global key management information diagram shows that the information of each PSU can be perceived by the upper-layer operation and maintenance software, and the CDU can also be obtained by the upper-layer operation and maintenance software in real time.

[0061] Once the spare part is provided to the user, maintenance personnel do not need to adjust the network topology or perform secondary adjustments to the RMC. The spare part RMC automatically synchronizes after installation, displaying a global key management information diagram. Simply install the new RMC3 into the vacant RMC slot; after installation, data will automatically synchronize between RMCs. Through redundant backup of data between the primary and secondary RMCs, information from PSUs, BMCs, and CDUs can be synchronized. The upper-level management software can obtain information to the maximum extent possible during this process, ensuring seamless operation and maintenance management within a controllable range.

[0062] In one implementation, such as Figure 2 This specification also provides a server management device applied to an RMC device. The RMC device is connected to a peer RMC device, the local RMC device is a slave RMC device, and the peer RMC device is a master RMC device. The master RMC device and the slave RMC device are respectively configured in a main power supply frame and a slave power supply frame. The master RMC device is connected to a thermal management device. The device includes: a first module for the autonomous RMC device to acquire and store asset information associated with the main power supply frame, the asset information associated with the main power supply frame including management information associated with the thermal management device; a second module for responding to a physical location migration event configured to the main power supply frame, and retrieving the stored asset information associated with the main power supply frame and the management information associated with the thermal management device; and a third module for managing the assets and thermal management device of the main power supply frame according to the retrieved information. The physical location migration event is triggered by the physical removal of the master RMC device, including physically migrating the slave RMC device to the main power supply frame to replace the physically removed master RMC device.

[0063] In one implementation, a fourth module is further included, configured to send asset information associated with the master power supply frame to the new RMC device in response to a physical installation event in which the new RMC device is configured in the slave power supply frame prior to the occurrence of the physical location migration event.

[0064] In one implementation, the asset information includes the associated power supply unit's BMC information and PSU information.

[0065] In one embodiment, the heat dissipation management device includes a coolant distribution unit.

[0066] The implementation methods of the apparatus are the same as or similar to the corresponding implementation methods, and will not be described again here.

[0067] In one embodiment, this specification provides an electronic device including a processor and a readable storage medium storing machine-executable instructions executable by the processor. The processor executes the machine-executable instructions to implement the aforementioned server management method. From a hardware perspective, a hardware architecture diagram can be found... Figure 3 As shown.

[0068] In one embodiment, this specification provides a readable storage medium storing machine-executable instructions that, when invoked and executed by a processor, cause the processor to implement the aforementioned server management method.

[0069] Here, a readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, a readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0070] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0071] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.

[0072] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification can take the form of a completely hardware implementation, a completely software implementation, or an implementation combining software and hardware aspects. Furthermore, embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0073] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments thereof. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0074] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0075] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0076] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification can take the form of a completely hardware implementation, a completely software implementation, or an implementation combining software and hardware aspects. Furthermore, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (which may include, but are not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0077] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A server management method, characterized in that, The method is applied to an RMC device, wherein the RMC device and a peer RMC device are connected via communication. The local RMC device is a slave RMC device, and the peer RMC device is a master RMC device. The master RMC device and the slave RMC device are respectively configured in a master power supply frame and a slave power supply frame. The master RMC device is connected via communication with a thermal management device. The method includes: The autonomous RMC device acquires and stores asset information associated with the main power supply frame, including management information associated with the thermal management device. In response to a physical location migration event configured to the main power supply frame, the stored asset information associated with the main power supply frame and the management information associated with the thermal management equipment are retrieved. Manage the main power supply unit's assets and thermal management equipment based on the information obtained from the call; The physical location migration event is triggered based on the physical removal of the primary RMC device, including physical migration from the RMC device to the main power box to replace the physically removed primary RMC device.

2. The method according to claim 1, characterized in that, include: In response to a physical installation event where a new RMC device is configured from the power supply box, send asset information associated with the main power supply box to the new RMC device; The physical installation event includes the physical installation and configuration of a new RMC device to the slave power box where the slave RMC device was located prior to the physical location migration event.

3. The method according to claim 1, characterized in that, The asset information includes the associated power supply unit's BMC information and PSU information.

4. The method according to claim 1, characterized in that, The heat dissipation management device includes a coolant distribution unit.

5. A server management device, characterized in that, An RMC device is used in which a communication connection is established between the local RMC device and a peer RMC device. The local RMC device is a slave RMC device, and the peer RMC device is a master RMC device. The master RMC device and the slave RMC device are respectively configured in a master power supply frame and a slave power supply frame. The master RMC device is communicatively connected to a thermal management device. The device includes: The first module is used for the autonomous RMC device to acquire and store asset information associated with the main power supply frame, wherein the asset information associated with the main power supply frame includes management information associated with the thermal management device. The second module is used to respond to a physical location migration event configured to the main power frame, and to call the stored asset information associated with the main power frame and the management information associated with the thermal management device. The third module is used to manage the assets and thermal management equipment of the main power supply frame based on the information obtained from the call; The physical location migration event is triggered based on the physical removal of the primary RMC device, including physical migration from the RMC device to the main power box to replace the physically removed primary RMC device.

6. The apparatus according to claim 5, characterized in that, include: The fourth module is used to send asset information associated with the main power supply frame to the new RMC device in response to a physical installation event from which a new RMC device is configured from the power supply frame. The physical installation event includes the physical installation and configuration of a new RMC device to the slave power box where the slave RMC device was located prior to the physical location migration event.

7. The apparatus according to claim 5, characterized in that, The asset information includes the associated power supply unit's BMC information and PSU information.

8. The apparatus according to claim 5, characterized in that, The heat dissipation management device includes a coolant distribution unit.

9. An electronic device, characterized in that, include: A processor and a readable storage medium storing machine-executable instructions that can be executed by the processor to implement the method of any one of claims 1-4.

10. A readable storage medium, characterized in that, The readable storage medium stores machine-executable instructions that, when invoked and executed by a processor, cause the processor to implement the method described in any one of claims 1-4.