Baseboard management function fault emergency system and method, server and storage medium
By introducing a fault emergency system for substrate management function in the server, and using the switching module and controller to realize automatic switching of management functions, the problems of increased costs, waste of resources and increased maintenance complexity caused by dual BMC configuration are solved, and high reliability and low-cost server management are achieved.
Patent Information
- Application Number
- CN202510238437.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, dual BMC configuration is adopted to improve the stability of server operation, and there are problems such as increased costs, waste of resources and increased maintenance complexity.
Provides a board management function fault emergency system, including a management module, a switching module and a controller. The management module manages states multiple devices on the substrate through multiple management functions, and switches the link between the management module and multiple devices on the substrate. When any management function of the management module is in a fault state, the controller controls the switching module to switch the link between the management function of the fault and the corresponding device, and connects the management module to the devices in multiple devices that can replace the fault management function of the multiple devices.
It realizes timely replacement of fault management functions, improves system reliability, reduces hardware costs, reduces the complexity of fault analysis and recovery, and improves the efficiency of fault analysis and recovery.
Smart Images

Figure CN120086072A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a substrate management function failure emergency system, method, server, and storage medium. Background Art
[0002] In today's rapidly developing technological field, servers, as the core devices for data processing and storage, their stability and reliability are crucial. The BMC (Baseboard Management Controller) is a key component in a server and undertakes important responsibilities such as monitoring the hardware status, managing the power supply, controlling the fan speed, and recording system logs. However, with the continuous expansion of the application scope of servers, the failure of the BMC has become a major hidden danger affecting the stable operation of the server.
[0003] In the related art, generally, through the dual BMC redundancy design method, the cost will increase when using two BMC chips. Troubleshooting requires analyzing the logs of both BMCs simultaneously, increasing the time cost. The standby BMC is in a low-load state for a long time and the hardware resources are not fully utilized. Moreover, in the actual execution process, the standby BMC can only be enabled when the main BMC completely fails, and it cannot respond in time to partial function failures, resulting in a decline in server performance or service interruption. Summary of the Invention
[0004] This application provides a substrate management function failure emergency system, method, server, and storage medium to at least solve the problems of increased cost, resource waste, and increased maintenance complexity in the related art when using a dual BMC configuration to improve the stability of server operation.
[0005] This application provides a substrate management function failure emergency system, including: a management module that performs status management on multiple devices on a substrate through multiple management functions; a switching module that switches the links between the management module and the multiple devices on the substrate; a controller that, when any management function of the management module is in a failure state, controls the switching module to switch the link between the failed management function and the corresponding device, and connects the management module to the device that can replace the failed management function among the multiple devices.
[0006] This application also provides a server, including: a substrate, multiple devices arranged on the substrate, and the above-mentioned substrate management function failure emergency system.
[0007] The present application also provides a method for emergency handling of substrate management function failures. This method is applied to the above-mentioned substrate management function failure emergency system and includes the following steps: identifying the operating state of the management function in the management module of the substrate; if the operating state is a failure state, controlling the switching module to switch the link between the failed management function and the corresponding device, and connecting the management module to the device that can replace the failed management function among the multiple devices; implementing the corresponding management function through the device that can replace the failed management function.
[0008] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-mentioned substrate management function failure emergency method.
[0009] With the present application, when any function of the management module fails, through the link switching of the switching module, the task of the failed function is re-assigned to the device that can replace the failed management function, and the corresponding management function is implemented through the device that can replace the failed management function. Thus, timely replacement of the failed management function can be achieved, improving the reliability of the system. Implementing the management function through the devices on the substrate without adding redundant management modules and using the devices that the substrate itself has effectively reduces the hardware cost. At the same time, the number of management modules is reduced, which can reduce the complexity of analysis during a failure and improve the efficiency of failure analysis and recovery. Thereby, the problems of increased cost, resource waste, and increased maintenance complexity in the related art when using a dual-BMC configuration to improve the stability of server operation are solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0011] Figure 1 It is a block diagram of a substrate management function failure emergency system provided by an embodiment of the present application;
[0012] Figure 2 It is a structural block diagram of a substrate management function failure emergency system provided by an embodiment of the present application;
[0013] Figure 3 It is a block diagram of a switching module provided by an embodiment of the present application;
[0014] Figure 4 It is a flowchart of a substrate management function failure emergency method provided by an embodiment of the present application;
[0015] Figure 5Flowchart of an emergency method for substrate management function failure provided by an embodiment of the present application. Detailed implementation manners
[0016] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present application.
[0017] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0018] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0019] Currently, in the operation system of a server, as a core management component, the working mechanism of the BMC plays a crucial role in the stable operation of the server. The BMC is usually directly connected to the server system and undertakes various important management responsibilities. During the operation of the server, the BMC may also crash and malfunction for some reason, and can no longer provide guarantee for the normal operation of the entire server system, resulting in the server being unable to operate normally, causing user data loss, and seriously affecting life and work. When the BMC chip fails, the BMC firmware cannot be remotely updated, and the staff can only go to the fault site to manually repair it using a burning tool. Therefore, when the BMC fails, a method to ensure the normal operation of each function of the server is extremely important. To further improve the reliability of the server, some servers adopt a dual-BMC backup redundancy design in related technologies. Under normal circumstances, the primary BMC (BMC0) undertakes all management work, while the standby BMC (BMC1) is in a standby state. Once BMC0 fails, the system will automatically switch to BMC1, and BMC1 will take over the management work of the server to ensure that the management functions of the server are not interrupted, and to ensure the continuous and stable operation of the server to the greatest extent. However, using two BMC chips will increase the cost, and troubleshooting requires analyzing the logs of both BMCs simultaneously, increasing the time cost. The standby BMC is in a low-load state for a long time and the hardware resources are not fully utilized. The biggest problem is that when a certain function of the primary BMC fails, such as the temperature monitoring fails, the standby BMC will not be immediately enabled. Only when the primary BMC completely crashes will the standby BMC be enabled, which causes problems with the server. And when the standby BMC also fails, the server will also become paralyzed.
[0020] To solve the above problems, an embodiment of the present application provides a substrate management function failure emergency system, as Figure 1 shown, including: a management module 101, a switching module 102, and a controller 103.
[0021] Among them, the management module 101 manages the states of multiple devices on the substrate through multiple management functions; the switching module 102 switches the link between the management module and multiple devices on the substrate; when any management function of the management module is in a fault state, the controller 103 controls the switching module to switch the link between the faulty management function and the corresponding device, and connects the management module to the device that can replace the faulty management function among the multiple devices.
[0022] Among them, the management functions may include temperature monitoring, power management, fan control, etc.
[0023] It can be understood that the management module 100 in the embodiments of the present application monitors and manages the states of multiple devices on the substrate through multiple management functions, and the switching module 102 switches the link connections between the management module 101 and each device on the substrate. Under normal circumstances, the management module 101 is kept directly connected to each device; once a certain management function is detected to fail, the switching module 102 can quickly change the connection path. During the actual execution process, when any management function in the management module 101 fails, the controller 103 intervenes and controls the operation of the switching module 102 to implement the switching of the link between the failed management function and the corresponding device. Specifically, the controller 103 will identify which management function has failed and direct the switching module 102 to transfer the tasks originally responsible for by the failed management function to other devices that can perform the same tasks. Thus, it can be ensured that even if some specific management functions have problems, the entire system can still operate continuously and stably, avoiding service interruption caused by a single failure point, and only one management module is adopted, and a switching module is added between the management module 101 and the system. Each functional bus of the management module 101 accesses this module, which not only guarantees the server functions but also greatly reduces the hardware cost. Moreover, when troubleshooting, only one BMC log needs to be analyzed, effectively saving time costs and improving the utilization rate of hardware resources.
[0024] In some embodiments, the switching module 102 includes multiple ports and multiple switches. Among them, the multiple ports are respectively connected to the management module and multiple devices, and the multiple switches are used to control the devices connected to the management module 101.
[0025] Among them, the multiple ports include at least one first port, at least one second port, and at least one third port. Among them, the first port is respectively connected to the management module 101 and one end of the switch, the third port is connected to the device corresponding to the management function, the second port is connected to the device that can replace the failed management function, and the other end of the switch is allowed to be connected to the second port or the third port.
[0026] The detailed block diagram of the substrate management function failure emergency system in the embodiments of the present application is as Figure 2 shown, where BCM is the management module 101, the Switch intelligent switching module is the switching module 102, and the CPU (Central Processing Unit) is the controller 103. For ease of understanding, in combination with Figure 2 and Figure 3 are elaborated in detail as follows:
[0027] Add a Switch intelligent switching module between the BMC and the system. The buses of each function of the BMC are connected to the first port ① of the Switch intelligent switching module. The third port ③ of the Switch intelligent switching module is connected to the system devices. The second port ② of the Switch intelligent switching module is connected to the devices that can replace the faulty management functions (such as Smart NIC (Smart Network Interface Card), FPGA (Field Programmable Gate Array), MCU (Microcontroller Unit), etc.). Among them, Smart NIC can replace the network management function of the BMC, FPGA can replace the functions of the BMC such as fan, temperature, power supply monitoring, and system reset, and MCU can replace the functions of the BMC such as event log collection and alarm. The devices in the system are then connected to Smart NIC, FPGA, and MCU. In addition, the other end of the switch in the embodiment of the present application can be connected to the second port or the third port. Specifically: The BMC is connected to the Switch intelligent switching module through the "bus base group", specifically connected to the "①" port of each switch. This means that the function signals of the BMC are first transmitted to these ports of the Switch intelligent switching module. One of the "②" ports of each switch is connected to the "system", and the other is respectively connected to the smart network card, FPGA, and MCU. When the switch is switched to the "②" port connection, the function signals originally connected to the system by the BMC will be switched to the corresponding alternative devices, realizing the transfer of functions; The CPU is connected to the Switch intelligent switching module through the "CTAL" signal line, which is used to control the switching action of the switch. When the CPU detects that the BMC fails, it will send a control signal to make the Switch intelligent switching module switch the function signals of the BMC to the corresponding alternative devices.
[0028] During the actual execution process, under normal circumstances, the switch of the Switch intelligent switching module is connected to the "①" port, and the BMC is directly connected to the system to execute its management functions. When the BMC fails, the CPU detects the fault signal and sends a control signal to the Switch intelligent switching module through the "CTAL" signal line. The Switch intelligent switching module switches the switch to the "②" port and connects the function signals of the BMC to the corresponding alternative devices (such as smart network card, FPGA, or MCU), and the alternative devices take over the corresponding functions of the BMC to ensure the normal operation of the server system. When the BMC returns to normal, the CPU can control the Switch intelligent switching module again to switch the connection back to the BMC, enabling it to resume its management functions.
[0029] It should be noted that in the embodiments of the present application, one BMC is used, and a Switch intelligent switching module is added between the BMC and the system. The buses of each function of the BMC are connected to the Switch intelligent switching module, and the selection and interconnection are performed by controlling the switch of the Switch intelligent switching module. The functions of the BMC are taken over by the existing devices on the board to ensure the normal operation of the system. Since only one BMC is used and there is no standby BMC, the hardware cost is reduced, and when troubleshooting, only the log of one BMC needs to be analyzed, reducing the time cost and the utilization rate of hardware resources.
[0030] In some embodiments, the controller 103 is configured to receive a first signal sent by the management module 101 and a second signal sent by a plurality of devices on the substrate, generate a target control instruction according to the first signal and the second signal, and use the target control instruction to control the switching module 102 to maintain the link unchanged or perform a link switching action.
[0031] It can be understood that the first signal can reflect the status of each management function, and the second signal can reflect the status of a plurality of devices on the substrate. The embodiments of the present application adopt a two-level judgment mechanism. As Figure 2 shown, the status of each function of the BMC is connected to the CPU through the fault / normal signal of each function of the BMC, and the device status in the system is also connected to the CPU through the fault / normal signal heartbeat packet of each function module. The CPU performs real-time monitoring. When the signals sent by the BMC and the system are detected simultaneously, the CPU will send a CTAL control signal to the Switch intelligent switching module to control the switching module 102 to maintain the link unchanged or perform a link switching action.
[0032] In some embodiments, the controller 103 is configured to: identify the signal types of the first signal and the second signal; if both the first signal and the second signal are normal signals, or the signal types of the first signal and the second signal are inconsistent, then use the target control instruction to control the switching module to maintain the link unchanged; if both the first signal and the second signal are fault signals, then use the target control instruction to control the switching module to perform a link switching action.
[0033] Wherein, the signal types of the first signal and the second signal both include fault signals and normal signals. The fault signal indicates that there is a problem with the corresponding component or function, and the normal signal indicates that the corresponding component or function is working as expected.
[0034] It can be understood that the controller 103 in the embodiments of the present application can dynamically adjust the link configuration by comprehensively analyzing the status types of the management module 101 and each device on the substrate, ensuring that the entire server can still operate smoothly even when some components fail.
[0035] Specifically, if the controller 103 receives normal signals from both the management module and the corresponding device simultaneously, it considers the current system state to be good, requires no operation, and maintains the existing link connection unchanged. If it receives one normal signal and one fault signal, it may be regarded as a false alarm or other abnormal situation. At this time, the controller 103 does not issue a control instruction and keeps the current link unchanged. If it receives fault signals from both the management module 101 and the corresponding device simultaneously, it confirms the existence of an actual fault. The controller 103 will generate a target control instruction and send it to the switching module 102. The switching module 102 adjusts the internal connection according to this instruction, reallocates the tasks of the faulty function to the standby or alternative device, ensures the continuous and stable operation of the work, and reduces the risk of service interruption.
[0036] In some embodiments, the controller 103 is configured to: when the faulty management function is restored, control the switching module to restore the link between the faulty management function and the corresponding table device.
[0037] It can be understood that the controller 103 can also monitor when the faulty function returns to normal and control the switching module to restore the original link configuration, ensuring automatic return to the optimal working state after the fault is repaired and maximizing the utilization of the functions of each component.
[0038] For example, when the temperature monitoring function of the management module (BMC) fails and sends a fault signal of the temperature monitoring function to the controller 103. The associated temperature sensor also reports a fault state. After the controller 103 confirms the existence of the fault, it commands the switching module to hand over the temperature monitoring task to the FPGA for execution. After a period of time, the BMC completes self-repair or solves the problem through remote upgrade and starts to send normal signals of temperature monitoring normally. The controller 103 recognizes that the temperature monitoring function of the BMC has returned to normal and the temperature sensor also returns to the normal state. The controller 103 can restore the original link configuration, generate a restoration instruction for the switching module 102, and the switching module 102, according to the restoration instruction, reassigns the temperature monitoring task from the FPGA back to the BMC, that is, the BMC is directly responsible for temperature monitoring.
[0039] In some embodiments, if multiple management functions of the management module fail, the device that can replace the faulty management function manages the management module.
[0040] It is understandable that in the embodiments of the present application, when multiple management functions of the management module fail simultaneously, other replaceable devices on the substrate will be used to take over and manage these faulty functions, ensuring that the server can still maintain high availability and stability in the face of multiple faults. It should be noted that the substrate is usually equipped with various types of devices that can replace faulty management functions. Specifically, the server substrate is usually equipped with various types of replaceable devices, such as SmartNIC, FPGA, and MCU. These devices each have different capabilities: SmartNIC has powerful network processing capabilities. When the network management functions of BMC (such as remote management access, network configuration, etc.) fail, SmartNIC can take over these functions. Through its own network communication interface and processing logic, it continues to provide stable network management services for the server, ensuring that remote administrators can access the server normally and perform relevant operations. FPGA has flexible programmable characteristics and can accurately monitor and control the hardware status. When the hardware monitoring-related functions of BMC, such as fan speed control, temperature monitoring, power status monitoring, and system reset, fail, FPGA can rely on the pre-written logic program inside to connect to the corresponding hardware sensors and control components, obtain hardware status information in real time, and perform regulation according to the set rules. For example, when it detects that the server temperature is too high, FPGA can control the fan to run faster to lower the temperature and ensure that the server hardware is in a safe operating state. MCU has certain data processing and storage capabilities and is mainly responsible for system event logging and alarm functions. When this part of the function of BMC fails, MCU can take over the work, record various events generated during the operation of the server in detail, such as hardware failures, software errors, user operations, etc., and analyze and judge abnormal events according to the preset rules. Once an abnormal situation is detected, MCU will send an alarm signal to the administrator in a timely manner by means of indicator light flashing, emitting a specific sound, or sending a notification to the remote management system, so that the administrator can quickly take measures to solve the problem.
[0041] Therefore, which device to choose to take over a specific management function is determined according to the capabilities of the device and the actual situation. First is the capability of the device, ensuring that the selected device has the hardware and software conditions to execute the corresponding management function. Second is the actual situation, such as the current load situation of the server and the working status of each device. If a certain device is currently in a high-load running state, even if it theoretically has the takeover ability, other relatively idle devices may be preferentially selected for function takeover to avoid new problems caused by overloading. Through such a comprehensive and flexible selection mechanism, the stable operation of the server system can be maximally guaranteed when multiple faults occur in BMC.
[0042] Secondly, an embodiment of the present application also provides a method for emergency handling of substrate management function failures. This method is applied to the above-mentioned substrate management function failure emergency system, as Figure 4 shown, and includes the following steps:
[0043] In step S101, identify the operating status of the management functions in the management module of the substrate.
[0044] It can be understood that the management module (BMC) of the substrate undertakes many key management functions, such as hardware monitoring (including fan speed, temperature of each device, power status, etc.), system event log collection, alarm, and remote management. Therefore, accurately identifying the operating status of the BMC management function is the basis for the stable operation of the entire system. Only by timely and accurately detecting the failures of the BMC can effective countermeasures be taken subsequently to avoid serious problems in the server system caused by BMC failures, such as hardware damage, data loss, and service interruption.
[0045] To identify the operating status of these management functions, the system will adopt a variety of monitoring means. The BMC itself will monitor the execution of its various functions in real time. For example, through internal sensors and monitoring circuits, it continuously obtains hardware-related data and compares it with the preset normal parameter range. At the same time, each functional module in the server system will also send its own working status information to the controller, and these information are sent periodically in the form of normal / fault signal heartbeat packets (for example, once per second). The BMC will also send the status information (normal or fault signal) of its own functions to the controller. The controller has intelligent detection capabilities. It determines the operating status of the BMC management function by receiving and analyzing the signals from the BMC and each functional module of the system. This process is based on the principles of signal transmission and logical judgment, and the CPU comprehensively processes signals from different sources to ensure the accuracy of the judgment. For example, when the CPU receives both a normal signal of a certain function sent by the BMC and a normal signal heartbeat packet sent by the corresponding system functional module, it will determine that the management function is in a normal operating state; conversely, if a fault signal is received, it will judge that the function may have a fault.
[0046] In step S102, if the operating status is a fault status, then control the switching module to switch the link between the faulty management function and the corresponding device, and connect the management module to the device that can replace the faulty management function among multiple devices.
[0047] It is understandable that by analyzing the signal and determining that a certain or certain management functions of the BMC are in a fault state, it will immediately send a CTAL control signal to the switching module. The switching module is a key component for realizing function switching in the entire system, and it contains multiple controllable switches inside. When receiving the CTAL signal from the CPU, the switching module will accurately control the corresponding switch actions according to the specific type of the fault, thereby changing the signal transmission link.
[0048] For example, if the network management function of the BMC fails, the switching module will switch the bus of the BMC network management function from the originally connected system port to the port corresponding to the SmartNIC intelligent network card, so that the SmartNIC establishes a connection with the network management function of the BMC; if the functions such as the fan, temperature, power supply monitoring, and system reset of the BMC fail, the relevant bus will be switched to the port corresponding to the FPGA to realize the connection between the BMC and the FPGA; when the functions such as the system event log and alarm of the BMC fail, the Switch intelligent switching module will switch the link to the port corresponding to the MCU, realizing the rapid transfer of the fault management function, ensuring that the server system can still maintain the normal operation of some key functions when the BMC fails, and avoiding the complete interruption of the server function caused by the BMC fault by timely switching the faulty function to the replaceable device.
[0049] In step S103, the corresponding management function is implemented by a device that can replace the faulty management function.
[0050] It is understandable that after the switching module completes the link switching, the device that can replace the faulty management function begins to play a role. Taking SmartNIC as an example, after taking over the network management function of the BMC, it will immediately undertake the task of remote management access to the server. For example, it allows the administrator to remotely log in to the server through the network to view the running status, configuration parameters, etc. of the server. At the same time, SmartNIC can also perform the operation of firmware upgrading the BMC to ensure that the software version of the BMC is always up-to-date to obtain better performance and stability. For the FPGA, after taking over the functions such as the fan, temperature, power supply monitoring, and system reset of the BMC, it will real-time monitor whether the rotation speed of the fan is normal, obtain the temperature data of each device through the temperature sensor, and monitor whether the output voltage and current of the power supply are within the normal range. Once an abnormal situation is found, the FPGA will perform corresponding processing according to the preset rules, such as adjusting the fan speed, sending an alarm signal, or performing a system reset operation, etc. After the MCU takes over the system event log and alarm functions of the BMC, it will record various events generated during the operation of the server in detail, including hardware failures, software errors, user operations, etc. At the same time, when an abnormal event is detected, the MCU will send an alarm signal in time to remind the administrator to pay attention through indicators, sounds, network notifications, etc.
[0051] Since the specific technical details have been described in detail in the previous system description, for the parts not elaborated, reference can be made to the above embodiments and will not be repeated here.
[0052] The following will Figure 2 describe in detail the working mechanism of the emergency method for substrate management function failure in the embodiments of the present application. As Figure 5 shown, the details are as follows:
[0053] Step 1: After the server board is powered on, the Switch intelligent switching switch connects port ① and port ③ of the Switch intelligent switching switch according to the default setting, enabling the BMC to directly establish communication links with each device in the server system. The BMC can immediately monitor the fan speed, temperature of each device, power supply status, etc. in real time, and at the same time start collecting system logs and preparing an alarm mechanism for possible failures. The standby functional devices such as Smart NIC, FPGA, and MCU are in a standby state during this stage and do not participate in the management work of the BMC.
[0054] It can be seen that this default connection method ensures that during the normal startup stage of the server, the BMC can quickly and efficiently perform its complete management functions and maintain the stable operation of the server system. Since the BMC is directly connected to the system devices, the data transmission path is short and the delay is low, enabling it to quickly obtain hardware status information, ensuring that all parameters of the server are within the normal range during the startup process, and laying a foundation for subsequent stable operation. At the same time, the standby devices do not intervene in the work of the BMC, reducing the system complexity and resource consumption, avoiding unnecessary interference, and improving the reliability and stability of the system;
[0055] Step 2: If all functions of the BMC are normal at this time, the BMC will send a normal signal to the CPU within a specified time (such as 1 second). And if the devices of the system are also working properly at this time, each device will also send a normal signal heartbeat packet to the CPU within the specified time. If the CPU receives normal signals from both the BMC and the system simultaneously, it is considered that the BMC and the system are normal, and the CPU sends a CTAL signal to the Switch intelligent switching module to control the switch to connect port ① and port ③, directly connecting the BMC and the system (if port ① and port ③ are already connected at this time, it remains unchanged). If the CPU receives one normal signal and one fault signal, the CPU determines it as a misoperation and does not send the CTAL signal, and the Switch intelligent switching module remains unchanged. This dual-signal confirmation mechanism greatly improves the accuracy of the system's judgment on the overall operating status of the BMC and the system. By having the BMC and the system devices send signals respectively and corroborate each other, the possibility of misjudgment of a single signal is reduced. When it is judged to be in a normal state, the direct connection between the BMC and the system is maintained, ensuring the efficiency and stability of data transmission, because the direct connection between the BMC and the system can quickly respond to system requirements and achieve precise management of the hardware. In the case of possible misjudgment, no switching operation is performed, effectively avoiding system instability or even failure caused by mis-switching, ensuring the continuous and stable operation of the server, and reducing the system maintenance cost and risk
[0056] Step 3: When some functions of the BMC are abnormal, the BMC will promptly send a function fault signal to the CPU within the specified time, notifying the CPU that there is a problem with itself. At the same time, the system devices affected by this abnormal function will also show corresponding abnormalities and send a function fault signal heartbeat packet to the CPU within the specified time. After receiving the fault signals from the BMC and the system, the CPU will immediately send a CTAL signal to the Switch intelligent switching module, instructing it to connect port ① and port ②. Taking the network management function fault of the BMC as an example, the Switch intelligent switching module will quickly connect the bus of the BMC network management to the SmartNIC
[0057] After that, the SmartNIC will fully take over the network management function of the BMC. It can not only continue to achieve remote management access to the server, but also perform maintenance operations such as firmware upgrade on the faulty BMC. It should be noted that during the switching process, if other functions of the BMC are still normal, the Switch intelligent switching module will not switch these normal functions, ensuring that the workload of each component of the server remains balanced and avoiding the impact on other functions caused by unnecessary switching. This precise switching mechanism for partial functional failures of the BMC effectively guarantees the normal operation of some functions of the server when some functions of the BMC are abnormal. By promptly switching the faulty function to the standby component, it ensures that the critical functions of the server are not interrupted. For example, the continuous availability of the network management function guarantees the smooth progress of remote management and avoids the server losing contact with the outside world due to BMC failure.
[0058] At the same time, only the faulty function is switched, and the original connections of other normal functions are maintained, avoiding large-scale adjustments to the entire system architecture, reducing the complexity and instability of the system, improving the fault tolerance and resource utilization rate of the system, reducing the waste of hardware resources, and extending the overall service life of the server.
[0059] Step 4: When all functions of the BMC fail, the situation is relatively critical. At this time, the Switch intelligent switching module will respond quickly, connect all port ① to port ②, enabling standby functional components such as Smart NIC, FPGA, and MCU to fully take over all functions of the BMC. The Smart NIC is responsible for network management and remote operation, the FPGA undertakes functions such as fan, temperature, power supply monitoring, and system reset, and the MCU is responsible for system event log collection and alarm. These standby components work together to ensure that the server system can continue to operate normally, maintain the basic functions of the server, and avoid server paralysis caused by complete BMC failure.
[0060] When the BMC is repaired or resumes normal function by itself, the CPU will detect the signal of the BMC function recovery and control the Switch intelligent switching module to switch the control right back to the BMC. The BMC resumes the management work of the server, and components such as SmartNIC, FPGA, and MCU return to the standby state again, waiting for the next possible fault switching requirement. Thus, in the extreme case of complete BMC failure, the mechanism of standby components fully taking over functions is the last line of defense for the server, which can ensure that the server will not be completely paralyzed, maximally guarantee the continuity of services on the server, avoid data loss and service interruption caused by server failure, and reduce economic losses.
[0061] When the BMC returns to normal, it can smoothly take over the control right, enabling the system to return to the optimal operating state, improving the reliability and maintainability of the system, providing a strong guarantee for the long-term stable operation of the server, enhancing the adaptability and flexibility of the entire server system, and meeting the usage requirements in different complex environments.
[0062] This application also provides a server, including: a substrate, a plurality of devices disposed on the substrate, and the above-mentioned substrate management function failure emergency system 10.
[0063] An embodiment of this application also provides a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is configured to execute the steps in any one of the above-mentioned substrate management function failure emergency method embodiments when running.
[0064] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0065] An embodiment of this application also provides a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-mentioned substrate management function failure emergency method embodiments.
[0066] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0067] The above has introduced in detail a substrate management function failure emergency method provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A baseboard management function failure emergency system, characterized in that: include: A management module manages the status of multiple devices on the substrate through multiple management functions; A switching module, switching the links between the management module and the plurality of devices on the substrate; The controller controls the switching module to switch the link between the failed management function and the corresponding device when any management function of the management module is in a faulty state, and connects the management module to a device among the multiple devices that can replace the failed management function.
2. The baseboard management function failure emergency system according to claim 1, characterized in that: The switching module includes a plurality of ports and a plurality of switches, wherein the plurality of ports are respectively connected to the management module and the plurality of devices, and the plurality of switches are used to control the devices connected to the management module.
3. The baseboard management function failure emergency system according to claim 2, characterized in that: The plurality of ports include at least one first port, at least one second port and at least one third port, wherein: The first port is connected to the management module and one end of the switch respectively, the second port is connected to a device that can replace the management function of the fault, the third port is connected to a device corresponding to the management function, and the other end of the switch allows connection to the second port or the third port.
4. The baseboard management function failure emergency system according to claim 1, characterized in that: The controller is used to: Receive a first signal sent by the management module and a second signal sent by multiple devices on the substrate, generate a target control instruction according to the first signal and the second signal, and use the target control instruction to control the switching module to maintain the link unchanged or perform a link switching action.
5. The baseboard management function failure emergency system according to claim 4, characterized in that: The signal types of the first signal and the second signal both include a fault signal and a normal signal.
6. The baseboard management function failure emergency system according to claim 5, characterized in that: The controller is used to: identifying signal types of the first signal and the second signal; If both the first signal and the second signal are the normal signals, or the signal types of the first signal and the second signal are inconsistent, controlling the switching module to maintain the link unchanged by using the target control instruction; If both the first signal and the second signal are fault signals, the target control instruction is used to control the switching module to perform a link switching action.
7. The baseboard management function failure emergency system according to claim 1, characterized in that: The controller is used to: When it is identified that the failed management function is restored, the switching module is controlled to restore the link between the failed management function and the corresponding table device.
8. A server, characterized in that: include: A substrate and a plurality of devices disposed on the substrate; The baseboard management function failure emergency system according to any one of claims 1 to 7.
9. An emergency method for baseboard management function failure, characterized in that: The method is applied to the baseboard management function failure emergency system according to any one of claims 1 to 7, and the method comprises the following steps: Identify the operating status of the management function in the management module of the baseboard; If the operation state is a fault state, controlling the switching module to switch the link between the faulty management function and the corresponding device, connecting the management module to a device among the multiple devices that can replace the faulty management function; The corresponding management function is implemented by the device that can replace the management function of the failure.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the baseboard management function failure emergency method as claimed in claim 9.
Citation Information
Cited By
Low-failure-rate switch, data transmission method and data transmission system
CN120281734A
Data acquisition system, data acquisition method, medium, and program
CN122470400A