Server cluster management system based on MCU
By introducing an MCU-based management system into the server cluster, the problems of complex server management and low management efficiency in the prior art are solved, and refined management and real-time monitoring of the server cluster are realized, which significantly improves management efficiency and stability.
Patent Information
- Application Number
- CN202510689804.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing server management system is complex and inconvenient in the operation and maintenance management process of multiple servers, resulting in low management efficiency and artificial misoperation can easily affect cluster stability.
A server cluster management system based on MCU is designed, including management nodes and server clusters. Each member node is configured with an MCU to monitor the server status and execute instructions. The main monitoring node is responsible for collecting and reporting status information and distributing instructions, while the management node processes and analyzes data and issues instructions to ensure the efficient and stable operation of the cluster.
Through real-time monitoring and unified management, the management efficiency of the server cluster is significantly improved, the risk of human misoperation is reduced, and the efficient and stable operation of the cluster is ensured.
Smart Images

Figure CN120223591A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology, and more specifically, to a server cluster management system based on MCU. Background Art
[0002] At present, the commonly used server management system is mainly based on the openBMC system. The status of a single server is managed through the BMC management network. In some large-scale application scenarios, multiple servers are generally equipped. The independence of each server makes the operation and maintenance management process complicated and inconvenient. In the process of deploying software environments on multiple servers, human errors will also affect the stability of the cluster. How to improve the efficiency of server cluster management is a technical problem that needs to be solved urgently in this field. Summary of the invention
[0003] In view of the defects of the prior art, the purpose of this application is to improve the efficiency of server cluster management.
[0004] To achieve the above object, the present application provides an MCU-based server cluster management system, comprising: a management node and a server cluster, the server cluster comprising a plurality of member nodes; Each member node is a server equipped with an MCU. The MCU communicates with one or more modules inside the server. The MCU is used to monitor the status of the server on this node and execute received instructions. At least one of the multiple member nodes is used as a monitoring node. The monitoring node is connected to each member node in communication. The MCU on the monitoring node is also used to obtain server status information of each member node, and to package the obtained server status information to obtain a server status data packet. The MCU on the monitoring node is also used to receive instructions issued by the management node. One monitoring node is used as the main monitoring node. The MCU on the main monitoring node is also used to send server status data packets to the management node. The MCU on the main monitoring node is also used to distribute instructions to the MCUs of the corresponding member nodes. The management node is used to receive the server status data packet sent by the main monitoring node, and the management node is also used to send instructions to the monitoring node.
[0005] In a possible implementation, the number of monitoring nodes is two, one monitoring node serves as a primary monitoring node, and the other monitoring node serves as a backup monitoring node; The standby monitoring node is used to determine whether the active monitoring node fails by monitoring the heartbeat signal of the active monitoring node; if it is determined that the active monitoring node fails, the node is converted into the active monitoring node.
[0006] In a possible implementation, the management node is further configured to determine whether a monitoring node has failed by monitoring the heartbeat signals of two monitoring nodes; if it is determined that a monitoring node has failed, a member node is configured as a new monitoring node.
[0007] In a possible implementation, the monitoring node is configured to: Pack the obtained server status information based on a linked list, where one node in the linked list is used to store the server status information of a member node.
[0008] In a possible implementation, the monitoring node is configured to: When a member node is added to the server cluster, add a corresponding node to the linked list; Or, when a member node is deleted from the server cluster, delete the corresponding node from the linked list.
[0009] In a possible implementation, the management node is configured to: Send a first instruction to the monitoring node, where the first instruction is used to instruct the reporting of server status information of a specified category.
[0010] In a possible implementation, the management node is configured to: Send a second instruction to the monitoring node, where the second instruction is used to instruct the execution of a firmware upgrade operation on a target member node according to a target firmware, and the target firmware is adapted to the server status of the target member node.
[0011] In a possible implementation, the management node is configured to: Send a third instruction to the monitoring node, where the third instruction is used to instruct the execution of a script update operation on a target member node according to a target script file, and the target script file is adapted to the server status of the target member node.
[0012] In a possible implementation, the MCU is configured with a status part thread and a task part thread; The status part thread is used to perform data transmission according to the short message communication method and provide the received instruction to the task part thread; The task part thread is used to perform corresponding operations according to the instruction provided by the status part thread.
[0013] In a possible implementation, the MCU is further configured to send abnormal status information to the monitoring node when the server status is abnormal; The management node is further configured to display the abnormal status information through an alarm interface when receiving the abnormal status information sent by the primary monitoring node.
[0014] Generally speaking, compared with the prior art, the above technical solution conceived by the present application has the following beneficial effects: The real-time monitoring and unified management of the cluster are realized through the management node and the member nodes (including at least one monitoring node) configured with an MCU. Among them, the primary monitoring node is responsible for collecting and reporting status information and distributing instructions, while the management node processes and analyzes data and issues instructions, jointly ensuring the efficient and stable operation of the cluster. Through this architecture, the server cluster management system realizes the refined management and real-time monitoring of the server cluster, significantly improving the management efficiency of the server cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic structural diagram of a server cluster management system based on an MCU provided by an embodiment of the present application; Figure 2 is a schematic diagram of the MCU communicating and connecting with other modules inside the server provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0017] The terms "first" and "second" in the description and claims of the present application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first instruction and the second instruction are used to distinguish different instructions, rather than to describe the specific order of the instructions.
[0018] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0019] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0020] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0021] Figure 1 is a schematic structural diagram of a server cluster management system based on an MCU provided by an embodiment of the present application, asFigure 1 As shown in the figure, the system includes: a management node (such as a PC) and a server cluster, and the server cluster includes multiple member nodes; Each member node is a server configured with an MCU (Microcontroller Unit), and the MCU is communicatively connected to one or more modules inside the server. The MCU is used to monitor the status of the server on this node, and the MCU is also used to execute the received instructions; At least one node among the multiple member nodes serves as a monitoring node. The monitoring node is communicatively connected to each member node. The MCU on the monitoring node is also used to obtain the server status information (such as device configuration information and device operating status information) of each member node, and pack (sort) the obtained server status information to obtain a server status data packet. The MCU on the monitoring node is also used to receive the instructions issued by the management node; One monitoring node serves as the primary monitoring node. The MCU on the primary monitoring node is also used to send the server status data packet to the management node, and the MCU on the primary monitoring node is also used to distribute the instructions to the MCU of the corresponding member node; The management node is used to receive the server status data packet sent by the primary monitoring node, and the management node is also used to issue instructions to the monitoring node.
[0022] Specifically, the system mainly consists of two parts: one is the management node, which can be a high-performance personal computer (PC) or other suitable devices; the other is the server cluster, which is composed of multiple member nodes.
[0023] Each member node is a server configured with an MCU. These MCUs are communicatively connected to one or more key modules inside the server to achieve real-time monitoring of the server status. The main functions of the MCU include: one is to monitor the running status of the server on this node in real time to ensure the stable and reliable operation of the server; the other is to execute the instructions received from the outside to achieve unified management and control of the server cluster.
[0024] Exemplarily, Figure 2 is a schematic diagram showing the communication connection between the MCU provided in the embodiment of the present application and other modules inside the server. As Figure 2 shown, the MCU uses bus communication (such as I2C bus communication) with key modules such as the CPU (Central Processing Unit) module, CPLD (Complex Programmable Logic Device) module, and storage module inside the server.
[0025] For example, the MCU can save key information to the FLASH in the way of a configuration table, such as the attributes of the server, IP address, serial number assigned by the management node, etc., so that after the network or power supply is restored, the member nodes can start and join the cluster network based on the key information in the FLASH to achieve a quick restoration of the cluster.
[0026] For example, the MCU supports system status reporting, such as processor model, memory frequency, memory capacity, graphics card model, system logs, etc.; the MCU supports whole-machine status reporting, such as the presence status of modules, the working status of modules, the presence status of storage hard disks, the temperature, current, voltage, and fan control of important chips, etc.
[0027] For example, the MCU can access the network by automatically obtaining an IP address. This network is used for communication between member nodes and can also be used for communication between member nodes and the management node.
[0028] For example, an optional way for a member node to join the cluster after power-on is as follows: after the member node is powered on, it establishes communication with the management node through broadcasting. After the management node configures the member node to join the cluster, it is managed by the monitoring node.
[0029] Among multiple member nodes, at least one node is designated as the monitoring node. The MCU of the monitoring node is not only responsible for monitoring the server status of its own node but also undertakes the task of communicating with other member nodes in the cluster. The MCU on the monitoring node regularly obtains the server status information of each member node, sorts and packages this information to form a server status data packet. In addition, the MCU on the monitoring node is also responsible for receiving instructions from the management node.
[0030] In the monitoring node, one node is set as the primary monitoring node. The MCU on the primary monitoring node also has the following functions: one is to send the sorted server status data packet to the management node so that the management node can grasp the running status of the entire cluster in real time; the other is to receive the instructions issued by the management node and distribute these instructions to the MCUs of the corresponding member nodes to achieve unified scheduling and management of the entire cluster.
[0031] The main functions of the management node include: one is to receive the server status data packets sent by the primary monitoring node, process and analyze these data to discover and solve problems in a timely manner; the other is to issue instructions to the monitoring node to achieve unified management and control of the entire server cluster through the monitoring node to ensure the efficient and stable operation of the cluster.
[0032] Optionally, as Figure 1 shown, a local area network can be provided for the cluster for communication between member nodes and between member nodes and the management node.
[0033] Therefore, real-time monitoring and unified management of the cluster are achieved through the management node and member nodes configured with MCUs (including at least one monitoring node). Among them, the primary monitoring node is responsible for collecting and reporting status information and distributing instructions, while the management node processes and analyzes data and issues instructions, jointly ensuring the efficient and stable operation of the cluster. Through this architecture, the server cluster management system realizes the refined management and real-time monitoring of the server cluster, significantly improving the management efficiency of the server cluster.
[0034] In a possible implementation, the number of monitoring nodes is two. One monitoring node serves as the primary monitoring node, and the other monitoring node serves as the standby monitoring node. The standby monitoring node is used to determine whether the primary monitoring node has failed by monitoring the heartbeat signal of the primary monitoring node; if it is determined that the primary monitoring node has failed, this node will be changed to the primary monitoring node.
[0035] Specifically, in this server cluster management system, the configuration of the monitoring nodes adopts a dual insurance mechanism, that is, it includes a primary monitoring node and a standby monitoring node. The standby monitoring node continuously monitors the heartbeat signal of the primary monitoring node to judge its operating state. If the standby monitoring node detects that the heartbeat signal of the primary monitoring node disappears or is abnormal, it indicates that the primary monitoring node may have failed. In this case, the standby monitoring node will automatically change to the primary monitoring node to ensure the continuity and reliability of the server cluster management system, thus avoiding the paralysis of the entire cluster management caused by a single point of failure.
[0036] In a possible implementation, the management node is also used to determine whether a monitoring node has failed by monitoring the heartbeat signals of the two monitoring nodes; if it is determined that a monitoring node has failed, a member node will be configured as a new monitoring node.
[0037] Specifically, in this server cluster management system, the management node is not only responsible for receiving and processing server status data packets, but also undertakes the important task of monitoring the operating states of the two monitoring nodes (the primary monitoring node and the standby monitoring node). The following is the detailed process of how the management node determines whether a monitoring node has failed by monitoring the heartbeat signal of the monitoring node and reconfigures the monitoring node when necessary.
[0038] (1) Heartbeat signal monitoring; The management node regularly receives heartbeat signals from the primary monitoring node and the standby monitoring node. These heartbeat signals are periodic messages sent by the monitoring nodes to prove that they are still operating normally.
[0039] (2) Fault detection; If the management node does not receive the heartbeat signal of a certain monitoring node within a predetermined time, or the received signal indicates that there may be a problem with the monitoring node, the management node will consider that the monitoring node may have failed.
[0040] (3)Fault confirmation; To avoid misjudgment, the management node may attempt to re - establish a connection with the suspected - failure monitoring node or send a test instruction to confirm the status of the monitoring node. If it is confirmed that the monitoring node cannot work properly, the management node will take actions.
[0041] (4)Re - configure the monitoring node; Once the management node confirms that the primary monitoring node has failed and the standby monitoring node has become the primary monitoring node, the system will lack a standby monitoring node; or, once the management node confirms that the standby monitoring node has failed, the system will also lack a standby monitoring node. At this time, the management node will select a suitable node from the existing member nodes (non - monitoring nodes) and configure it as the new standby monitoring node.
[0042] When selecting a new monitoring node, the management node will consider various factors, such as the current load of the node, network location, hardware resources, etc., to ensure that the new monitoring node can effectively undertake the monitoring task.
[0043] (5)Configuration process; The management node will send a configuration instruction to the selected member node to configure the node as the standby monitoring node.
[0044] (6)System recovery; After the new standby monitoring node is configured, the system will return to the state of dual - monitoring nodes, and continue to ensure the stable operation and management of the server cluster.
[0045] Therefore, through this mechanism, the server cluster management system can not only quickly respond and restore the monitoring function when a monitoring node fails, but also, by configuring a new monitoring node, enable the system to continuously have the ability of self - repair, thus ensuring the continuous monitoring and management of the entire cluster.
[0046] In a possible implementation, the monitoring node is used for: Packaging the obtained server status information based on a linked list, where one node in the linked list is used to store the server status information of a member node.
[0047] Specifically, a linked list is a commonly used data structure, which consists of a series of nodes, and each node contains a data part and a pointer pointing to the next node. In the monitoring node, the linked list is used to organize and manage the server status information.
[0048] The server status information of each member node is encapsulated in a linked list node. This linked list node usually contains the following information: the identifier of the member node (such as IP address or unique number); the real-time status of the server (such as CPU usage, memory usage, disk space, network traffic, etc.); the health status of the server (such as whether a failure has occurred, error logs, etc.); the timestamp (recording the time when the information was collected), etc.
[0049] Packing process: When the MCU of the monitoring node collects the server status information from each member node, it fills the collected information into the data part of the corresponding linked list node. Once the status information of all member nodes is added to the linked list, the MCU of the monitoring node packs the entire linked list into a server status data packet. This data packet is usually in a structured data format, such as JSON, XML, or a custom binary format.
[0050] Packing the server status information using a linked list structure has the following advantages. (1) Flexibility: The linked list allows nodes to be added or deleted dynamically to adapt to changes in the cluster size. (2) High efficiency: The linked list can effectively utilize memory and only occupies the memory space required for the actual data. (3) Easy management: The structure of the linked list makes it easier to update and maintain the server status information.
[0051] Therefore, in this way, the monitoring node can efficiently collect and organize the server status information, and the primary monitoring node can provide the management node with the real-time cluster operation status, thereby realizing the refined management and real-time monitoring of the server cluster.
[0052] In a possible implementation, the monitoring node is used to: When a member node is added to the server cluster, add a corresponding node to the linked list; Or, when a member node is deleted from the server cluster, delete the corresponding node from the linked list.
[0053] Specifically, the monitoring node is responsible for maintaining the linked list of the member node status information in the server cluster management system to ensure that the linked list always reflects the current configuration of the cluster. The following is a specific introduction to how the monitoring node updates the linked list when a member node is added or deleted from the server cluster.
[0054] When adding a member node. (1) Detecting a new node: When a new member node is added to the server cluster, the monitoring node will detect this change through a certain mechanism (such as network discovery protocol, configuration update notification, etc.). (2) Creating a new linked list node: The monitoring node will create a new linked list node for the newly added member node. This node will contain the identifier and initial status information of the new member node. (3) Inserting into the linked list: The monitoring node inserts the newly created linked list node into the linked list. The insertion position can be determined according to specific sorting rules, such as in the order of node identifiers or according to the joining time of the nodes. (4) Updating the linked list: After inserting the new node, the monitoring node will update the structure of the linked list to ensure the integrity and correctness of the linked list. This may include updating the pointers pointing to the new node and / or the end pointer of the linked list.
[0055] When deleting a member node. (1) Detecting node removal: When a certain member node in the server cluster is removed, the monitoring node will also receive a notification of this change through a certain mechanism. (2) Locating the linked list node: The monitoring node will search for the linked list node corresponding to the deleted member node in the linked list. This is usually done through the identifier of the node. (3) Deleting the linked list node: Once the corresponding linked list node is found, the monitoring node will delete the node from the linked list. This involves updating the pointer of the previous node to point to the next node of the deleted node, or if the deleted node is the last one, updating the end pointer of the linked list. (4) Releasing resources: After deleting the linked list node, the monitoring node will release the resources associated with this node to avoid memory leaks. (5) Updating the linked list status: The monitoring node will update the status of the linked list to ensure that the linked list reflects the actual configuration of the current cluster.
[0056] In a possible implementation, the management node is used to: Send a first instruction to the monitoring node, and the first instruction is used to instruct to report server status information of a specified category (processor information, driver information, system logs, etc.).
[0057] Specifically, the process of the management node sending instructions. (1) Instruction formulation: The management node formulates specific instructions according to the current monitoring requirements or system maintenance plan. These instructions can be periodic regular checks or investigations for specific events or problems. (2) Instruction content: The first instruction, that is, "report server status information of a specified category", will clearly indicate the categories of server status information that the monitoring node needs to collect and report. These categories may include processor information, driver information, system logs, etc. (3) Instruction sending: The management node sends the first instruction to the primary monitoring node.
[0058] Accordingly, after receiving the first instruction from the management node, the primary monitoring node will parse the instruction content and prepare to execute it. Furthermore, according to the requirements of the instruction, it will collect server status information of specified categories from the MCUs of each member node. Then, it will organize and package the collected information to form a structured data packet and send it to the management node for the management node to analyze and process.
[0059] It can be understood that by issuing the above first instruction, the management node can actively obtain the latest information on the server status of specified categories, assisting the administrator to timely understand the real-time status of the cluster, so as to quickly identify and solve potential problems. The advantage of doing so is to improve the initiative and predictability of system monitoring.
[0060] In a possible implementation, the management node is used to: Issue a second instruction to the monitoring node. The second instruction is used to instruct to perform a firmware upgrade operation on the target member node according to the target firmware, and the target firmware is adapted to the server status of the target member node.
[0061] Specifically, the purpose of the second instruction is to instruct the monitoring node to perform a firmware upgrade operation on the specified member node. Such an operation is usually to repair known vulnerabilities, improve hardware performance or add new functions.
[0062] The second instruction may include information on the target firmware and the identifier of the target member node that needs to perform the firmware upgrade. The target firmware is adapted to the server status of the target member node to ensure the compatibility and security of the upgrade operation.
[0063] After the management node issues the second instruction to the monitoring node, the primary monitoring node receives and parses the second instruction to determine the task details of the firmware upgrade. Then, the primary monitoring node transfers the target firmware to the target member node according to the instruction and guides the MCU on the target member node to execute the firmware upgrade process. After the upgrade is completed, the primary monitoring node reports the upgrade execution result to the management node.
[0064] Exemplarily, the firmware of the BIOS (Basic Input / Output System), CPLD module, and ARM (Advanced RISC Machines) module can be upgraded.
[0065] Therefore, by issuing the second instruction, the management node can control the member nodes in the cluster to perform firmware upgrades, thereby realizing the automated management and maintenance of the cluster firmware, improving the operation and maintenance efficiency, and at the same time reducing the risk of human operation errors.
[0066] In a possible implementation, the management node is used to: Send the third instruction to the monitoring node. The third instruction is used to instruct to perform a script update operation on the target member node according to the target script file, and the target script file is adapted to the server state of the target member node.
[0067] Specifically, the third instruction is used to instruct the monitoring node to perform a script update operation on the target member node. This is usually to automate specific configuration changes or maintenance tasks. For example, to deploy a software environment in a cluster with one click.
[0068] The third instruction includes the target script file and the identifier of the target member node on which the script needs to be executed. The target script file is customized according to the server state of the target member node to ensure the accuracy and effectiveness of script execution.
[0069] The management node sends the third instruction to the monitoring node; then the primary monitoring node receives and parses the third instruction to understand the specific requirements for script update; then the primary monitoring node sends the target script file to the target member node and guides the MCU of the target member node to perform the script update operation; after the script update is completed, the primary monitoring node collects the execution results of the script update operation and reports them to the management node.
[0070] Therefore, by sending the third instruction, the management node can control the member nodes in the cluster to perform script updates, thereby realizing the automated management and maintenance of cluster scripts, improving the operation and maintenance efficiency, and at the same time reducing the risk of human operation errors.
[0071] Optionally, if the firmware upgrade of the member node requires a new script file to be adapted to it, then the second instruction and the third instruction can be sent simultaneously. This means that when the member node receives a new firmware update, the relevant script files will be adjusted accordingly to adapt to the changes at the firmware level. For example, if the firmware update introduces new commands or changes the setting method of certain parameters, the script files will be updated accordingly to ensure that these new commands can be correctly called or the new parameter format can be used when executed. Vice versa, when the script files are updated to introduce new configuration policies or optimize the existing operation process, they can run smoothly with the new functions and support provided by the firmware. This two-way adaptation mechanism forms a close collaborative relationship between the firmware and the script files, so that the system can maintain the best operating state after each update, avoiding potential problems caused by version mismatches, and enhancing the stability and reliability of the system.
[0072] In a possible implementation, the MCU is configured with a status part thread and a task part thread; The status part thread is used to perform data transmission according to the short message communication method and provide the received instruction to the task part thread; The task part threads are used to perform corresponding operations according to the instructions provided by the status part threads.
[0073] The main function of the status part threads is to perform data transmission in the short message communication mode. Short message communication is an efficient and concise data exchange method that allows the MCU to send and receive information with the minimum bandwidth and the fastest speed. The status part threads will monitor the key parameters of the server in real time, such as CPU load, memory usage, disk I / O status, etc., and send these status information to the monitoring nodes in the form of short messages, or receive instructions from the monitoring nodes. This communication method ensures the real-time and reliability of data. Even in the case of poor network conditions, stable communication can be maintained.
[0074] After receiving the instructions, the status part threads will pass these instructions to the task part threads. The task part threads will then perform corresponding operations according to the instructions provided by the status part threads. These operations may include but are not limited to server restart, configuration update, software deployment, security policy implementation, etc. The execution process of the task part threads is highly optimized to ensure that the instructions can be executed quickly and accurately, so as to respond to the needs of cluster management.
[0075] The working process of the task part threads is as follows. (1) Receive instructions: The task part threads obtain instructions from the status part threads. These instructions come from the management node and are relayed through the primary monitoring node. (2) Parse instructions: The task part threads parse the instructions to determine the specific operations to be performed. (3) Execute operations: According to the instruction content, the task part threads will perform corresponding operations, such as adjusting server configuration, starting or stopping services, loading or unloading modules, etc. (4) Feedback results: After the execution is completed, the task part threads will feedback the operation results to the status part threads, and the status part threads will send the result information back to the monitoring nodes through short messages, so that the management node can understand the status of instruction execution and the overall situation of the cluster.
[0076] Therefore, through the cooperation between the status part threads and the task part threads, it is possible to monitor real-time data while also being able to execute instructions in a timely manner. The close combination of real-time data monitoring and immediate instruction execution can effectively ensure the real-time and reliability of data transmission and instruction execution.
[0077] Optionally, a state machine mechanism and a communication mechanism with responses are adopted for data processing.
[0078] A state machine is an abstract machine that can change to the next state based on the current state and input signals. A state machine consists of the following basic components. (1) State: The situation of the system at a certain moment. (2) Input: External or internal events that trigger state transitions. (3) Transition: The change from one state to another. (4) Action: Operations that may be performed during state transitions.
[0079] The communication mechanism with acknowledgment means that during the data communication process, after the sender sends data, it requires the receiver to return an acknowledgment signal (ACK) to confirm that the data has been correctly received. If the sender does not receive an acknowledgment within the predetermined time, it will resend the data or take other error handling measures.
[0080] In a possible implementation, the MCU is also used to send abnormal status information to the monitoring node when the server status is abnormal; The management node is also used to display the abnormal status information through the alarm interface when it receives the abnormal status information sent by the primary monitoring node.
[0081] Specifically, when the MCU detects that a key module on the server is abnormal, such as overheating, hardware failure, system crash, or performance degradation, it will immediately start the emergency response mechanism. The MCU will quickly collect specific information about the abnormality, including error codes, timestamps of the abnormality occurrence, affected modules, and possible causes of the failure, and send this abnormal status information to the monitoring node. This process ensures that the monitoring node can know the health status of the member nodes in real time, so that problems will not expand due to information delay.
[0082] Optionally, the MCU can implement device status alerts by using a timer plus a software interrupt.
[0083] After receiving the abnormal status information, the monitoring node will immediately process it. If it is the primary monitoring node, it will further send this information to the management node. When the management node receives the abnormal status information sent by the primary monitoring node, it will display this information in a graphical or list form through the alarm interface. This information may include server node identification, abnormal type, abnormal description, recommended solution steps, and possible impact scope.
[0084] Optionally, when an abnormal status occurs, the management node can trigger the alarm interface by means of an interrupt.
[0085] Through the alarm interface, operation and maintenance personnel can quickly take actions, such as remotely restarting the server, adjusting the load balance, starting the standby server, or dispatching technicians for on-site repair. This timely alarm and response mechanism significantly shortens the fault handling time, reduces potential losses, and ensures the continuous and stable operation of the server cluster. In addition, the management node can record all abnormal events and corresponding handling measures, providing data support for subsequent problem analysis and system optimization.
[0086] It can be understood that through the management system of the server cluster provided by this application, functions such as synchronizing the status information of member nodes to the management node, remotely upgrading the CPLD / BIOS firmware of the cluster, monitoring the health status of the whole machine / module, and obtaining system logs can be realized, breaking the traditional point-to-point management, reducing the impact caused by human misoperations, and significantly improving the management efficiency of the server cluster.
[0087] It can be understood that the various digital numbers involved in the embodiments of this application are only for convenient distinction in description and are not used to limit the scope of the embodiments of this application.
[0088] Those skilled in the art can easily understand that the above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this application shall be included in the protection scope of this application.
Claims
1. A server cluster management system based on MCU, characterized in that, Including: A management node and a server cluster, where the server cluster includes multiple member nodes; Each member node is a server configured with an MCU, and the MCU is communicatively connected to one or more modules inside the server. The MCU is used to monitor the status of the server on this node, and the MCU is also used to execute the received instructions; At least one node among the multiple member nodes serves as a monitoring node. The monitoring node is communicatively connected to each member node. The MCU on the monitoring node is also used to obtain the server status information of each member node, package the obtained server status information to obtain a server status data packet, and the MCU on the monitoring node is also used to receive the instructions issued by the management node; One monitoring node serves as the primary monitoring node. The MCU on the primary monitoring node is also used to send the server status data packet to the management node, and the MCU on the primary monitoring node is also used to distribute the instructions to the MCUs of the corresponding member nodes; The management node is used to receive the server status data packet sent by the primary monitoring node, and the management node is also used to issue instructions to the monitoring node.
2. The MCU-based server cluster management system according to claim 1, characterized in that, The number of monitoring nodes is two. One monitoring node serves as the primary monitoring node, and the other monitoring node serves as the standby monitoring node; The standby monitoring node is used to determine whether the primary monitoring node has failed by monitoring the heartbeat signal of the primary monitoring node; If it is determined that the primary monitoring node has failed, then change this node to the primary monitoring node.
3. The MCU-based server cluster management system according to claim 2, wherein The management node is also used to determine whether the monitoring node has failed by monitoring the heartbeat signals of the two monitoring nodes; if it is determined that one monitoring node has failed, then configure a member node as a new monitoring node.
4. The MCU-based server cluster management system according to claim 1, wherein The monitoring node is used for: Packaging the obtained server status information based on a linked list, where one node in the linked list is used to store the server status information of one member node.
5. The MCU-based server cluster management system according to claim 4, wherein, The monitoring node is used for: When a member node is added to the server cluster, adding a corresponding node to the linked list; Or, when a member node is deleted from the server cluster, deleting the corresponding node from the linked list.
6. The MCU-based server cluster management system according to claim 1, characterized in that, The management node is used for: Issuing a first instruction to the monitoring node, and the first instruction is used to instruct to report the server status information of a specified category.
7. The MCU-based server cluster management system according to claim 6, wherein, The management node is used for: Issuing a second instruction to the monitoring node, and the second instruction is used to instruct to perform a firmware upgrade operation on the target member node according to the target firmware, and the target firmware is adapted to the server status of the target member node.
8. The MCU-based server cluster management system according to claim 6, characterized in that, The management node is used for: Issuing a third instruction to the monitoring node, and the third instruction is used to instruct to perform a script update operation on the target member node according to the target script file, and the target script file is adapted to the server status of the target member node.
9. The MCU-based server cluster management system according to claim 1, wherein The MCU is configured with a status part thread and a task part thread; The status part thread is used to perform data transmission according to the short message communication method and provide the received instructions to the task part thread; The task part thread is used to perform corresponding operations according to the instructions provided by the status part thread.
10. The MCU-based server cluster management system according to any one of claims 1-9, characterized in that, The MCU is also used to send abnormal status information to the monitoring node when the server status is abnormal; The management node is also used to display the abnormal status information through an alarm interface when receiving the abnormal status information sent by the primary monitoring node.
Citation Information
Patent Citations
Vehicle hierarchical control network system and control method
CN103929476A
Server monitoring system
CN106941431A
High-availability cluster system
CN107948017A
Method and device for large-scale cluster management and real-time service monitoring
CN114285772A
Electromechanical management MCU firmware burning method and system
CN115328505A