Server management system
By combining a time-triggered Ethernet (TTE) switch and a dual-board management controller, high reliability and real-time performance of the server management system are achieved, solving the problem of low reliability in traditional systems and realizing microsecond-level business continuity and resource optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-03
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional server management systems have low reliability and cannot meet the requirements of high availability and real-time performance.
The system employs a time-triggered Ethernet TTE switch, a main baseboard management controller, and a backup baseboard management controller, connected via TTE network cards, to achieve global clock synchronization and traffic scheduling. Combined with dynamic load balancing and AI-driven proactive fault prediction, it ensures deterministic transmission of heartbeat messages and dynamic adjustment of tasks.
It improves the reliability and real-time performance of the server management system, achieves microsecond-level business continuity and seamless switching, and enhances resource utilization and security.
Smart Images

Figure CN121644306A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a server management system. Background Technology
[0002] Server hardware management technology is a crucial area in computer systems, particularly in scenarios with extremely high reliability and real-time requirements, such as high-availability servers, edge computing nodes, and aerospace avionics. Traditional server management systems suffer from relatively low reliability.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a server management system to at least address the problem of low reliability in server management systems in related technologies.
[0005] This application provides a server management system, including: a time-triggered Ethernet (TTE) switch, a main baseboard management controller, and a backup baseboard management controller, wherein the main baseboard management controller is connected to the TTE switch via a first TTE network interface card (NIC); and the backup baseboard management controller is connected to the TTE switch via a second TTE NIC.
[0006] In one exemplary embodiment, the method further includes: after the TTE switch is started, the TTE switch periodically broadcasts Protocol Control (PCF) frames; the primary baseboard management controller and the backup baseboard management controller calibrate their local clocks according to the received PCF frames, so that the clocks of the TTE switch are synchronized with those of the primary baseboard management controller and the backup baseboard management controller.
[0007] In one exemplary embodiment, the system further includes: the TTE switch dividing the network communication cycle into fixed time slots according to a predetermined scheduling table.
[0008] In one exemplary embodiment, the fixed time slot includes at least one of the following: a dedicated time-triggered (TT) time slot, a rate-limited (RC) time slot, and a best-effort (BE) time slot.
[0009] In one exemplary embodiment, after the TTE switch divides the network communication cycle into fixed time slots according to a predetermined scheduling table, the method further includes: triggering the TT time slot at a dedicated time, and the main baseboard management controller sending a heartbeat message to the backup baseboard management controller through the first TTE network card.
[0010] In an exemplary embodiment, after the main baseboard management controller sends a heartbeat message to the backup baseboard management controller through the first TTE network card, the method further includes: the backup baseboard management controller detecting the heartbeat message within the TT time slot; and confirming heartbeat loss if the heartbeat message is not received within a preset window.
[0011] In one exemplary embodiment, after the TTE switch divides the network communication cycle into fixed time slots according to a predetermined scheduling table, the method further includes transmitting non-critical data in the RC time slot or the BE time slot.
[0012] In one exemplary embodiment, the system further includes: the main baseboard management controller reading a first weight allocation table from a configuration file, wherein the first weight allocation table records the initial allocation ratio of various tasks in the main baseboard management controller; and the backup baseboard management controller reading a second weight allocation table from a configuration file, wherein the second weight allocation table records the initial allocation ratio of various tasks in the backup baseboard management controller.
[0013] In an exemplary embodiment, after the main substrate management controller and the backup substrate management controller read the initial weight allocation table from the configuration file, the method further includes: the main substrate management controller executing various tasks sequentially according to the priority of various tasks in the first weight allocation table; and the backup substrate management controller executing various tasks sequentially according to the priority of various tasks in the second weight allocation table.
[0014] In one exemplary embodiment, the system further includes: a first dynamic engine connected to the mainboard management controller via a first interface; and a second dynamic engine connected to the mainboard management controller via a second interface.
[0015] In one exemplary embodiment, the system further includes: the first dynamic engine or the first process monitoring the load of the main baseboard management controller in real time; and the second dynamic engine or the second process monitoring the load of the backup baseboard management controller in real time.
[0016] In one exemplary embodiment, the method further includes: when the load of the main baseboard management controller is greater than or equal to a preset load threshold, adjusting the allocation ratio of the target tasks in the first weight allocation table and the second weight allocation table.
[0017] In an exemplary embodiment, adjusting the allocation ratio of target tasks in the first weight allocation table and the second weight allocation table includes: determining tasks with a priority lower than a first preset value as the target tasks; reducing the allocation ratio of the target tasks in the first weight allocation table by a second preset value; and increasing the allocation ratio of the target tasks in the second weight allocation table by the second preset value.
[0018] In one exemplary embodiment, the method further includes: the main board management controller collecting performance parameters of each local hardware component; determining risk parameters based on the performance parameters; and performing a primary / backup switch if the risk parameters are greater than or equal to a third preset value.
[0019] In an exemplary embodiment, performing a primary / backup switch includes: the primary baseboard management controller freezing tasks with a priority higher than a fourth preset value; and the primary baseboard management controller synchronizing the frozen tasks to the backup baseboard management controller via an RC time slot.
[0020] In an exemplary embodiment, the main baseboard management controller determines risk parameters based on the performance parameters, including: inputting the performance parameters into a neural network model to obtain the risk parameters; wherein the neural network model is deployed on the main baseboard management controller or a remote server.
[0021] In one exemplary embodiment, the performance parameters include at least one of the following: CPU core temperature, CPU core temperature change slope, memory error count growth rate, jitter history of preset duration heartbeat messages, and watchdog timeout count.
[0022] This application also provides a time-triggered Ethernet (TTE) switch, which is connected to the main baseboard management controller via a first TTE network card and to the backup baseboard management controller via a second TTE network card, including: periodic broadcast protocol control PCF frames to synchronize the clock of the TTE switch with the main baseboard management controller and the backup baseboard management controller.
[0023] In one exemplary embodiment, the method further includes: dividing the network communication cycle into fixed time slots according to a predetermined scheduling table, wherein the fixed time slots include at least one of the following: dedicated time-triggered (TT) time slot, rate-limited (RC) time slot, and best-effort (BE) time slot. In the dedicated time-triggered (TT) time slot, the primary baseboard management controller sends a heartbeat message to the backup baseboard management controller through a first TTE network card, and transmits non-critical data in the RC time slot or the BE time slot.
[0024] This application also provides a main baseboard management controller, which is connected to the main baseboard management controller via a first TTE network card and to a backup baseboard management controller via a second TTE network card, comprising: reading a first weight allocation table from a configuration file, wherein the first weight allocation table records the initial allocation ratio of various tasks in the main baseboard management controller; and executing various tasks sequentially according to the priority of various tasks in the first weight allocation table.
[0025] In one exemplary embodiment, the method further includes: monitoring the load of the main board management controller in real time through a first process, and adjusting the allocation ratio of the target task in the first weight allocation table when the load of the main board management controller is greater than or equal to a preset load threshold.
[0026] This application provides a server management system, including: a time-triggered Ethernet (TTE) switch, a main baseboard management controller, and a backup baseboard management controller. The main baseboard management controller is connected to the TTE switch via a first TTE network interface card (NIC); the backup baseboard management controller is connected to the TTE switch via a second TTE NIC. By integrating the deterministic communication capabilities of TTE with the dual baseboard management controllers, the reliability of the server management system is improved. Therefore, it solves the technical problem of low reliability in related technologies for server management systems. Attached Figure Description
[0027] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A schematic diagram of a server management system provided in an embodiment of this application;
[0029] Figure 2 A time slot diagram provided for an embodiment of this application;
[0030] Figure 3 This is a schematic diagram of task allocation weights provided for embodiments of this application. Detailed Implementation
[0031] The keywords are explained below:
[0032] The Baseboard Management Controller (BMC) is used for remote management, monitoring, and maintenance of servers.
[0033] Time-Triggered Ethernet (TTE) is a deterministic network protocol that achieves highly reliable communication through time synchronization and time slot scheduling.
[0034] Protocol Control Frame (PCF) is a specific frame used for clock synchronization in TTE networks.
[0035] Time-triggered, or TT for short, refers to deterministic traffic transmitted within a predetermined time window.
[0036] Rate-Constrained (RC) and Best-Effort (BE) are both nondeterministic network traffic types in TTE (Traffic Technology).
[0037] Long Short-Term Memory (LSTM) is a type of recurrent neural network (RNN) commonly used for time series forecasting.
[0038] A Field-Programmable Gate Array (FPGA) is a type of reconfigurable integrated circuit.
[0039] Dynamic Load Splitting, or DLS for short.
[0040] The Neural Fault-Tolerance Analyzer, or Neuro-FTA for short.
[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0042] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0043] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] This application provides a server management system, including: a time-triggered Ethernet (TTE) switch, a main baseboard management controller, and a backup baseboard management controller, wherein the main baseboard management controller is connected to the TTE switch via a first TTE network interface card (NIC); and the backup baseboard management controller is connected to the TTE switch via a second TTE NIC.
[0045] like Figure 1 This is a system architecture diagram, mainly composed of a primary BMC, a backup BMC, TTE network interface cards (including the aforementioned first and second TTE network interface cards), a TTE switch, and an FPGA dynamic policy engine (optional, which can be replaced by a software module). The primary BMC and the backup BMC are connected to the TTE switch through independent TTE physical links. The TTE switch is responsible for global clock synchronization and traffic scheduling.
[0046] The primary and backup BMCs are BMC chips that support PCIe interfaces. A TTE network interface card is installed in the PCIe slot of each BMC.
[0047] TTE switches support event-triggered scheduling in managed switches. Configure the switch as a central master clock (CMC) responsible for generating and distributing global synchronization clock signals.
[0048] The primary BMC and the backup BMC are each connected directly to two different ports of the TTE switch via their respective TTE network interface cards (NICs) using independent twisted-pair cables or fiber optic cables. This link is dedicated to communication between BMCs and is physically isolated from the service network.
[0049] If the real-time requirements for load scheduling are extremely high, an FPGA chip can be added. The FPGA is connected to the main BMC through a PCIe or SPI interface to run the load scheduling algorithm and achieve fast task allocation at the hardware level.
[0050] In one exemplary embodiment, the method further includes: after the TTE switch is started, the TTE switch periodically broadcasts Protocol Control (PCF) frames; the primary baseboard management controller and the backup baseboard management controller calibrate their local clocks according to the received PCF frames, so that the clocks of the TTE switch are synchronized with those of the primary baseboard management controller and the backup baseboard management controller.
[0051] After the system powers on, the TTE switch starts up and acts as the central master clock (CMC), beginning to periodically broadcast protocol control frames. The TTE network cards on the primary BMC and backup BMC receive the PCF frames and use them to calibrate their local clocks, achieving nanosecond-level time synchronization with the switch and with each other.
[0052] In one exemplary embodiment, the system further includes: the TTE switch dividing the network communication cycle into fixed time slots according to a predetermined scheduling table.
[0053] After synchronization is complete, the TTE switch divides the network communication cycle into fixed time slots according to the predetermined schedule, as shown in the attached table. Figure 2 As shown, the fixed time slot includes at least one of the following: dedicated time-triggered (TT) time slot, rate-limited (RC) time slot, and best-effort (BE) time slot.
[0054] PCF time slot: Used to transmit protocol control frames, achieving nanosecond-level global clock synchronization. TT frame dedicated time slot: Exclusively used for transmitting heartbeat messages, ensuring they are not interfered with by any other traffic, reducing heartbeat latency and jitter to the microsecond level. RC / BE dynamic time slot: Used to transmit non-critical data such as sensor data and logs.
[0055] In one exemplary embodiment, after the TTE switch divides the network communication cycle into fixed time slots according to a predetermined scheduling table, the method further includes: triggering the TT time slot at a dedicated time, and the main baseboard management controller sending a heartbeat message to the backup baseboard management controller through the first TTE network card.
[0056] During the dedicated TT time slot of each communication cycle, the primary BMC sends a heartbeat message containing its own status (such as CPU load and memory usage) to the backup BMC through its TTE network card. Since this time slot is exclusive, this heartbeat message is transmitted without delay or jitter.
[0057] In an exemplary embodiment, after the main baseboard management controller sends a heartbeat message to the backup baseboard management controller through the first TTE network card, the method further includes: the backup baseboard management controller detecting the heartbeat message within the TT time slot; and confirming heartbeat loss if the heartbeat message is not received within a preset window.
[0058] The backup BMC listens for heartbeat messages within the expected TT time slot. If no heartbeat is received within the specified window, the heartbeat loss can be confirmed within microseconds.
[0059] In one exemplary embodiment, after the TTE switch divides the network communication cycle into fixed time slots according to a predetermined scheduling table, the method further includes transmitting non-critical data in the RC time slot or the BE time slot.
[0060] Other non-critical data (such as large amounts of sensor logs) are transmitted within the RC / BE time slot, and even if such data traffic is sudden, it will not affect the determinism of the TT heartbeat time slot.
[0061] In one exemplary embodiment, the system further includes: the main baseboard management controller reading a first weight allocation table from a configuration file, wherein the first weight allocation table records the initial allocation ratio of various tasks in the main baseboard management controller; and the backup baseboard management controller reading a second weight allocation table from a configuration file, wherein the second weight allocation table records the initial allocation ratio of various tasks in the backup baseboard management controller.
[0062] When the system starts, it reads a task weight allocation table from a configuration file (such as an XML or JSON file). This table defines the initial allocation ratio of various tasks on the primary and backup BMCs. The primary BMC reads the first weight allocation table mentioned above, and the backup BMC reads the second weight allocation table mentioned above.
[0063] Main BMC: Based on the first weight table, it mainly performs high-priority real-time tasks, such as sensor data acquisition multiple times per second.
[0064] Backup BMC: Simultaneously, based on the second weight table, execute low-priority background tasks, such as compressing and storing log files.
[0065] In an exemplary embodiment, after the main substrate management controller and the backup substrate management controller read the initial weight allocation table from the configuration file, the method further includes: the main substrate management controller executing various tasks sequentially according to the priority of various tasks in the first weight allocation table; and the backup substrate management controller executing various tasks sequentially according to the priority of various tasks in the second weight allocation table.
[0066] In one exemplary embodiment, the system further includes: a first dynamic engine connected to the mainboard management controller via a first interface; and a second dynamic engine connected to the mainboard management controller via a second interface.
[0067] In one exemplary embodiment, the system further includes: the first dynamic engine or the first process monitoring the load of the main baseboard management controller in real time; and the second dynamic engine or the second process monitoring the load of the backup baseboard management controller in real time.
[0068] The first dynamic engine mentioned above is the FPGA dynamic engine connected to the main BMC in the above embodiment, and the second backup dynamic engine mentioned above is the FPGA dynamic engine connected to the backup BMC in the above embodiment.
[0069] A policy engine running in the BMC user space or on the FPGA monitors the load of the primary and backup BMCs in real time. If the primary BMC is found to be overloaded, it can dynamically adjust the weights in the first weight allocation table, for example, by migrating some sensor polling tasks to the backup BMC to achieve load balancing.
[0070] The Dynamic Load Sharing (DLS) mechanism allows both primary and backup BMCs to participate in computational tasks, rather than through traditional cold backups, significantly improving resource utilization. The FPGA dynamic policy engine (or central control software) dynamically adjusts task allocation weights based on real-time system load.
[0071] The primary BMC prioritizes tasks with high real-time requirements (such as sensor polling and watchdog control), while the backup BMC handles non-real-time tasks (such as log storage and firmware update verification). The FPGA dynamic policy engine (or central control software) dynamically adjusts task allocation weights based on the real-time system load, as shown in the example below. Figure 3 .
[0072] In one exemplary embodiment, the method further includes: when the load of the main baseboard management controller is greater than or equal to a preset load threshold, adjusting the allocation ratio of the target tasks in the first weight allocation table and the second weight allocation table.
[0073] In an exemplary embodiment, adjusting the allocation ratio of target tasks in the first weight allocation table and the second weight allocation table includes: determining tasks with a priority lower than a first preset value as the target tasks; reducing the allocation ratio of the target tasks in the first weight allocation table by a second preset value; and increasing the allocation ratio of the target tasks in the second weight allocation table by the second preset value.
[0074] In one exemplary embodiment, the method further includes: the main board management controller collecting performance parameters of each local hardware component; determining risk parameters based on the performance parameters; and performing a primary / backup switch if the risk parameters are greater than or equal to a third preset value.
[0075] The Neuro-FTA agent deployed on each BMC continuously collects local hardware performance metrics to form time-series data. The performance parameters include at least one of the following: CPU core temperature, CPU core temperature change slope, memory error count growth rate, jitter history of preset duration heartbeat messages, and watchdog timeout count.
[0076] The collected data is fed into a pre-trained LSTM neural network model for inference. This model has been trained offline and is capable of identifying potential patterns that lead to BMC failures.
[0077] If the model outputs a risk score current_risk_score > 0.8 (threshold is configurable), the policy engine will determine that the main BMC is about to fail.
[0078] The system continuously collects and analyzes time-series data such as the slope of BMC kernel temperature changes, memory ECC error rate, and heartbeat message jitter history. It predicts fault risks and outputs a risk value (0-1). When the risk value exceeds a preset threshold (e.g., 0.8), the system determines that the primary BMC has a potential fault risk. Before the heartbeat timeout, the system smoothly migrates real-time tasks from the primary BMC to the backup BMC, achieving proactive fault avoidance and preventing service interruption.
[0079] The system then triggers an active failover process: the policy engine instructs the primary BMC to instantly freeze the status of all high-priority tasks under its responsibility and quickly synchronize this status to the backup BMC via RC time slots. Upon receiving the status, the backup BMC immediately takes over all responsibilities of the primary BMC, while the old primary BMC is downgraded to a standby node. The entire process is completed without the user's awareness, resulting in zero service interruption.
[0080] Once the primary BMC, which was downgraded due to a failure, is repaired or restarted, it will automatically rejoin the system as a standby node. It first synchronizes all status information with the current primary BMC, then gradually begins to take over background tasks, and finally the system returns to the normal operating mode of dynamic load balancing.
[0081] In an exemplary embodiment, the main baseboard management controller determines risk parameters based on the performance parameters, including: inputting the performance parameters into a neural network model to obtain the risk parameters; wherein the neural network model is deployed on the main baseboard management controller or a remote server.
[0082] If cost is a primary concern or real-time requirements are less stringent, the FPGA chip can be omitted. The dynamic load scheduling algorithm can be run as a high-priority daemon directly in the Linux user space of the main BMC. Although the scheduling latency is slightly higher than the FPGA solution, it still works effectively.
[0083] LSTM models can be deployed locally on the BMC or on a remote server. Deploying locally results in lower latency and higher reliability; deploying remotely facilitates unified model updates and centralized analysis across a large number of nodes.
[0084] This application deeply integrates Time-Triggered Networking (TTE) technology with a dual-BMC redundancy architecture, combined with AI-driven proactive prediction and dynamic load management. This not only achieves orders-of-magnitude improvements in reliability, real-time performance, and resource efficiency, but also brings a qualitative leap in security, maintainability, and business continuity, comprehensively addressing the inherent shortcomings of traditional I²C / IPMI-based dual-BMC solutions. Compared to existing technologies, the dual-BMC redundancy control method and system based on Time-Triggered Networking (TTE) proposed in this patent offer the following significant advantages: extremely high reliability and determinism, business continuity and seamless switching, significant resource optimization and efficiency improvement, enhanced security and isolation, and flexible scalability and maintainability.
[0085] This application applies Time Triggered Network (TTE) to dual BMC redundancy control. Through the innovative combination of TTE time slot isolated heartbeat, Dynamic Load Segmentation (DLS), and AI-based proactive fault prediction (Neuro-FTA), it solves the fundamental defects of traditional I²C / IPMI solutions and achieves a qualitative leap.
[0086] This application can be used not only for server BMC redundancy management, but its core TTE deterministic communication, dynamic load balancing, and AI proactive predictive fault tolerance framework can also be applied to other distributed systems with extremely high reliability requirements, such as redundancy control of autonomous driving computing units, industrial control systems (ICS), and 5G baseband units (BBU). In an exemplary embodiment, performing primary / standby switchover includes: the primary baseboard management controller freezing tasks with a priority higher than a fourth preset value; and the primary baseboard management controller synchronizing the frozen tasks to the standby baseboard management controller via RC time slots.
[0087] This application also provides a time-triggered Ethernet (TTE) switch, which is connected to the main baseboard management controller via a first TTE network card and to the backup baseboard management controller via a second TTE network card, including: periodic broadcast protocol control PCF frames to synchronize the clock of the TTE switch with the main baseboard management controller and the backup baseboard management controller.
[0088] In one exemplary embodiment, the method further includes: dividing the network communication cycle into fixed time slots according to a predetermined scheduling table, wherein the fixed time slots include at least one of the following: dedicated time-triggered (TT) time slot, rate-limited (RC) time slot, and best-effort (BE) time slot. In the dedicated time-triggered (TT) time slot, the primary baseboard management controller sends a heartbeat message to the backup baseboard management controller through a first TTE network card, and transmits non-critical data in the RC time slot or the BE time slot.
[0089] This application also provides a main baseboard management controller, which is connected to the main baseboard management controller via a first TTE network card and to a backup baseboard management controller via a second TTE network card, comprising: reading a first weight allocation table from a configuration file, wherein the first weight allocation table records the initial allocation ratio of various tasks in the main baseboard management controller; and executing various tasks sequentially according to the priority of various tasks in the first weight allocation table.
[0090] In one exemplary embodiment, the method further includes: monitoring the load of the main board management controller in real time through a first process, and adjusting the allocation ratio of the target task in the first weight allocation table when the load of the main board management controller is greater than or equal to a preset load threshold.
[0091] This application aims to provide a dual BMC redundancy control system and method based on a time-triggered communication architecture (TTE), which solves the problems of high center-slip latency, low resource utilization, and passive fault tolerance in traditional solutions, achieving microsecond-level switching, dynamic load balancing, and proactive fault prediction. The core is the deep integration of TTE's deterministic communication capabilities with dual BMC redundancy control, constructing a highly reliable redundancy architecture through global clock synchronization, time-slot isolation scheduling, and AI-driven load management.
[0092] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0093] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described server management system embodiments.
[0094] Embodiments of this application also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described server management system embodiments when run.
[0095] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0096] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described server management system embodiments.
[0097] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above-described server management system embodiments.
[0098] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0099] The server management system provided in this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A server management system, characterized by, Comprising: a time-triggered Ethernet (TTE) switch, a primary baseboard management controller, a backup baseboard management controller, wherein, the primary baseboard management controller is connected to the TTE switch through a first TTE network card; and the backup baseboard management controller is connected to the TTE switch through a second TTE network card.
2. The system of claim 1, wherein, Further comprising: after the TTE switch is started, the TTE switch periodically broadcasts a protocol control frame (PCF) frame; the primary baseboard management controller and the backup baseboard management controller calibrate local clocks according to the received PCF frame, so that the TTE switch is clock-synchronized with the primary baseboard management controller and the backup baseboard management controller.
3. The system of claim 1, wherein, Further comprising: the TTE switch divides a network communication period into fixed time slots according to a predetermined schedule table.
4. The system of claim 3, wherein, The fixed time slots include at least one of a time-triggered (TT) time slot, a rate-constrained (RC) time slot, and a best-effort (BE) time slot.
5. The system of claim 4, wherein, After the TTE switch divides the network communication period into the fixed time slots according to the predetermined schedule table, further comprising: in the TT time slot, the primary baseboard management controller sends a heartbeat packet to the backup baseboard management controller through the first TTE network card.
6. The system of claim 5, wherein, After the primary baseboard management controller sends the heartbeat packet to the backup baseboard management controller through the first TTE network card, further comprising: the backup baseboard management controller detects the heartbeat packet in the TT time slot; in a case where the heartbeat packet is not received within a preset window, it is determined that the heartbeat is lost.
7. The system of claim 4, wherein, After the TTE switch divides the network communication period into the fixed time slots according to the predetermined schedule table, further comprising: non-critical data is transmitted in the RC time slot or the BE time slot.
8. The system of claim 1, wherein, Further comprising: the primary baseboard management controller reads a first weight distribution table from a configuration file, wherein the first weight distribution table records initial distribution proportions of various tasks in the primary baseboard management controller; the backup baseboard management controller reads a second weight distribution table from a configuration file, wherein the second weight distribution table records initial distribution proportions of various tasks in the backup baseboard management controller.
9. The system of claim 8, wherein, After the primary baseboard management controller and the backup baseboard management controller read the initial weight distribution table from the configuration file, further comprising: the primary baseboard management controller executes various tasks in the first weight distribution table according to priorities of the various tasks; the backup baseboard management controller executes various tasks in the second weight distribution table according to priorities of the various tasks.
10. The system of claim 8, wherein, Further comprising: a first dynamic engine connected to the primary baseboard management controller through a first interface; a second dynamic engine connected to the primary baseboard management controller through a second interface.
11. The system of claim 10, wherein, Further comprising: the first dynamic engine or a first process monitors a load of the primary baseboard management controller in real time; the second dynamic engine or a second process monitors a load of the backup baseboard management controller in real time.
12. The system of claim 11, wherein, Further comprising: in a case where the load of the primary baseboard management controller is greater than or equal to a preset load threshold, distribution proportions of target tasks in the first weight distribution table and the second weight distribution table are adjusted.
13. The system of claim 12, wherein, Adjusting the allocation proportion of a target task in the first weight distribution table and the second weight distribution table comprises: determining a task with a priority lower than a first preset value as the target task; decreasing the allocation proportion of the target task in the first weight distribution table by a second preset value and increasing the allocation proportion of the target task in the second weight distribution table by the second preset value.
14. The system of claim 1, wherein, Further comprising: the main baseboard management controller collects performance parameters of each piece of hardware locally; determining a risk parameter according to the performance parameters; in the case where the risk parameter is greater than or equal to a third preset value, performing a master-backup switchover.
15. The system of claim 14, wherein, Performing a master-backup switchover comprises: the main baseboard management controller freezes a task with a priority higher than a fourth preset value; the main baseboard management controller synchronizes the frozen task to the backup baseboard management control through an RC time slot.
16. The system of claim 14, wherein, The main baseboard management controller determines a risk parameter according to the performance parameters, comprising: inputting the performance parameters into a neural network model to obtain the risk parameter; wherein the neural network model is deployed on the main baseboard management controller or a remote server.
17. The system of claim 14, wherein, The performance parameters comprise at least one of the following: central processing unit core temperature, central processing unit core temperature change slope, memory error count growth rate, heartbeat packet jitter history within a preset time length, watchdog timeout number.
18. A time-triggered Ethernet (TTE) switch, characterized by connected to the main baseboard management controller through a first TTE network card and connected to the backup baseboard management controller through a second TTE network card, comprising: periodically broadcast a protocol control PCF frame to synchronize the TTE switch with the main baseboard management controller and the backup baseboard management controller clock.
19. The TTE switch of claim 18, wherein, Further comprising: dividing network communication periods into fixed time slots according to a predetermined schedule, wherein the fixed time slots comprise at least one of the following: a dedicated time-triggered TT time slot, a rate-constrained RC time slot, and a best-effort BE time slot, in the dedicated time-triggered TT time slot, the main baseboard management controller sends a heartbeat packet to the backup baseboard management controller through the first TTE network card, and non-critical data is transmitted in the RC time slot or the BE time slot.
20. A host baseboard management controller, comprising: connected to the main baseboard management controller through a first TTE network card and connected to the backup baseboard management controller through a second TTE network card, comprising: reading a first weight distribution table from a configuration file, wherein the first weight distribution table records the initial allocation proportion of each type of task in the main baseboard management controller; executing each type of task in turn according to the priority of each type of task in the first weight distribution table.
21. The host baseboard management controller of claim 20, wherein, Further comprising: real-time monitoring the load of the main baseboard management controller through a first process, and in the case where the load of the main baseboard management controller is greater than or equal to a preset load threshold, adjusting the allocation proportion of a target task in the first weight distribution table.
Citation Information
Patent Citations
Timed task execution optimization method and system
CN121210072A
Server system and redundant management method thereof
US20150067084A1