A distributed, centralized management system based on SoC array servers
By employing a distributed and centralized management system in the SoC array server, combined with management network topology and election strategies, the high cost, complexity, and thermal power consumption issues of the BMC management module are resolved, achieving an efficient and reliable management solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-25
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies that manage SoC array servers by improving the performance of the BMC management module suffer from high hardware costs, increased complexity, reduced reliability, and increased thermal power consumption, making it difficult to simultaneously resolve the contradictions between performance, cost, reliability, and heat dissipation.
A distributed and centralized management system is adopted. The management network is designed through BMC management board and node baseboard. The management interface is expanded by using gigabit Ethernet bus interface and network switching chip. Combined with election strategy, Master and Slave are elected to realize regional management of SoC array card and reduce the workload of BMC management system.
It reduces the hardware cost and software complexity of the BMC management system, improves reliability, reduces processor power consumption, improves chassis heat dissipation, and resolves the contradiction between performance, cost, and reliability.
Smart Images

Figure CN116633928B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server technology, specifically to a distributed and centralized management system based on a SoC array server. Background Technology
[0002] SoC array servers are currently the most suitable underlying hardware infrastructure devices for cloud phones and cloud gaming (cloud mobile games) applications. A single SoC array server typically integrates dozens or even hundreds of SoC array cards. For server products, due to their deployment in data centers, which makes local maintenance difficult and requires 24 / 7 uninterrupted operation, remote operation and maintenance management capabilities are just as important as business processing capabilities. Servers have much higher standards for system maintainability than home PCs and more stringent requirements for operational stability; therefore, general-purpose servers need to combine high performance, high availability, and high reliability.
[0003] The general-purpose server uses a dedicated BMC management module (baseboard management controller) to ensure that the server can be effectively managed during operation, to diagnose faults in a timely manner, and to report the collected management information to the upper-level operation and maintenance network management system in a timely manner, which plays a vital role in the back-end protection of the server system.
[0004] However, for special servers such as SoC array servers, there are usually dozens or even hundreds of service boards (SoC array cards). The BMC management module concurrently manages such a large number of service boards, which places more stringent requirements on the processor performance of the BMC management module and the reliability of the BMC management system.
[0005] Existing technologies all achieve this by improving the performance of the BMC management chip and using higher-performance multi-core ARM or x86 processors. The drawback is:
[0006] 1. First, it will increase the hardware cost of the BMC management module.
[0007] 2. Secondly, if the performance of the BMC management system is improved simply by increasing the performance of the BMC management chip, the design complexity of the BMC management module hardware and software will increase significantly. However, the complexity of the hardware and software design is positively correlated with the reliability of the BMC management module, which will lead to a decrease in the reliability of the BMC management system.
[0008] 3. Improvements in the hardware performance of the BMC management module processor and the increased workload of the management system software mean that low-power ARM processors can no longer be used, necessitating the adoption of multi-core ARM or x86 processors. This, in turn, leads to increased thermal power consumption of the BMC management module, consequently increasing the complexity of the chassis cooling system design.
[0009] In summary, compared with general-purpose x86 and big-core ARM servers, SoC array servers place more stringent requirements on the selection of BMC management chips and the design of hardware and software systems. It is difficult to resolve the contradictions between performance, cost, reliability and heat dissipation by simply selecting the processor chip for the BMC management module. Summary of the Invention
[0010] This invention provides a distributed and centralized management system based on a SoC array server. It adopts a server management hardware and software framework that combines distributed and centralized management within the server chassis, which can effectively solve the aforementioned technical problems.
[0011] To achieve the above objectives, the present invention provides the following technical solution: a distributed and centralized management system based on a SoC array server, comprising a BMC management board and node baseboards. Both the BMC management board and the node baseboards are designed with management networks. On the BMC management board, the BMC management chip extends one gigabit Ethernet bus interface through a network interface card chip or a network physical layer chip; and further extends M gigabit Ethernet bus interfaces through a network switching chip for interactive management information communication with M node baseboards.
[0012] On the node baseboard, a gigabit Ethernet switching chip is used to expand one gigabit Ethernet bus interface as an uplink network port to connect to the BMC management board as one management network connection; and expands N gigabit Ethernet bus interfaces as downlink network ports to implement interactive management information communication with N SoC array cards.
[0013] The N SoC RAID cards on each node blade can form a regional management group. The N SoC RAID card members in each regional management group can communicate with each other through the local area network extended by the backplane management network switching chip. One SoC RAID card is elected as the management SoC RAID card Master from the N SoC RAID cards through an election strategy, and the other N-1 SoC RAID cards are member SoC RAID card Slaves. The Master is responsible for reporting the election results of the regional management group to the BMC management system. The BMC management system and each SoC RAID card will record the election results.
[0014] The Master can collect hardware status information from N-1 Slaves. The Master is responsible for periodically reporting the collected hardware status information of the entire node to the BMC management system.
[0015] Preferably, the position number of the SoC array card on the baseboard is used as the election strategy.
[0016] Preferably, the SoC array cards are deployed sequentially on the node backplane. The SoC array card closest to the chassis backplane is numbered 1, and the subsequent SoC array cards are numbered 2, 3...N. The recommended election strategy is to elect the SoC array card with the largest position number as the Master. After the node powers on, each SoC array card reads its position information on the chassis and backplane, encapsulates its management information, and broadcasts it within the management area network. Each SoC array card on the same node compares the management information received from other SoC array cards with its own position information and running status information. If it finds an SoC array card with a larger position number than itself that is running normally, it marks its master / slave attribute as "slave"; if it finds that its position number is the largest and its running status is normal, it marks its master / slave attribute as "master".
[0017] Preferably, when the node is running normally, the board with the largest position number will be elected as the "master" by default. If the board with position number N is not in place or is in an abnormal state, the SoC array card with the smaller position number will have the opportunity to be elected as the master. Each SoC array card on the same node needs to record its own and the management information of the other N-1 SoC array cards as management group information, and perform dynamic maintenance of management group information: every certain period of time, it sends its own management information to the other N-1 SoC array cards through heartbeat. If each SoC array card finds that its own or other SoC array cards' management information has changed, it updates its recorded management group information.
[0018] Preferably, if the Master fails during node operation, that is, the Slave does not receive heartbeat packets from the Master for a certain period of time, or the Slave finds that the running status in the Master's management information is abnormal, the election strategy will be re-triggered to elect a new Master.
[0019] Compared with the prior art, the beneficial effects of the present invention are:
[0020] By adopting a management framework solution that combines distributed and centralized management, the concurrent instruction execution volume of the BMC management system operating the SoC array card is reduced, greatly alleviating the hardware and software workload of the BMC management module.
[0021] 1. It can effectively reduce the hardware cost of BMC processor chips and their peripherals;
[0022] 2. It can effectively enhance the hardware and software reliability of the BMC management module;
[0023] 3. The heat dissipation of the chassis system can be improved by reducing the processor power consumption of the BMC management board. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of the network hardware topology architecture in a distributed and centralized management system based on a SoC array server according to the present invention.
[0026] Figure 2 This is a schematic diagram of the SoC array card in a distributed and centralized management system based on a SoC array server according to the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0029] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] Definitions:
[0031] BMC: Baseboard Management Controller.
[0032] SoC: System on Chip.
[0033] This invention provides a distributed and centralized management system based on a SoC array server, including a BMC management board and a node baseboard.
[0034] I. Network Hardware Management Solution:
[0035] Assuming the total number of blade slots in a SoC array server node is M (e.g., M=12), and the number of SoC array cards per blade in each node is N (e.g., N=6), then the total number of slots in the entire system is N. M (e.g., 12) (6=72) SoC array cards.
[0036] Both the BMC management board (the hardware carrier of the BMC management module) and the node baseboard (the SoC array card deployed on the node baseboard) are designed with a dedicated management network (the dedicated management network is used to achieve separation from the service network and ensure that they do not interfere with each other):
[0037] 1.1 On the BMC management board, the BMC management chip expands one gigabit Ethernet bus interface through a network card chip (or network physical layer chip) (if the management network bandwidth needs to be increased, a 10Gb network card or a 2.5Gb network card can be used, or several network cards can be bonded together to expand the network bandwidth); and then expands M gigabit Ethernet bus interfaces through a network switching chip (if the management network bandwidth needs to be increased, it can be expanded to 2.5Gb network), which are used to implement interactive management information communication with M node baseboards.
[0038] 1.2 On the node baseboard, a gigabit Ethernet switching chip is used to expand one gigabit Ethernet bus interface (if the management network bandwidth needs to be increased, it can be expanded to 2.5Gb network) as an uplink network port to connect to one management network port of the BMC management board; N gigabit Ethernet bus interfaces are expanded as downlink network ports to implement interactive management information communication with N SoC array cards.
[0039] The above-mentioned management network hardware topology is as follows: Figure 1 As shown.
[0040] II. Distributed + Centralized Management System Solution:
[0041] 2.1. The N SoC array cards on each node blade can be used as a regional management group, so the whole machine can be divided into M regional management groups.
[0042] 2.2 The N SoC RAID cards in each regional management group can interconnect and communicate via a local area network extended by the backplane management network switching chip. From the N SoC RAID cards, a certain election strategy is used to elect one SoC RAID card as the management SoC RAID card (Master), while the other (N-1) SoC RAID cards serve as member SoC RAID cards (Slaves). The Master is responsible for reporting the election results of its regional management group to the BMC management system. The BMC management system and each SoC RAID card record the election results.
[0043] 2.3 The Master can collect hardware status information (such as temperature, power supply, power consumption, presence status, and operating status) from (N-1) Slaves. The Master is responsible for periodically reporting the collected hardware status information of the entire node to the BMC management system. In this way, the BMC management system does not need to query the hardware status information of each SoC array card, but only needs to query the management SoC array card (Master) in each node, which greatly reduces the workload of the BMC management system (the SoC array card can also actively report the information).
[0044] III. Election Strategies and Maintenance Plans:
[0045] The election strategy mentioned above can have several options, such as: based on the order in which the SoC RAID cards complete their power-on boot process, based on the position number of the SoC RAID cards on the backplane, based on the MAC address value of the SoC RAID card's management network, based on the IP address value of the SoC RAID card's management network, based on the SN serial number value of the SoC RAID card, etc. Here, we recommend using the position number of the SoC RAID cards on the backplane as the election strategy. The advantage is that under normal conditions, the position of the Master on each node is consistent on its node's backplane and will not change when the SoC RAID card is replaced. The election strategy is as follows:
[0046] 3.1 SoC array cards are deployed sequentially on the node backplane. The SoC array card closest to the chassis backplane is numbered 1, and the subsequent SoC array cards are numbered 2, 3...N (or the position numbers can be reversed from largest to smallest). The recommended election strategy is to elect the SoC array card with the largest position number (N) as the Master. The advantage is that this SoC array card is closer to the chassis air intake, resulting in a lower SoC processor chip temperature and higher operational stability, which is beneficial for this SoC array card to undertake node management tasks. Figure 2 As shown.
[0047] After a node powers on, each SoC array card reads its location information on the chassis and backplane (refer to patent CN2022109528898), encapsulates its management information (location information + IP address information + operating status information (normal or faulty)) and broadcasts it within the management area network. Each SoC array card on the same node compares the received management information (IP address + location information + operating status information) from other SoC array cards with its own location information + operating status information. If a SoC array card with a higher location number than its own is present and operating normally, it marks its own master / slave attribute as "slave"; if its own location number is the highest and its operating status is normal, it marks its own master / slave attribute as "master". In summary, under normal node operation, the card with the highest location number (N) will be elected as the "master" by default. If the card with location number N is not present or is in an abnormal state, the SoC array card with the lower location number will have the opportunity to be elected as the master. Each SoC array card on the same node needs to record its own and the management information (location information + IP address information + running status information) of the other N-1 SoC array cards as management group information, and perform dynamic maintenance of management group information: every certain period of time, it sends its own management information to the other N-1 SoC array cards through heartbeat. If each SoC array card finds that its own or other SoC array cards' management information has changed, it updates its recorded management group information.
[0048] 3.2 If the Master fails during node operation, i.e. the Slave does not receive heartbeat packets from the Master for a certain period of time, or the Slave finds that the Master's running status in the management information is abnormal, the election strategy will be re-triggered to elect a new Master. The election process and strategy are the same as described above.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] By adopting a management framework solution that combines distributed and centralized management, the concurrent instruction execution volume of the BMC management system operating the SoC array card is reduced, greatly alleviating the hardware and software workload of the BMC management module.
[0051] 1. It can effectively reduce the hardware cost of BMC processor chips and their peripherals;
[0052] 2. It can effectively enhance the hardware and software reliability of the BMC management module;
[0053] 3. The heat dissipation of the chassis system can be improved by reducing the processor power consumption of the BMC management board.
[0054] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A distributed, centrally managed system based on an array of SoC servers, characterized in that, The BMC management board and the node backboard are designed with a management network, on the BMC management board, a BMC management chip expands one gigabit Ethernet bus interface through a network card chip or a network physical layer chip; Through a network switching chip, M gigabit Ethernet bus interfaces are expanded for interactive management information communication with M node backboards; On the node backboard, a gigabit Ethernet switching chip is used to expand one gigabit Ethernet bus interface as an upper network interface to connect with the one management network of the BMC management board; N gigabit Ethernet bus interfaces are expanded as lower network interfaces to implement interactive management information communication with N SoC array cards; The N SoC array cards on each node blade can be a regional management group, the N SoC array card members in each regional management group can be interconnected and communicated through the local area network expanded by the backboard management network switching chip, one SoC array card is elected as a management SoC array card Master from the N SoC array cards through an election strategy, and the other N-1 SoC array cards are member SoC array cards Slave; the Master is responsible for reporting the election result of the regional management group to the BMC management system; the BMC management system and each SoC array card record the election result; Through the Master, the hardware state information of the N-1 Slaves can be collected, and the Master is responsible for reporting the collected hardware state information of the entire node to the BMC management system regularly.
2. The distributed, centralized management system based on the SoC array server according to claim 1, characterized in that: The position number of the SoC array card on the backboard is used as the election strategy.
3. The distributed, centralized management system based on SoC array server according to claim 2, characterized in that: The SoC array cards are deployed in sequence on the node backboard, the SoC array card close to the backboard of the case is 1, and the position numbers of the SoC array cards behind it are 2, 3, …, N in turn, the recommended election strategy elects the SoC array card with the largest position number as the Master, after the node is powered on and runs, each SoC array card reads the position information of the board card on the case and the backboard, and encapsulates the management information of the board card and broadcasts it in the management local area network, each SoC array card on the same node compares the received management information of other SoC array cards with the position information and running state information in its own management information, if a SoC array card with a larger position number than itself exists and is running normally, the master-slave attribute of the SoC array card is marked as "Slave"; If the position number of the SoC array card is the largest and the running state of the SoC array card is normal, the master-slave attribute of the SoC array card is marked as "Master".
4. The distributed, centralized management system based on SoC array servers according to claim 3, characterized in that: In normal operation, the card with the largest position number is elected as the master by default. If the card with position number N is not in place or is abnormal, the SoC array card with a smaller position number will have a chance to be elected as the master. Each SoC array card on the same node needs to record its own and other N-1 SoC array cards' management information as management group information and dynamically maintain the management group information: the management information of each SoC array card is sent to other N-1 SoC array cards through a heartbeat mode at a certain time interval, and each SoC array card updates the recorded management group information if it finds that the management information of itself or other SoC array cards has changed.
5. The distributed, centralized management system based on SoC array servers according to claim 4, characterized in that: If the master fails during the operation of the node, i.e., the slave cannot receive the heartbeat packet from the master within a certain time interval or the slave finds that the running state in the management information of the master is abnormal, the election strategy is triggered again to select a new master.
Citation Information
Patent Citations
Multi-interface management network architecture for cloud server
CN104486130A
Node election method and device
CN107579860A