Edge server system and its control method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-15
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]本发明实施例提供一种边缘服务器系统,旨在解决现有技术的机架式服务器中所存在的体积庞大、部署环境苛刻、功耗与噪音高、网络配置复杂的技术问题
1、紧凑与家居化:采用多块核心板构成核心板集群,替代机架式服务器中的处理器模块,系统的整体体积相比现有2U服务器较小,可直接放置于桌面、电视柜等家用或办公环境,实现“即插即用”,无需专业机房;
Smart Images

Figure CN122569693A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of server system technology, and in particular relates to an edge server system and its control method. Background Technology
[0002] Existing edge computing or cloud phone services typically rely on large rack-mount servers (such as 2U servers). These servers have built-in dedicated Baseboard Management Controller (BMC) chips for centralized monitoring and management of dozens or even hundreds of computing nodes (such as power on / off and status queries). However, the existing server solutions have the following significant drawbacks: 1. Large size and demanding deployment environment: It requires professional server racks and is equipped with facilities such as air conditioning and UPS. It cannot be deployed in non-professional environments such as homes or ordinary offices. 2. High power consumption and noise: The power consumption of the whole machine usually exceeds 300W, and it relies on high-speed fans for forced heat dissipation, which is very noisy and not suitable for quiet environments; 3. Complex network configuration: The computing nodes and the baseboard management controller chip are usually located on the same management network, which poses security risks and has a complex network path. Summary of the Invention
[0003] This invention provides an edge server system designed to address the technical problems of existing rack servers, such as large size, demanding deployment environment, high power consumption and noise, and complex network configuration.
[0004] This invention is implemented as follows: an edge server system includes: A chassis, wherein a PCB board is provided in the chassis, and the PCB board is provided with multiple interfaces; A heat dissipation module disposed in the chassis and electrically connected to the PCB board; A network module disposed on the PCB board, the network module being used for external communication; and A core board cluster is disposed on the PCB board, the core board cluster including a management core board electrically disposed on the PCB board and multiple computing core boards; the management core board is communicatively connected to the network module; the multiple computing core boards are all communicatively connected to the management core board and are evenly distributed on at least two sides of the management core board; the number of computing core boards is 2N, N≥1; The management core board and the multiple computing core boards are located in different network subnets. The management core board is used to distribute address information, management instructions and external data input through the network module to the multiple computing core boards. The computing core boards are used to perform corresponding operations and calculations according to the management instructions and the external data.
[0005] In one embodiment, the housing includes a front panel and a rear panel disposed opposite to each other, the front panel having a first ventilation hole structure and the rear panel having a second ventilation hole structure; The projection of the first ventilation hole structure in the direction from the front panel to the rear panel covers the core plate cluster; and / or The projection of the second ventilation hole structure in the direction from the rear panel to the front panel covers the core plate cluster.
[0006] In one embodiment, the heat dissipation module includes multiple cooling fans and multiple heat sinks. The multiple cooling fans are at least disposed on the inner side of the front panel to create a cooling airflow to be output to the outside of the chassis. The multiple heat sinks are respectively disposed on the top of the multiple computing core boards and cover the corresponding computing core boards.
[0007] In one embodiment, the PCB board is provided with multiple slots, and the management core board and multiple computing core boards are detachably disposed on the multiple slots.
[0008] In one embodiment, the network module is a physical network port integrated on the rear side of the management core board, which communicates with multiple computing core boards through an internal network topology.
[0009] The present invention also provides a control method for an edge server system, applied to the aforementioned edge server system, the control method comprising: The management core board is controlled to start first and load a simplified management firmware. After a first set time, multiple computing core boards are controlled to start. The management core board is controlled to discover the communication protocol through internal network broadcast, automatically identify all online computing core boards, and allocate internal address information and basic configuration files to each computing core board. The simplified management firmware includes network and basic I / O drivers. The management core board is controlled to continuously listen to the heartbeat packets periodically sent by each computing core board, and actively poll the key indicators of each computing core board, including but not limited to computing indicators, storage indicators and temperature indicators. The control system sends management commands to one or more computing core boards according to an internal address mapping table, employs different broadcast methods depending on the number of computing core boards communicating, and controls the computing core boards receiving the management commands to respond to them. The internal address mapping table is used to record the correspondence between the MAC addresses and physical locations of each computing core board; and When a computing core board fails, the management core board is controlled to mark the failed computing core board as offline, troubleshoot the fault to determine the cause of the fault, and perform corresponding processing operations based on the cause of the fault.
[0010] In one embodiment, the control of the management core board automatically identifies information about all online computing core boards by broadcasting a communication protocol over an internal network, and assigns internal address information and basic configuration files to each computing core board, including: The management core board is controlled to scan the presence signals of all slots used to install the computing core board through the I2C bus of the PCB board, identify the number of inserted computing core boards and physical slot IDs in real time, and read the hardware identity information of each computing core board. The management core board is controlled to dynamically generate an internal network topology based on the identification results and assign fixed MAC addresses to each computing core board. During network initialization, the management core board is controlled to actively detect IP conflicts and automatically reassign IP addresses to each computing core board when a conflict occurs; and The management core board controls the batch distribution of basic configuration files to each computing core board through the internal network topology.
[0011] In one embodiment, the step of using different broadcast methods based on the number of communicating computing core boards, and controlling the computing core boards that receive management commands to respond to the management commands, includes: When the management core board communicates with a specific computing core board, the response is handled using unicast. When the management core board communicates with multiple computing core boards to enable collaborative operation of the multiple computing core boards, the management core board is controlled to process responses using a multi-threaded concurrent approach; and The computing core board that controls the communication responds to the management instructions sent by the management core board and executes the corresponding operations. After the operations are completed, it sends an acknowledgment response to the management core board.
[0012] In one embodiment, troubleshooting to determine the cause of the fault and performing corresponding processing operations based on the cause of the fault includes: If the cause of the failure is a temporary process failure, the management core board is controlled to send a remote command through the internal network topology to control each computing core board to automatically restart the currently failed business process a set number of times, with a second set time interval between each restart. The failure event is recorded after the business process is successfully restored. If the fault is a system-level fault, and restarting the business process is ineffective, the management core board is controlled to trigger a soft restart of each of the computing core boards; and If the cause of the fault is a hardware failure, the management core board will isolate the faulty computing core board from the network and mark it as a hardware failure, prompting the administrator to query the specific faulty computing core board through a specified command.
[0013] In one embodiment, the control method for the edge server system further includes: The management core board and each of the computing core boards perform hardware-level signature verification during startup; and All management traffic between the management core board and each of the computing core boards is transmitted using a lightweight encryption protocol.
[0014] The edge server system of this invention has the following beneficial effects: 1. Compact and Home-Friendly: It uses multiple core boards to form a core board cluster, replacing the processor module in the rack server. The overall size of the system is smaller than that of the existing 2U server. It can be placed directly on the desktop, TV cabinet and other home or office environments, achieving "plug and play" without the need for a professional data center. 2. Significantly Reduced Power Consumption and Noise: Since the management core board is only responsible for management tasks, its power consumption and heat generation are low. The computing core boards are distributed around the management core board. This layout disperses the high-heat-generating computing units and places the low-heat-generating control units in the center, creating an optimized heat dissipation airflow that promotes overall thermal balance. This structural layout results in lower typical system power consumption and even lower standby power consumption. It allows for the use of low-requirement cooling modules combined with optimized airflow, generating significantly less noise than the high-power fans of existing rack-mount servers, meeting the requirements for a quiet environment. 3. Cost advantage: Replacing the dedicated BMC chip in existing rack servers with a standard management core board effectively reduces hardware costs; 4. Security and Network Simplification: By defining management core boards and computing core boards, strict network isolation and routing are achieved, ensuring that computing nodes are not directly exposed to the external network, thus improving security. Furthermore, with only one network module serving as the data interface, network configuration on the user side is greatly simplified. 5. High reliability design: The functions of the management core board and the computing core board are separated. When the management core board fails, it will not cause service interruption (only the computing core board will be unmanaged). The reasonable layout and heat dissipation design ensure the stability of long-term operation. Attached Figure Description
[0015] Figure 1 This is a three-dimensional schematic diagram of the edge server system provided in an embodiment of the present invention; Figure 2 This is another perspective view of the edge server system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the PCB board and core board cluster provided in an embodiment of the present invention; Figure 4 This is an assembly diagram of the chassis and cooling fan provided in an embodiment of the present invention; Figure 5 This is a flowchart illustrating the control method of the edge server system provided in an embodiment of the present invention.
[0016] Explanation of key component symbols: Edge server system-100; Chassis-10; Front panel-11; First ventilation hole structure-111; Rear panel-12; Second ventilation hole structure-121; PCB board-20; Cooling fan-30; Core board cluster-40; Management core board-41; Computing core board-42. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the invention, and should not be construed as limiting the invention. Furthermore, it should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0018] In the description of this invention, it should be understood that the orientation or positional relationship indicated in the description of direction and positional relationship is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing this invention and simplifying the description, and is not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0019] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0020] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention.
[0021] Furthermore, the present invention may repeat reference numerals and / or reference letters in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or settings discussed. In addition, the present invention provides examples of various specific processes and materials, but those skilled in the art will recognize the application of other processes and / or the use of other materials.
[0022] Please see Figures 1 to 4 The edge server system 100 of this embodiment includes: The enclosure 10 contains a PCB board 20, which has multiple interfaces. A heat dissipation module is installed in the chassis 10 and electrically connected to the PCB board 20; A network module is mounted on PCB board 20, and the network module is used for external communication; and A core board cluster 40 is disposed on PCB board 20. The core board cluster 40 includes a management core board 41 electrically disposed on PCB board 20 and multiple computing core boards 42. The management core board 41 is communicatively connected to the network module. The multiple computing core boards 42 are all communicatively connected to the management core board 41 and are evenly distributed on at least two sides of the management core board 41. The number of computing core boards 42 is 2N, where N≥1. The management core board 41 and the multiple computing core boards 42 are located in different network subnets. The management core board 41 is used to distribute address information, management commands and external data input through the network module to the multiple computing core boards 42. The computing core boards 42 are used to perform corresponding operations and calculations according to the management commands and external data.
[0023] The edge server system 100 of this invention has the following beneficial effects: 1. Compact and Home-Friendly: The system uses multiple core boards to form a core board cluster 40, replacing the processor module in the rack server. The overall size of the system is smaller than that of the existing 2U server, and it can be placed directly on the desktop, TV cabinet and other home or office environments, achieving "plug and play" without the need for a professional server room. 2. Significantly Reduced Power Consumption and Noise: Since the management core board 41 is only responsible for management tasks, its power consumption and heat generation are low. The computing core boards 42 are distributed around the management core board 41. This layout disperses the high-heat-generating computing units and places the low-heat-generating control units in the center, forming an optimized heat dissipation airflow that promotes overall thermal balance. Through this structural layout, the system's typical power consumption is lower, and its standby power consumption is even lower. It can use low-requirement cooling modules combined with optimized airflow, resulting in noise levels far lower than the high-power fans of existing rack-mount servers, meeting the requirements for a quiet environment. 3. Cost advantage: The adoption of the standard management core board 41 replaces the dedicated BMC chip in the existing rack server, effectively reducing hardware costs; 4. Security and Network Simplification: By defining the management core board 41 and the computing core board 42, strict network isolation and routing are achieved, ensuring that computing nodes are not directly exposed to the external network, thus improving security. Furthermore, with only one network module serving as the data interface, network configuration on the user side is greatly simplified. 5. High reliability design: The functions of the management core board 41 and the computing core board 42 are separated. When the management core board 41 fails, it will not cause service interruption (only the computing core board 42 will be out of control). The reasonable layout and heat dissipation design ensure the stability of long-term operation.
[0024] Specifically, in this embodiment of the invention, the chassis 10 is a relatively flat metal box, which makes the edge server system 100 roughly in a relatively flat box shape, effectively controlling its overall size. Therefore, there is no need to configure a professional server room rack for the edge server system 100, and it can be deployed in non-professional environments such as homes or ordinary offices.
[0025] Moreover, the chassis 10 is designed for quick disassembly. Its top cover and the front, back, left and right panels are detachably connected by screws. Only a few screws need to be unscrewed to open the cover and directly expose the management core board 41 and computing core board 42 inside, which is convenient for assembly, disassembly and maintenance.
[0026] In this embodiment of the invention, the PCB board 20 can be fixed to the bottom plate of the housing 10 by screws, and multiple interfaces can be arranged in a row at the rear edge of the PCB board 20. Please see Figure 3 On the right side, for example, multiple interfaces may include a power interface (12V power supply), an LED light interface, a MASKROM interface, a LOADER interface, an HDMI interface, a BMC debugging port, a Type-C interface, dual RJ45 interfaces, a USB 3.0 interface, an Audio Codec interface, and two dual USB 2.0 interfaces, etc.
[0027] The edge server system 100 of this embodiment is only a fraction of the size of a traditional 2U server, and the typical power consumption of the whole machine is about 130W. It is relatively small in size and relatively low in power consumption. Therefore, the heat dissipation module may include multiple small-sized cooling fans 30 to replace the large cooling fans 30 in traditional servers, which can effectively reduce noise and reduce the space occupied by the chassis 10.
[0028] In this embodiment of the invention, the PCB board 20 is provided with multiple slots, and the management core board 41 and multiple computing core boards 42 are detachably disposed on the multiple slots. The PCB board 20, the management core board 41 and the multiple computing core boards 42 are integrated into a hardware architecture.
[0029] Specifically, multiple slots are used to install board-to-board connectors, that is, the management core board 41 and multiple computing core boards 42 are plugged into the board-to-board connectors, so as to be electrically and detachably installed on the PCB board 20, which facilitates the assembly, disassembly and maintenance of the core board cluster 40.
[0030] In addition, each slot has a corresponding physical slot ID, which is associated with the type and serial number of each computing core board 42, ensuring that the management core board 41 can accurately identify the computing core board 42.
[0031] In one embodiment, all core boards adopt the ARM architecture. One core board is defined as the management core board 41 (i.e., the BMC board), which may be model RK3588S. The remaining core boards are defined as the computing core boards 42, which may be the same model as the management core board 41 or other models.
[0032] In another embodiment, the management core board 41 may adopt a different architecture than the computing core board 42, such as an x86 architecture, and be dedicated to management routing functions.
[0033] The management core board 41 and each computing core board 42 are interconnected through an internal gigabit Ethernet switching network to ensure data transmission speed. Each computing core board 42 has its own IP address and MAC address (i.e., address information assigned by the management core board 41), and also its own number, such as 1, 2, 3, 4... or a, b, c, d... etc. For example, it can be called computing core board 1, computing core board 2, computing core board 3, computing core board 4...
[0034] The operation of the edge server system 100 is managed and controlled by the management core board 41. Management commands and external data are received and executed by the computing core board 42. The management core board 41 and the multiple computing core boards 42 are logically located in two different network subnets. All computing core boards 42 cannot directly access the external network. External data must enter through the network module - management core board 41. The management core board 41 performs routing forwarding and firewall filtering to achieve network isolation and improve security.
[0035] In one embodiment, the network module is a physical network port integrated on the rear side of the management core board 41, which communicates with multiple computing core boards 42 through the internal network topology.
[0036] The network module is used to connect the edge server system 100 to the external network, and the management core board 41 acts as a router. It can assign IP addresses to all computing core boards 42 and forward data through the internal network topology. Since the network module is the only network exit point of the edge server system 100, it can greatly simplify the network configuration on the user side.
[0037] In this embodiment of the invention, the management core board 41 is located in the middle of the PCB board 20, and multiple computing core boards 42 are distributed at least to both sides of the management core board 41. The core board cluster 40 is a physical layout design with the management core board 41 in the center and multiple computing core boards 42 surrounding it. This disperses the computing units with high heat sources and places the control units with low heat sources in the center, which is beneficial to the overall thermal balance.
[0038] The number of core boards 42 is calculated to be 2N, where N≥1, for example: When N=1, the two computing core boards 42 are respectively set on the left and right sides of the management core board 41; When N=2, two computing core boards 42 are respectively set on the left and right sides of the management core board 41, or four computing core boards 42 are respectively set around the management core board 41. When N=3, the three computing core boards 42 are respectively set on the left and right sides of the management core board 41...; When N≥2, the computing core boards 42 on the same side can be arranged in a row, a column, or a compact arrangement.
[0039] Please see Figure 3 In one embodiment, N equals 4, and the number of computing core boards 42 is 8. Four computing core boards 42 are evenly distributed on both sides of the management core board 41. The four computing core boards 42 on the same side are arranged in a grid pattern to ensure the compactness of the core board cluster 40, thereby controlling the size of the edge server system 100.
[0040] It is understandable that a total of 8 computing core boards 42 can ensure that the system has sufficiently high computing performance while controlling the system's hardware cost, and at the same time, the system's heat dissipation burden is also within a controllable range.
[0041] Of course, the number of computing core boards 42 is not limited to 8. It can be adjusted to 4, 6 or 12 depending on the size of the chassis 10, the system heat dissipation capacity and the system computing performance requirements. The specific settings can be configured according to specific needs.
[0042] Furthermore, the edge server system 100 also includes a power module located on the rear side of the PCB board 20, which can be input using a standard household three-prong AC power outlet. After the power module converts AC to DC, it independently supplies power to the management core board 41, multiple computing core boards 42, and heat dissipation modules through the PCB board 20.
[0043] Please see Figure 3 and Figure 4 In one embodiment, the housing 10 includes a front panel 11 and a rear panel 12 disposed opposite to each other. The front panel 11 is provided with a first ventilation hole structure 111, and the rear panel 12 is provided with a second ventilation hole structure 121. The projection of the first ventilation hole structure 111 in the direction from the front panel 11 to the rear panel 12 covers the core plate cluster 40; and / or The projection of the second ventilation hole structure 121 in the direction from the rear panel 12 to the front panel 11 covers the core plate cluster 40.
[0044] Specifically, in the width direction of the chassis 10, the first ventilation hole structure 111 and the second ventilation hole structure 121 cover the core board cluster 40. The first ventilation hole structure 111 and the second ventilation hole structure 121 include a plurality of densely arranged ventilation holes to ensure that the position of the core board cluster 40 can be completely covered, and an effective heat dissipation airflow channel is established. When the airflow flows in the chassis 10, effective heat dissipation is achieved.
[0045] Please see Figure 3 and Figure 4 In one embodiment, the heat dissipation module includes multiple cooling fans 30 and multiple heat sinks. The multiple cooling fans 30 are at least disposed on the inner side of the front panel 11 to create a cooling airflow to be output to the outside of the chassis 10. The multiple heat sinks are respectively disposed on the top of multiple computing core boards 42 and cover the corresponding computing core boards 42.
[0046] Specifically, multiple cooling fans 30 are arranged in a row on the inner side of the front panel 11, and the area covered is roughly equal to the area covered by the first ventilation hole structure 111. The structural layout of the core board cluster 40, combined with the multiple front-mounted cooling fans 30, forms a "center-radial" heat dissipation airflow.
[0047] Cool air is drawn in from the rear panel 12 of the chassis 10 by the cooling fan 30, which in turn washes over the peripheral computing core board 42 and the central management core board 41. Finally, hot air is exhausted from the front panel 11 of the chassis 10, thus achieving effective heat dissipation for the system.
[0048] It is understandable that since the computing core board 42 is used for the specific calculation process of the system, its heat generation is higher than that of the management core board 41. Therefore, a removable heat sink is set on the top of the computing core board 42. After the heat sink absorbs heat, it dissipates heat through the heat dissipation channel, which can improve the heat dissipation effect of the computing core board 42 and further ensure the stability of the system.
[0049] In one embodiment, several cooling fans 30 can also be added to the inner side of the left and right panels of the chassis 10, so that the three sides of the chassis 10 can dissipate heat, further improving the heat dissipation performance of the system.
[0050] In one embodiment, the PCB board 20 may also be equipped with several temperature sensors for detecting the temperature of the core board cluster 40. These temperature sensors are electrically connected to the management core board 41 to transmit the collected temperature data. Based on the real-time temperature feedback from the temperature sensors, the management core board 41 can dynamically adjust the speed of different fans among the multiple cooling fans 30, making the operation of the cooling fans 30 more closely match the system's temperature changes, further optimizing noise and energy consumption.
[0051] Please see Figures 1 to 5 The control method for the edge server system of this invention, applied to the edge server system 100 in any of the above embodiments, includes the following steps: S10: The control and management core board starts up first and loads the simplified management firmware. After a first set time, it controls multiple computing core boards to start up. The control and management core board discovers the communication protocol through internal network broadcast, automatically identifies all online computing core boards, and assigns internal address information and basic configuration files to each computing core board. The simplified management firmware includes network and basic I / O drivers.
[0052] Step S10 is the system startup and self-test process. Specifically, when the system is powered on and starts working, the management core board is powered on to start up and load the simplified management firmware. Since the simplified management firmware only includes network and basic I / O drivers, the management core board can speed up the startup speed while ensuring the availability of basic functions and quickly ensure the availability of management channels.
[0053] After the management core board starts up for the first set duration (e.g., 150ms), the management channel is now available, and multiple compute core boards are then powered on to start. The management core board discovers the communication protocol via internal network broadcast and begins communication, thereby identifying all online compute core boards. To ensure that the management core board can accurately identify and locate each compute core board, it assigns internal address information (e.g., IP address and MAC address) to each compute core board. To ensure that the compute core boards can be quickly put into use, a basic configuration file is sent to the compute core boards to guarantee their basic functions, enabling plug-and-play functionality.
[0054] S20: The control and management core board continuously listens to the heartbeat packets sent periodically by each computing core board and actively polls the key indicators of each computing core board. Key indicators include, but are not limited to, computing indicators, storage indicators, and temperature indicators.
[0055] Step S20 is the system status monitoring process. The calculation indicators may include CPU / memory load, process health, etc. The storage indicators may include the remaining lifespan of the eMMC / SD cards of the management core board and the computing core board, read / write error count, etc. The temperature indicators may include the temperature of the onboard sensors, etc.
[0056] The management core board continuously monitors the heartbeat packets of the computing core boards and actively polls their key indicators to ensure a clear understanding of their working and usage status, facilitating accurate control over each computing core board.
[0057] S30: The control management core board sends management commands to one or more computing core boards according to the internal address mapping table. It adopts different broadcast methods according to the number of computing core boards being communicated, and controls the computing core boards that receive the management commands to respond to the management commands. The internal address mapping table is used to record the correspondence between the MAC address and physical location of each computing core board.
[0058] Step S30 is the system's instruction execution flow, including broadcast mode and response mechanism. Specifically, during normal operation, the management core board maintains an internal address mapping table so that communication can be performed quickly in a short time when needed. The internal address mapping table is generated by the management core board based on the generated MAC address of each computing core board and the obtained physical location of each computing core board (i.e., the location of the slot where each computing core board is located).
[0059] It's understandable that not all compute core boards need to be operational during normal system operation. The management core board communicates with one or more compute core boards depending on the task. Some tasks may only require a single compute core board to work. In this case, the management core board can communicate with a single target compute core board via a single broadcast to reduce network traffic. Other tasks may require multiple compute core boards to work collaboratively. In this case, the management core board can communicate with multiple target compute core boards via parallel broadcasts to improve management efficiency.
[0060] Moreover, only the computing core board that communicates with the management core board to perform its work will respond to management commands and execute operations, effectively reducing the traffic load on the internal network.
[0061] S40: When a computing core board fails, the control and management core board marks the failed computing core board as offline, performs troubleshooting to determine the cause of the failure, and performs corresponding processing operations based on the cause of the failure.
[0062] Step S40 is the system fault handling process. The fault causes may include temporary process faults, system-level faults, and hardware faults. The handling operations include process restart, core board restart, core board maintenance and replacement, and alarms.
[0063] It should be noted that a failure of the management core board itself will not affect the services on the computing core board that is already running normally, but it will lose the ability to control the computing core board and interrupt its access to the external network. Therefore, when the management core board detects its own failure, it can alert the administrator through the system's management interface to manually intervene and restore the system to normal operation as soon as possible.
[0064] In one embodiment, the PCB board may also be provided with several temperature sensors for detecting the temperature of the core board cluster. The several temperature sensors are electrically connected to the management core board to transmit the collected temperature to the management core board.
[0065] Based on the above structure, the control method may further include the following steps: the control management core board dynamically adjusts the speed of different fans among multiple cooling fans according to real-time temperature feedback from several temperature sensors. This makes the operation of the cooling fans more closely match the temperature changes of the system, further optimizing noise and energy consumption.
[0066] The control method of the edge server system in this embodiment of the invention is implemented by the management core board, which controls the operation of itself, all computing core boards and other functional modules of the system.
[0067] In one embodiment, step S10, which involves the control management core board automatically identifying information about all online computing core boards by broadcasting a communication protocol over the internal network, and allocating internal address information and basic configuration files to each computing core board, includes the following steps: S101: The control and management core board scans the presence signals of all slots used to install computing core boards through the I2C bus of the PCB board, identifies the number of inserted computing core boards and their physical slot IDs in real time, and reads the hardware identity information of each computing core board.
[0068] It is understandable that when the computing core board is installed on the PCB, there will be a corresponding presence signal on the corresponding slot to indicate that the computing core board has been inserted. The management core board uses this to identify the number of computing core boards inserted and the ID of the slot. The hardware identity information may include the core board type, serial number, etc. The physical slot ID is associated with the core board type, serial number, etc., rather than simply identifying the number of computing core boards. This is how to achieve accurate identification of the location of the computing core board.
[0069] S102: The control and management core board dynamically generates the internal network topology based on the identification results and assigns fixed MAC addresses to each computing core board.
[0070] The fixed MAC address is generated based on the hardware ID of each computing core board to ensure that it remains unchanged after the computing core board is restarted.
[0071] S103: During network initialization, the control management core board actively detects IP conflicts and automatically reassigns IP addresses to each computing core board when a conflict occurs.
[0072] By managing the core board to detect whether there are IP address conflicts among the various computing core boards, and reallocating IP addresses when conflicts occur, the system ensures accurate identification of each computing core board and thus guarantees the orderly operation of each computing core board.
[0073] S104: The control and management core board distributes basic configuration files to each computing core board in batches through the internal network topology.
[0074] By distributing basic configuration files to all computing core boards through the management core board, the computing core boards are guaranteed to be available immediately, ensuring basic computing capabilities and achieving "plug and play".
[0075] In one embodiment, step S30, which involves using different broadcast methods based on the number of communicating computing core boards and controlling the computing core boards that receive management commands to respond to the management commands, includes the following steps: S301: When the management core board communicates with a specific computing core board, the response is handled using unicast.
[0076] When the management command issued by the management core board is only for a specific computing core board (such as "restart computing core board 5"), the unicast method is used to process the response, which can reduce network traffic by 85% and effectively reduce system power consumption and load.
[0077] S302: When the management core board communicates with multiple computing core boards to enable the multiple computing core boards to work together, the control management core board uses a multi-threaded concurrent approach to process responses.
[0078] When the management commands issued by the management core board require the collaboration of multiple computing core boards (such as "synchronizing time"), the management core board uses a multi-threaded concurrent approach to process the responses, which greatly improves management efficiency.
[0079] S303: The computing core board that controls communication responds to the management commands sent by the management core board and executes the corresponding operations. After execution, it sends an acknowledgment response to the management core board.
[0080] Specifically, only the target computing core board will respond to management commands and execute operations. After execution, it will send an acknowledgment response to the management core board, which can significantly reduce the management traffic load of the internal network.
[0081] Meanwhile, the control method of this invention also supports setting timed tasks and management strategies, and the management core board executes tasks on all computing core boards according to the task schedule.
[0082] In one embodiment, step S40, which involves troubleshooting to determine the cause of the fault and performing corresponding processing operations based on the cause, includes the following steps: S401: If the cause of the fault is a temporary process failure, the control management core board sends a remote command through the internal network topology to control each computing core board to automatically restart the faulty service process a set number of times, with a second set time interval between each restart. After the service process is successfully restored, the fault event is recorded.
[0083] Specifically, if the management core board detects a system failure such as process freezing or a sudden high load, it determines the cause to be a temporary fault. The management core board then sends a remote command to each computing core board via the internal network topology, triggering the computing core boards to automatically restart the currently faulty service process a set number of times (e.g., 3 times), with each restart spaced at a second set interval (e.g., 200ms). This attempts to repair the temporary process failure. If the service process restarts successfully and recovers, the fault event is recorded, indicating the fault is resolved and no alarm is needed in the management interface. However, if the service process restart fails, an alarm is triggered in the management interface to alert the administrator, and the process proceeds to the next step.
[0084] S402: If the cause of the fault is a system-level fault, when restarting the business process is ineffective, the control management core board will trigger a soft restart of each computing core board.
[0085] Specifically, if the management core board detects a system failure such as an operating system crash or memory leak, it determines that the failure is a system-level failure. If restarting the business process in step S401 cannot resolve the current failure, the management core board controls all computing core boards to perform a soft restart in order to preserve memory data and prevent the loss or damage of memory data.
[0086] S403: If the cause of the fault is a hardware fault, the control and management core board will isolate the faulty computing core board network and mark it as a hardware fault, reminding the administrator to query the specific faulty computing core board through a specified command.
[0087] Specifically, if the management core board detects a fault such as eMMC damage or hardware error in the system, it determines that the cause of the fault is a hardware failure. At this time, the management core board isolates the faulty computing core board through the internal switching chip ACL rules to prevent it from affecting other computing core boards. At the same time, it activates the system's fault indicator lights or alarms to remind the administrator to query the specific faulty computing core board through the SSH command and guide the administrator to quickly replace the faulty computing core board through the management interface.
[0088] In one embodiment, the control method for the edge server system further includes the steps of: S50: The control and management core board and each computing core board perform hardware-level signature verification during startup.
[0089] Step S50 is the system's security assurance process. Through hardware-level signature authentication, it ensures that the firmware running on the management core board and the computing core board has not been tampered with, thereby enhancing security.
[0090] S60: All management traffic between the management core board and each computing core board is transmitted via a lightweight encryption protocol.
[0091] Step S60 is the system's security assurance process. Management traffic may include heartbeat packets, control commands, and operation logs. Lightweight encryption protocols may include DTLS, MQTT, and AES-128 protocols. By transmitting management traffic through lightweight encryption protocols, security is ensured while minimizing resource overhead, preventing internal network eavesdropping, and guaranteeing the security of system data transmission.
[0092] In this specification, the terms "in one embodiment," "in another embodiment," etc., refer to specific features, structures, materials, or characteristics described in connection with embodiments or examples that are included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiments or examples. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0093] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An edge server system, characterized in that, include: A chassis, wherein a PCB board is provided in the chassis, and the PCB board is provided with multiple interfaces; A heat dissipation module disposed in the chassis and electrically connected to the PCB board; A network module disposed on the PCB board, the network module being used for external communication; and A core board cluster is disposed on the PCB board, the core board cluster including a management core board electrically disposed on the PCB board and multiple computing core boards; the management core board is communicatively connected to the network module; the multiple computing core boards are all communicatively connected to the management core board and are evenly distributed on at least two sides of the management core board; the number of computing core boards is 2N, N≥1; The management core board and the multiple computing core boards are located in different network subnets. The management core board is used to distribute address information, management instructions and external data input through the network module to the multiple computing core boards. The computing core boards are used to perform corresponding operations and calculations according to the management instructions and the external data.
2. The edge server system according to claim 1, characterized in that, The housing includes a front panel and a rear panel arranged opposite to each other. The front panel is provided with a first ventilation hole structure, and the rear panel is provided with a second ventilation hole structure. The projection of the first ventilation hole structure in the direction from the front panel to the rear panel covers the core plate cluster; and / or The projection of the second ventilation hole structure in the direction from the rear panel to the front panel covers the core plate cluster.
3. The edge server system according to claim 2, characterized in that, The heat dissipation module includes multiple cooling fans and multiple heat sinks. The multiple cooling fans are at least located on the inner side of the front panel to create a cooling airflow to the outside of the chassis. The multiple heat sinks are respectively located on the top of the multiple computing core boards and cover the corresponding computing core boards.
4. The edge server system according to claim 1, characterized in that, The PCB board has multiple slots, and the management core board and multiple computing core boards are detachably mounted on the multiple slots.
5. The edge server system according to claim 1, characterized in that, The network module is a physical network port integrated on the rear side of the management core board, which communicates with multiple computing core boards through an internal network topology.
6. A control method for an edge server system, applied to the edge server system according to any one of claims 1 to 5, characterized in that, The control method for the edge server system includes: The management core board is controlled to start first and load a simplified management firmware. After a first set time, multiple computing core boards are controlled to start. The management core board is controlled to discover the communication protocol through internal network broadcast, automatically identify all online computing core boards, and allocate internal address information and basic configuration files to each computing core board. The simplified management firmware includes network and basic I / O drivers. The management core board is controlled to continuously listen to the heartbeat packets periodically sent by each computing core board, and actively poll the key indicators of each computing core board, including but not limited to computing indicators, storage indicators and temperature indicators. The control system sends management commands to one or more computing core boards according to an internal address mapping table, employs different broadcast methods depending on the number of computing core boards communicating, and controls the computing core boards receiving the management commands to respond to them. The internal address mapping table is used to record the correspondence between the MAC addresses and physical locations of each computing core board; and When a computing core board fails, the management core board is controlled to mark the failed computing core board as offline, troubleshoot the fault to determine the cause of the fault, and perform corresponding processing operations based on the cause of the fault.
7. The control method for the edge server system according to claim 6, characterized in that, The control management core board discovers communication protocols via internal network broadcast, automatically identifies information about all online computing core boards, and assigns internal address information and basic configuration files to each computing core board, including: The management core board is controlled to scan the presence signals of all slots used to install the computing core board through the I2C bus of the PCB board, identify the number of inserted computing core boards and physical slot IDs in real time, and read the hardware identity information of each computing core board. The management core board is controlled to dynamically generate an internal network topology based on the identification results and assign fixed MAC addresses to each computing core board. During network initialization, the management core board is controlled to actively detect IP conflicts and automatically reassign IP addresses to each computing core board when a conflict occurs; and The management core board controls the batch distribution of basic configuration files to each computing core board through the internal network topology.
8. The control method for the edge server system according to claim 6, characterized in that, The method of using different broadcast methods according to the number of the communicating computing core boards, and controlling the computing core boards that receive management commands to respond to the management commands, includes: When the management core board communicates with a specific computing core board, the response is handled using unicast. When the management core board communicates with multiple computing core boards to enable collaborative operation of the multiple computing core boards, the management core board is controlled to process responses using a multi-threaded concurrent approach; and The computing core board that controls the communication responds to the management instructions sent by the management core board and executes the corresponding operations. After the operations are completed, it sends an acknowledgment response to the management core board.
9. The control method for the edge server system according to claim 6, characterized in that, The process of troubleshooting to determine the cause of the fault and performing corresponding processing operations based on the cause includes: If the cause of the failure is a temporary process failure, the management core board is controlled to send a remote command through the internal network topology to control each computing core board to automatically restart the currently failed business process a set number of times, with a second set time interval between each restart. The failure event is recorded after the business process is successfully restored. If the fault is a system-level fault, and restarting the business process is ineffective, the management core board is controlled to trigger a soft restart of each of the computing core boards; and If the cause of the fault is a hardware failure, the management core board will isolate the faulty computing core board from the network and mark it as a hardware failure, prompting the administrator to query the specific faulty computing core board through a specified command.
10. The control method for the edge server system according to claim 6, characterized in that, The control method for the edge server system also includes: The management core board and each of the computing core boards perform hardware-level signature verification during startup; and All management traffic between the management core board and each of the computing core boards is transmitted using a lightweight encryption protocol.