Edge server management method and device, electronic equipment and storage medium

Through the master-slave control unit architecture and multi-link communication, the problem of insufficient adaptability of BMC management chips is solved, and the efficient, stable operation and fault self-healing capabilities of edge servers are achieved. It is suitable for a variety of CPU architectures and has the advantages of energy saving.

CN120407349AInactive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510872967.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing multi-node edge servers, the programmability of the BMC management chip is not high and cannot be adapted to CPUs other than Intel X86 architecture, resulting in frequent operational failures and cannot meet the requirements of high efficiency and high real-time.

Method used

The master-slave control unit architecture adopts the master control unit, and the master control unit communicates with multiple slave control units through multiple links and supplies power independently. The master control unit sends data acquisition tasks to the slave control unit. If the response feedback information is not received, the fault slave control unit will be taken over by the adjacent slave control unit, collect data and report it, and layered management and fault alarm will be realized.

Benefits of technology

It realizes adaptation to various CPU architectures, ensures the communication stability and efficient operation of edge servers, has significant energy saving effects, and can maintain normal operation when node failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407349A_ABST
    Figure CN120407349A_ABST
Patent Text Reader

Abstract

The invention discloses an edge server management method and device, electronic equipment and a storage medium, an edge server comprises a master control unit and a plurality of target nodes, each target node comprises a processor and a slave control unit, and the master control unit and the slave control unit communicate through various links and are independently powered; the slave control unit in each target node can send a response request to the slave control unit in the adjacent target node, and if the first slave control unit does not receive response feedback information corresponding to the response request sent by the second slave control unit within the first preset duration, it is indicated that the second slave control unit may have a fault. The first slave control unit can take over the second slave control unit, collect and report the data of the processor corresponding to the second slave control unit, and give a fault alarm at the same time, thereby achieving the hierarchical management of each node and the processor in each node, and guaranteeing the normal operation of each node and the stable communication of an edge server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of server management, and in particular to an edge server management method, device, electronic device, and storage medium. Background Art

[0002] In a multi-node edge server, each node has a separate Central Processing Unit (CPU), and each CPU also corresponds to a management chip for monitoring and controlling the CPU. At the same time, there is also a Chassis Management Controller (CMC) chip that communicates with the management chip corresponding to each CPU to manage and control the entire multi-node edge server. Currently, most multi-node edge servers use the Baseboard Management Controller (BMC) as the management chip for the CPU and the CMC chip. However, the low programmability of the BMC results in it being only compatible with CPUs of some architectures, so operation failures may occur during operation and maintenance, which cannot meet the requirements of high efficiency and high real-time performance of multi-node edge servers. Summary of the Invention

[0003] This application provides an edge server management method, device, electronic device, and storage medium to at least solve the problem that the control unit set in the current edge server can only be compatible with CPUs of some architectures, may malfunction during operation, and cannot meet the requirements of high efficiency and high real-time performance of multi-node edge servers.

[0004] This application provides an edge server management method. The edge server includes a main control unit and multiple target nodes, and each target node includes a processor and a slave control unit. The main control unit and the slave control unit communicate through multiple links; the main control unit and the slave control unit are independently powered; the method includes: sending data collection tasks to multiple slave control units respectively through the main control unit; In response to the data collection task, sending response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node; If the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within the first preset duration, collecting the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit; Sending the real-time feedback data and a fault warning message to the main control unit through the first slave control unit, and the fault warning message is used to indicate that the second slave control unit has a fault.

[0005] The present application also provides an edge server management device. The edge server includes a main control unit and multiple target nodes. Each target node includes a processor and a slave control unit. The main control unit and the slave control unit communicate through multiple links; the main control unit and the slave control unit are independently powered; the device includes: a transceiver module, configured to send data collection tasks to multiple slave control units respectively through the main control unit; The transceiver module is further configured to, in response to the data collection task, send response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node; An acquisition module, configured to, if the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within the first preset duration, collect the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit; The transceiver module is further configured to send the real-time feedback data and a fault warning message to the main control unit through the first slave control unit, and the fault warning message is used to indicate that the second slave control unit has a fault.

[0006] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above-mentioned edge server management methods when executing the computer program.

[0007] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned edge server management methods are implemented.

[0008] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned edge server management methods are implemented.

[0009] In this application, the edge server includes a main control unit and multiple target nodes. Each target node includes a processor and a slave control unit. The main control unit and the slave control units communicate through multiple links; the main control unit and the slave control units are independently powered; the main control unit can send data collection tasks to multiple slave control units respectively. At this time, in response to the data collection task, the slave control unit in each target node can send a response request to the slave control unit in an adjacent target node. If the first slave control unit does not receive the response feedback information corresponding to the response request from the second slave control unit within the first preset time period, it indicates that the second slave control unit may have failed. Therefore, the first slave control unit can take over the second slave control unit, collect the data of the processor corresponding to the second slave control unit, report it, and issue a fault alarm at the same time. In this solution, through the master-slave hierarchical setting of the control units in the edge server, hierarchical management of each node and the processors in each node is realized, which can be adapted to processors of various architectures and can achieve the effect of energy saving; in addition, when a certain slave control unit fails, the slave control unit of an adjacent node can take over and report data, which can ensure the normal operation of each node. The communication stability of the edge server can be effectively ensured through multi-link communication and independent power supply. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0011] Figure 1 It is the architecture diagram of the edge server provided by the embodiment of the present application; Figure 2 It is the flowchart of a method for managing an edge server provided by the embodiment of the present application Figure 1 ; Figure 3 It is the flowchart of a method for managing an edge server provided by the embodiment of the present application Figure 2 ; Figure 4 It is the schematic diagram of the control unit link provided by the embodiment of the present application; Figure 5 It is the module diagram of the edge server provided by the embodiment of the present application; Figure 6 It is the structural diagram of a device for managing an edge server provided by the embodiment of the present application; Figure 7 It is the structural diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0013] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0015] In an edge multi-node server, each node has a separate CPU, and each CPU also corresponds to a management chip for monitoring and controlling the CPU. At the same time, there is also a CMC chip for communicating with the management chip corresponding to each CPU to manage and control the entire edge server.

[0016] Currently, most multi-node servers use BMC as the management chip for the CPU and the CMC chip. A small number use Field Programmable Gate Array (FPGA) chips for communication and programming optimization. For edge multi-node servers, the programming of BMC is too rigid and the functions are too redundant, lacking flexibility. Generally speaking, BMC can only be applied to CPUs with the Intel series X86 architecture. In an edge server, each node not only has a CPU with the X86 architecture, but also a CPU or GPU with the ARM architecture. Since the programmability of BMC is not high and it cannot be applied to edge servers with the ARM architecture, there is no suitable management chip for controlling the CPU at this time. When special scenarios such as BMC restart occur, since the restart time of BMC is very slow, generally taking 1-2 minutes, it is difficult to meet the requirements of edge servers for high efficiency and high real-time performance.

[0017] To enable those skilled in the art of this technical field to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0018] As Figure 1 shown, Figure 1 is the architecture diagram of the edge server provided by the embodiment of this application. It can be seen that the edge server includes a main control unit and multiple target nodes, and each target node includes a processor and a slave control unit. Figure 1 The number N of target nodes in the edge server is not limited. In addition, the main control unit and the slave control unit can communicate through multiple links. Figure 1 There can be multiple links for communication between the main control unit and each slave control unit. The link can include: SPI, Uart, CAN, I2C, etc., which can be selected according to the actual scenario requirements; and, in addition to the unified power supply of the edge server, the main control unit and the slave control unit are independently powered respectively.

[0019] As Figure 2 shown, Figure 2 is the flowchart of the edge server management method provided by the embodiment of this application. The method can include the following steps: 201. Send data acquisition tasks to multiple slave control units respectively through the main control unit.

[0020] In the embodiment of this application, both the main control unit and the slave control unit can be a microcontroller unit (MCU), and data can be transmitted between the main control unit and each slave control unit. Since during the operation of the edge server, the main control unit needs to timely understand the operation status and operation data of the processors in each node, the main control unit can send data acquisition tasks to multiple slave control units respectively. Since each target node includes a slave control unit and a processor, after the slave control unit receives the data acquisition task, it can start to collect the relevant data of its corresponding processor.

[0021] 202. In response to the data acquisition task, send response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node.

[0022] In the embodiment of this application, as Figure 1As shown, since data transmission may also be required between the slave control units in each target node, there are also communication links between the respective slave control units. After receiving the data acquisition task, the slave control unit can start collecting relevant data of the processor in the target node where it is located and report it to the master control unit. At the same time, since some of the slave control units may malfunction, these malfunctioning slave control units may be unable to collect data and report it to the master control unit. Therefore, to ensure that the master control unit can receive the relevant data of the processors in each target node, the slave control units in each target node can send response requests to the slave control units in adjacent target nodes. This response request can be used to request the slave control units in adjacent target nodes to provide response feedback, so that the slave control units in each target node can determine whether the slave control units in adjacent target nodes are malfunctioning based on whether they receive the corresponding feedback.

[0023] Exemplarily, as Figure 1 shown, the edge processor includes N target nodes arranged in the illustrated manner. Then, the slave control unit 1 in target node 1 can send a response request to the slave control unit 2 in target node 2. Similarly, the slave control unit 2 in target node 2 can send response requests to both the slave control unit 1 in target node 1 and the slave control unit 3 in target node 3. Similarly, the slave control unit 3 in target node 3 can send response requests to both the slave control unit 2 in target node 2 and the slave control unit 4 (not shown in the figure) in target node 4 (not shown in the figure). The slave control unit N in target node N can send a response request to the slave control unit N - 1 (not shown in the figure) in target node N - 1 (not shown in the figure). Of course, the N target nodes in the edge processor can also be arranged in other ways, which is not specifically limited.

[0024] 203. If the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within the first preset duration, the first slave control unit collects the real-time feedback data of the processor corresponding to the second slave control unit.

[0025] In the embodiment of the present application, after the slave control units in each target node respectively send response requests to the slave control units in adjacent target nodes, they can wait for the response feedback from the slave control units in adjacent target nodes. If the response feedback information corresponding to the response request can be received, it indicates that the slave control unit in the adjacent target node is operating normally, and no other processing is required. If the response feedback information corresponding to the response request is not received, it can be indicated that the slave control unit in the adjacent target node is not operating normally and may have malfunctioned.

[0026] Specifically, the first slave control unit can be the slave control unit in any target node, and the second slave control unit can be the slave control unit in the target node adjacent to the target node where the first slave control unit is located. Additionally, after sending the response request, it is possible that the adjacent target node is too busy to immediately provide a response feedback. Therefore, a certain waiting duration can be set. That is, if the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within the first preset duration, it indicates that the second slave control unit has a fault and is unable to collect the relevant data of the processor corresponding to the second slave control unit and report it to the master control unit. Therefore, the first slave control unit can actively take over the second slave control unit and replace the second slave control unit to collect the real-time feedback data of the processor corresponding to the second slave control unit. Among them, the first preset duration can be set by itself, such as 5s, 10s, etc.

[0027] In some embodiments, the real-time feedback data may include processor data and sensor data. The sensor data may include temperature sensor data, voltage sensor data, current sensor data, etc. The processor data may include CPU load rate, CPU power, CPU turbo frequency status, CPU hardware fault status, memory load, etc.

[0028] 204. Send the real-time feedback data and fault warning information to the master control unit through the first slave control unit.

[0029] In the embodiment of the present application, after the first slave control unit collects the real-time feedback data of the processor corresponding to the second slave control unit, it can report the real-time feedback data of the processor corresponding to the second slave control unit to the master control unit, and at the same time, it can also send a fault warning information, which is used to indicate that the second slave control unit has a fault, so that the master control unit knows that the second slave control unit is currently in a fault state and unable to transmit data.

[0030] In some embodiments, when sending the real-time feedback data through the first slave control unit, the collected real-time feedback data can be packaged. For example, it can be packaged and sent in the form of "header frame + data + check code".

[0031] In some embodiments, when sending the real-time feedback data through the first slave control unit, the collected real-time feedback data can be preprocessed first. Specifically, noise reduction filtering, outlier removal, and pre-calculation can be performed. The pre-calculation can include calculating the temperature change rate based on multiple frames of collected temperature sensor data and calculating the voltage fluctuation amplitude based on multiple frames of collected voltage sensor data. Only the temperature change rate and the voltage fluctuation amplitude are sent to the master control unit, which can reduce the data transmission volume by more than 50%.

[0032] In the embodiment of the present application, the edge server includes a main control unit and multiple target nodes. Each target node includes a processor and a slave control unit. The main control unit and the slave control unit communicate through multiple links; the main control unit and the slave control unit are independently powered; the main control unit can send data collection tasks to multiple slave control units respectively. At this time, in response to the data collection task, the slave control unit in each target node can send a response request to the slave control unit in the adjacent target node. If the first slave control unit does not receive the response feedback information corresponding to the response request from the second slave control unit within the first preset time period, it indicates that the second slave control unit may have failed. Therefore, the first slave control unit can take over the second slave control unit, collect the data of the processor corresponding to the second slave control unit and report it while giving a fault alarm. In this solution, through the master-slave hierarchical setting for the control unit in the edge server, hierarchical management of each node and the processors in each node can be achieved, which can be adapted to processors of various architectures and can achieve the effect of energy saving; in addition, when a certain slave control unit fails, the slave control unit of the adjacent node can take over and report the data, so as to ensure the normal operation of each node. The communication stability of the edge server can be effectively ensured through multi-link communication and independent power supply.

[0033] As Figure 3 shown, Figure 3 FIG. is another flowchart of the edge server management method provided by the embodiment of the present application. The method may include the following steps: 301. After the edge server is powered on, initialize the main control unit and the slave control unit respectively.

[0034] In the embodiment of the present application, after the edge server is powered on, each control unit needs to be initialized before running. The initialization may specifically include: clock frequency initialization, communication interface initialization, General Purpose Input / Output (GPIO) interface initialization, Web service initialization, etc.

[0035] 302. The main control unit monitors the real-time bus load rate and real-time error rate between the main control unit and each slave control unit in real time.

[0036] In the embodiment of the present application, the main control unit and each slave control unit can communicate through multiple links, and each link is applicable to different scenarios. In order to select a suitable link for communication, it can be determined according to the current data transmission requirements. Therefore, the main control unit can monitor the real-time bus load rate and real-time error rate between the main control unit and each slave control unit in real time.

[0037] It should be noted that since the data to be transmitted between the master control unit and each slave control unit may be different, and the scenario requirements may also be different, it is necessary to monitor the real-time bus load rate and real-time error rate for the links between the master control unit and each slave control unit.

[0038] In some embodiments, the real-time bus load rate can be determined according to the actual transmission rate and the bus bandwidth. Specifically, the real-time bus load rate = actual transmission rate / bus bandwidth * 100%; the real-time error rate can be determined according to the number of error data frames and the total number of data frames. Specifically, the real-time error rate = number of error data frames / total number of data frames * 100%.

[0039] 303. Determine the corresponding target bus according to the real-time bus load rate and the real-time error rate.

[0040] In the embodiments of the present application, since the characteristics of different links are different and the applicable scenarios are also different, the current scenario can be determined according to the real-time bus load rate and the real-time error rate, so as to determine the corresponding target bus.

[0041] In some embodiments, common current links may include SPI, Uart, CAN, I2C, etc. As Figure 4 shown, SPI can be applicable to high-speed scenarios, Uart can be applicable to low-power scenarios, CAN can be applicable to strong anti-interference scenarios, and I2C can be applicable to conventional low-speed scenarios. Therefore, the slave control units can transmit data through I2C. Generally speaking, most scenarios default to transmit data through the SPI bus; in addition, if a large amount of data files are transmitted, such as log files stored in the MCU, firmware upgrade files sent by the CMC, etc., the current real-time bus load rate will increase (there is a probability of exceeding 70%), and the real-time error rate may also increase. Then, in order to ensure sequential transmission, a link with strong anti-interference ability is required, so communication can be carried out through the CAN bus; in addition, if it is a low-power scenario and not too many resources are wanted to be consumed, then the Uart bus can be selected; in addition, if the CMC obtains small amounts of data such as the firmware version number from the MCU, the I2C bus can be selected.

[0042] 304. Through the target bus, respectively construct communication links between the master control unit and multiple slave control units.

[0043] In the embodiments of the present application, after determining the target bus according to the current actual scenario requirements, the communication links between the master control unit and each slave control unit can be constructed through the target bus.

[0044] It should be noted that the communication link between the main control unit and each slave control unit is determined according to the real-time data transmission scenario requirements between the main control unit and the slave control unit, and can be adjusted at any time according to the real-time scenario changes. For example, if there is no need to transmit a large amount of data currently, the SPI bus can be adjusted to the I2C bus.

[0045] In the embodiment of the present application, an appropriate link can be selected for transmission according to the real-time data transmission scenario, which can avoid the situation that the data channel does not support transmission or consumes channel resources, effectively save resource consumption, and ensure the normal transmission of data.

[0046] 305. Send data acquisition tasks to multiple slave control units respectively through the main control unit.

[0047] 306. In response to the data acquisition task, send response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node.

[0048] In the embodiment of the present application, for the description of steps 305 to 306, please refer to the detailed description of steps 201 to 202 in the above embodiment, and the embodiments of the present invention will not be repeated here.

[0049] In some embodiments, sending response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node may specifically include: sending response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node according to a preset period.

[0050] It should be noted that in order to know the operating status of the slave control units in adjacent target nodes, the slave control unit may send response requests to the slave control units in adjacent target nodes according to a preset period. For example, the first slave control unit sends a response request to the second slave control unit every 10s or every 30s.

[0051] It should be noted that the slave control units in each target node may send response requests to the slave control units in adjacent target nodes according to a preset period in response to the data acquisition task; when no data acquisition task is received, they may also send response requests to the slave control units in adjacent target nodes according to a preset period.

[0052] In some embodiments, sending response requests from the slave control units in each target node to the slave control units in adjacent target nodes respectively may specifically include: if the master control unit does not receive the feedback message from the slave control unit in the first target node within the second preset duration, the master control unit sends a takeover instruction to the slave control unit in the second target node, where the first target node and the second target node are adjacent nodes; according to the takeover instruction, the slave control unit in the second target node sends a response request to the slave control unit in the first target node.

[0053] It should be noted that the master control unit can also command the slave control unit to take over the slave control unit in its adjacent target node. After the master control unit issues a data collection task, it will receive real-time feedback data reported by multiple slave control units. If it has not received the feedback message from the slave control unit in a certain target node for a certain period of time, then the master control unit can consider that the slave control unit in that target node has failed. At this time, in order to obtain the data of the processor corresponding to the failed slave control unit, it can command the slave control unit in the adjacent target node to take over the failed slave control unit. Among them, the second preset duration can be set by itself, such as: 5s, 10s, etc.

[0054] In the embodiments of the present application, by sending response requests from the slave control units in each target node to the slave control units in adjacent target nodes respectively, it can enable the slave control units in each target node to timely know whether there is a failure in the slave control units in adjacent target nodes, and sending response requests according to a preset period or according to the instruction of the master control unit can take over and report the failure in time when there is a failure in the slave control units in adjacent target nodes, reduce the losses of adjacent target nodes caused by the failure, and ensure the stability of each target node in the edge server.

[0055] 307. If the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within the first preset duration, the first slave control unit collects the real-time feedback data of the processor corresponding to the second slave control unit.

[0056] 308. The first slave control unit sends the real-time feedback data and the fault alarm information to the master control unit.

[0057] In the embodiments of the present application, for the descriptions of steps 307-308, please refer to the detailed descriptions of steps 203-204 in the above embodiments, and the embodiments of the present invention will not be elaborated here.

[0058] 309. According to the real-time feedback data received by the master control unit, determine the processor load information of each target node.

[0059] 310. Dynamically adjust the processor fan speed according to the processor load information.

[0060] In the embodiments of the present application, the main control unit can obtain the load information of the processors in each target node from the received real-time feedback data. Since there are fans in the target nodes to cool the processors, the processor fan speed can be dynamically adjusted according to the processor load information. For example, if the processor load information indicates that the current load of the processor is large, then the temperature of the corresponding processor will increase. Therefore, the processor fan speed can be synchronously increased to cool the processor.

[0061] In some embodiments, as [[ID= shown in the schematic diagram of the edge sensor module, the slave control unit is connected to the sensor and the processor to implement data acquisition and control, and the slave control unit is also connected to the fan. That is to say, the fan in the target node is controlled by the slave control unit. Therefore, after the main control unit determines the processor load information, it can first determine the corresponding processor fan speed according to the processor load information, and then send the processor fan speed to the slave control unit so that the slave control unit makes dynamic adjustments according to the processor fan speed.

[0062] In some embodiments, the slave control unit can control the fan speed through PWM.

[0063] In some embodiments, as ​ shown, the slave control unit is connected to and controls the fan. At the same time, the slave control unit also monitors the fan bearing life, which specifically includes: obtaining the actual fan speed and the current temperature value of the processor through the first slave control unit; determining the corresponding duty cycle according to the current temperature value; determining the current theoretical speed according to the duty cycle and the maximum fan speed of the processor; and outputting the fan life prompt information of the processor according to the actual fan speed and the current theoretical speed.

[0064] It should be noted that the maximum fan speed is set when the fan leaves the factory. When the fan is running, it will not exceed the maximum fan speed of the fan. And as the number of uses and the duration increase, the fan bearing life will decrease, and the difference between the fan speed and the fan theoretical speed will become larger and larger. The fan theoretical speed is obtained by multiplying the duty cycle of the fan by the maximum fan speed of the fan. The duty cycle is related to the real-time temperature. Therefore, the duty cycle corresponding to the current temperature value can be determined according to the current temperature value, and then multiplied by the maximum fan speed of the fan to obtain the current theoretical speed of the fan. Then, it is compared with the actual fan speed, and the fan bearing life is determined according to the difference between the actual fan speed and the current theoretical speed, and the prompt information is output.

[0065] For example: when the difference between the actual rotation speed and the current theoretical rotation speed of the fan is less than or equal to 10% of the current theoretical rotation speed, it can be considered that the life of the fan bearing is good; when the difference between the actual rotation speed and the current theoretical rotation speed of the fan is greater than 10% of the current theoretical rotation speed, it can be considered that the loss of the fan bearing is large and maintenance or replacement is required. Among them, 10% can be a value set by oneself.

[0066] It should be noted that the slave control unit can control the rotation speed of the fan in the target node through PWM, and obtain the current actual rotation speed of the fan by detecting the Tach signal. From the currently detected temperature value of the slave control unit, the required rotation speed of the fan can be calculated through a specific PID regulation algorithm for automatic control and life monitoring of the fan.

[0067] In the embodiment of the present application, by adjusting the rotation speed and monitoring the life of the fan in the target node where the slave control unit is located, the normal heat dissipation and operation of the processor can be ensured, and it can prevent the heat dissipation effect from being affected due to the excessive deviation between the fan rotation speed and the theory, thereby affecting the service life of the target node and the fan itself.

[0068] 311. Parse and store the real-time feedback data through the main control unit.

[0069] In the embodiment of the present application, after receiving the real-time feedback data sent by the slave control unit, the main control unit can store it, and can also parse it to analyze whether the real-time feedback data indicates that there is an abnormal situation in the processor and record the abnormal situation in the log.

[0070] In some embodiments, before parsing and storing, the main control unit can also verify the real-time feedback data, and perform parsing and storing in the case of passing the verification, which can ensure data security and server security.

[0071] 312. If the first slave control unit receives the response feedback information corresponding to the response request sent by the second slave control unit, stop collecting the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit.

[0072] In the embodiments of the present application, when the second slave control unit fails, the first slave control unit can take over the second slave control unit for data acquisition and reporting, while the second slave control unit can be restarted and reset. After the second slave control unit can continue data acquisition and reporting after reset, there is no need for the first slave control unit to take over anymore. Therefore, the response feedback information corresponding to the response request sent by the second slave control unit to the first slave control unit can be used. When the first slave control unit receives the response feedback information corresponding to the response request sent by the second slave control unit, the first slave control unit can confirm that the second slave control unit has been restored from the fault and can operate normally. Therefore, the first slave control unit can stop collecting the real-time feedback data of the processor corresponding to the second slave control unit.

[0073] In some embodiments, the reset operation of the faulty slave control unit can be completed through the watchdog mechanism. The watchdog mechanism (Watchdog Timer) is a hardware or software mechanism used to monitor the running state of the system and automatically recover when the system encounters an exception. It is widely used in fields such as embedded systems, computer systems, and network devices. Its core purpose is to improve the reliability and stability of the system and prevent the system from falling into an infinite loop, crashing, or becoming unresponsive due to software failures, hardware anomalies, or external interference.

[0074] In some embodiments, both the master control unit and the slave control units can be independently powered. Specifically, both the master control unit and the slave control units are respectively configured with independent capacitors to achieve independent power supply. When there is an abnormal power supply for the master control unit and / or the slave control units, continuous power supply is carried out through the respectively configured independent capacitors.

[0075] It should be noted that the master control unit and the slave control units are equipped with independent supercapacitors as backup power supplies. Since the required supply voltage of the control unit is not high, a capacitor with a capacity of 500 mF can meet the requirement. When there is a power outage or abnormal power-off such as a power failure, the supercapacitor can continue to supply power to the control unit for at least more than 10 seconds, maintaining the continuous operation of the control unit, ensuring the storage of critical status data (such as registers) at the moment of power-off. In case of an abnormal power-off, the cause of the power-off can also be recorded to help quickly locate the problem.

[0076] In the embodiments of the present application, both the master control unit and each slave control unit are respectively configured with independent capacitors for independent power supply. In this way, in the case of a sudden power-off, power supply can also be continued through the independent capacitors, so that the control unit will not suddenly stop running, avoiding situations such as data loss and transmission failures, and ensuring the operation stability of the server.

[0077] In some embodiments, the main control unit obtains the operation data of the power supply module of each target node in real time; according to the real-time feedback data received by the main control unit, the operation data of the power supply module is dynamically adjusted.

[0078] It should be noted that as ​ shown, the main control unit can be connected to the power supply module (Power Supply Unit, PSU) in each target node, and can be connected through the SMBus bus for monitoring the operation data of the PSU, such as input and output voltage, current, power, etc., so as to calculate the power consumption of each target node. The main control unit can also control the output efficiency of the PSU according to the power consumption, thereby saving power. The main control unit can determine the processor load information according to the received real-time feedback data, and thus dynamically adjust the operation data of the power supply module according to the processor load information.

[0079] In some embodiments, the main control unit receives the control instructions sent by the cloud network platform; the main control unit sends the control instructions to multiple slave control units respectively.

[0080] It should be noted that the control unit can implement cloud network services. The cloud network platform sends control instructions such as restart and shutdown to the main control unit. After the main control unit receives the control instructions, it then sends the corresponding control instructions to the slave control units in the communication task. The slave control units complete the control of a single target node.

[0081] In some embodiments, as ​ shown, both the main control unit and the slave control unit are connected to the phy chip through the RMII protocol for network communication, and can be connected to the upper computer through the network port. The upper computer can directly access the Web services of the main control unit and the slave control unit through the IP address, and can intuitively monitor and control each target node of the edge server in the Web.

[0082] In the embodiments of the present application, both the main control unit and the slave control unit can adopt a real-time operating system (RTOS). The main control unit uses FreeRTOS to implement multitask management, mainly including communication tasks, data processing tasks, remote control tasks, energy-saving management tasks, firmware upgrade tasks, multi-bus dynamic switching tasks, etc. The communication task can communicate with the slave control unit, receive status information and send control instructions; the data processing task can verify, parse and store the received status information, and determine whether there is an abnormal situation in the server; the remote control task can communicate with a remote terminal or a cloud network, receive remote control instructions and forward them to the slave control unit; the energy-saving management task can dynamically adjust the fan speed, PSU output voltage and current, and the working frequency of the control unit according to the load conditions of each target node; the firmware upgrade task can upgrade its own firmware and the firmware of the slave control unit. The slave control unit also uses FreeRTOS, mainly including sensor acquisition tasks, communication tasks, firmware upgrade tasks and control tasks. The acquisition task periodically controls the sensor module to acquire status information; the communication task can communicate with the main control unit, receive status acquisition instructions and control instructions, and send status information and operation results; the firmware upgrade task completes the firmware upgrade of itself; the control task performs corresponding operations according to the control instructions of the main control unit.

[0083] The edge server disclosed in the embodiments of the present application can greatly reduce the software implementation cost and PCB layout cost compared with the traditional edge server implemented by BMC. The overall cost is saved by at least 60%, and it is applicable to all CPUs, different from the traditional BMC solution which is only applicable to Intel's X86 architecture CPUs. When the main control unit and the slave control unit communicate with each other, not only a lightweight neural network engine is built in the slave control unit for data preprocessing, but also a multi-bus dynamic switching mechanism is implemented to dynamically switch the bus to cope with different transmission scenarios, ensuring the communication reliability in a complex electromagnetic environment while accelerating the transmission rate. When dealing with node failure problems, a distributed self-healing mechanism is adopted. When a target node fails, the slave control unit of an adjacent node can temporarily take over important tasks and continue to take over after the reset and recovery are completed. It can also dynamically adjust the PSU output voltage according to the CPU load conditions, and adjust the main frequencies of the main control unit and the slave control unit, so as to achieve the purpose of saving power. When the power supply is abnormally cut off, the key status data can also be stored, and the status data can be restored after the power is restored, ensuring the stability and reliability of the server operation.

[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0085] As ​ shown, an embodiment of the present application further provides an edge server management device. The edge server includes a main control unit and a plurality of target nodes. Each target node includes a processor and a slave control unit. The main control unit and the slave control unit communicate through multiple links; the main control unit and the slave control unit are independently powered; the edge server management device may include: A transceiver module 601, configured to send data collection tasks to a plurality of slave control units respectively through the main control unit; The transceiver module 601 is further configured to, in response to the data collection task, send response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node; An acquisition module 602, configured to, if the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within a first preset duration, collect the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit; The transceiver module 601 is further configured to send the real-time feedback data and a fault alarm message to the main control unit through the first slave control unit, and the fault alarm message is used to indicate that the second slave control unit has a fault.

[0086] In some embodiments, the transceiver module 601 is specifically configured to send response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node at a preset period.

[0087] In some embodiments, the transceiver module 601 is specifically configured to, if the main control unit does not receive the feedback message from the slave control unit in the first target node within a second preset duration, send a takeover instruction to the slave control unit in the second target node through the main control unit, where the first target node and the second target node are adjacent nodes; The transceiver module 601 is specifically configured to send a response request to the slave control unit in the first target node through the slave control unit in the second target node according to the takeover instruction.

[0088] In some embodiments, the acquisition module 602 is further configured to monitor the real-time bus load rate and the real-time error rate between the main control unit and each slave control unit in real time through the main control unit; The edge server management device may further include: A processing module 603, configured to determine a corresponding target bus according to the real-time bus load rate and the real-time error rate; The processing module 603 is further configured to construct communication links between the main control unit and a plurality of slave control units respectively through the target bus.

[0089] In some embodiments, the processing module 603 is further configured to perform continuous power supply through separately configured independent capacitors when there is an abnormal power supply in the main control unit and / or the slave control unit.

[0090] In some embodiments, the obtaining module 602 is further configured to obtain, in real time through the main control unit, the operation data of the power supply module of each target node; The processing module 603 is further configured to dynamically adjust the operation data of the power supply module according to the real-time feedback data received by the main control unit.

[0091] In some embodiments, the processing module 603 is further configured to determine the processor load information of each target node according to the real-time feedback data received by the main control unit; The processing module 603 is further configured to dynamically adjust the rotation speed of the processor fan according to the processor load information.

[0092] In some embodiments, the processing module 603 is further configured to initialize the main control unit and the slave control unit respectively after the edge server is powered on.

[0093] In some embodiments, the obtaining module 602 is further configured to obtain the actual rotation speed of the fan of the processor and the current temperature value through the first slave control unit; The processing module 603 is further configured to determine the corresponding duty cycle according to the current temperature value; The processing module 603 is further configured to determine the current theoretical rotation speed according to the duty cycle and the maximum rotation speed of the fan of the processor; The processing module 603 is further configured to output a fan life prompt message for the processor according to the actual rotation speed of the fan and the current theoretical rotation speed.

[0094] In some embodiments, the processing module 603 is further configured to parse and store the real-time feedback data through the main control unit.

[0095] In some embodiments, the obtaining module 602 is further configured to stop collecting the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit if the first slave control unit receives the response feedback information corresponding to the response request sent by the second slave control unit.

[0096] In some embodiments, the transceiver module 601 is further configured to receive, through the main control unit, a control instruction sent by the cloud network platform; The transceiver module 601 is further configured to send the control instruction to multiple slave control units respectively through the main control unit.

[0097] In the embodiments of the present application, for the description of the features in the corresponding embodiments of the edge server management device, reference may be made to the relevant descriptions in the corresponding embodiments of the edge server management method, which will not be elaborated herein one by one.

[0098] As ​ shown, an embodiment of the present application further provides an electronic device, including a memory 701 and a processor 702. A computer program is stored in the memory 701, and the processor 702 is configured to run the computer program to execute the steps in any of the above-described embodiments of the edge server management method.

[0099] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the edge server management method when running.

[0100] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0101] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the edge server management method.

[0102] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the edge server management method.

[0103] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0104] The above has introduced in detail the process monitoring of a storage system provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for managing an edge server, characterized in that, The edge server includes a main control unit and multiple target nodes. Each target node includes a processor and a slave control unit. The main control unit and the slave control units communicate through multiple links; the main control unit and the slave control units are independently powered; the method includes: Sending data collection tasks to multiple slave control units respectively through the main control unit; In response to the data collection tasks, sending response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node; If the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within a first preset time period, collecting real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit; Sending the real-time feedback data and a fault alarm message to the main control unit through the first slave control unit, where the fault alarm message is used to indicate that the second slave control unit has a fault.

2. The method according to claim 1, wherein The step of sending response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node includes: Sending the response requests to the slave control units in the adjacent target nodes respectively through the slave control units in each target node according to a preset period.

3. The method according to claim 1, wherein The step of sending response requests to the slave control units in adjacent target nodes respectively through the slave control units in each target node includes: If the main control unit does not receive a feedback message from the slave control unit in the first target node within a second preset time period, sending a takeover instruction to the slave control unit in the second target node through the main control unit, where the first target node and the second target node are adjacent nodes; According to the takeover instruction, sending the response request to the slave control unit in the first target node through the slave control unit in the second target node.

4. The method according to claim 1, characterized in that Before sending data collection tasks to multiple slave control units respectively through the main control unit, the method further includes: The main control unit monitors the real-time bus load rate and real-time error rate between the main control unit and each slave control unit in real time; Determining a corresponding target bus according to the real-time bus load rate and the real-time error rate; Constructing communication links between the main control unit and the multiple slave control units respectively through the target bus.

5. The method according to claim 1, characterized in that, The method further includes: When there is a power supply anomaly in the main control unit and / or the slave control units, continuous power supply is performed through independently configured capacitors.

6. The method according to claim 1, characterized in that The method further includes: The main control unit obtains the operation data of the power module of each target node in real time; Dynamically adjusting the operation data of the power module according to the real-time feedback data received by the main control unit.

7. The method according to claim 1, characterized in that, The method further includes: Determining the processor load information of each target node according to the real-time feedback data received by the main control unit; Dynamically adjusting the rotation speed of the processor fan according to the processor load information.

8. The method according to claim 1, characterized in that, Before sending data collection tasks to multiple slave control units respectively through the main control unit, the method further includes: After the edge server is powered on, initialize the main control unit and the slave control unit respectively.

9. The method according to claim 1, wherein The method further includes: Obtain the actual fan speed and the current temperature value of the processor through the first slave control unit; Determine the corresponding duty cycle according to the current temperature value; Determine the current theoretical speed according to the duty cycle and the maximum fan speed of the processor; Output the fan life prompt information of the processor according to the actual fan speed and the current theoretical speed.

10. The method according to claim 1, wherein After sending the real-time feedback data and the fault alarm information to the main control unit through the first slave control unit, the method includes: Parse and store the real-time feedback data through the main control unit.

11. The method according to claim 1, characterized in that, If the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within the first preset time period, after collecting the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit, the method further includes: If the first slave control unit receives the response feedback information corresponding to the response request sent by the second slave control unit, stop collecting the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit.

12. The method according to claim 1, wherein The method further includes: Receive a control instruction sent by the cloud network platform through the main control unit; Send the control instruction to the multiple slave control units respectively through the main control unit.

13. An edge server management device, characterized in that, The edge server includes a main control unit and multiple target nodes. Each target node includes a processor and a slave control unit. The main control unit and the slave control unit communicate through multiple links; the main control unit and the slave control unit are independently powered; the device includes: A transceiver module, configured to send data collection tasks to multiple slave control units respectively through the main control unit; The transceiver module is further configured to, in response to the data collection task, send a response request to the slave control unit in an adjacent target node respectively through the slave control unit in each target node; An acquisition module, configured to collect the real-time feedback data of the processor corresponding to the second slave control unit through the first slave control unit if the first slave control unit does not receive the response feedback information corresponding to the response request sent by the second slave control unit within the first preset time period; The transceiver module is further configured to send the real-time feedback data and the fault alarm information to the main control unit through the first slave control unit, and the fault alarm information is used to indicate that the second slave control unit has a fault.

14. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the edge server management method according to any one of claims 1 to 12 when executing the computer program.

15. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the edge server management method according to any one of claims 1 to 12 when executed by a processor.

Citation Information

Patent Citations

  • State information processing method and device, server and readable storage medium

    CN110493328A

  • Server, node equipment information acquisition method and device thereof, equipment and medium

    CN113872796A

  • Multi-node server power supply module state control method and device

    CN115756138A

  • Abnormality processing method and device

    CN118945037A