Firmware refreshing method and electronic equipment
By using a power consumption prediction model to dynamically adjust the firmware transmission duration and sequence in multi-cascaded servers, the peak load and hardware risks of firmware flashing in multi-cascaded servers are resolved, achieving an efficient and reliable firmware flashing process and ensuring the safe and stable operation of the server.
Patent Information
- Application Number
- CN202511432816.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
AI Technical Summary
How to efficiently and reliably perform firmware updates on multiple slave servers in a multi-cascaded server architecture, avoiding risks such as peak load, power overload, and hardware overheating caused by concurrent updates.
By obtaining the total power consumption of the master server and slave servers, the future power consumption is predicted using a power consumption prediction model. When the predicted total power consumption exceeds a threshold, the firmware transmission duration and sending order of each slave server are dynamically adjusted to ensure that each slave server performs firmware updates within an independent time period.
It effectively avoids power consumption fluctuations and hardware risks during firmware flashing, improves the stability and reliability of multi-cascaded servers, and ensures safe and stable operation.
Smart Images

Figure CN120909625A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a firmware refreshing method and electronic equipment. BACKGROUND
[0002] With the rapid development of cloud computing, big data, data mining and deep learning, higher requirements are put forward for the computing performance, resource scalability and hardware synergy efficiency of servers. In order to meet the demand of large-scale parallel computing in this application scenario, multi-cascaded servers emerge as the times require. The multi-cascaded server is composed of a master server and multiple slave servers. The master server is connected with the slave servers through a Peripheral Component Interconnect Express (PCIe) switch or a network device. This connection mode helps the master server to quickly discover all slave servers and realize point-to-point firmware updating. How to realize efficient and reliable firmware updating for the architecture characteristics of multi-cascaded servers is the current focus. SUMMARY
[0003] The present application provides a firmware refreshing method and electronic equipment to at least solve the problem of how the master server efficiently and reliably refreshes the firmware of multiple slave servers.
[0004] The present application provides a firmware refreshing method applied to a master server, wherein the master server is connected with each of multiple slave servers; the method comprises: obtaining a total power consumption value within a preset historical time period back to the current time, wherein the total power consumption value is the power consumption value consumed by the master server and all slave servers together; predicting a predicted total power consumption value within a preset future time period based on the current time according to the total power consumption value and a pre-constructed power consumption prediction model; when the predicted total power consumption value is greater than a preset power consumption threshold, determining a firmware transmission time length corresponding to each slave server according to the total number of all slave servers; in accordance with a preset sending sequence, sequentially sending the firmware corresponding to each slave server to the slave server within the firmware transmission time length corresponding to each slave server, so as to refresh the firmware of each slave server.
[0005] The present application also provides a firmware refreshing device, comprising: an obtaining module configured to obtain a total power consumption value within a preset historical time period back to the current time, wherein the total power consumption value is the power consumption value consumed by the master server and all slave servers together; The prediction module is configured to predict a predicted total power consumption value in a preset time period in the future based on the total power consumption value and a pre-constructed power consumption prediction model; The determination module is configured to determine a firmware transmission time length corresponding to each slave server respectively based on the total number of all slave servers when the predicted total power consumption value is greater than the preset power consumption threshold. The sending module is configured to send the firmware corresponding to each slave server to the slave server in turn within the firmware transmission time length corresponding to each slave server respectively according to a preset sending sequence, so that each slave server performs firmware refreshing respectively.
[0006] The application further provides an electronic device, including a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of any of the firmware refreshing methods.
[0007] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of any of the firmware refreshing methods.
[0008] The application further provides a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps of any of the firmware refreshing methods.
[0009] Through the application, in the process of refreshing firmware of a plurality of slave servers by a master server, the total power consumption value in a future time period is predicted based on the total power consumption value of the master server and the slave servers in a historical time period, and the firmware refreshing of the slave servers is dynamically adjusted when the predicted total power consumption value is greater than a preset power consumption threshold, thereby effectively avoiding the risk of power overload, hardware overheating and the like caused by the peak load due to concurrent refreshing in the firmware refreshing process, and ensuring the safe and stable operation of the multi-level server including the master server and the plurality of slave servers. In the second aspect, the firmware transmission time length of each slave server is dynamically determined based on the total number of the slave servers, thereby avoiding the problem of transmission timeout or resource limitation of some slave servers when the fixed time length is allocated, and the instant peak power consumption caused by parallel firmware refreshing of the plurality of slave servers is reduced by sequentially transmitting the firmware according to the preset sending sequence. In the application, the power consumption value is combined with the firmware refreshing, which not only solves the power consumption fluctuation problem of the multi-level server in the firmware refreshing process, but also improves the reliability and efficiency of the firmware refreshing by allocating the transmission time length to each slave server and sequentially sending the firmware, thereby providing a basis for the stable operation of the multi-level server. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below only illustrate some of the embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0011] Figure 1 A flow chart of a firmware refreshing method provided by the embodiments of the present application; Figure 2 A refreshing schematic diagram of firmware of components with different priorities provided by the embodiments of the present application; Figure 3 An information interaction diagram of the master server sending firmware information of each slave server provided by the embodiments of the present application; Figure 4 An application scenario schematic diagram of a firmware refreshing system provided by the embodiments of the present application; Figure 5 A structure schematic diagram of a firmware refreshing device provided by the embodiments of the present application; Figure 6 A structure schematic diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0012] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the protection scope of the present application.
[0013] It should be noted that, in the description of the present application, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0014] In order to make the skilled in the art better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0015] First, the application scenario of the embodiments of the present application is exemplarily introduced.
[0016] With the rapid development of cloud computing, big data, data mining and deep learning, higher requirements are put forward for the computing performance, resource scalability and hardware synergy efficiency of servers. To meet the demand of large-scale parallel computing in this application scenario, device pooling servers emerge as the times require, forming a multi-cascaded server architecture composed of one master server (including a central processing unit (CPU) server) and one to eight slave server graphics processing unit boxes (GPU BOX). In this architecture, the slave server serves as a dedicated computing resource carrier, only mounting graphics processing units (GPUs) and supporting four GPU deployments per slave server, and the master server is directly connected to each slave server through a peripheral component interconnect express (PCIe) switch or a network switch in a star or tree topology. Unlike the chain architecture of traditional slave servers in series, this point-to-point connection method not only quickly identifies all slave servers with the help of link layer discovery protocols (such as the link layer discovery protocol (LLDP)), but also lays a foundation for efficient execution of hardware management operations.
[0017] In the whole life cycle operation and maintenance of the multi-cascaded server, the version stability and functional compatibility of the firmware directly determine the running efficiency and reliability of the core hardware such as the GPU. For example, updating the GPU firmware can repair hardware compatibility vulnerabilities and optimize computing power scheduling logic. However, due to the large number of slave servers mounted on the master server in this architecture, and each slave server containing four GPU components, firmware flashing involves a large number of hardware nodes and complex component types (covering multiple firmware such as GPUs and baseboard management controllers (BMCs)). If sudden conditions such as slave server communication timeout and power consumption mutation occur during the slave server flashing process, the lack of targeted exception handling mechanism will further increase the risk of firmware flashing failure, and even affect the service availability of the entire multi-cascaded server. Therefore, how to realize efficient and reliable firmware updating according to the architecture characteristics of the multi-cascaded server is the focus of attention.
[0018] Therefore, the embodiments of the present application provide a firmware flashing method to solve the problem of how the master server efficiently and reliably flashes the firmware of multiple slave servers.
[0019] It should be noted that the execution subject of the firmware updating method provided by the embodiment of the present application can be a firmware updating device, which can be realized by software, hardware or a combination of software and hardware to become part or all of an electronic device, wherein the electronic device can be a server or a terminal, wherein the server in the embodiment of the present application can be a server or a server cluster composed of multiple servers, and the terminal in the embodiment of the present application can be a smart phone, a personal computer, a tablet computer, a wearable device, a smart robot and other smart hardware devices. In the following method embodiment, the execution subject is taken as an example of the electronic device.
[0020] According to the embodiment of the present application, a firmware updating method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0021] Figure 1 is a flowchart of a firmware updating method provided by the embodiment of the present application, which is executed by a master server, wherein the master server is connected with each of a plurality of slave servers, and the master server and the plurality of slave servers constitute a multi-level server cluster. The master server is responsible for global scheduling, including topology discovery, cluster power consumption prediction, time-sharing mechanism triggering, refresh sequence generation, virtual channel management, security authentication and rollback control, etc. The slave server receives the instruction of the master server and is responsible for local refresh operation, including firmware receiving, writing, resource constraint setting, temperature monitoring, etc. As shown in Figure 1 , the flowchart includes: S101, obtaining a total power consumption value backtracking from a preset historical time period based on a current time.
[0022] The total power consumption value is the power consumption value consumed by the master server and all slave servers.
[0023] The total power consumption value backtracking from a preset historical time period based on a current time means that the sum of the power consumption of the master server and all slave servers in a historical time period before the current time as the end point. For example, the current time is 10:00, and the preset historical time period is the past 1 hour (i.e. 9:00-10:00).
[0024] S102, predicting a predicted total power consumption value in a future preset time period based on the current time according to the total power consumption value and a pre-constructed power consumption prediction model.
[0025] Specifically, the power consumption prediction model can be a Kalman filter model, a linear regression model, etc. For example, a Kalman filter model can be constructed based on the energy consumption fingerprint data (such as the maximum power consumption value, the thermal resistance coefficient) corresponding to the master server and each slave server, so as to predict the predicted total power consumption value in a future preset time period based on the total power consumption value by using the Kalman filter model. It should be noted that the constructed Kalman filter model will be described in subsequent embodiments, which will not be described here.
[0026] S103, when the predicted total power consumption value is greater than the preset power consumption threshold, determining the firmware transmission time length corresponding to each slave server according to the total number corresponding to all slave servers.
[0027] Specifically, the preset power consumption threshold can be set according to the power supply capacity corresponding to the master server and the slave server, and the hardware tolerance, which is not limited in the embodiments of the present application. For example, the preset power consumption threshold is determined according to the maximum output power of the power supply module of each server, the maximum heat dissipation capacity of the heat dissipation system, etc.
[0028] For example, the preset power consumption threshold is determined by the following formula:
[0029] Wherein, P is the preset power consumption threshold, is a preset safety factor (0 <1), is the power distribution unit capacity, which is the power supply of the master server and the slave server. In this way, it is ensured that the safety range of the power distribution unit capacity is not exceeded during the firmware refreshing process, avoiding the overloading of the power supply or the sudden increase of power consumption caused by refreshing multiple slave servers in parallel.
[0030] The firmware transmission time length corresponding to the slave server refers to the time length of sending firmware from the master server to a single slave server.
[0031] Since there are physical limits of power supply, heat dissipation system, etc., when the total power consumption value is greater than the preset power consumption threshold, it may cause server hardware damage and other faults. If the master server transmits firmware to all slave servers in parallel, the instantaneous total power consumption will increase sharply, and at this time the predicted total power consumption value may be greater than the preset power consumption threshold. Therefore, it is necessary to dynamically adjust the transmission of firmware from the master server to each slave server, to allocate separate transmission time periods for each slave server, stagger the firmware transmission time of each slave server, and disperse the instantaneous power consumption increment, so as to reduce the total power consumption of the multi-level server. To improve the safety and stability of each server during the firmware refreshing process.
[0032] S104, according to the preset sending sequence, sequentially sending the firmware corresponding to each slave server to the slave server in the firmware transmission duration corresponding to each slave server, so that each slave server performs firmware refreshing respectively.
[0033] Specifically, the preset sending sequence is the order of sending the firmware of each slave server according to the service priority, hardware state and the like of the slave server. For example, the master server can preferentially refresh the slave server with high service priority (such as a slave server processing core data), and then refresh the slave server with low service priority (such as a slave server processing non-core data).
[0034] In the embodiments of the present application, in the process of refreshing the firmware of the plurality of slave servers by the master server, the total power consumption value in the future time period is predicted according to the total power consumption value of the master server and the slave servers in the historical time period, and the firmware refreshing of each slave server is dynamically adjusted in the case that the predicted total power consumption value is greater than the preset power consumption threshold, thereby effectively avoiding the peak load caused by concurrent refreshing in the firmware refreshing process, and the risks of power overload, hardware overheating and the like caused by the peak load, and guaranteeing the safe and stable operation of the multi-level server including the master server and the plurality of slave servers. In the second aspect, the firmware transmission duration of each slave server is dynamically determined based on the total number of slave servers, thereby avoiding the problem of transmission timeout or resource limitation of some slave servers when the fixed duration is allocated, and sequentially transmitting each firmware through the preset sending sequence, that is, transmitting the firmware corresponding to only one slave server in one firmware transmission duration, thereby reducing the instantaneous peak power consumption caused by parallel firmware refreshing of the plurality of slave servers. In the present application, the power consumption value is combined with the firmware refreshing, which not only solves the power consumption fluctuation problem of the multi-level server in the firmware refreshing process, but also improves the reliability and efficiency of the firmware refreshing by allocating the transmission duration to each slave server and sequentially sending each firmware, thereby providing a basis for the stable operation of the multi-level server.
[0035] In some embodiments, on the basis of the foregoing embodiments, the power consumption prediction module is a Kalman filter model. In the model, the Kalman filter algorithm is used to dynamically predict the predicted total power consumption value of the master server when refreshing the plurality of slave servers. The model is represented as follows:
[0036] wherein, is the predicted total power consumption value in the future preset time period based on the current time, is the total power consumption value in the preset historical time period based on the current time, A is a state transition matrix, and B is a control input matrix, is a control input, is the process noise. A and B can be calibrated offline by system identification method (e.g. least square method) based on server power consumption history, or updated online adaptively. may represent the current active refresh server number or refresh load factor. represents the error caused by factors not considered in the model. In the model, the filter state initialization and constraint setting are based on the maximum power consumption value and thermal resistance coefficient of each server, which can be determined by conducting energy consumption baseline test on each slave server.
[0037] In some embodiments, on the basis of the foregoing embodiments, when the predicted total power consumption value is greater than the preset power consumption threshold, the firmware transmission duration corresponding to each slave server is determined by the following formula:
[0038] wherein, is the firmware transmission duration, a is a configurable coefficient, is the total number of slave servers. In the embodiments of the present application, the logarithmic function is used to slow down the firmware transmission duration as the cluster size grows.
[0039] Each slave server is allocated an independent time period, and the firmware transmission processes of the slave servers are not overlapping:
[0040] wherein, , are the time period for transmitting firmware of the i-th slave server and the time period for transmitting firmware of the j-th slave server, respectively.
[0041] This ensures that only one slave server is performing firmware transmission at any time, thereby controlling the total power consumption.
[0042] In some embodiments, before sending the firmware corresponding to the first slave server to the first slave server within the firmware transmission duration corresponding to the first slave server, the method provided by the embodiments of the present application further includes the following content: First, the reference power consumption value corresponding to the first slave server is obtained, and the historical power consumption value corresponding to the first slave server during the firmware transmission process of the second slave server is obtained.
[0043] wherein, the first slave server and the second slave server are each a server in the plurality of slave servers, and in the preset sending order, the second slave server is the previous slave server of the first slave server. For example, the second slave server is the first slave server to which the master server transmits firmware, and the first slave server is the second slave server to which the master server transmits firmware.
[0044] Specifically, the reference power consumption value can be defined according to the actual situation of the first slave server. For example, when the power consumption value of the first slave server exceeds the reference power consumption value, the first slave server can have hardware damage and the like. The historical power consumption value refers to the power consumption data generated by the first slave server during the transmission of the firmware of the second slave server.
[0045] Then, according to the reference power consumption value and the historical power consumption value, the interval duration after the transmission of the firmware corresponding to the second slave server to the second slave server is determined. The interval duration is used to indicate that after a time period corresponding to the interval duration, the firmware corresponding to the first slave server is sent to the first slave server.
[0046] Specifically, the interval duration refers to the time interval from the completion of the firmware transmission of the second slave server to the start of the firmware transmission to the first slave server. The time period corresponding to the time interval is used as the power consumption cooling window of the first slave server, so as to wait for the power consumption of the first slave server to return to the stable state before transmitting the firmware to the first slave server, thereby avoiding the transmission failure of the master server in the process of transmitting the firmware to the first slave server, or the abnormality of the hardware of the first slave server and the like.
[0047] In a possible implementation, the interval duration is determined according to the difference between the historical power consumption value and the reference power consumption value.
[0048] Exemplarily, the interval duration has a positive correlation with the historical power consumption value corresponding to the slave server.
[0049] In the embodiment of the present application, the formula for determining the interval duration is as follows:
[0050] wherein, is the interval duration, is the historical power consumption value, is the reference power consumption value, and b is a cooling coefficient, which can be set according to the actual situation. Through the interval duration, it is ensured that the temperature and power consumption of the slave server return to the safe level when the next transmission period starts.
[0051] In the embodiment of the present application, during the process in which the master server transmits the firmware to the second slave server, the first slave server may have a temporary increase in power consumption at runtime. At this time, if the firmware is immediately transmitted to the first slave server, it will cause the power consumption of the first slave server to continuously increase, and even break through the power consumption threshold, resulting in hardware damage, firmware transmission failure, and the like. Therefore, after the master server completes the transmission of the firmware of the second slave server, a time period corresponding to the waiting interval is waited for to ensure that the hardware state of the first slave server returns to stable, and then the firmware corresponding to the first slave server is transmitted to the first slave server, thereby improving the reliability of firmware transmission and avoiding further impact of firmware transmission on the first slave server.
[0052] In some embodiments, on the basis of any of the foregoing embodiments, the first slave server includes at least one component, and the firmware corresponding to the first slave server includes component firmware corresponding to each component.
[0053] Specifically, the components in the slave server refer to hardware functional modules constituting the slave server. For example, the components can be a BMC, a GPU, a field-programmable gate array (FPGA), and the like. The component firmware refers to a firmware program adapted to the component, used to implement the basic functions of the component (such as a BMC firmware to implement power consumption detection, a GPU firmware to implement computing power scheduling, and the like). The component firmware can be stored in a built-in storage chip of the component.
[0054] During the firmware transmission time period of the first slave server, the firmware corresponding to the first slave server is transmitted to the first slave server, specifically including the following steps: a1, obtaining the dependency relationship between all component firmwares.
[0055] Specifically, different component firmware refreshes have a logical relationship, and if the firmware refresh of component A needs to depend on the firmware refresh of component B first, it is called that component A depends on component B. At this time, under the condition that both component A and component B need firmware refresh, during the refresh, the component firmware corresponding to component B needs to be refreshed first, and then the component firmware corresponding to component A needs to be refreshed. For example, the PCIe interface firmware depends on the flashing channel provided by the BMC.
[0056] In a possible implementation manner, each component firmware includes at least one function, and the dependency relationship between all component firmwares is the dependency relationship between functions in all component firmwares.
[0057] Specifically, the function in the component firmware refers to the smallest function code block constituting the firmware, which can implement a specific sub-function of the component, such as data receiving, data verification, hardware driving, etc. There can be a calling and called dependency relationship between different functions, such as function A needs to call the result of function B to execute. For example, the component firmware corresponding to the PCIe port contains function 1 for establishing a PCIe link, and the component firmware corresponding to the BMC contains function 2 for initializing the flashing interface, and function 1 depends on the execution of function 2, that is, after the BMC initializes the flashing interface through function 2, the PCIe port can establish a link through function 1. At this time, function 1 depends on function 2.
[0058] a2, determine the sending order corresponding to each component firmware according to the dependency relationship between all component firmwares.
[0059] In this way, the sending order of each component firmware determined based on the dependency relationship between the component firmwares can avoid the problem that the component firmware corresponding to the upper component cannot be flashed when the underlying component is not ready, and reduce the firmware flashing failure caused by incorrect order.
[0060] In a possible implementation, the sending order corresponding to each component firmware is determined according to the dependency relationship between all component firmwares, and specifically includes the following steps: First, a directed graph corresponding to all component firmwares is constructed according to the dependency relationship between functions in all component firmwares.
[0061] Each component firmware is a vertex in the directed graph, and the dependency relationship between functions is an edge in the directed graph.
[0062] Specifically, in the constructed directed graph, each vertex uniquely corresponds to a component firmware, which is a basic node unit of the directed graph, and is used to identify the component firmware to be sorted. The vertex can include core attributes: component firmware name (such as BMC firmware), belonging to slave server, etc., so as to determine the function dependency and sorting later. Each edge in the directed graph uniquely corresponds to the dependency relationship between a pair of functions. For example, if function 1 depends on function 2, an edge between the component firmware corresponding to function 1 and the component firmware corresponding to function 2 is added in the directed graph.
[0063] Then, the sending order corresponding to each component firmware is determined according to the directed graph.
[0064] For example, the sending order corresponding to each component firmware is determined by using the directed graph through a topological sorting algorithm.
[0065] In the embodiments of the present application, the main server extracts the dependent function symbol set in the firmware by inversely resolving the export symbol table of the binary file corresponding to the firmware, and constructs a directed graph according to the dependent function symbol set, and then performs topological sorting on the directed graph to generate the firmware update sequence.
[0066] a3, within the firmware transmission duration corresponding to the first slave server, sequentially sending all component firmwares to the first slave server according to the sending order corresponding to each component firmware.
[0067] In addition, the component firmwares with the dependency relationship can be sent to the first slave server according to the sending order, and the component firmwares without the dependency relationship can be sent to the first slave server in parallel.
[0068] In some embodiments, on the basis of the foregoing embodiments, the present application provides a method further comprising the following content: b1, obtaining the power consumption change rate corresponding to the first slave server in the process of sending all component firmwares to the first slave server.
[0069] Specifically, the power consumption change rate refers to the change rate of power per unit time of the main server in the process of sending component firmwares to the first slave server. The power consumption change rate is used to judge whether the first slave server appears a sudden increase in power consumption in the transmission process. Here, the sudden increase in power consumption may be caused by the first slave server receiving component firmwares, or may be caused by the first slave server processing business data, etc.
[0070] b2, when the power consumption change rate is greater than a preset change rate threshold, or when the main server does not send all component firmwares to the first slave server within the firmware transmission duration, obtaining at least one remaining component firmware that is not sent and a preconfigured bandwidth corresponding to the first slave server.
[0071] Specifically, the preset change rate threshold can be a power consumption change rate safety upper limit set according to the hardware tolerance of the first slave server, such as the maximum output power of the power module in the first slave server, the heat dissipation capacity, etc. If the actual power change rate of the first slave server exceeds the preset change rate threshold, it may cause the first slave server hardware to overheat or power fluctuation, and it is necessary to further split the transmission of the component firmwares to reduce the impact of firmware transmission on the first slave server.
[0072] Assuming that the sending order of the component firmware in the first slave server is component firmware 1, component firmware 2, and component firmware 3 in sequence. If the firmware transmission time length is reached, component firmware 1 and component firmware 2 have completed transmission, and component firmware 3 has not been transmitted, then component firmware 3 is the remaining component firmware. If the firmware transmission time length is not reached, but the power consumption rate of the first slave server is greater than the preset change rate threshold when transmitting component firmware 2, then the remaining component firmware is component firmware 2 and component firmware 3.
[0073] b3, according to the size of the first remaining component firmware and the preconfigured bandwidth, splitting the first remaining component firmware to determine at least one sub-firmware corresponding to the first remaining component firmware.
[0074] Wherein, the first remaining component firmware is any one of the at least one remaining component firmware.
[0075] Specifically, the size of the first remaining component firmware refers to the data volume of the remaining component firmware, and the unit is usually MB or GB.
[0076] The preconfigured bandwidth refers to the bandwidth preset between the master server and the first slave server. Each sub-firmware obtained by splitting the remaining component firmware is a small volume firmware fragment, so the transmission power consumption of each sub-firmware is small, and compared with transmitting the component firmware, transmitting the sub-firmware in sequence helps to further reduce the power consumption value of the first slave server. The preconfigured bandwidth can be set according to the actual situation of the first slave server, which is not limited here.
[0077] In a possible implementation, at least one sub-firmware corresponding to the first remaining component firmware is determined by the following method: First, according to the size of the first remaining component firmware and the preconfigured bandwidth, the firmware splitting specification is determined.
[0078] Specifically, the firmware splitting specification refers to the size of each sub-firmware, which is used as a basis for dividing the remaining component firmware.
[0079] For example, the firmware splitting specification is determined by the following formula:
[0080] Wherein, is the firmware splitting specification, is the size of the first remaining component firmware, is the preconfigured bandwidth, is the frequency reduction factor, In this way, by introducing the frequency reduction factor, the transmission bandwidth is further reduced, the hardware resource occupation is reduced, and the hardware dynamic state is adapted.
[0081] Then, the first remaining component firmware is divided according to the firmware splitting specification to determine at least one sub-firmware.
[0082] b4, when determining the sub-firmware corresponding to each remaining component firmware respectively, according to the preset sending rule, all sub-firmwares are sent to the first slave server in turn.
[0083] Specifically, the preset sending rule refers to the preset sending order of the sub-firmwares after splitting. For example, small volume sub-firmwares can be sent first, and after a preset time interval, large volume sub-firmwares are sent.
[0084] In one possible implementation, when the sub-firmware corresponding to each remaining component firmware is determined, all sub-firmwares are sent to the first slave server in turn according to the preset sending rule, which specifically includes the following steps: First, according to the sending order of each remaining component firmware and the position of each sub-firmware in the remaining component firmware, the sending order of each sub-firmware is determined.
[0085] Specifically, the position of the sub-firmware in the remaining component firmware can be indicated by the segment number of the sub-firmware after the remaining component firmware is split. The position of each sub-firmware in the remaining component firmware can be used to identify the order of the sub-firmwares in the complete remaining component firmware, to ensure that the sub-firmwares belonging to the same remaining component firmware are transmitted in order, avoiding data disorder when the first slave server receives and splices the sub-firmwares.
[0086] Then, according to the sending order of each sub-firmware, the sub-firmwares are sent to the first slave server in turn.
[0087] For example, the remaining component firmware is the component firmware corresponding to BMC and the component firmware corresponding to the fan. The component firmware corresponding to BMC is split into three sub-firmwares: sub-firmware 1, sub-firmware 2 and sub-firmware 3. The component firmware corresponding to the fan is split into sub-firmware 4 and sub-firmware 5. Among them, sub-firmware 1 is located in front of sub-firmware 2 in the component firmware corresponding to BMC, sub-firmware 2 is located in front of sub-firmware 3 in the component firmware corresponding to BMC. Sub-firmware 5 is located in front of sub-firmware 4 in the component firmware corresponding to BMC. And the sending order of the component firmware corresponding to BMC is earlier than that of the component firmware corresponding to the fan. Therefore, the sending order of the sub-firmwares transmitted by the master server is sub-firmware 1, sub-firmware 2, sub-firmware 3, sub-firmware 5, sub-firmware 4.
[0088] In this way, when the power consumption change rate is greater than the preset change rate threshold, the transmission of each sub-firmware in sequence can further reduce the impact of firmware transmission on the power consumption of the first slave server, and when the master server does not transmit all component firmwares to the first slave server within the firmware transmission time length, the remaining component firmwares are split into sub-firmwares, and since the transmission time of each sub-firmware is short, the transmission of each sub-firmware can be completed quickly, reducing the timeout probability caused by long-time occupation of the link. In addition, according to the transmission order corresponding to each remaining component firmware and the position of each sub-firmware in the remaining component firmware, the transmission order of each sub-firmware is determined, which ensures the orderliness of sub-firmware transmission while ensuring that each remaining component firmware is transmitted in the transmission order of each remaining component, and avoids that the first slave server can splice in order after receiving the sub-firmware.
[0089] In some embodiments, on the basis of any of the preceding embodiments, the master server is connected with each slave server through at least one network device.
[0090] Specifically, the network device is used to support data forwarding between the master server and the slave server, and is an intermediate node for communication between the master server and the slave server. For example, the network device can be an Ethernet switch, a router, a PCIe switch, etc.
[0091] Different combinations of network devices between the master server and the slave server form multiple transmission paths. For example, the master server can communicate with the slave server 1 through the network device a, and can also communicate with the slave server 1 through the network device b, thereby forming two transmission paths of "master server-network device a-slave server 1" and "master server-network device b-slave server 1".
[0092] In a possible implementation, the method provided by the embodiment of the application further includes the following steps: First, at least one transmission path between the master server and the ith slave server is determined according to the connection relationship among the ith slave server, the network device corresponding to the ith slave server, and the master server.
[0093] Wherein, i is a positive integer.
[0094] For example, the connection relationship among the ith slave server, the network device corresponding to the ith slave server, and the master server can be determined by a link layer discovery protocol.
[0095] For example, all possible transmission paths are enumerated according to the connection relationship among the ith slave server, the network device corresponding to the ith slave server, and the master server.
[0096] Then, the bandwidth information, the delay information and the security level information corresponding to all transmission paths are acquired.
[0097] Specifically, the bandwidth information can be the maximum data amount (bandwidth) that can be transmitted in a unit time of the transmission path, reflecting the data transmission capability of the transmission path. Exemplarily, the bandwidth information can be collected in real time by a network speed testing tool or read from a network device configuration.
[0098] The delay information refers to the time required for data to be sent from the master server to the i-th slave server through the transmission path. The delay information can reflect the data response speed of the transmission path. Exemplarily, the delay information includes network device forwarding delay, link transmission delay, etc.
[0099] The security level information is used to reflect the security protection capability of the transmission path for data transmission. Exemplarily, the security level information can be three levels of high, medium and low according to the strength of the protection measures. For example, the transmission path with high security level needs to support data encryption, access control, integrity verification, and the transmission path with low security level can only support basic verification.
[0100] Finally, according to the bandwidth information, the delay information and the security level information corresponding to all transmission paths respectively, a target transmission path is determined from all transmission paths, so as to transmit the firmware corresponding to the i-th slave server from the target transmission path to the i-th slave server when the sending order corresponding to the i-th slave server is determined according to the preset sending order.
[0101] Specifically, the target transmission path is the optimal path selected from all transmission paths by comprehensively considering the bandwidth, delay and security level information, which can meet the rate requirement, real-time requirement and security requirement of firmware transmission.
[0102] Exemplarily, according to the bandwidth information, the delay information and the security level information of each transmission path, each transmission path is quantitatively scored, and the transmission path with the highest score is taken as the target transmission path.
[0103] In the embodiments of the present application, the target transmission path is determined from all transmission paths according to the bandwidth information, the delay information and the security level information corresponding to all transmission paths, and specifically includes the following steps: Firstly, according to the bandwidth information of the first transmission path and a preset bandwidth scoring rule, a first score corresponding to the bandwidth information is determined. According to the delay information of the first transmission path and a preset delay scoring rule, a second score corresponding to the delay information is determined. According to the security level information of the first transmission path and a preset security scoring rule, a third score corresponding to the security level information is determined.
[0104] Exemplarily, in a case that the bandwidth information comprises a bandwidth, a preset bandwidth scoring rule can be that, when the bandwidth is greater than or equal to 10 GB / s, a corresponding first score is 100; when the bandwidth is less than 10 GB / s and greater than 5 GB / s, the corresponding first score is 60; and when the bandwidth is less than or equal to 5 GB / s, the corresponding first score is 30.
[0105] Exemplarily, in a case that the delay information comprises a delay duration, a preset delay scoring rule can be that, when the delay duration is greater than or equal to 1 ms, a corresponding second score is 30; when the delay duration is less than 1 ms and greater than 10 us, the corresponding second score is 60; and when the bandwidth is less than or equal to 10 us, the corresponding second score is 100.
[0106] Exemplarily, in a case that the security level information comprises a security level, a preset security scoring rule can be that, when the security level is a high security level, a corresponding third score is 100; when the security level is a medium security level, the corresponding third score is 60; and when the security level is a low security level, the corresponding third score is 30.
[0107] Then, a total score corresponding to the first transmission path is determined according to the first score, the second data and the third score.
[0108] Exemplarily, the first score, the second score and the third score are weighted and summed to obtain the total score corresponding to the first transmission path. Here, weights corresponding to the first score, the second data and the third score can be set according to actual conditions, which are not limited here.
[0109] Finally, when the total score corresponding to each transmission path is determined respectively, a target transmission path is determined according to the total score corresponding to each transmission path respectively.
[0110] In the embodiments of the present application, the transmission paths are selected through the bandwidth information and the delay information of the transmission paths, so that the firmware transmission rate and real-time performance can be guaranteed and timeout can be avoided. Meanwhile, in combination with the security level information of the transmission paths, the data security requirement of the firmware can be met, the firmware can be prevented from being leaked or tampered in the transmission process, and the server security can be guaranteed. In addition, the target transmission path can be flexibly selected according to the differentiated requirements of the server, and the flexibility of the firmware transmission process can be improved.
[0111] In a possible implementation, a plurality of logical channels are created on the physical link between the master server and each slave server to distinguish the traffic corresponding to different slave servers (such as firmware 1 of slave server 1, firmware 2 of slave server 2, and management data of the master server to each slave server).
[0112] Exemplarily, a Virtual Local Area Network (VLAN) technology is used to create N_vlan virtual local area network channels in the network constructed by the master server and the slave servers, where N_vlan is not less than the number of the slave servers. Each channel corresponds to a slave server or the management data of the master server to the slave server, and the independent transmission performance of each channel is guaranteed by isolating the propagation domain. The manager also implements a Quality of Service (QoS) policy for the transmission of different firmware types.
[0113] In some embodiments, on the basis of any of the preceding embodiments, the method provided by the embodiments of the present application further includes the following content: First, the device type corresponding to the i-th slave server is acquired.
[0114] Specifically, the device type corresponding to the i-th slave server can be a BMC server, an FPGA server, etc. Exemplarily, the master server can read the hardware interface information of each slave server through the LLDP protocol to determine the device type corresponding to each slave server.
[0115] Then, according to the device type, the data transmission protocol corresponding to the i-th slave server is determined, so as to subsequently send the firmware corresponding to the i-th slave server to the master server in compliance with the data transmission protocol.
[0116] Specifically, the data transmission protocol refers to the standardized rules for data interaction between the master server and the i-th slave server, including but not limited to data frame format, transmission rate control, error checking mechanism, etc., which needs to be compatible with the hardware interface of the slave server to ensure that the firmware can be correctly transmitted and parsed.
[0117] Exemplarily, according to the device type and the preset mapping relationship between the device type and the data transmission protocol, the data transmission protocol corresponding to the i-th slave server is determined.
[0118] In the embodiments of the present application, the master server establishes a multi-protocol overlay network at the Open Systems Interconnection (OSI) data link layer, and dynamically selects a data transmission protocol (also referred to as a tunneling protocol) according to the device type corresponding to the slave server. For example, the User Datagram Protocol (UDP) is used for encapsulated transmission for the FPGA slave server, and the Transmission Control Protocol (TCP) is used for encapsulated transmission for the BMC slave server. This is because UDP is connectionless and has small overhead, and is suitable for the scenario where the firmware corresponding to the FPGA slave server has high real-time requirements and can tolerate a small amount of packet loss, thereby improving transmission efficiency. The TCP protocol is a reliable connection protocol, and can ensure the integrity and order of the firmware corresponding to the BMC slave server during transmission, thereby avoiding the failure of the BMC corresponding component firmware in the BMC slave server due to packet loss (the BMC is a key management component).
[0119] Considering that if the data transmission protocol is incompatible with the hardware of the slave server, the data cannot be received, and the firmware transmission fails, therefore, the master server accurately matches the data transmission protocol for the slave server according to the device type corresponding to the slave server, optimizes the transmission performance of the transmission path, and improves the reliability of the firmware transmission.
[0120] In some embodiments, on the basis of any of the preceding embodiments, the method provided by the embodiments of the present application further includes: In the process of communication between the master server and each slave server, the master server encrypts the firmware corresponding to each slave server through a preset encryption rule, and sends the encrypted firmware to the first slave server.
[0121] For example, the master server and each slave server correspond to a temporary key pair when communicating, which is used for two-factor authentication of a session, and the validity period of the temporary integer is less than 300 seconds, thereby enhancing security. In this way, even if the integer is compromised, its available time is extremely short, and it does not pose a long-term threat to the multi-level server.
[0122] For example, the two-factor authentication is implemented based on the Fast Identity Online Universal 2nd Factor (FIDO U2F) standard.
[0123] In some embodiments, on the basis of any of the preceding embodiments, different virtual local area networks are constructed in the transmission path between the master server and the slave server to implement transmission of different component firmwares.
[0124] The method provided by the embodiments of the present application further includes the following content: First, a firmware type corresponding to the first firmware is acquired.
[0125] Specifically, the firmware type can be a Basic Input / Output System (BIOS) firmware, an FPGA firmware, or a NIC firmware.
[0126] Then, a bandwidth corresponding to the first firmware is determined according to the firmware type.
[0127] For the BIOS firmware, the bandwidth is greater than or equal to 50 Mbps, for the FPGA firmware, the bandwidth is greater than or equal to 100 Mbps, and for the NIC firmware, the bandwidth is greater than or equal to 30 Mbps.
[0128] In some embodiments, on the basis of any of the foregoing embodiments, the server calculates resource constraints for each slave server and issues the constraints to the slave server, and the slave server configures a local Control Group (cgroup) according to the constraints.
[0129] In a possible implementation, the method provided by the embodiment of the application further includes the following content: First, a first number of central processor cores included in a master server is acquired.
[0130] The master server is any one of all slave servers.
[0131] Specifically, the first number of central processor cores is used to determine the total amount of CPU resources in the slave server and determine the basis for the number of cores for firmware writing.
[0132] Then, a second number of central processor cores used by the master server when writing firmware is determined according to the first number.
[0133] For example, the first number of the preset proportion can be used as the second number.
[0134] Finally, the second number is sent to the master server.
[0135] For example, a CPU quota constraint is set for the firmware writing process of the first slave server, and the CPU quota is set as follows:
[0136] In addition, a memory constraint can also be set for the firmware writing process of the first slave server, for example, an upper limit of the memory is set as follows:
[0137] In which, The number of CPU cores allocated to the first slave server for firmware refreshing, that is, the second number, a first number of CPU cores in the first slave server, a size of memory allocated for the first slave server, a size of firmware corresponding to the first slave server, so as to ensure that the memory usage of the refreshing process does not exceed three times the size of the firmware plus a fixed buffer.
[0138] In the embodiments of the present application, by setting the number of central processor cores corresponding to the slave server, on the one hand, sufficient CPU computing power is ensured for firmware writing, and the refreshing efficiency is ensured. On the other hand, the slave server is prevented from occupying business CPU resources during firmware writing, and the stability of the core business of the slave server is ensured.
[0139] In a possible implementation, after the slave server receives the firmware, in order to ensure that the firmware writing process is isolated from other system services and to control resource usage, the firmware writing process is isolated by a container / control group.
[0140] For example, based on the Extended Berkeley Packet Filter (eBPF) technology, the firmware writing process is isolated in an independent Control Group (cgroups). Through the isolation and resource control mechanism of cgroups and eBPF, it can be effectively ensured that the firmware refreshing does not interfere with other functions of the slave server.
[0141] In a possible implementation, the master server is further configured to detect the temperature of the slave server when the firmware is written after the slave server receives the corresponding firmware. When the temperature of the slave server is greater than a preset temperature threshold, the master server sends a firmware writing pause instruction to the slave server and starts a cooling process. After pausing for a preset time, the master server sends a firmware writing start instruction to the slave server to instruct the slave server to continue the firmware writing.
[0142] In a possible implementation, the entire refreshing process of the slave server includes a firmware transmission process in which the master server sends firmware to the slave server, and a firmware writing process in which the slave server writes the firmware after receiving the firmware. In addition, the refreshing process can also include a rollback process.
[0143] For example, the rollback process includes a three-stage rollback process: In the first stage, when the master server and the first slave server fail to verify the firmware digital signature, the master server terminates the refresh process and generates an alarm signal; In the second stage, when the first slave server exceeds the preset write time for firmware writing, the first slave server performs a local backup firmware recovery operation; In the third stage, when the first slave server fails to write firmware continuously, the master server sends the firmware corresponding to the first slave server to the slave server corresponding to the first slave server, so that the slave server corresponding to the first slave server can refresh the firmware after receiving the firmware. At this time, the slave server corresponding to the first slave server is a mirrored redundant node of the first slave server.
[0144] The overall process of the firmware flashing method is described below. The master server initiates the flashing task, obtaining network topology information and slave server identities from the multi-cascaded servers via a link-layer discovery protocol. Then, based on a Kalman filter model, it determines the predicted total power consumption. When the predicted total power consumption exceeds a preset power consumption threshold, the firmware transmission duration for each slave server is determined according to the total number of slave servers, thus achieving time-sharing firmware transmission for each slave server. Figure 2 As shown, after receiving the firmware from the server, if high-priority component firmware exists in the firmware, the slave server uses cgroups and eBPF technology to isolate and refresh the high-priority component firmware, ensuring that firmware writing does not interfere with other functions of the device. If no high-priority component firmware (such as BMC firmware) exists in the firmware, the slave server refreshes each component firmware in the order of receiving the component firmware (also known as writing each component firmware). When the firmware writing is successful, the master server receives a firmware refresh completion signal from the slave server, and the firmware refresh of the slave server ends. When the firmware writing fails, the slave server rolls back the firmware to the previous version.
[0145] Figure 3 This is a diagram showing the interaction between the master server and multiple slave servers, where each server sends its own firmware information. Figure 3 In this process, the master server sends its respective firmware to slave servers A, B, and C sequentially according to the firmware transmission duration. Specifically, the master server first sends firmware 1 to slave server A within the firmware transmission duration corresponding to server A, while servers B and C wait. After sending to slave server A, the master server sends firmware 2 to slave server B within the firmware transmission duration corresponding to server B, while slave server C waits. After sending to slave server B, the master server sends firmware 3 to slave server C within the firmware transmission duration corresponding to server C.
[0146] This application also provides a firmware flashing system. The system includes: a topology discovery module, an energy consumption awareness engine, a flashing sequence generator, a virtual channel manager, a security authentication module, and a three-stage rollback controller.
[0147] The topology discovery module is configured to acquire the network topology information and slave server identities of the multi-server cascade through a link layer discovery protocol. The module periodically transmits and receives LLDP data packets, identifies new or changed slave server connection conditions, and provides the topology structure information to the master server for use.
[0148] The energy consumption-aware engine is configured to perform dynamic energy consumption prediction based on a Kalman filtering algorithm in the Kalman prediction unit to determine a predicted total power consumption value. The engine updates the Kalman filter state using the collected energy consumption fingerprint data and real-time measurement values to estimate the power consumption change during the refresh process, thereby supporting power consumption threshold judgment.
[0149] In the system, when the predicted total power consumption value is greater than the preset power consumption threshold, a time-sharing polling refresh mechanism is started. The mechanism includes dividing the refresh window into multiple time slices, and the length of each time slice is the firmware refresh duration described above. Each slave server is allocated an independent time slice, and the time slices satisfy a time slice mutual exclusion condition. At the same time, a power consumption cooling window is inserted between adjacent time slices, and the cooling window is a time period corresponding to the interval duration described above. In this way, after starting the time-sharing mechanism, the refresh window is divided into multiple mutually exclusive time slices, and a cooling time slot is inserted between the time slices, which can balance power consumption and refresh throughput.
[0150] The refresh sequence generator is configured to parse firmware identification (FID) obtained from each slave server, extract dependency function symbols, and construct a firmware dependency graph (i.e., the directed graph described above). Then, the graph is topologically sorted to generate a firmware refresh sequence matrix. The sequence design in the matrix ensures that components with dependency relationships are executed in the order of dependency; for components that are independent of each other, they can be refreshed simultaneously in parallel channels to speed up the total refresh speed.
[0151] When a slave server refresh times out or power consumption suddenly changes, the firmware corresponding to the slave server is fragmented and transmitted at a reduced frequency. The firmware is divided into sub-firmwares of the firmware splitting specification size described above.
[0152] The virtual channel manager is configured to create multiple logical channels on a physical link to distinguish different traffic. Specifically, N_vlan virtual local area network channels can be created in a switching network using VLAN technology, and N_vlan is not less than the number of slave servers. Each channel corresponds to a slave server or a service type, and the propagation domain is isolated to ensure independent transmission performance of each channel. The manager also implements quality of service (QoS) policies for the transmission of different firmware types.
[0153] For critical firmware such as the BIOS, the highest priority and lowest bandwidth guarantee are assigned to ensure timely updates even during network congestion. This manager uses a dynamic path selection algorithm to choose the highest-scoring path among multiple available paths. Path scores are weighted by bandwidth, latency, and security level, for example:
[0154] in, The total score is represented by Bandwidth, path bandwidth, Latency (normalized path latency), and Security (path security level). This multi-metric scoring mechanism comprehensively considers QoS indicators such as bandwidth and latency, as well as security attributes, to ensure that the selected transmission path is both efficient and secure.
[0155] The security authentication module implements two-factor authentication and key management for firmware refresh sessions. Internally, it integrates a Physically Unclonable Function (PUF) chip to generate a unique identifier for each slave server. At the start of each refresh session, this unit dynamically generates a temporary session key based on the slave server's PUF identifier and session conditions using an HMAC-based Key Derivation Function (HKDF), and uses this key in conjunction with a short-term digital certificate for key exchange and verification. This HKDF-based key derivation method securely generates encryption keys from a shared secret, and the short-term certificate used has a validity period of less than 300 seconds, resisting replay attacks.
[0156] The three-stage rollback controller monitors the status of the refresh process and executes corresponding recovery strategies based on different error types, specifically including three stages of response: Phase 1: Immediately terminate the refresh process and generate an alarm signal to notify the management system if firmware digital signature verification fails. This prevents unauthorized or tampered firmware from being written to the slave server.
[0157] Second stage: If the refresh process exceeds the preset timeout period, it will roll back to the local backup firmware and continue to ensure the functionality of the slave server through the recovery operation; at the same time, the failure information will be recorded for subsequent analysis.
[0158] Phase 3: If multiple refresh attempts fail consecutively, a redundancy strategy is triggered, switching the refresh task to a mirrored redundant slave server in the cluster to ensure the reliability and availability of the update operation from a hardware perspective.
[0159] In the system, on the one hand, by fusing the energy consumption fingerprint collection and the power consumption prediction algorithm based on Kalman filtering, the refresh period power consumption trend of each slave server can be accurately predicted, and the refresh plan can be dynamically adjusted. In this way, the peak load caused by concurrent refresh is significantly reduced, the power supply redundancy design and power module overload are avoided, the energy consumption resources are effectively saved, and the overall energy efficiency ratio of the system is improved. On the other hand, by introducing the firmware interface descriptor (FID) mechanism and the three-stage rollback control process, the compatibility and fault controllability of the refresh process in the server cluster of multiple vendors and heterogeneous architectures are ensured. The FID supports semantic analysis and dynamic verification of the firmware structure, and can quickly recover through the rollback mechanism in the case of system exception, effectively reducing the risk of business interruption caused by refresh failure. On the other hand, by constructing the refresh window time-sharing scheduling algorithm and the virtual channel QoS policy table, on the basis of link layer topology discovery and resource perception, the refresh order and time slice of each slave server in the multi-cascaded architecture are dynamically planned. Combined with the eBPF resource isolation mechanism, the refresh isolation scheduling of high-priority services can be realized, the refresh behavior is guaranteed to have zero disturbance to the core services, and the intelligence and controllability of system operation and maintenance are improved.
[0160] Figure 4 An application scenario diagram of a firmware refresh system. In Figure 4 the main server in the multi-cascaded server cluster discovers the network topology information of the multi-cascaded server and the slave server identity through the topology discovery module; creates multiple logical channels on the physical link to distinguish different traffic through the virtual channel manager; determines the predicted total power consumption value based on the Kalman filtering algorithm in the Kalman prediction unit through the energy perception engine; generates the firmware refresh order matrix through the refresh sequence generator, so as to determine the sending order of each component firmware of each slave server. During the transmission of the sub-firmware, the main server realizes the two-factor authentication and key management of the firmware refresh session through the security authentication module. In addition, the main server monitors the state of the refresh process through the three-stage rollback controller, and executes the corresponding recovery strategy according to different error types.
[0161] Through the description of the above implementation, those skilled in the art can clearly understand that the method according to the above embodiment can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better implementation.
[0162] In the embodiments of the present application, a firmware refresh device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, which have been described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware is also possible and contemplated.
[0163] The embodiment provides a firmware refreshing device, which comprises the following steps of: Figure 5 As shown in the figure, comprising: An acquisition module 501 is configured to acquire a total power consumption value in a preset historical time period back to a current time, wherein the total power consumption value is a power consumption value consumed by a master server and all slave servers together. A prediction module 502 is configured to predict a predicted total power consumption value in a future preset time period based on the total power consumption value and a pre-constructed power consumption prediction model. A determination module 503 is configured to determine a firmware transmission time length corresponding to each slave server according to a total number of all slave servers when the predicted total power consumption value is greater than a preset power consumption threshold. A sending module 504 is configured to sequentially send firmware corresponding to each slave server to the slave server in the firmware transmission time length corresponding to each slave server according to a preset sending sequence, so that each slave server performs firmware refreshing.
[0164] In a possible implementation, before the firmware corresponding to the first slave server is sent to the first slave server in the firmware transmission time length corresponding to the first slave server, the acquisition module 501 is further configured to acquire a reference power consumption value corresponding to the first slave server and a historical power consumption value corresponding to the first slave server in a firmware transmission process of a second slave server, wherein the first slave server and the second slave server are both one of the plurality of slave servers, and the second slave server is a slave server before the first slave server in the preset sending sequence. The determination module 503 is further configured to determine an interval time length after the firmware corresponding to the second slave server is transmitted to the second slave server according to the reference power consumption value and the historical power consumption value, wherein the interval time length indicates that the firmware corresponding to the first slave server is sent to the first slave server after a time period corresponding to the interval time length.
[0165] In a possible implementation, the first slave server comprises at least one component, the firmware corresponding to the first slave server comprises component firmware corresponding to each component, and the sending module 504 is specifically configured to acquire a dependency relationship between all component firmwares. The sending module 504 is specifically configured to determine a sending sequence corresponding to each component firmware according to the dependency relationship between all component firmwares. The sending module 504 is specifically configured to sequentially send all component firmwares to the first slave server according to the sending sequence corresponding to each component firmware in the firmware transmission time length corresponding to the first slave server.
[0166] In a possible implementation, the obtaining module 501 is further configured to obtain a power consumption change rate corresponding to the first slave server in the process of sending all the component firmwares to the first slave server; when the power consumption change rate is greater than a preset change rate threshold, or when the master server does not send all the component firmwares to the first slave server within a firmware transmission time length, obtain at least one remaining component firmware that is not sent and a preconfigured bandwidth corresponding to the first slave server; The determining module 503 is further configured to split the first remaining component firmware according to a size of the first remaining component firmware and the preconfigured bandwidth, to determine at least one sub-firmware corresponding to the first remaining component firmware, where the first remaining component firmware is any one of the at least one remaining component firmware. The sending module 504 is further configured to split the first remaining component firmware according to a size of the first remaining component firmware and the preconfigured bandwidth, to determine at least one sub-firmware corresponding to the first remaining component firmware, where the first remaining component firmware is any one of the at least one remaining component firmware. In a possible implementation, the determining module 503 is specifically configured to determine a firmware splitting specification according to a size of the first remaining component firmware and the preconfigured bandwidth. The first remaining component firmware is divided according to the firmware splitting specification, to determine at least one sub-firmware.
[0167] In a possible implementation, the sending module 504 is specifically configured to determine a sending order corresponding to each sub-firmware according to a sending order corresponding to each remaining component firmware and a position of each sub-firmware in the remaining component firmware. The sub-firmware is sent to the first slave server in sequence according to the sending order corresponding to each sub-firmware.
[0168] In a possible implementation, each component firmware includes at least one function, and a dependency relationship among all the component firmwares is a dependency relationship among functions in all the component firmwares. In a possible implementation, the determining module 503 is specifically configured to construct a directed graph corresponding to all the component firmwares according to the dependency relationship among the functions in all the component firmwares, where each component firmware is used as a vertex in the directed graph, and the dependency relationship among the functions is used as an edge in the directed graph. The sending order corresponding to each component firmware is determined according to the directed graph.
[0169] In a possible implementation, the master server is connected to each slave server through at least one network device; and the determining module 503 is further configured to connect the master server to each slave server through the at least one network device. In a possible implementation, the acquisition module 501 is further configured to acquire a device type corresponding to the i-th slave server. The determination module 503 is further configured to determine a data transmission protocol corresponding to the i-th slave server according to the device type, so as to subsequently send the firmware corresponding to the i-th slave server to the master server in a manner of following the data transmission protocol.
[0170] Through the device provided by the embodiment of the present application, first, in the process of firmware refreshing of the master server to the multiple slave servers, the total power consumption value of the future time period is predicted according to the total power consumption value of the master server and the slave servers in the historical time period, and in the case that the predicted total power consumption value is greater than the preset power consumption threshold, the firmware refreshing of each slave server is dynamically adjusted, which effectively avoids the risk of power supply overload, hardware overheating and the like caused by the peak load due to the concurrent refreshing in the firmware refreshing process, thereby guaranteeing the safe and stable operation of the multi-level server including the master server and the multiple slave servers. Second, the firmware transmission duration of each slave server is dynamically determined based on the total number of the slave servers, which avoids the problem of transmission timeout or resource limitation of some slave servers when the fixed duration is allocated, and meanwhile, each firmware is sequentially transmitted through the preset sending order, thereby reducing the instantaneous peak power consumption caused by the parallel firmware refreshing of the multiple slave servers. In the present application, by combining the power consumption value with the firmware refreshing, both the power consumption fluctuation problem of the multi-level server in the firmware refreshing process and the reliability and efficiency of the firmware refreshing are improved by allocating the transmission duration to each slave server and sequentially sending each firmware, thereby providing a basis for the stable operation of the multi-level server.
[0171] The description of the features in the embodiment of the firmware refreshing device can be referred to the related description of the embodiment of the firmware refreshing method, which will not be repeated here.
[0172] The embodiment of the present application further provides an electronic device, as shown in the figure, which comprises a memory 10 and a processor 20, the memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above firmware refreshing method embodiments. Figure 6 The embodiment of the present application further provides an electronic device, as shown in the figure, which comprises a memory 10 and a processor 20, the memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above firmware refreshing method embodiments.
[0173] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above firmware refreshing method embodiments when running.
[0174] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0175] Embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program, when executed by a processor, implements the steps in any of the above-described firmware updating method embodiments.
[0176] Embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps in any of the above-described firmware updating method embodiments.
[0177] The skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in a general manner. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0178] The above describes in detail a firmware updating method and an electronic device provided by the present application. The principles and implementation manners of the present application are described by applying specific examples in this paper, and the above description of the examples is only applicable to help understand the method of the present application and its core idea. It should be pointed out that, for the ordinary skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A firmware refresh method, characterized by, The application is applied to a master server, wherein the master server is connected with each of a plurality of slave servers; the method comprises: obtaining a total power consumption value in a preset historical time period back to a current time, wherein the total power consumption value is a power consumption value consumed by the master server and all the slave servers; predicting a predicted total power consumption value in a future preset time period based on the total power consumption value and a pre-constructed power consumption prediction model; when the predicted total power consumption value is greater than a preset power consumption threshold, determining a firmware transmission time length corresponding to each of the slave servers according to a total number corresponding to all the slave servers; in the firmware transmission time length corresponding to each of the slave servers, sequentially sending firmware corresponding to each of the slave servers to the slave server in a preset sending order, so that each of the slave servers performs firmware refreshing.
2. The method of claim 1, wherein, Before the firmware corresponding to the first slave server is sent to the first slave server in the firmware transmission time length corresponding to the first slave server, the method further comprises: obtaining a reference power consumption value corresponding to the first slave server and a historical power consumption value corresponding to the first slave server in the firmware transmission process of the second slave server, wherein the first slave server and the second slave server are both one of the plurality of slave servers, and the second slave server is a previous slave server of the first slave server in the preset sending order; determining an interval time length after the firmware corresponding to the second slave server is transmitted to the second slave server according to the reference power consumption value and the historical power consumption value, wherein the interval time length indicates that the firmware corresponding to the first slave server is sent to the first slave server after a time period corresponding to the interval time length.
3. The method of claim 2, wherein, The first slave server comprises at least one component, and the firmware corresponding to the first slave server comprises component firmware corresponding to each of the components; in the firmware transmission time length corresponding to the first slave server, the firmware corresponding to the first slave server is sent to the first slave server, comprising: obtaining a dependency relationship between all the component firmware; determining a sending order corresponding to each of the component firmware according to the dependency relationship between all the component firmware; in the firmware transmission time length corresponding to the first slave server, sequentially sending all the component firmware to the first slave server according to the sending order corresponding to each of the component firmware.
4. The method of claim 3, wherein, The method further comprises: obtaining a power consumption change rate corresponding to the first slave server in the process of sending all the component firmware to the first slave server; when the power consumption change rate is greater than a preset change rate threshold, or the master server does not send all the component firmware to the first slave server in the firmware transmission time length, obtaining at least one remaining component firmware that is not sent and a preconfigured bandwidth corresponding to the first slave server; splitting the first residual component firmware according to a size of the first residual component firmware and the preconfigured bandwidth, to determine at least one sub-firmware corresponding to the first residual component firmware, wherein the first residual component firmware is any one of the at least one residual component firmware; after determining the sub-firmware corresponding to each of the residual component firmware respectively, sequentially transmitting all the sub-firmwares to the first slave server according to a preset transmission rule.
5. The method of claim 4, wherein, The splitting the first residual component firmware according to a size of the first residual component firmware and the preconfigured bandwidth, to determine at least one sub-firmware corresponding to the first residual component firmware, comprises: determining a firmware splitting specification according to the size of the first residual component firmware and the preconfigured bandwidth; dividing the first residual component firmware according to the firmware splitting specification, to determine at least one sub-firmware.
6. The method according to claim 4 or 5, characterized in that, The after determining the sub-firmware corresponding to each of the residual component firmware respectively, sequentially transmitting all the sub-firmwares to the first slave server according to a preset transmission rule, comprises: determining a transmission order corresponding to each of the sub-firmwares according to the transmission order corresponding to each of the residual component firmware respectively and a position of each of the sub-firmwares in the residual component firmware; sequentially transmitting the sub-firmwares to the first slave server according to the transmission order corresponding to each of the sub-firmwares.
7. The method according to any one of claims 3-5, characterized in that, Each of the component firmwares comprises at least one function, and a dependency relationship among all the component firmwares is a dependency relationship among the functions in all the component firmwares. The determining the transmission order corresponding to each of the component firmwares according to the dependency relationship among all the component firmwares, comprises: constructing a directed graph corresponding to all the component firmwares according to the dependency relationship among the functions in all the component firmwares, wherein each of the component firmwares is a vertex in the directed graph, and the dependency relationship among the functions is an edge in the directed graph; determining the transmission order corresponding to each of the component firmwares according to the directed graph.
8. The method according to any one of claims 1-5, characterized in that, The master server is connected with each of the slave servers through at least one network device; the method further comprises: determining at least one transmission path between the master server and the ith slave server according to a connection relationship among the ith slave server, a network device corresponding to the ith slave server and the master server, wherein i is a positive integer; obtaining bandwidth information, delay information and security level information corresponding to all the transmission paths respectively; determining a target transmission path from all the transmission paths according to the bandwidth information, the delay information and the security level information corresponding to all the transmission paths respectively, so as to transmit the firmware corresponding to the ith slave server from the target transmission path to the ith slave server when a transmission order corresponding to the ith slave server is determined according to the preset transmission order.
9. The method of claim 8, wherein, The method further comprises: obtaining a device type corresponding to the ith slave server; According to the device type, a data transmission protocol corresponding to the i-th slave server is determined, so that the firmware corresponding to the i-th slave server is sent to the master server in subsequent following the data transmission protocol.
10. An electronic device, comprising: Comprising: a memory for storing a computer program; a processor for implementing the steps of the firmware updating method as claimed in any one of claims 1-9 when executing the computer program.
Citation Information
Patent Citations
Firmware updating method, electronic device and control system
CN109933352A
Intelligent scheduling method and system for monitoring power consumption load of server
CN117421131A
Server cooperative control method, storage medium and electronic equipment
CN119917350A
Methods and systems for scene driven content creation
US10540699B1
Automated firmware updates based on device driver version in a server device
US20230205508A1