Firmware updating method and electronic device
By predicting power consumption values and dynamically adjusting firmware transmission duration and sequence in multi-cascaded servers, the peak load and hardware risks of firmware flashing in multi-cascaded servers are resolved, achieving an efficient and reliable firmware flashing process and ensuring the safe and stable operation of the server.
Patent Information
- Application Number
- CN202511432816.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-09
AI Technical Summary
How to efficiently and reliably perform firmware updates on multiple slave servers in a multi-cascaded server architecture, avoiding risks such as peak load, power overload, and hardware overheating caused by concurrent updates.
By obtaining the total power consumption of the master server and slave servers, the future power consumption is predicted using a power consumption prediction model. The firmware transmission duration of each slave server is dynamically adjusted, and the firmware is sent one by one according to the preset sending order to avoid instantaneous peak power consumption.
It effectively avoids power consumption fluctuations and hardware risks during firmware flashing, improves the stability and reliability of multi-cascaded servers, and ensures safe and stable operation.
Smart Images

Figure CN120909625B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a firmware flashing method and an electronic device. Background Technology
[0002] With the rapid development of cloud computing, big data, data mining, and deep learning, higher demands are being placed on server computing performance, resource scalability, and hardware collaboration efficiency. To meet the needs of large-scale parallel computing in these application scenarios, multi-cascaded servers have emerged. A multi-cascaded server consists of a master server and multiple slave servers. The master server connects to the slave servers via a high-speed Peripheral Component Interconnect Express (PCIe) switch or network device. This connection method helps the master server quickly discover all slave servers and achieve point-to-point firmware updates. How to achieve efficient and reliable firmware updates tailored to the architectural characteristics of multi-cascaded servers is currently a key focus. Summary of the Invention
[0003] This application provides a firmware flashing method and electronic device to at least solve the problem of how a master server can efficiently and reliably flash firmware on multiple slave servers.
[0004] This application provides a firmware flashing method applied to a master server, wherein the master server is connected to each of a plurality of slave servers; the method includes:
[0005] Get the total power consumption value within a preset historical time period based on the current time. The total power consumption value is the power consumption value consumed by the master server and all slave servers.
[0006] Based on the total power consumption value and the pre-built power consumption prediction model, predict the total power consumption value within a future preset time period with the current time as the reference.
[0007] When the predicted total power consumption is greater than the preset power consumption threshold, the firmware transmission time for each slave server is determined based on the total number of all slave servers.
[0008] According to the preset sending order, within the firmware transmission time corresponding to each slave server, the firmware corresponding to each slave server is sent to the slave server in sequence, so that each slave server can perform firmware refresh.
[0009] This application also provides a firmware flashing device, including:
[0010] The acquisition module is used to acquire the total power consumption value within a preset historical time period based on the current time. The total power consumption value is the power consumption value consumed by the master server and all slave servers.
[0011] The prediction module is used to predict the total power consumption value within a preset future time period based on the current time, according to the total power consumption value and the pre-built power consumption prediction model.
[0012] The determination module is used to determine the firmware transmission duration for each slave server based on the total number of all slave servers when the predicted total power consumption value is greater than the preset power consumption threshold.
[0013] The sending module is used to send the firmware corresponding to each slave server to the slave server in a preset sending order within the firmware transmission time corresponding to each slave server, so that each slave server can perform firmware refresh.
[0014] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above firmware flashing methods.
[0015] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described firmware flashing methods.
[0016] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described firmware flashing methods.
[0017] This application addresses two key aspects. First, during firmware updates from a master server to multiple slave servers, it predicts the total power consumption for future periods based on the historical total power consumption of the master and slave servers. If the predicted total power consumption exceeds a preset power consumption threshold, it dynamically adjusts the firmware updates for each slave server. This effectively avoids peak loads caused by concurrent updates, preventing power overload and hardware overheating risks, thus ensuring the safe and stable operation of the cascaded server system, including the master server and multiple slave servers. Second, it dynamically determines the firmware transmission duration for each slave server based on the total number of slave servers. This avoids timeouts or resource limitations for some slave servers when using fixed duration allocations. Furthermore, by transmitting firmware sequentially in a preset order, it reduces instantaneous peak power consumption caused by parallel firmware updates from multiple slave servers. In this application, by combining power consumption with firmware updates, it solves the power fluctuation problem during firmware updates in cascaded servers and improves the reliability and efficiency of firmware updates by allocating transmission durations to each slave server and transmitting firmware sequentially, providing a foundation for the stable operation of cascaded servers. Attached Figure Description
[0018] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a firmware flashing method provided in this application embodiment;
[0020] Figure 2 A schematic diagram illustrating the firmware refresh of components with different priorities provided in the embodiments of this application;
[0021] Figure 3 This application provides an example of an information exchange diagram showing how a master server sends its firmware to multiple slave servers.
[0022] Figure 4 This is a schematic diagram illustrating an application scenario of a firmware flashing system provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the structure of a firmware flashing device provided in an embodiment of this application;
[0024] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0026] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0027] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] First, the application scenarios of the embodiments of this application will be introduced by way of example.
[0029] With the rapid development of cloud computing, big data, data mining, and deep learning, higher demands are being placed on server computing performance, resource scalability, and hardware collaboration efficiency. To meet the needs of large-scale parallel computing in this application scenario, device pooling servers have emerged, forming a multi-cascaded server architecture consisting of one master server (including a Central Processing Unit (CPU) server) and 1-8 slave server Graphics Processing Unit Boxes (GPU BOXes). In this architecture, slave servers serve as dedicated computing resource carriers, only equipped with Graphics Processing Units (GPUs), with each slave server supporting a deployment of up to four GPUs. The master server directly connects to each slave server via a high-speed Peripheral Component Interconnect Express (PCIe) switch or network switch, in a star or tree topology. Unlike the traditional chain architecture of serialized slave servers, this point-to-point connection method not only enables rapid identification of all slave servers through link layer discovery protocols (such as Link Layer Discovery Protocol, LLDP), but also lays the foundation for efficient execution of hardware management operations.
[0030] In the entire lifecycle operation and maintenance of multi-cascaded servers, firmware version stability and functional compatibility directly determine the operating efficiency and reliability of core hardware such as GPUs. For example, GPU firmware updates can fix hardware compatibility vulnerabilities and optimize computing power scheduling logic. However, due to the large number of slave servers mounted on the master server in this architecture, and each slave server containing four GPU components, firmware updates involve a large number of hardware nodes and complex component types (covering various firmware for GPUs, Baseboard Management Controllers (BMCs), etc.). If unexpected situations such as slave server communication timeouts or sudden power consumption changes occur during slave server updates, the lack of targeted exception handling mechanisms will further increase the risk of firmware update failures and may even affect the service availability of the entire multi-cascaded server. Therefore, how to achieve efficient and reliable firmware updates tailored to the architectural characteristics of multi-cascaded servers is a current focus.
[0031] In view of this, embodiments of this application provide a firmware flashing method to solve the problem of how a master server can efficiently and reliably flash firmware for multiple slave servers.
[0032] It should be noted that the firmware flashing method provided in this embodiment of the invention can be executed by a firmware flashing device. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. The electronic device can be a server or a terminal. In this embodiment, the server can be a single server or a server cluster composed of multiple servers. The terminal can be a smartphone, personal computer, tablet computer, wearable device, or other intelligent hardware device such as a smart robot. The following method embodiments will use an electronic device as the execution subject for illustration.
[0033] According to an embodiment of the present invention, a firmware flashing method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] Figure 1This is a flowchart illustrating a firmware flashing method according to an embodiment of the present invention. The method is executed by a master server, which is connected to each of a plurality of slave servers, forming a multi-cascaded server cluster. The master server is responsible for global scheduling, including topology discovery, cluster power consumption prediction, time-sharing mechanism triggering, flashing sequence generation, virtual channel management, security authentication, and rollback control. Slave servers receive instructions from the master server and are responsible for local flashing operations, including firmware reception, writing, resource constraint setting, and temperature monitoring. Figure 1 As shown, the process includes:
[0035] S101, obtain the total power consumption value within a preset historical time period based on the current time.
[0036] The total power consumption value represents the power consumption consumed by the master server and all slave servers.
[0037] The total power consumption value within a preset historical time period based on the current time refers to the sum of the power consumption of the master server and all slave servers within a historical time period ending at the current time. For example, if the current time is 10:00, the preset historical time period is the past hour (i.e., 9:00-10:00).
[0038] S102, based on the total power consumption value and the pre-built power consumption prediction model, predicts the total power consumption value within a future preset time period with the current time as the reference.
[0039] Specifically, the power consumption prediction model can be a Kalman filter model, a linear regression model, etc. For example, a Kalman filter model can be constructed based on the energy consumption fingerprint data (such as maximum power consumption and thermal resistance coefficient) corresponding to the master server and each slave server, so as to predict the total power consumption value within a preset time period based on the total power consumption value. It should be noted that the constructed Kalman filter model will be described in subsequent embodiments and will not be elaborated here.
[0040] S103, when the predicted total power consumption value is greater than the preset power consumption threshold, determine the firmware transmission time corresponding to each slave server according to the total number of all slave servers.
[0041] Specifically, the preset power consumption threshold can be set based on the power supply capabilities of the master and slave servers, as well as the hardware tolerance. This application embodiment does not impose specific limitations on this. For example, the preset power consumption threshold can be determined based on the maximum output power of each server's power module and the maximum heat dissipation capacity of the cooling system.
[0042] For example, the preset power consumption threshold is determined by the following formula:
[0043]
[0044] Where P is the preset power consumption threshold. Preset safety factor (0 < <1), The power distribution unit capacity is set to supply power to both the master and slave servers. This ensures that the firmware update process does not exceed the safe range of the power distribution unit capacity, avoiding power overload or sudden power surges caused by updating multiple slave servers simultaneously.
[0045] The firmware transmission time for a slave server refers to the time it takes for the master server to send firmware to a single slave server.
[0046] Due to physical limitations of power supplies and cooling systems, exceeding a preset power consumption threshold can potentially cause server hardware damage. If the master server transmits firmware to all slave servers in parallel, the instantaneous total power consumption will surge, potentially exceeding the preset threshold. Therefore, it's necessary to dynamically adjust the master server's firmware transmission to each slave server. By allocating separate transmission time slots to each slave server and staggering their firmware transmission times, the instantaneous power consumption increase is distributed, reducing the overall power consumption of the multi-cascaded servers. This improves the security and stability of each server during the firmware flashing process.
[0047] S104, according to the preset sending order, sequentially sends the firmware corresponding to each slave server to the slave server within the firmware transmission time corresponding to each slave server, so that each slave server can perform firmware refresh.
[0048] Specifically, the preset sending order is the order in which each slave server firmware is sent based on the slave server's service priority, hardware status, and other settings. For example, the master server can prioritize refreshing slave servers with higher service priority (such as slave servers that process core data) and then refresh slave servers with lower service priority (such as slave servers that process non-core data).
[0049] In this application embodiment, firstly, during the firmware update process of the master server on multiple slave servers, the total power consumption value for the future time period is predicted based on the total power consumption value of the master server and slave servers in the historical time period. If the predicted total power consumption value is greater than a preset power consumption threshold, the firmware update of each slave server is dynamically adjusted. This effectively avoids the risks of power overload and hardware overheating caused by peak loads during concurrent firmware updates, ensuring the safe and stable operation of the multi-cascaded server including the master server and multiple slave servers. Secondly, the firmware transmission duration for each slave server is dynamically determined based on the total number of slave servers. This avoids the problem of some slave servers experiencing transmission timeouts or resource limitations when allocating fixed durations. Simultaneously, each firmware is transmitted sequentially through a preset sending order, meaning only the firmware corresponding to one slave server is transmitted within a single firmware transmission duration, reducing the instantaneous peak power consumption generated by parallel firmware updates on multiple slave servers. In this application, by combining power consumption values with firmware updates, the power consumption fluctuation problem during firmware updates in multi-cascaded servers is solved. Furthermore, by allocating transmission durations to each slave server and sending firmware sequentially, the reliability and efficiency of firmware updates are improved, providing a foundation for the stable operation of multi-cascaded servers.
[0050] In some embodiments, based on the foregoing embodiments, the power consumption prediction module is a Kalman filter model. In this model, the Kalman filter algorithm is used to dynamically predict the total power consumption of the master server when refreshing multiple slave servers. The model is represented as follows:
[0051]
[0052] in, This represents the predicted total power consumption over a future preset time period based on the current time. To backtrack the total power consumption over a preset historical time period based on the current moment, A is the state transition matrix, and B is the control input matrix. To control the input, This is process noise. A and B can be calibrated offline using system identification methods (such as the least squares method) or updated adaptively online based on historical server power consumption data. This can be represented as the number of currently active refresh slave servers or the refresh load factor. This indicates errors caused by factors not considered in the model. In this model, the filter state is initialized and constraints are set based on the maximum power consumption and thermal resistance of each server. The maximum power consumption and thermal resistance of each server can be determined by performing baseline energy consumption tests on each slave server.
[0053] In some embodiments, based on the foregoing embodiments, when the predicted total power consumption value is greater than a preset power consumption threshold, the firmware transmission duration corresponding to each slave server is determined by the following formula:
[0054]
[0055] in, Where 'a' is the firmware transmission time, and 'a' is a configurable coefficient. This represents the total number of servers. In this embodiment, a logarithmic function is used to gradually increase the firmware transmission time as the cluster size grows.
[0056] Each slave server is assigned an independent time period, and the firmware transmission processes of each slave server do not overlap:
[0057]
[0058] in, , These represent the time periods for transmitting the firmware of the i-th slave server and the time periods for transmitting the firmware of the j-th slave server, respectively.
[0059] This ensures that only one server is transmitting firmware at any given time, thereby controlling total power consumption.
[0060] In some embodiments, before sending the firmware corresponding to the first slave server to the first slave server within the firmware transmission time corresponding to the first slave server, the method provided in this application embodiment further includes the following:
[0061] First, obtain the baseline power consumption value corresponding to the first slave server, and the historical power consumption value corresponding to the first slave server during the firmware transmission process of the second slave server.
[0062] In this configuration, both the first and second slave servers are one of a plurality of slave servers. In a preset transmission order, the second slave server precedes the first slave server. For example, the second slave server is the first slave server whose firmware is transmitted by the master server, and the first server is the second slave server whose firmware is transmitted by the master server.
[0063] Specifically, the baseline power consumption value can be limited based on the actual situation of the first slave server. For example, if the power consumption of the first slave server exceeds the baseline power consumption value, the first slave server may have hardware damage, etc. The historical power consumption value refers to the power consumption data generated by the first slave server during operation while the second slave server is transmitting firmware.
[0064] Then, based on the baseline power consumption value and the historical power consumption value, the interval length after transmitting the firmware corresponding to the second slave server to the second slave server is determined. The interval length is used to indicate that after a time period corresponding to the interval length, the firmware corresponding to the first slave server will be sent to the first slave server again.
[0065] Specifically, the interval refers to the time interval from the completion of firmware transmission from the second slave server to the start of sending firmware to the first slave server. The time interval is used as a power cooling window for the first slave server, so that the master server waits for the power consumption of the first slave server to recover to a stable state before transmitting firmware to the first slave server, thus avoiding abnormal situations such as transmission failure or hardware failure of the first slave server during the firmware transmission process.
[0066] In one possible implementation, the interval length is determined based on the difference between the historical power consumption value and the reference power consumption value.
[0067] For example, the interval duration is positively correlated with the historical power consumption value corresponding to the server.
[0068] In this embodiment of the application, the formula for determining the interval duration is as follows:
[0069]
[0070] in, The interval duration Historical power consumption value Here, 'b' represents the baseline power consumption value, and 'b' represents the cooling coefficient, which can be set according to actual conditions. The interval duration ensures that the server's temperature and power consumption return to safe levels before the next transmission period begins.
[0071] In this embodiment, during the process of the master server sending firmware to the second slave server, the first slave server may experience a brief increase in power consumption during operation. If the firmware is immediately transmitted to the first slave server at this time, the power consumption of the first slave server will continue to increase, or even exceed the power consumption threshold, resulting in hardware damage, firmware transmission failure, etc. Therefore, after the master server completes the firmware transmission to the second slave server, it waits for a period of time corresponding to the interval to ensure that the hardware state of the first slave server recovers to a stable state before sending the firmware corresponding to the first slave server. This improves the reliability of firmware transmission and avoids further impact of firmware transmission on the first slave server.
[0072] In some embodiments, based on any of the foregoing embodiments, the first slave server includes at least one component, and the firmware corresponding to the first slave server includes component firmware corresponding to each component.
[0073] Specifically, components in a slave server refer to the hardware functional modules that make up the slave server. Examples of components include a BMC, GPU, and Field-Programmable Gate Array (FPGA). Component firmware refers to the firmware program adapted to the component, used to implement the component's basic functions (such as BMC firmware for power consumption detection, GPU firmware for computing power scheduling, etc.). Component firmware can be stored in the component's built-in storage chip.
[0074] Within the firmware transmission time corresponding to the first slave server, the firmware corresponding to the first slave server is sent to the first slave server, specifically including the following steps:
[0075] a1 retrieves the dependencies between all component firmware.
[0076] Specifically, there is a logical sequence to firmware updates for different components. If the firmware update of component A requires the firmware update of component B to be performed first, this is called component A depending on component B. In this case, when both component A and component B need to be updated, the firmware of component B must be updated first, followed by the firmware of component A. For example, PCIe interface firmware depends on the flashing channel provided by BMC.
[0077] In one possible implementation, each component firmware includes at least one function, and the dependencies between all component firmwares are the dependencies between the functions in all component firmwares.
[0078] Specifically, functions in component firmware refer to the smallest functional code blocks that constitute the firmware, capable of implementing specific sub-functions of the component, such as data reception, data verification, and hardware drivers. Different functions may have calling or being called dependencies; for example, function A requires the result of calling function B to execute. For instance, the component firmware corresponding to the PCIe port contains function 1 for establishing a PCIe link, while the component firmware corresponding to the BMC contains function 2 for initializing the flashing interface. Function 1 depends on the execution of function 2; that is, after the BMC initializes the flashing interface through function 2, the PCIe port can only establish a link through function 1. In this case, function 1 depends on function 2.
[0079] a2 determines the sending order of each component firmware based on the dependencies between all component firmware.
[0080] In this way, the sending order of each component firmware, determined by the dependencies between component firmware, can avoid the problem that the corresponding component firmware of the upper-level component cannot be flashed when the lower-level component is not ready, and reduce firmware flashing failures caused by incorrect order.
[0081] In one possible implementation, the sending order of each component firmware is determined based on the dependencies between all component firmware, specifically including the following steps:
[0082] First, based on the dependencies between functions in all component firmware, a directed graph corresponding to all component firmware is constructed.
[0083] In this system, each component firmware is represented as a vertex in the directed graph, and the dependencies between functions are represented as edges in the directed graph.
[0084] Specifically, in the constructed directed graph, each vertex uniquely corresponds to a component firmware and is the basic node unit of the directed graph, used to identify the component firmware to be sorted. Vertices can contain core attributes: component firmware name (e.g., BMC firmware), belonging slave server, etc., to facilitate subsequent determination of function dependencies and sorting. Each edge in the directed graph uniquely corresponds to a dependency relationship between a pair of functions. For example, if function 1 depends on function 2, then an edge is added to the directed graph to establish the dependency relationship between the component firmware corresponding to function 1 and the component firmware corresponding to function 2.
[0085] Then, based on the directed graph, the sending order corresponding to each component firmware is determined.
[0086] For example, a topology sorting algorithm is used to determine the transmission order of each component firmware using a directed graph.
[0087] In this embodiment of the application, the main server reverse-parses the exported symbol table of the binary file corresponding to the firmware, extracts the set of dependent function symbols in the firmware, constructs a directed graph based on the set of dependent function symbols, and then performs topological sorting on the directed graph to generate a firmware refresh sequence.
[0088] a3, within the firmware transmission time corresponding to the first slave server, sequentially sends all component firmware to the first slave server according to the sending order corresponding to each component firmware.
[0089] Furthermore, firmware for components with dependencies can be sent to the first slave server in the order of transmission, while firmware for components without dependencies can be sent to the first slave server in parallel.
[0090] In some embodiments, based on the foregoing embodiments, the method provided in this application further includes the following:
[0091] b1, obtain the power consumption change rate of the first slave server during the process of sending all component firmware to the first slave server.
[0092] Specifically, the power consumption change rate refers to the rate of change in power per unit time during the process of the master server sending component firmware to the first slave server. The power consumption change rate is used to determine whether the first slave server experiences a sudden increase in power consumption during transmission. This sudden increase in power consumption could be due to the first slave server receiving component firmware, or it could be due to the first slave server processing business data, etc.
[0093] b2, when the power consumption change rate is greater than the preset change rate threshold, or when the master server does not send all component firmware to the first slave server within the firmware transmission time, obtain at least one remaining component firmware that has not been sent, as well as the pre-configured bandwidth corresponding to the first slave server.
[0094] Specifically, the preset power change rate threshold can be set based on the hardware tolerance of the first slave server, such as the maximum output power of the power module and its heat dissipation capacity, to establish a safe upper limit for the power consumption change rate. If the actual power change rate of the first slave server exceeds the preset power change rate threshold, it may cause the hardware of the first slave server to overheat or experience power fluctuations. In such cases, it may be necessary to further split the transmitted component firmware for transmission to reduce the impact of firmware transmission on the first slave server.
[0095] Assume that in the component firmware of the first slave server, the transmission order is component firmware 1, component firmware 2, and component firmware 3. If, when the firmware transmission time is reached, component firmware 1 and component firmware 2 have been transmitted, but component firmware 3 has not yet been transmitted, then component firmware 3 is the remaining component firmware. If, before the firmware transmission time is reached, but during the transmission of component firmware 2, the power consumption change rate of the first slave server is greater than a preset change rate threshold, then the remaining component firmware consists of component firmware 2 and component firmware 3.
[0096] b3, based on the size of the first remaining component firmware and the pre-configured bandwidth, split the first remaining component firmware to determine at least one sub-firmware corresponding to the first remaining component firmware.
[0097] The first remaining component firmware is any one of at least one remaining component firmware.
[0098] Specifically, the size of the first remaining component firmware refers to the data volume of the remaining component firmware, usually in MB or GB.
[0099] The pre-configured bandwidth refers to the bandwidth pre-set between the master server and the first slave server. Each sub-firmware obtained after splitting the remaining component firmware is a small firmware fragment, thus each sub-firmware consumes less power during transmission. Compared to transmitting the component firmware sequentially, transmitting the sub-firmware helps further reduce the power consumption of the first slave server. The pre-configured bandwidth can be set according to the actual situation of the first slave server; no specific limitations are made here.
[0100] In one possible implementation, at least one sub-firmware corresponding to the first remaining component firmware is determined in the following manner:
[0101] First, the firmware splitting specification is determined based on the size of the first remaining component firmware and the pre-configured bandwidth.
[0102] Specifically, firmware splitting specifications refer to the size of each sub-firmware, which serves as the basis for dividing the remaining component firmware.
[0103] For example, the firmware split specification is determined by the following formula:
[0104]
[0105] in, For firmware split specifications, The size of the firmware for the first remaining component. For pre-configured bandwidth, As a frequency reduction factor, In this way, by introducing a frequency reduction factor, the transmission bandwidth is further reduced, the hardware resource consumption is reduced, and the system adapts to dynamic hardware conditions.
[0106] Then, according to the firmware splitting specifications, the first remaining component firmware is divided to determine at least one sub-firmware.
[0107] b4. After determining the sub-firmware corresponding to each remaining component firmware, send all sub-firmware to the first slave server in sequence according to the preset sending rules.
[0108] Specifically, the preset sending rule refers to the preset sending order of the split sub-firmware. For example, smaller sub-firmware can be sent first, and larger sub-firmware can be sent after a preset interval.
[0109] In one possible implementation, after determining the sub-firmware corresponding to each remaining component firmware, all sub-firmware are sequentially sent to the first slave server according to a preset sending rule, specifically including the following steps:
[0110] First, based on the transmission order corresponding to each remaining component firmware and the position of each sub-firmware in the remaining component firmware, the transmission order corresponding to each sub-firmware is determined.
[0111] Specifically, the position of a sub-firmware within the remaining component firmware can be indicated by its segment number after being split from the remaining component firmware. The position of each sub-firmware within the remaining component firmware can be used to identify the order of sub-firmware within the complete remaining component firmware, ensuring that sub-firmware belonging to the same remaining component firmware are transmitted in the correct order, thus avoiding data corruption when the first slave server reassembles the sub-firmware after receiving them.
[0112] Then, according to the sending order corresponding to each sub-firmware, the sub-firmware is sent to the first slave server in sequence.
[0113] For example, the remaining component firmware consists of the firmware corresponding to the BMC and the firmware corresponding to the fan. The firmware corresponding to the BMC is further divided into three sub-firmware: sub-firmware 1, sub-firmware 2, and sub-firmware 3. The firmware corresponding to the fan is divided into sub-firmware 4 and sub-firmware 5. Sub-firmware 1 is located earlier than sub-firmware 2 in the firmware corresponding to the BMC, sub-firmware 2 is located earlier than sub-firmware 3, and sub-firmware 5 is located earlier than sub-firmware 4. Furthermore, the firmware corresponding to the BMC is sent before the firmware corresponding to the fan. Therefore, the main server transmits the sub-firmware in the following order: sub-firmware 1, sub-firmware 2, sub-firmware 3, sub-firmware 5, and sub-firmware 4.
[0114] In this way, when the power consumption change rate exceeds a preset threshold, the impact of firmware transmission on the power consumption of the first slave server can be further reduced by sequentially sending each sub-firmware. Furthermore, if the master server does not send all component firmware to the first slave server within the firmware transmission time, the remaining component firmware can be split into sub-firmware. Since sub-firmware transmission takes less time, the transmission of a single sub-firmware can be completed quickly, reducing the probability of timeouts caused by prolonged link occupation. In addition, the transmission order of each sub-firmware is determined based on the transmission order of the remaining component firmware and the position of each sub-firmware within the remaining component firmware. This ensures that the remaining component firmware follows the transmission order of the remaining components while maintaining the orderliness of sub-firmware transmission, preventing the first slave server from being unable to orderly assemble the received sub-firmware.
[0115] In some embodiments, based on any of the foregoing embodiments, the master server connects to each slave server through at least one network device.
[0116] Specifically, network devices are used to support data forwarding between master and slave servers, acting as intermediary nodes in communication between them. Examples of network devices include Ethernet switches, routers, and PCIe switches.
[0117] The different network devices between the master server and the slave server form multiple transmission paths. For example, the master server can communicate with slave server 1 through network device a, and it can also communicate with slave server 1 through network device b, thus forming two transmission paths: "master server - network device a - slave server 1" and "master server - network device b - slave server 1".
[0118] In one possible implementation, the method provided in this application embodiment further includes the following steps:
[0119] First, based on the connection relationship between the i-th slave server, the network device corresponding to the i-th slave server, and the master server, at least one transmission path between the master server and the i-th slave server is determined.
[0120] Where i is a positive integer.
[0121] For example, the connection relationship between the i-th slave server, the network device corresponding to the i-th slave server, and the master server can be determined by the link layer discovery protocol.
[0122] For example, based on the connection relationship between the i-th slave server, the network device corresponding to the i-th slave server, and the master server, all possible transmission paths are enumerated.
[0123] Then, obtain the bandwidth information, latency information, and security level information corresponding to each transmission path.
[0124] Specifically, bandwidth information can be defined as the maximum amount of data that can be transmitted per unit time along a transmission path (bandwidth), reflecting the data transmission capacity of the transmission path. For example, bandwidth information can be collected in real time using network speed test tools or read from network device configurations.
[0125] Latency information refers to the time required for data to travel from the master server to the i-th slave server via the transmission path. Latency information reflects the data response speed of the transmission path. For example, latency information includes network device forwarding latency and link transmission latency.
[0126] Security level information reflects the security protection capabilities of a transmission path for data transmission. For example, security level information can be categorized into high, medium, and low levels based on the strength of the protection measures. A high-security transmission path needs to support data encryption, access control, and integrity verification, while a low-security transmission path may only support basic verification.
[0127] Finally, based on the bandwidth, latency, and security level information corresponding to each transmission path, a target transmission path is determined from all transmission paths. This ensures that when the transmission order corresponding to the i-th slave server is determined according to the preset transmission order, the firmware corresponding to the i-th slave server is transmitted from the target transmission path to the i-th slave server.
[0128] Specifically, the target transmission path is the optimal path selected from all transmission paths by comprehensively considering bandwidth, latency, and security level information, which can meet the firmware transmission rate requirements, real-time requirements, and security requirements.
[0129] For example, based on the bandwidth, latency, and security level information of each transmission path, each transmission path is quantitatively scored, and the transmission path with the highest score is selected as the target transmission path.
[0130] In this embodiment of the application, the target transmission path is determined from all transmission paths based on the bandwidth information, latency information, and security level information corresponding to each transmission path. Specifically, this includes the following steps:
[0131] First, based on the bandwidth information of the first transmission path and a preset bandwidth scoring rule, a first score corresponding to the bandwidth information is determined. Then, based on the delay information of the first transmission path and a preset delay scoring rule, a second score corresponding to the delay information is determined. Finally, based on the security level information of the first transmission path and a preset security scoring rule, a third score corresponding to the security level information is determined.
[0132] For example, when the bandwidth information includes bandwidth, the preset bandwidth scoring rules can be as follows: when the bandwidth is greater than or equal to 10GB / s, the corresponding first score is 100 points; when the bandwidth is less than 10GB / s but greater than 5GB / s, the corresponding first score is 60 points; and when the bandwidth is less than or equal to 5GB / s, the corresponding first score is 30 points.
[0133] For example, when the latency information includes latency duration, the preset latency scoring rules can be as follows: when the latency duration is greater than or equal to 1ms, the corresponding second score is 30 points; when the latency duration is less than 1ms but greater than 10us, the corresponding second score is 60 points; and when the bandwidth is less than or equal to 10us, the corresponding second score is 100 points.
[0134] For example, when the security level information includes the security level, the preset security scoring rules can be that when the security level is high, the corresponding third score is 100 points; when the security level is medium, the corresponding third score is 60 points; and when the security level is low, the corresponding third score is 30 points.
[0135] Then, based on the first score, the second data, and the third score, the total score corresponding to the first transmission path is determined.
[0136] For example, the first score, second score, and third score are weighted and summed to obtain the total score corresponding to the first transmission path. Here, the weights corresponding to the first score, the second score, and the third score can be set according to the actual situation, and are not limited here.
[0137] Finally, after determining the total score for each transmission path, the target transmission path is determined based on the total score for each transmission path.
[0138] In this embodiment, selecting transmission paths based on bandwidth and latency information ensures firmware transmission speed and real-time performance, preventing timeouts. Simultaneously, combining the security level information of each transmission path meets the firmware's data security requirements, preventing leakage or tampering during transmission and ensuring server security. Furthermore, the target transmission path can be flexibly selected based on the server's differentiated needs, enhancing the flexibility of the firmware transmission process.
[0139] In one possible implementation, multiple logical channels are created on the physical links between the master server and each slave server to distinguish the traffic corresponding to different slave servers (such as firmware 1 of slave server 1, firmware 2 of slave server 2, management data of the master server for each slave server, etc.).
[0140] For example, using Virtual Local Area Network (VLAN) technology, N_vlan VLAN channels are created in the network constructed between the master server and each slave server, where N_vlan is not less than the number of slave servers. Each channel corresponds to a slave server or the master server's management data for the slave server, and the independent transmission performance of each channel is ensured by isolating propagation domains. The manager also implements Quality of Service (QoS) policies for transmissions of different firmware types.
[0141] In some embodiments, based on any of the foregoing embodiments, the method provided in this application further includes the following:
[0142] First, obtain the device type corresponding to the i-th slave server.
[0143] Specifically, the device type corresponding to the i-th slave server can be a BMC server, FPGA server, etc. For example, the master server can read the hardware interface information of each slave server through the LLDP protocol to determine the device type corresponding to each slave server.
[0144] Then, based on the device type, the data transmission protocol corresponding to the i-th slave server is determined so that the firmware corresponding to the i-th slave server can be sent to the master server in accordance with the data transmission protocol.
[0145] Specifically, the data transmission protocol refers to the standardized rules for data interaction between the master server and the i-th slave server, including but not limited to data frame format, transmission rate control, error checking mechanism, etc. It needs to be compatible with the slave server hardware interface to ensure that the firmware can transmit and parse correctly.
[0146] For example, the data transmission protocol corresponding to the i-th slave server is determined based on the device type and the preset mapping relationship between device type and data transmission protocol.
[0147] In this embodiment, the master server establishes a multi-protocol overlay network at the Open Systems Interconnection (OSI) data link layer, dynamically selecting the data transmission protocol (also known as the tunneling protocol) based on the device type corresponding to the slave server. For example, the FPGA slave server uses User Datagram Protocol (UDP) for encapsulation and transmission, while the BMC slave server uses Transmission Control Protocol (TCP) for encapsulation and transmission. This is because UDP is connectionless and has low overhead, making it suitable for scenarios where the firmware corresponding to the FPGA slave server has high real-time requirements and can tolerate a small amount of packet loss, thus improving transmission efficiency. TCP, as a reliable connection protocol, ensures the integrity and order of the firmware corresponding to the BMC slave server during transmission, preventing firmware refresh failures of the BMC component (a critical management component) due to packet loss.
[0148] Considering that incompatibility between the data transmission protocol and the slave server's hardware could lead to data reception failure and firmware transmission failure, the master server determines the data transmission protocol by precisely matching the slave server's device type, thereby optimizing the transmission path's performance and improving firmware transmission reliability.
[0149] In some embodiments, based on any of the foregoing embodiments, the method provided in this application further includes:
[0150] During communication with each slave server, the master server encrypts the firmware corresponding to each slave server according to the preset encryption rules and sends the encrypted firmware to the first slave server.
[0151] For example, the master server and each slave server have a corresponding temporary key pair during communication for two-factor authentication of the session. The validity period of the temporary integer is less than 300 seconds, thereby enhancing security. In this way, even if the integer is leaked, its usability time is extremely short and will not pose a long-term threat to the multi-level servers.
[0152] For example, two-factor authentication is implemented based on the Fast Identity Online Universal 2nd Factor (FIDO U2F) standard.
[0153] In some embodiments, based on any of the foregoing embodiments, different component firmware can be transmitted by constructing different virtual local area networks in the transmission path between the master server and the slave server.
[0154] The method provided in this application embodiment also includes the following:
[0155] First, obtain the firmware type corresponding to the first firmware.
[0156] Specifically, the firmware type can be Basic Input / Output System (BIOS) firmware, FPGA firmware, or NIC firmware.
[0157] Then, based on the firmware type, determine the bandwidth corresponding to the first firmware.
[0158] For BIOS firmware, the bandwidth is ≥50Mbps; for FPGA firmware, the bandwidth is ≥100Mbps; and for NIC firmware, the bandwidth is ≥30Mbps.
[0159] In some embodiments, based on any of the foregoing embodiments, the server calculates resource constraints for each slave server and issues them to the slave server, and the slave server configures a local control group (cgroup) according to the constraints.
[0160] In one possible implementation, the method provided in this application embodiment further includes the following:
[0161] First, obtain the initial number of central processing unit cores contained in the main server.
[0162] The master server is any one of the slave servers.
[0163] Specifically, the first number of central processing unit cores is used to determine the total amount of CPU resources in the server and to determine the number of cores to write the firmware to.
[0164] Then, based on the first number, a second number of CPU cores used by the master server when writing firmware is determined.
[0165] For example, a first quantity with a preset ratio can be used as a second quantity.
[0166] Finally, the second quantity is sent to the main server.
[0167] For example, a CPU quota constraint is set for the firmware writing process of the first slave server, setting the CPU quota:
[0168]
[0169] In addition, memory constraints can be set for the firmware writing process of the first slave server, for example, setting a memory limit:
[0170]
[0171] in, The number of CPU cores allocated to the first slave server for firmware flashing, i.e., the second number. The first number refers to the number of CPU cores in the first slave server, i.e., the first quantity. The size of the memory allocated to the first slave server. The size of the firmware corresponding to the first slave server is set to ensure that the memory usage of the flashing process does not exceed three times the firmware size plus a fixed buffer.
[0172] In this embodiment, by setting the number of CPU cores corresponding to the slave server, on the one hand, sufficient CPU computing power is ensured for firmware writing, guaranteeing refresh efficiency. On the other hand, it prevents the slave server from preempting business CPU resources during firmware writing, ensuring the stability of the slave server's core business.
[0173] In one possible implementation, after receiving the firmware from the server, the firmware writing process is isolated by a container / control group to ensure isolation between the firmware writing process and other system services, and to control resource usage.
[0174] For example, the firmware writing process is isolated in separate control groups (cgroups) based on Extended Berkeley Packet Filter (eBPF) technology. Through the isolation and resource control mechanisms of cgroups and eBPF, it can be effectively ensured that firmware flashing does not interfere with other functions of the server.
[0175] In one possible implementation, the master server is also used to detect the temperature of each slave server during firmware writing after receiving its corresponding firmware. If the temperature of a slave server exceeds a preset temperature threshold, a firmware writing pause command is sent to each slave server, and a cooling process is initiated. After a preset pause duration, a firmware writing start command is sent to the slave server to instruct it to continue writing the firmware.
[0176] In one possible implementation, the entire refresh process on the server includes a firmware transfer process where the master server sends the firmware to the slave server, and a firmware writing process on the slave server after receiving the firmware. Furthermore, the refresh process may also include a rollback process.
[0177] For example, the rollback process includes a three-stage rollback process: In the first stage, when the master server and the first slave server fail to verify the firmware digital signature, the master server terminates the refresh process and generates an alarm signal; In the second stage, when the first slave server exceeds the preset write time for firmware writing, the first slave server performs a local backup firmware recovery operation; In the third stage, when the first slave server fails to write firmware continuously, the master server sends the firmware corresponding to the first slave server to the slave server corresponding to the first slave server, so that the slave server corresponding to the first slave server can refresh the firmware after receiving the firmware. At this time, the slave server corresponding to the first slave server is a mirrored redundant node of the first slave server.
[0178] The overall process of the firmware flashing method is described below. The master server initiates the flashing task, obtaining network topology information and slave server identities from the multi-cascaded servers via a link-layer discovery protocol. Then, based on a Kalman filter model, it determines the predicted total power consumption. When the predicted total power consumption exceeds a preset power consumption threshold, the firmware transmission duration for each slave server is determined according to the total number of slave servers, thus achieving time-sharing firmware transmission for each slave server. Figure 2 As shown, after receiving the firmware from the server, if high-priority component firmware exists in the firmware, the slave server uses cgroups and eBPF technology to isolate and refresh the high-priority component firmware, ensuring that firmware writing does not interfere with other functions of the device. If no high-priority component firmware (such as BMC firmware) exists in the firmware, the slave server refreshes each component firmware in the order of receiving the component firmware (also known as writing each component firmware). When the firmware writing is successful, the master server receives a firmware refresh completion signal from the slave server, and the firmware refresh of the slave server ends. When the firmware writing fails, the slave server rolls back the firmware to the previous version.
[0179] Figure 3 This is a diagram showing the interaction between the master server and multiple slave servers, where each server sends its own firmware information. Figure 3In this process, the master server sends its respective firmware to slave servers A, B, and C sequentially according to the firmware transmission duration. Specifically, the master server first sends firmware 1 to slave server A within the firmware transmission duration corresponding to server A, while servers B and C wait. After sending to slave server A, the master server sends firmware 2 to slave server B within the firmware transmission duration corresponding to server B, while slave server C waits. After sending to slave server B, the master server sends firmware 3 to slave server C within the firmware transmission duration corresponding to server C.
[0180] This application also provides a firmware flashing system. The system includes: a topology discovery module, an energy consumption awareness engine, a flashing sequence generator, a virtual channel manager, a security authentication module, and a three-stage rollback controller.
[0181] The topology discovery module is used to obtain network topology information and slave server identities of multi-cascaded servers through the Link Layer Discovery Protocol (LLDP). This module periodically sends and receives LLDP packets, identifies new or changed slave server connections, and provides the topology information to the master server.
[0182] The energy consumption sensing engine performs dynamic energy consumption prediction based on the Kalman filter algorithm in the Kalman prediction unit to determine the predicted total power consumption. This engine uses collected energy consumption fingerprint data and real-time measurements to update the Kalman filter state, estimating power consumption changes during the refresh process and thus supporting power consumption threshold judgment.
[0183] In this system, when the predicted total power consumption exceeds a preset power consumption threshold, a time-sharing refresh mechanism is activated. This mechanism includes: dividing the refresh window into multiple time slices, the length of each time slice being the firmware refresh duration; allocating an independent time slice to each slave server, ensuring mutual exclusion between time slices; and inserting a power cooling window between adjacent time slices, the cooling window being the time period corresponding to the aforementioned interval duration. Thus, by activating the time-sharing mechanism, dividing the refresh window into multiple mutually exclusive time slices, and inserting cooling time slots between time slices, power consumption and refresh throughput can be balanced.
[0184] A refresh sequence generator parses the firmware interface descriptors (FIDs) obtained from each slave server, extracts dependency function symbols, and constructs a firmware dependency graph (i.e., the directed graph mentioned above). The graph is then topologically sorted to generate a firmware refresh order matrix. The sequence design in the matrix ensures that components with dependencies are refreshed in the order of dependency; for independent components, they can be refreshed simultaneously in parallel channels, accelerating the overall refresh speed.
[0185] When the server refresh times out or the power consumption changes abruptly, the firmware corresponding to the server is fragmented and transmitted at a reduced frequency, dividing the firmware into sub-firmware of the aforementioned firmware splitting specifications.
[0186] The Virtual Channel Manager is used to create multiple logical channels on physical links to differentiate between different traffic types. Specifically, it utilizes VLAN technology to create N_vlan virtual LAN channels in the switched network, where N_vlan is no less than the number of slave servers. Each channel corresponds to a slave server or service type, and independent transmission performance is ensured for each channel through isolated propagation domains. The manager also implements Quality of Service (QoS) policies for transmissions of different firmware types.
[0187] For critical firmware such as the BIOS, the highest priority and lowest bandwidth guarantee are assigned to ensure timely updates even during network congestion. This manager uses a dynamic path selection algorithm to choose the highest-scoring path among multiple available paths. Path scores are weighted by bandwidth, latency, and security level, for example:
[0188]
[0189] in, The total score is represented by Bandwidth, path bandwidth, Latency (normalized path latency), and Security (path security level). This multi-metric scoring mechanism comprehensively considers QoS indicators such as bandwidth and latency, as well as security attributes, to ensure that the selected transmission path is both efficient and secure.
[0190] The security authentication module implements two-factor authentication and key management for firmware refresh sessions. Internally, it integrates a Physically Unclonable Function (PUF) chip to generate a unique identifier for each slave server. At the start of each refresh session, this unit dynamically generates a temporary session key based on the slave server's PUF identifier and session conditions using an HMAC-based Key Derivation Function (HKDF), and uses this key in conjunction with a short-term digital certificate for key exchange and verification. This HKDF-based key derivation method securely generates encryption keys from a shared secret, and the short-term certificate used has a validity period of less than 300 seconds, resisting replay attacks.
[0191] The three-stage rollback controller monitors the status of the refresh process and executes corresponding recovery strategies based on different error types, specifically including three stages of response:
[0192] Phase 1: Immediately terminate the refresh process and generate an alarm signal to notify the management system if firmware digital signature verification fails. This prevents unauthorized or tampered firmware from being written to the slave server.
[0193] Second stage: If the refresh process exceeds the preset timeout period, it will roll back to the local backup firmware and continue to ensure the functionality of the slave server through the recovery operation; at the same time, the failure information will be recorded for subsequent analysis.
[0194] Phase 3: If multiple refresh attempts fail consecutively, a redundancy strategy is triggered, switching the refresh task to a mirrored redundant slave server in the cluster to ensure the reliability and availability of the update operation from a hardware perspective.
[0195] In this system, on the one hand, by integrating energy consumption fingerprinting with a Kalman filter-based power consumption prediction algorithm, the system can accurately predict the power consumption trend of each slave server during the refresh period and dynamically adjust the refresh plan. This significantly reduces the peak load caused by concurrent refreshes, avoids redundant power supply design and power module overload, effectively saves energy resources, and improves the overall energy efficiency ratio of the system. On the other hand, by introducing a Firmware Interface Descriptor (FID) mechanism and a three-stage rollback control process, the system ensures high compatibility and fault controllability of the refresh process in multi-vendor, heterogeneous server clusters. FID supports semantic parsing and dynamic verification of the firmware structure, and can quickly recover through the rollback mechanism in case of system anomalies, effectively reducing the risk of service interruption caused by refresh failures. Furthermore, by constructing a refresh window time-sharing scheduling algorithm and a virtual channel QoS policy table, the system dynamically plans the refresh order and time slice of each slave server in a multi-cascaded architecture based on link layer topology discovery and resource awareness. Combined with the eBPF resource isolation mechanism, it can achieve isolated scheduling of high-priority services during refresh, ensuring zero interference of refresh behavior with core services, thereby improving the intelligence and controllability of system operation and maintenance.
[0196] Figure 4 This is a schematic diagram illustrating an application scenario for a firmware flashing system. Figure 4 In this multi-cascaded server cluster, the master server obtains network topology information and slave server identities through a topology discovery module and discovery protocol. It creates multiple logical channels on the physical link to differentiate between different traffic flows through a virtual channel manager. An energy consumption awareness engine determines the predicted total power consumption based on the Kalman filter algorithm in the Kalman prediction unit. A refresh sequence generator generates a firmware refresh order matrix, thus determining the sending order of firmware for each component on each slave server. During sub-firmware transmission, the master server implements two-factor authentication and key management for the firmware refresh session through a security authentication module. Furthermore, the master server monitors the refresh process status through a three-stage rollback controller and executes corresponding recovery strategies based on different error types.
[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0198] This application also provides a firmware flashing device for implementing the above embodiments and preferred embodiments, which will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0199] This embodiment provides a firmware flashing device, such as... Figure 5 As shown, it includes:
[0200] The acquisition module 501 is used to acquire the total power consumption value within a preset historical time period based on the current time. The total power consumption value is the power consumption value consumed by the master server and all slave servers.
[0201] The prediction module 502 is used to predict the total power consumption value within a future preset time period based on the current time, according to the total power consumption value and the pre-built power consumption prediction model.
[0202] The determination module 503 is used to determine the firmware transmission duration for each slave server based on the total number of all slave servers when the predicted total power consumption value is greater than the preset power consumption threshold.
[0203] The sending module 504 is used to send the firmware corresponding to each slave server to the slave server in a preset sending order within the firmware transmission time corresponding to each slave server, so that each slave server can perform firmware refresh.
[0204] In one possible implementation, before sending the firmware corresponding to the first slave server to the first slave server during the firmware transmission time corresponding to the first slave server, the acquisition module 501 is further used to acquire the reference power consumption value corresponding to the first slave server, and the historical power consumption value corresponding to the first slave server during the firmware transmission process of the second slave server, wherein the first slave server and the second slave server are both one of a plurality of slave servers, and in a preset transmission order, the second slave server is the slave server preceding the first slave server.
[0205] The determining module 503 is further configured to determine, based on the reference power consumption value and the historical power consumption value, the interval duration after transmitting the firmware corresponding to the second slave server to the second slave server, the interval duration being used to indicate that after a time period corresponding to the interval duration, the firmware corresponding to the first slave server will be sent to the first slave server again.
[0206] In one possible implementation, the first slave server includes at least one component, and the firmware corresponding to the first slave server includes the component firmware corresponding to each component. The sending module 504 is specifically used to obtain the dependency relationship between all component firmware.
[0207] Based on the dependencies between all component firmware, determine the sending order corresponding to each component firmware;
[0208] Within the firmware transmission time corresponding to the first slave server, all component firmware are sent to the first slave server in sequence according to the sending order corresponding to each component firmware.
[0209] In one possible implementation, the acquisition module 501 is further configured to acquire the power consumption change rate of the first slave server during the process of sending all component firmware to the first slave server; when the power consumption change rate is greater than a preset change rate threshold, or when the master server does not send all component firmware to the first slave server within the firmware transmission time, the module acquires at least one remaining component firmware that has not been sent, as well as the pre-configured bandwidth corresponding to the first slave server.
[0210] The determining module 503 is further configured to split the first remaining component firmware according to the size of the first remaining component firmware and the pre-configured bandwidth, and determine at least one sub-firmware corresponding to the first remaining component firmware, wherein the first remaining component firmware is any one of the at least one remaining component firmware;
[0211] The sending module 504 is further configured to split the first remaining component firmware according to the size of the first remaining component firmware and the pre-configured bandwidth, and determine at least one sub-firmware corresponding to the first remaining component firmware, wherein the first remaining component firmware is any one of the at least one remaining component firmware;
[0212] In one possible implementation, the determining module 503 is specifically used to determine the firmware splitting specification based on the size of the first remaining component firmware and the pre-configured bandwidth;
[0213] Based on the firmware splitting specifications, the first remaining component firmware is divided to determine at least one sub-firmware.
[0214] In one possible implementation, the sending module 504 is specifically used to determine the sending order of each sub-firmware according to the sending order corresponding to each remaining component firmware and the position of each sub-firmware in the remaining component firmware.
[0215] According to the sending order of each sub-firmware, the sub-firmware is sent to the first slave server in sequence.
[0216] In one possible implementation, each component firmware includes at least one function, and the dependencies between all component firmwares are the dependencies between the functions in all component firmwares.
[0217] In one possible implementation, the determining module 503 is specifically used to construct a directed graph corresponding to all component firmware based on the dependencies between functions in all component firmware, wherein each component firmware is a vertex in the directed graph and the dependencies between functions are edges in the directed graph.
[0218] Based on the directed graph, determine the transmission order corresponding to each component firmware.
[0219] In one possible implementation, the master server connects to each slave server through at least one network device; the determining module 503 is also used for the master server to connect to each slave server through at least one network device.
[0220] In one possible implementation, the acquisition module 501 is also used to acquire the device type corresponding to the i-th slave server;
[0221] The determination module 503 is also used to determine the data transmission protocol corresponding to the i-th slave server according to the device type, so that the firmware corresponding to the i-th slave server can be sent to the master server in accordance with the data transmission protocol.
[0222] The apparatus provided in this application, in its first aspect, during the firmware update process of the master server on multiple slave servers, predicts the total power consumption of the future time period based on the total power consumption of the master server and slave servers in historical time periods. If the predicted total power consumption exceeds a preset power consumption threshold, the firmware update of each slave server is dynamically adjusted. This effectively avoids peak loads caused by concurrent updates during firmware update, thus preventing risks such as power overload and hardware overheating, and ensuring the safe and stable operation of the multi-cascaded server, including the master server and multiple slave servers. In its second aspect, the firmware transmission duration of each slave server is dynamically determined based on the total number of slave servers. This avoids the problem of some slave servers experiencing timeouts or resource limitations when using fixed duration allocation. Simultaneously, by transmitting each firmware in a preset sending order, the instantaneous peak power consumption generated by parallel firmware updates on multiple slave servers is reduced. In this application, by combining power consumption values with firmware updates, the power consumption fluctuation problem during firmware update in multi-cascaded servers is solved. Furthermore, by allocating transmission duration to each slave server and transmitting each firmware in sequence, the reliability and efficiency of firmware updates are improved, providing a foundation for the stable operation of multi-cascaded servers.
[0223] For a description of the features in the corresponding embodiment of the firmware flashing device, please refer to the relevant description of the corresponding embodiment of the firmware flashing method, which will not be repeated here.
[0224] Embodiments of this application also provide an electronic device, such as... Figure 6 As shown, it includes a memory 10 and a processor 20. The memory 10 stores a computer program, and the processor 20 is configured to run the computer program to perform the steps in any of the firmware flashing method embodiments described above.
[0225] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the firmware flashing method embodiments described above when running.
[0226] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0227] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the firmware flashing method embodiments described above.
[0228] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the firmware flashing method embodiments described above.
[0229] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0230] The firmware flashing method and electronic device provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A firmware flashing method, characterized in that, Applied to a master server, wherein the master server is connected to each of a plurality of slave servers; the method includes: Obtain the total power consumption value within a preset historical time period based on the current time, wherein the total power consumption value is the power consumption value consumed by the master server and all the slave servers; Based on the total power consumption value and the pre-built power consumption prediction model, predict the total power consumption value within a future preset time period with the current time as the reference. When the predicted total power consumption value is greater than the preset power consumption threshold, the firmware transmission duration corresponding to each of the slave servers is determined according to the total number of all the slave servers. According to the preset sending order, within the firmware transmission time corresponding to each of the slave servers, the firmware corresponding to each of the slave servers is sent to the slave server in sequence, so that each of the slave servers can perform firmware refresh respectively; Before sending the firmware corresponding to the first slave server to the first slave server within the firmware transmission time corresponding to the first slave server, the method further includes: Obtain the baseline power consumption value corresponding to the first slave server, and the historical power consumption value corresponding to the first slave server during the firmware transmission process of the second slave server, wherein the first slave server and the second slave server are both one of the multiple slave servers, and in the preset transmission order, the second slave server is the slave server preceding the first slave server; Based on the baseline power consumption value and the historical power consumption value, the interval duration after transmitting the firmware corresponding to the second slave server to the second slave server is determined. The interval duration is used to indicate that after a time period corresponding to the interval duration, the firmware corresponding to the first slave server is sent to the first slave server again. The first slave server includes at least one component, and the firmware corresponding to the first slave server includes component firmware corresponding to each of the components. During the firmware transmission time corresponding to the first slave server, sending the firmware corresponding to the first slave server to the first slave server includes: Obtain the dependencies between all the component firmware; Based on the dependencies between all the component firmware, determine the sending order corresponding to each component firmware; Within the firmware transmission time corresponding to the first slave server, all component firmwares are sent to the first slave server in sequence according to the sending order corresponding to each component firmware. The method further includes: Obtain the power consumption change rate of the first slave server during the process of sending all the component firmware to the first slave server; When the power consumption change rate is greater than the preset change rate threshold, or when the master server fails to send all the component firmware to the first slave server within the firmware transmission time, the master server obtains at least one remaining component firmware that has not been sent, as well as the pre-configured bandwidth corresponding to the first slave server. Based on the size of the first remaining component firmware and the pre-configured bandwidth, the first remaining component firmware is split to determine at least one sub-firmware corresponding to the first remaining component firmware, wherein the first remaining component firmware is any one of at least one of the remaining component firmware; Once the sub-firmware corresponding to each of the remaining component firmwares is determined, all the sub-firmwares are sent to the first slave server in sequence according to the preset sending rules. The step of splitting the first remaining component firmware according to the size of the first remaining component firmware and the pre-configured bandwidth, and determining at least one sub-firmware corresponding to the first remaining component firmware, includes: The firmware splitting specification is determined based on the size of the first remaining component firmware and the pre-configured bandwidth; According to the firmware splitting specifications, the first remaining component firmware is divided to determine at least one of the sub-firmware; Firmware split specifications are determined by the following formula: in, For firmware split specifications, The size of the firmware for the first remaining component. For pre-configured bandwidth, This is the frequency reduction factor.
2. The method according to claim 1, characterized in that, After determining the sub-firmware corresponding to each of the remaining component firmwares, the step of sequentially sending all the sub-firmwares to the first slave server according to a preset sending rule includes: The transmission order of each sub-firmware is determined based on the transmission order corresponding to each of the remaining component firmwares and the position of each sub-firmware in the remaining component firmwares. According to the sending order corresponding to each of the sub-firmware, the sub-firmware is sent to the first slave server in sequence.
3. The method according to claim 1, characterized in that, Each component firmware includes at least one function, and the dependencies between all component firmwares are the dependencies between each function in all component firmwares; The step of determining the transmission order of each component firmware according to the dependencies between all the component firmware includes: Based on the dependencies between the functions in all the component firmware, a directed graph corresponding to all the component firmware is constructed, wherein each component firmware is a vertex in the directed graph, and the dependencies between the functions are edges in the directed graph; Based on the directed graph, the transmission order corresponding to each component firmware is determined.
4. The method according to claim 1, characterized in that, The master server is connected to each of the slave servers via at least one network device; the method further includes: Based on the connection relationship between the i-th slave server, the network device corresponding to the i-th slave server, and the master server, at least one transmission path between the master server and the i-th slave server is determined, where i is a positive integer; Obtain the bandwidth information, latency information, and security level information corresponding to all the transmission paths; Based on the bandwidth information, latency information, and security level information corresponding to all the transmission paths, a target transmission path is determined from all the transmission paths so that when the transmission order corresponding to the i-th slave server is determined according to the preset transmission order, the firmware corresponding to the i-th slave server is transmitted from the target transmission path to the i-th slave server.
5. The method according to claim 4, characterized in that, The method further includes: Obtain the device type corresponding to the i-th slave server; Based on the device type, a data transmission protocol corresponding to the i-th slave server is determined so that the firmware corresponding to the i-th slave server can be sent to the master server in accordance with the data transmission protocol.
6. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the firmware flashing method as described in any one of claims 1-5 when executing the computer program.
Citation Information
Patent Citations
Firmware updating method, electronic device and control system
CN109933352A