Server cluster maintenance method and device, electronic equipment and storage medium

Through the server cluster maintenance method of automated grouping upgrades and business migration, the server cluster maintenance problems are solved, and efficient and reliable server cluster management is achieved, suitable for large-scale data centers and cloud computing scenarios.

CN120276753APending Publication Date: 2025-07-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510421214.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the maintenance efficiency of server clusters is low and easy to miss, and firmware upgrades are prone to compatibility conflicts and downtime accidents, lack of collaborative management mechanisms, insufficient security, and cannot meet the requirements of high security demand scenarios.

Method used

Through firmware management module, multi-node collaboration module and security management module, automated group upgrades and business traffic migration of server clusters are realized, combining AI predictive maintenance and blockchain evidence storage to ensure the security and reliability of firmware upgrades.

Benefits of technology

It improves the maintenance efficiency of server clusters, reduces the impact of firmware upgrades on business, ensures the normal business capabilities of server clusters, enhances security and reliability, and is suitable for large-scale data centers and cloud computing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276753A_ABST
    Figure CN120276753A_ABST
Patent Text Reader

Abstract

The invention discloses a server cluster maintenance method and device, electronic equipment and a storage medium, and relates to the technical field of server intelligent manufacturing and operation maintaining.The server cluster maintenance method comprises the steps that after a firmware upgrading request is received, servers in a server cluster can be automatically divided into two groups, firmware upgrading is conducted in one group, and then firmware upgrading is conducted in the other group; and when it is monitored that the firmware of the servers in the group is successfully upgraded, firmware upgrading of the servers in the other group is carried out, so that centralized maintenance and automatic firmware upgrading of the servers in the server cluster are realized, a user does not need to manually carry out firmware upgrading operation on the servers in the server cluster one by one, and the user experience is improved. Therefore, the technical problem that the maintenance efficiency of the server cluster is low can be solved, and the technical effect of improving the maintenance efficiency of the server cluster is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of server intelligent manufacturing and operation and maintenance, and particularly relates to a maintenance method, device, electronic device and storage medium for a server cluster. Background Art

[0002] With the continuous development of information technology, server clusters have become an important infrastructure for enterprises and organizations to process large amounts of data and provide critical services. As the scale of server clusters continues to expand, their management and maintenance work has become increasingly complex.

[0003] Currently, in related technologies, the maintenance of server clusters is usually carried out on a per-server basis. For example, the firmware of servers is upgraded by manually operating each server one by one. This maintenance method is time-consuming and prone to omissions, resulting in low maintenance efficiency. Summary of the Invention

[0004] This application provides a maintenance method, device, electronic device and storage medium for a server cluster to at least solve the problem of low efficiency in maintaining servers one by one in related technologies.

[0005] This application provides a maintenance method for a server cluster, including:

[0006] Responding to receiving a firmware upgrade request, determining a target firmware version to be upgraded;

[0007] Based on the status information and current load information of servers in the server cluster, dividing the servers in the server cluster into a first group and a second group;

[0008] Sending the target firmware corresponding to the target firmware version to the servers in the first group for firmware upgrade;

[0009] When it is monitored that the firmware upgrade of the servers in the first group is successful, switching the service traffic of the servers in the second group to the first group, and sending the target firmware to the servers in the second group for firmware upgrade.

[0010] This application also provides a maintenance device for a server cluster, including:

[0011] A version determination module, configured to respond to receiving a firmware upgrade request and determine a target firmware version to be upgraded;

[0012] A grouping module, configured to divide the servers in the server cluster into a first group and a second group based on the status information and current load information of servers in the server cluster;

[0013] A sending module, configured to send the target firmware corresponding to the target firmware version to the servers in the first group for firmware upgrade;

[0014] A service migration module, configured to switch the service traffic of the servers in the second group to the first group when it is detected that the firmware upgrade of the servers in the first group is successful;

[0015] The sending module is further configured to send the target firmware to the servers in the second group for firmware upgrade.

[0016] This application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above server cluster maintenance methods when executing the computer program.

[0017] This application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above server cluster maintenance methods are implemented.

[0018] This application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above server cluster maintenance methods are implemented.

[0019] Through this application, after receiving a firmware upgrade request, the servers in the server cluster can be automatically divided into two groups. First, firmware upgrade is performed in one group, and when it is detected that the firmware upgrade of the servers in this group is successful, the firmware upgrade of the servers in the other group is then performed. Thus, centralized maintenance and automatic firmware upgrade of the servers in the server cluster are achieved. There is no need for the user to manually perform firmware upgrade operations on each server in the server cluster one by one, improving the server maintenance efficiency. Moreover, by switching the service traffic of these servers to the previously upgraded first group before performing the firmware upgrade on the servers in the other group, the impact of the firmware upgrade on the service can be reduced, ensuring the normal service capabilities of the server cluster. Therefore, the technical problem of low server cluster maintenance efficiency can be solved, and the technical effect of improving the server cluster maintenance efficiency can be achieved. Description of the Drawings

[0020] To more clearly illustrate the embodiments of this application, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1Schematic diagram of the maintenance system architecture for implementing the maintenance method of the server cluster provided by an exemplary embodiment of the present application;

[0022] Figure 2 Flow chart of a maintenance method for a server cluster provided by an embodiment of the present application;

[0023] Figure 3 Shows the schematic diagram of the end-to-end secure link implementation process of an exemplary embodiment of the present application;

[0024] Figure 4 Shows the schematic diagram of the process of multi-node collaborative fault tolerance upgrade of an exemplary embodiment of the present application;

[0025] Figure 5 Flow chart of another maintenance method for a server cluster provided by an embodiment of the present application;

[0026] Figure 6 Shows the schematic diagram of the firmware release process of an exemplary embodiment of the present application;

[0027] Figure 7 Flow chart of yet another maintenance method for a server cluster provided by an embodiment of the present application;

[0028] Figure 8 Shows the schematic diagram of the predictive maintenance process based on the AI model of an exemplary embodiment of the present application;

[0029] Figure 9 Schematic diagram of the structure of a maintenance device for a server cluster provided by an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0031] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0032] Currently, in related technologies, the maintenance of servers mainly revolves around server management tools (such as Intelligent Platform Management Interface (IPMI), the open standard data model and protocol Redfish developed by the Distributed Management Task Force (DMTF)), and traditional operation and maintenance processes, mainly including the following aspects:

[0033] I. Basic monitoring function

[0034] (1) IPMI: A standardized hardware management interface used to monitor the physical state of servers, such as temperature, voltage, fan speed, etc. IPMI supports remote monitoring and control of servers through Out-of-Band (OOB);

[0035] (2) Redfish: A server management protocol based on the Representational State Transfer (REST) Application Programming Interface (API), used for remote management and monitoring of servers. Redfish provides a more modern interface than IPMI and supports more flexible management functions.

[0036] II. Firmware upgrade and management

[0037] (1) Firmware upgrades such as the Basic Input Output System (BIOS) and Baseboard Manager Controller (BMC) usually require one-by-one operation, which is time-consuming and prone to omission, and firmware updates may cause compatibility conflicts and lead to downtime accidents.

[0038] III. Hardware changes and software configuration:

[0039] (1) Hardware and software configurations (such as virtualization platforms) such as the replacement of the Central Processing Unit (CPU) and memory modules require multiple cross-team coordinations, and the processes are fragmented. Currently, there is a lack of a collaborative management mechanism, resulting in low efficiency.

[0040] IV. Maintenance process

[0041] (1) The existing maintenance process lacks the ability to dynamically adjust and cannot optimize the maintenance strategy according to the server type (such as rack-mounted, blade) or business load.

[0042] V. Security Shortcomings

[0043] (1) Remote maintenance lacks end-to-end encryption and operation auditing, posing a risk of data leakage.

[0044] The current server maintenance solutions mainly face the following problems:

[0045] I. Low Maintenance Efficiency

[0046] (1) Firmware upgrade requires operation on each server one by one, which is time-consuming and prone to omission. The related technologies lack the ability of batch operation and automated management;

[0047] (2) Coordination among multiple teams is required multiple times for hardware changes and software configurations. The process is fragmented, and the related technologies lack a collaborative management mechanism, resulting in low efficiency.

[0048] II. High Maintenance Risks

[0049] (1) Firmware update may lead to compatibility conflicts and cause downtime accidents. The related technologies lack an effective compatibility verification mechanism;

[0050] (2) When maintaining multiple servers in a server cluster, manual operation is prone to cause cascading failures.

[0051] III. Weak Iteration Ability

[0052] (1) Importing new hardware requires re-adapting drivers and firmware, and the verification cycle is as long as several weeks;

[0053] (2) Lack of real-time health monitoring and predictive maintenance, and fault repair depends on post-event response;

[0054] (3) Lack of a closed-loop optimization mechanism, and historical operation and maintenance data is not effectively used for firmware version iteration or hardware compatibility improvement.

[0055] IV. Existence of Security Shortcomings

[0056] (1) Remote maintenance lacks end-to-end encryption and operation auditing, posing a risk of data leakage. The related technologies have obvious deficiencies in terms of security and cannot meet the requirements of high-security demand scenarios (such as finance and government affairs).

[0057] In summary, the server maintenance tools and processes in the related technologies have obvious limitations in mass production server maintenance, especially in scenarios such as large-scale data centers, cloud computing server clusters, and edge computing nodes.

[0058] In view of the above problems, the present application provides a systematic collaborative management solution for server mass production maintenance. Through firmware driver maintenance, multi-node collaborative scheduling, and security enhancement mechanisms, the problems of low efficiency, high risk, and slow iteration in server cluster maintenance are solved, realizing the efficient and reliable operation and maintenance of large-scale server clusters. It is particularly applicable to the mass production maintenance and diagnosis scenarios of large-scale data centers, cloud computing server clusters, and edge computing nodes, and can meet the requirements of high reliability, multi-node collaboration, and remote management.

[0059] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0060] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the server cluster maintenance method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0061] Figure 1 The schematic diagram of the maintenance system architecture for implementing the server cluster maintenance method provided in an exemplary embodiment of the present application is shown in Figure 1As shown in the figure, the maintenance system mainly includes the following core modules: firmware management module, multi-node coordination module, predictive maintenance module, and security management module. Among them, the firmware management module supports batch signing, version verification, and gray release of firmware such as BIOS, BMC, and Complex Programmable Logic Device (CPLD). It has a built-in compatibility rule library (such as firmware version X is only compatible with CPU model Y), and dynamically associates with hardware changes (Engineering Change, EC). For example, adding a Non-Volatile Memory Express (NVMe) hard drive triggers an automatic driver adaptation test; the multi-node coordination module realizes parallel maintenance of server clusters based on a distributed task engine; the predictive maintenance module collects hardware health data (such as temperature, power consumption, Error Correction Code (ECC) errors, etc.) in real time through the BMC or IPMI interface, combines with an Artificial Intelligence (AI) model to predict faults, and supports the Secure Shell (SSH) and OOB protocols. All operation records are encrypted and stored to generate audit logs; the security management module signs firmware packages and maintenance instructions using national cryptographic algorithms to prevent malicious tampering, and restricts maintenance permissions through Role-Based Access Control (RBAC). Multi-factor authentication (MFA) is required for critical operations.

[0062] such as Figure 1As shown in the figure, the maintenance system can interact with user terminals and server clusters. Users can send maintenance instructions (such as firmware upgrades) through user terminals. The security management module performs national cryptography signature and Quantum Key Distribution (QKD) encryption on the maintenance instructions to ensure transmission security. In addition, all operation records during the maintenance process can be stored on the blockchain through the security management module to meet compliance audit requirements such as Equal Protection 2.0 and the General Data Protection Regulation (GDPR). The multi-node collaboration module decomposes the maintenance tasks into parallel subtasks and dynamically selects maintenance nodes (such as the blue-green deployment mode) based on the current health status and business load of the server cluster. When the firmware management module performs a firmware upgrade, it generates an incremental update package according to the task and interacts with the servers in the server cluster to send the incremental update package to the servers for firmware upgrade. The server cluster executes the firmware update operation and reports health data to the predictive maintenance module in real time through BMC or IPMI. The predictive maintenance module analyzes the data. If a high-risk fault is found, it triggers an alarm and optimizes the maintenance strategy (such as giving priority to replacing spare parts). In addition, in this system, the predictive maintenance module supports an adaptive learning mechanism. The maintenance results and fault data are fed back to the predictive maintenance module to update the parameters of the AI model of the predictive maintenance module, thereby improving the prediction accuracy of the AI model.

[0063] It should be noted that Figure 1 the architecture shown is only used as an example to explain this application and cannot be used as a limitation to this application.

[0064] The maintenance system provided by this application, in response to the high-reliability, large-scale, and security requirements in the server field, provides an intelligent and secure mass production maintenance solution for data center and cloud computing scenarios through firmware drive, AI prediction, and multi-node collaboration technologies.

[0065] Embodiments of this application provide a method for maintaining a server cluster. In combination with the execution process of the method for maintaining a server cluster, the method will be described in detail.

[0066] Figure 2 It is a schematic flowchart of a method for maintaining a server cluster provided by an embodiment of this application. This method can be executed by the device for maintaining a server cluster provided by an embodiment of this application and can be integrated in an electronic device.

[0067] As Figure 2 shown, the method for maintaining this server cluster includes the following steps:

[0068] Step 101, in response to receiving a firmware upgrade request, determine the target firmware version to be upgraded.

[0069] Exemplarily, a firmware upgrade request can be initiated by a user. For example, the user can initiate a firmware upgrade request through the user terminal in their possession. Among them, the server model can be carried in the firmware upgrade request. Then, according to the adaptation relationship between different firmware versions and the server model, the firmware version adapted to the server model carried in the firmware upgrade request can be determined as the target firmware version to be upgraded.

[0070] Optionally, there may be more than one firmware version that can adapt to the same server model. To ensure that the firmware version of the server is the latest, the latest firmware version can be selected from multiple firmware versions that adapt to the server model as the target firmware version. Or, to ensure the success rate of the firmware upgrade, according to the firmware version currently deployed in the server of this server model, the firmware version that is closest to the currently deployed firmware version and newer than the currently deployed firmware version can be selected from multiple firmware versions that adapt to the server model as the target firmware version.

[0071] Step 102: Based on the status information and current load information of the servers in the server cluster, divide the servers in the server cluster into a first group and a second group.

[0072] Among them, the status information and current load information of the servers can be obtained through BMC or IPMI. The status information can be information reflecting the hardware health status of the servers, and the current load information can be, but is not limited to, CPU occupancy, hard disk occupancy, etc.

[0073] In this embodiment, when performing firmware upgrade on the servers in the server cluster, the servers can be grouped first according to the status information and current load information of each server in the server cluster, and the servers in the server cluster are divided into a first group and a second group.

[0074] Exemplarily, the servers with better status information (such as no hardware failures) and smaller current load information can be divided into the first group, and the remaining servers can be divided into the second group. Additionally, optionally, if the number of servers with good status and small load is large, a certain number of servers can be selected from them and divided into the first group, and the remaining servers can be divided into the second group to ensure that the number of servers in the first group is small and reduce the impact of firmware upgrade on the business.

[0075] Step 103: Send the target firmware corresponding to the target firmware version to the servers in the first group for firmware upgrade.

[0076] In this embodiment, after determining the first group and the second group, the target firmware corresponding to the target firmware version can be sent to the servers in the first group, and the servers in the first group can perform firmware upgrade after receiving the target firmware.

[0077] Optionally, in order to reduce the size of the firmware package and the transmission time, an incremental update package can be determined according to the difference between the firmware currently deployed on the server in the first group and the target firmware, and the incremental update package is sent to the servers in the first group for firmware upgrade.

[0078] Step 104, when it is monitored that the firmware upgrade of the servers in the first group is successful, switch the service traffic of the servers in the second group to the first group, and send the target firmware to the servers in the second group for firmware upgrade.

[0079] In this embodiment, after sending the target firmware to the servers in the first group for firmware upgrade, the servers in the first group can be monitored, and when it is monitored that the firmware upgrades of all the servers in the first group are successful, the service traffic of each server in the second group is switched to the servers in the first group, and then the target firmware is sent to the servers in the second group for firmware upgrade.

[0080] Exemplarily, a health check can be performed on the servers in the first group. If the health check passes, a feedback of successful upgrade will be received, so that it can be determined that the firmware upgrade of the server is successful; if the health check fails, a feedback of failed health check will be received, so that it can be determined that the firmware upgrade of the server fails.

[0081] For the method for maintaining a server cluster according to the embodiment of the present application, after receiving a firmware upgrade request, the servers in the server cluster can be automatically divided into two groups. First, the firmware upgrade is performed in one group, and when it is monitored that the firmware upgrade of the servers in this group is successful, the firmware upgrade of the servers in the other group is performed. Thus, centralized maintenance and automatic firmware upgrade of the servers in the server cluster are achieved. There is no need for the user to manually perform firmware upgrade operations on each server in the server cluster one by one, which improves the server maintenance efficiency. Moreover, by switching the service traffic of these servers to the previously upgraded first group before performing the firmware upgrade on the servers in the other group, the impact of the firmware upgrade on the service can be reduced, and the normal service capacity of the server cluster can be ensured. Therefore, the technical problem of low server cluster maintenance efficiency can be solved, and the technical effect of improving the server cluster maintenance efficiency can be achieved.

[0082] In an alternative embodiment of the present application, when sending the target firmware to the server, such as when sending the target firmware to the servers in the first group and / or when sending the target firmware to the servers in the second group, the integrity of the target firmware corresponding to the target firmware version can be verified first. For example, the target firmware can be verified by calculating the hash value of the firmware package of the target firmware. If the verification passes, it is determined that the integrity verification passes. Then, in the case where the verification passes, a key is obtained to encrypt the target firmware to obtain an encrypted firmware, where the key can be a key generated using a preset encryption algorithm. For example, the QKD algorithm can be used to generate the key. Furthermore, the encrypted firmware is sent to the servers in the first group for firmware upgrade.

[0083] In this embodiment, by verifying the integrity of the target firmware, the integrity of the firmware package sent to the server for upgrade can be ensured, thereby reducing the probability of firmware upgrade failure caused by an incomplete firmware package; after the integrity verification passes, the target firmware is encrypted using a key and then sent to the servers in the first group for firmware upgrade. By encrypting the transmission of the target firmware, the security of the target firmware transmission is ensured, and malicious attacks can be effectively avoided.

[0084] Figure 3 The figure shows a schematic diagram of the end-to-end secure link implementation process of an exemplary embodiment of the present application, and the target firmware can be securely transmitted to the server for firmware upgrade. As Figure 3 shown, before transmission, the quantum key is first generated using QKD technology, and then the hash value of the firmware package of the target firmware is calculated for hash verification to determine whether the target firmware passes the integrity verification. If the verification fails, the firmware upgrade process is interrupted, and the corresponding security event is recorded for maintenance personnel to trace. If the verification is successful, the generated quantum key is used to encrypt the firmware package to obtain the encrypted firmware, and the encrypted firmware is sent to the target node (i.e., each server in the first group) through the secure link. After receiving the encrypted firmware, the target node decrypts it in the Trusted Execution Environment (TEE) to obtain the decrypted firmware package, and verifies the integrity of the decrypted firmware package. If the verification fails, the firmware upgrade process is interrupted, and the corresponding security event is recorded; if the verification is successful, the obtained firmware package is deployed on the target node to perform the firmware upgrade operation. Then, all operations are recorded on the blockchain to generate an immutable audit log. Thus, through firmware package integrity verification, encryption, decryption in the trusted execution environment, and integrity verification, the integrity and security of the firmware are ensured, which helps to improve the success rate of firmware upgrade, and by recording the operations on the blockchain, it is ensured that the operation records cannot be tampered with, guaranteeing data security.

[0085] Among them, the specific implementation code for firmware encryption and blockchain evidence storage is as follows:

[0086] from cryptography.hazmat.primitives import hashes

[0087] from blockchain import Blockchain

[0088] def encrypt_firmware(firmware,qkd_key):

[0089] cipher = AES.new(qkd_key,AES.MODE_GCM)

[0090] encrypted_data,tag = cipher.encrypt_and_digest(firmware)

[0091] return encrypted_data,cipher.nonce,tag

[0092] def log_to_blockchain(operation,user):

[0093] blockchain = Blockchain()

[0094] block = blockchain.create_block({

[0095] 'operation':operation,

[0096] 'user':user,

[0097] 'timestamp':datetime.now()

[0098] })

[0099] return block.hash

[0100] # Call example:

[0101] encrypted_fw,nonce,tag = encrypt_firmware(firmware_data,qkd_key)

[0102] block_hash = log_to_blockchain("Firmware upgrade","admin")

[0103] Servers of the same model may not use exactly the same hardware. Therefore, the target firmware may not be fully compatible in servers of the same model, resulting in the failure of firmware upgrade based on the target version. In an alternative embodiment of the present application, in response to this phenomenon, after sending the target firmware to the servers in the first group for firmware upgrade, if it is detected that the firmware upgrade of the first server in the first group fails, the first server is controlled to execute a preset rollback script to roll back to the firmware version before the upgrade, where the first server can be any server in the first group. If it is detected that the firmware upgrade of a certain server in the first group fails, an automatic rollback is triggered based on the rollback script to roll back the firmware of the server to the version before the upgrade. Moreover, when the firmware upgrade of the first server fails, the hardware configuration information of the first server can also be obtained. The hardware configuration information may include the version information of at least one piece of hardware in the first server. Then, based on the hardware configuration information and the target firmware version, the compatibility rule library is updated. For example, the incompatibility relationship between the target firmware version and the hardware configuration information of the first server can be recorded in the compatibility rule library, so that developers or maintenance personnel can develop firmware adapted to the first server according to the content recorded in the compatibility rule library.

[0104] Figure 4 shows a schematic flowchart of multi-node collaborative fault-tolerant upgrade according to an exemplary embodiment of the present application, as Figure 4 shown. When performing collaborative upgrade on multiple servers (i.e., nodes) in a server cluster, the servers in the server cluster can be divided into blue-green groups (i.e., the first group and the second group). Among them, the blue group (i.e., the second group) maintains the operation of the current production environment (i.e., the current firmware version), and the green group (i.e., the first group) is prepared for upgrade and runs the new firmware version environment. Then, based on the new firmware (i.e., the target firmware in the foregoing embodiment), the firmware is upgraded in parallel in the green group, and the business metrics are monitored in real time to determine whether the upgrade is successful. If successful, the business traffic is gradually switched to the green group, the green group goes online, and the servers in the blue group continue to be upgraded with the new firmware. If the firmware upgrade of a certain server in the green group fails, an automatic rollback is triggered to roll back the firmware to the version before the upgrade, and the failure reason is recorded. The compatibility rule library is updated according to the hardware configuration information of the server where the upgrade fails. Thereby, the firmware upgrade efficiency can be greatly improved, shortening the firmware upgrade cycle of thousands of servers to within 1 hour and reducing manual operations by 90%.

[0105] Among them, the specific implementation code for distributed task scheduling and rollback in multi-node firmware upgrade is as follows:

[0106] from celery import Celery

[0107] app = Celery('tasks', broker='redis: / / localhost')

[0108] @app.task

[0109] def deploy_firmware(node_id, firmware_version):

[0110] try:

[0111] node = get_node(node_id)

[0112] node.upgrade(firmware_version)

[0113] if node.health_check():

[0114] return "Upgrade successful"

[0115] else:

[0116] raise Exception("Health check failed")

[0117] except Exception as e:

[0118] rollback(node_id) # Automatically roll back to the previous version

[0119] log_error(f"Failed to upgrade node {node_id}: {str(e)}")

[0120] # Trigger upgrade tasks in batches

[0121] for node in green_group:

[0122] deploy_firmware.delay(node.id, "bios_v2.0")

[0123] In this embodiment, by rolling back to the firmware version before the upgrade when it is detected that the firmware upgrade of the first server in the first group fails, it is possible to enable the server with the failed upgrade to still operate normally, avoid a downtime accident caused by a compatibility conflict due to firmware update, and by updating the compatibility rule library according to the hardware configuration information of the server with the failed upgrade, it provides a reference for developers to develop subsequent firmware to develop new firmware adapted to the first server.

[0124] To ensure the success rate of server firmware upgrade in a server cluster, before a new firmware version is released, it is usually necessary to test the firmware of the version to be released, and use it for the firmware upgrade of the server cluster after passing the small-scale test. Thus, in an optional implementation manner of the present application, as Figure 5 shown, based on the foregoing embodiment, the method for maintaining a server cluster of the present application may further include the following steps:

[0125] Step 201, in response to receiving a firmware release request, obtain the correspondence between the firmware version of the firmware to be released and the server model, where the firmware release request carries the correspondence and the firmware package of the firmware to be released.

[0126] In this embodiment, when a developer develops a new version of the firmware, a firmware release request may be initiated to release the new version of the firmware after passing the test.

[0127] Among them, the firmware release request carries the firmware package of the firmware to be released, and the correspondence between the firmware version of the firmware to be released and the server model.

[0128] Step 202, determine candidate servers that match the firmware to be released according to the correspondence.

[0129] In this embodiment, according to the obtained correspondence between the firmware version of the firmware to be released and the server model, servers that are adapted to the firmware version of the firmware to be released (referred to as candidate servers) can be determined.

[0130] Step 203, determine target servers based on the hardware configuration information and load information of the candidate servers.

[0131] Among them, the hardware configuration information may be, but is not limited to, information such as CPU, memory, and hard disk, and the load information may be, but is not limited to, the current occupancy rate of the CPU, the remaining memory capacity, etc.

[0132] In this embodiment, for the determined candidate servers, some servers can be selected from them for the first batch of firmware upgrades according to the hardware configuration information and load information of each candidate server. The selected part of the servers is called the target server.

[0133] For example, for candidate servers, at least one server with different hardware configuration information can be selected as the target server based on the hardware configuration information of each server. And if there are multiple servers with the same hardware configuration, the server with the smallest load can be selected as the target server according to the load information of the multiple servers. Thus, the configuration differences among the servers in the target server can be enriched, the diversity of the target server can be improved, so as to realize the firmware upgrade test for servers of the same model but different configurations at the same time, quickly detect the compatibility between the firmware to be released and servers with different hardware configurations, and by preferentially selecting servers with small loads for firmware upgrade, it is beneficial to reduce the probability of firmware upgrade failure caused by excessive server load and improve the success rate of firmware upgrade.

[0134] Step 204: Send the firmware package of the firmware to be released to the target server for firmware upgrade.

[0135] In this embodiment, after the target server is determined, the firmware package of the firmware to be released can be sent to the target server, and the target server performs firmware upgrade based on this firmware package.

[0136] Step 205: When there is no abnormality within a preset duration after it is monitored that the target server is performing firmware upgrade, send the firmware package to the remaining servers in the candidate servers for firmware upgrade.

[0137] Among them, the preset duration can be set according to actual needs. For example, the preset duration can be set to 24 hours.

[0138] In this embodiment, after the firmware upgrade of the target server is successful, it is further monitored whether an abnormality occurs within the preset duration for the target server. If there is no abnormality, the firmware package is sent to the remaining servers in the candidate servers for firmware upgrade, and after the firmware upgrade of the remaining servers is successful, the firmware package is released, so that the firmware package can be used for the firmware upgrade of the corresponding model servers in the server cluster. If an abnormality occurs in a server among the target servers, the firmware version of the server with the abnormality is rolled back to the version before the upgrade, an alarm is triggered, and the relationship that the hardware configuration information of this server is not compatible with the version to be released is recorded in the compatibility rule library. In this case, for the remaining servers, the firmware package is no longer sent to the servers in the remaining servers that have the same hardware configuration information as the target server with the abnormality.

[0139] The maintenance method of the server cluster according to the embodiments of the present application, when receiving a firmware release request, determines candidate servers that match the to-be-released firmware according to the correspondence between the firmware version of the to-be-released firmware and the server model, and then determines target servers from the candidate servers for upgrading to the to-be-released firmware according to the hardware configuration information and load information of the candidate servers. When it is detected that there is no abnormality within a preset time period after the target server performs firmware upgrade, a firmware package is sent to the remaining servers in the candidate servers for firmware upgrade. Thus, by first selecting some servers from the servers that match the to-be-released firmware for firmware upgrade according to the hardware configuration information and load information, the impact of firmware upgrade on the business can be minimized, and when the target server upgrade is successful and there is no abnormality within the preset time period, the remaining servers are then upgraded, which can improve the firmware upgrade success rate and enhance the reliability of firmware upgrade.

[0140] Figure 6 shows a schematic diagram of the firmware release process according to an exemplary embodiment of the present application. As Figure 6 shown, when a new version of firmware is released, first, a developer submits a firmware release request. In response to this request, the system automatically signs the firmware package and generates a release plan to ensure the security of the firmware package. For example, a national cryptography algorithm can be used for signing. A certain number of nodes (i.e., servers) are selected for firmware upgrade first, and low-load nodes are preferentially selected. The firmware package of the to-be-released firmware is deployed to the selected nodes, and it is monitored whether there is no abnormality within 24 hours. If an abnormality occurs in a certain node, the firmware upgrade of this node fails, and this node automatically rolls back to the old version and triggers an alarm. At the same time, compatibility problems are recorded and the compatibility rule library is updated; if it is monitored that there is no abnormality after 24 hours, the firmware package of the to-be-released firmware is deployed to the remaining nodes in batches, and the remaining nodes complete the upgrade in batches. Thus, through the small-scale firmware upgrade test and automatic rollback mechanism before firmware release, the firmware upgrade success rate is increased to 99.9%, and the reliability is enhanced.

[0141] Among them, the specific implementation code for firmware signing and release is as follows:

[0142] class FirmwareManager:

[0143] def__init__(self):

[0144] self.compatibility_rules = load_rules_from_db()# Load compatibility rules from the database

[0145] def sign_firmware(self,firmware_file,private_key):

[0146] # Sign the firmware using the national cryptographic algorithm

[0147] signature = sm2_sign(firmware_file, private_key)

[0148] return firmware_file + ".signed", signature

[0149] def gray_release(self, nodes, firmware_version):

[0150] # Intelligently select the first batch of gray nodes (low-load nodes first)

[0151] selected_nodes = filter_low_load(nodes)

[0152] # Call example

[0153] manager = FirmwareManager()

[0154] signed_fw = manager.sign_firmware("bios_v2.0.bin", "private_key.pem")

[0155] manager.gray_release(all_nodes, signed_fw)

[0156] Regarding the problem in the related technology that there is a lack of real-time health monitoring and predictive maintenance, and the fault repair depends on the ex post response, the maintenance method of the server cluster provided by this application also supports remote diagnosis and fault prediction of the server cluster, can timely detect possible faults in the server cluster, and realize the health monitoring and predictive maintenance of the server cluster. Thus, in an alternative embodiment of this application, as Figure 7 shown, based on the foregoing embodiment, the maintenance method of the server cluster of this application may further include the following steps:

[0157] Step 301, obtain the hardware status data of the second server in the server cluster.

[0158] Among them, the second server may be any server in the server cluster, and the hardware status data may be, but is not limited to, the hardware health data of the server, such as temperature, power consumption, data monitored through Self-Monitoring Analysis and Report Technology (SMART), etc.

[0159] In this embodiment, for each server in the server cluster, the hardware status data of the server can be periodically collected at a preset time interval.

[0160] Step 302: Input the hardware status data into a pre-trained hardware health diagnosis model for hardware diagnosis, and obtain the diagnosis result output by the hardware health diagnosis model.

[0161] Among them, the hardware health diagnosis model is pre-trained. The hardware health diagnosis model can be trained by collecting the historical failure data of the server. For example, the historical failure data of the Graphics Processing Unit (GPU) of the server can be collected to analyze the failure law of the GPU radiator, and these data can be used to train the AI model to obtain a hardware health diagnosis model that can predict whether the GPU radiator fails. Among them, the AI model can be, for example, Long Short-term Memory Networks (LSTM).

[0162] In this embodiment, the collected hardware status data is input into a pre-trained hardware health diagnosis model. The hardware health diagnosis model analyzes the input hardware status data and outputs the corresponding diagnosis result. For example, the output diagnosis result can be the remaining life, remaining capacity, etc. of the hardware.

[0163] Exemplarily, taking the hard disk life prediction based on the LSTM model as an example, the specific implementation code is as follows:

[0164] class HealthPredictor:

[0165] def __init__(self):

[0166] self.model = tf.keras.Sequential([

[0167] tf.keras.layers.LSTM(64, input_shape=(30, 5)), # Input: SMART data within 30 days (5-dimensional features)

[0168] tf.keras.layers.Dense(1, activation='sigmoid') # Output:

[0169] Percentage of remaining life in total life )

[0171] self.model.compile(optimizer='adam', loss='mse')

[0172] def train(self, smart_data, labels):

[0173] self.model.fit(smart_data, labels, epochs=50)

[0174] def predict(self, new_data):

[0175] return self.model.predict(new_data)

[0176] # Call example

[0177] predictor = HealthPredictor()

[0178] predictor.train(historical_smart_data, historical_labels)

[0179] predicted_life = predictor.predict(current_smart_data)

[0180] if predicted_life < 0.1: # Remaining life < 10%

[0181] trigger_maintenance_alert()

[0182] Step 303, in the case where the diagnosis result indicates that there is hardware to be maintained, generate a maintenance work order based on the diagnosis result.

[0183] In this embodiment, after obtaining the diagnosis result, it is possible to determine whether there is hardware that needs to be maintained according to the diagnosis result. For example, if the remaining life of a certain hardware in the diagnosis result is 5% of the total life, which is less than the preset threshold of 10%, it is determined that the hardware needs to be maintained. When the diagnosis result indicates that there is hardware to be maintained, a maintenance work order is generated according to the relevant data of the hardware to be maintained in the diagnosis result. For example, if the remaining life of a certain hardware in the diagnosis result is 5% of the total life and needs to be maintained, the generated maintenance work order may include [unique identifier of the second server, name and model of the hardware to be maintained, diagnosis result: remaining life is 5%, maintenance suggestion: replacement].

[0184] Step 304, feedback the maintenance work order to the user.

[0185] In this embodiment, the generated maintenance work order can be automatically sent to the user so that the user can allocate corresponding hardware for maintenance according to the maintenance work order.

[0186] In an alternative embodiment of the present application, when a user performs maintenance on a second server with a hardware failure in the server cluster, a hardware maintenance instruction can be initiated. For example, the hardware maintenance instruction can be sent through a user terminal, and the maintenance window, that is, the time period of this maintenance, can be carried in the hardware maintenance instruction. After the system receives the hardware maintenance instruction, it can determine a third server adjacent to the second server according to the distribution information of the servers in the server cluster. During the maintenance window period, the service traffic of the second server is migrated to the third server, and after the service traffic of the second server is migrated to the third server, the maintenance of the hardware to be maintained in the second server is started. For example, when detecting a hard disk failure of a server in the server cluster, it is detected that the SMART error rate of the hard disk of a certain server has increased, and the remaining life is predicted to be 7 days, which is less than the preset threshold. Then the system automatically sends a maintenance work order to the warehouse, so that the warehouse staff can allocate spare hard disks and plan the maintenance window according to the maintenance work order. During the maintenance period, the service load of this server is migrated to the adjacent server, realizing zero-downtime replacement. In this embodiment, before maintaining the server that requires hardware maintenance, the service traffic on this server is migrated to the adjacent server, thereby ensuring that the service of the server is not affected by the hardware maintenance operation and ensuring the continuity of the service during the maintenance process.

[0187] In an alternative embodiment of the present application, during the maintenance of the second server, a maintenance operation record can be obtained, where the maintenance operation record manually input by the user can be obtained, or the automatically generated maintenance operation record can be obtained, and the maintenance operation record is stored in the blockchain to achieve blockchain evidence storage of the maintenance operation, ensuring that the operation record cannot be tampered with and guaranteeing security.

[0188] The maintenance method of the server cluster in the embodiment of the present application obtains the hardware status data of the servers in the server cluster, inputs the hardware status data into the hardware health diagnosis model for hardware diagnosis to obtain a diagnosis result, generates a maintenance work order based on the diagnosis result and feeds it back to the user, so that the user can prepare the required hardware in advance or propose a fault response strategy according to the maintenance work order, thereby being able to detect possible faults in the server cluster in time and perform maintenance, realizing the health monitoring and predictive maintenance of the server cluster, and solving the problem in the related technology that the fault repair depends on the ex post response, resulting in untimely fault maintenance.

[0189] Figure 8 Shows a schematic diagram of the predictive maintenance process based on the AI model in an exemplary embodiment of the present application, as Figure 8As shown in the figure, when performing predictive maintenance, the hardware health data of the servers in the server cluster is collected in real time, and the hardware health data is input into a pre-trained AI model for hardware detection. It is determined whether the diagnosis result includes a hardware fault. If there is no fault, the hardware health data of each server is continuously monitored; if there is a hardware fault, a maintenance work order is generated and fed back to the user, so that the user can allocate spare hardware according to the maintenance work order and plan the maintenance window, and migrate the business load of the server to be maintained to other adjacent servers within the maintenance window, and then perform maintenance operations, such as hard disk replacement, and record the maintenance operations to update the maintenance log. By performing predictive maintenance on the servers in the server cluster, maintenance is carried out before a failure occurs, reducing the hardware failure rate by 40% and the spare parts inventory cost by 30%.

[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0191] The embodiment of the present application also provides a maintenance device for a server cluster. Figure 9 As shown in the structure schematic diagram of a maintenance device for a server cluster provided by an embodiment of the present application, Figure 9 As shown, the maintenance device 50 of the server cluster includes: a version determination module 510, a grouping module 520, a sending module 530, and a service migration module 540.

[0192] Among them, the version determination module 510 is used to determine the target firmware version to be upgraded in response to receiving a firmware upgrade request;

[0193] The grouping module 520 is used to divide the servers in the server cluster into a first group and a second group based on the status information and current load information of the servers in the server cluster;

[0194] The sending module 530 is used to send the target firmware corresponding to the target firmware version to the servers in the first group for firmware upgrade;

[0195] The service migration module 540 is used to switch the service traffic of the servers in the second group to the first group when it is detected that the firmware upgrade of the servers in the first group is successful;

[0196] The sending module 530 is further used to send the target firmware to the servers in the second group for firmware upgrade.

[0197] Optionally, the sending module 530 is further used to:

[0198] Perform integrity verification on the target firmware corresponding to the target firmware version;

[0199] When the verification passes, obtain a key pair to encrypt the target firmware to obtain an encrypted firmware;

[0200] Send the encrypted firmware to the servers in the first group for firmware upgrade.

[0201] Optionally, the maintenance device 50 of the server cluster further includes: a rollback module and an update module;

[0202] The rollback module is used to control the first server to execute a preset rollback script to roll back to the firmware version before the upgrade when it is detected that the firmware upgrade of the first server in the first group fails;

[0203] The update module is used to obtain the hardware configuration information of the first server; update the compatibility rule library based on the hardware configuration information and the target firmware version.

[0204] Optionally, the maintenance device 50 of the server cluster further includes: a firmware release module; the firmware release module is used for:

[0205] In response to receiving a firmware release request, obtain the correspondence between the firmware version of the firmware to be released and the server model, where the firmware release request carries the correspondence and the firmware package of the firmware to be released;

[0206] Determine candidate servers that match the firmware to be released according to the correspondence;

[0207] Determine the target server based on the hardware configuration information and load information of the candidate servers;

[0208] Send the firmware package of the firmware to be released to the target server for firmware upgrade;

[0209] When it is detected that there is no abnormality within a preset time period after the target server performs a firmware upgrade, send the firmware package to the remaining servers in the candidate servers for firmware upgrade.

[0210] Optionally, the maintenance device 50 of the server cluster further includes: a prediction and diagnosis module; the prediction and diagnosis module is used for:

[0211] Obtain the hardware status data of the second server in the server cluster;

[0212] Input the hardware status data into a pre-trained hardware health diagnosis model for hardware diagnosis, and obtain the diagnosis result output by the hardware health diagnosis model;

[0213] When the diagnosis result indicates that there is hardware to be maintained, generate a maintenance work order based on the diagnosis result;

[0214] Feed back the maintenance work order to the user.

[0215] Optionally, the maintenance device 50 of the server cluster further includes: a maintenance module; the maintenance module is configured to:

[0216] In response to receiving a hardware maintenance instruction, determine a third server adjacent to the second server according to the distribution information of the servers in the server cluster;

[0217] After migrating the service traffic of the second server to the third server, start the maintenance of the hardware to be maintained in the second server.

[0218] Optionally, the maintenance device 50 of the server cluster further includes: a storage module; the storage module is configured to:

[0219] Obtain a maintenance operation record;

[0220] Store the maintenance operation record in the blockchain.

[0221] For the description of the features in the corresponding embodiments of the maintenance device of the server cluster, reference can be made to the relevant descriptions in the corresponding embodiments of the maintenance method of the server cluster, which will not be elaborated here one by one.

[0222] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the maintenance method of the server cluster.

[0223] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the maintenance method of the server cluster when running.

[0224] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0225] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and the steps in any one of the above embodiments of the maintenance method of the server cluster are implemented when the computer program is executed by a processor.

[0226] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the steps in any one of the above embodiments of the maintenance method of the server cluster are implemented when the computer program is executed by a processor.

[0227] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered as exceeding the scope of this application.

[0228] The above has introduced in detail a method, apparatus, electronic device, and storage medium for maintaining a server cluster provided in this application. Specific examples have been used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A maintenance method for a server cluster, characterized in that, Including: Upon receiving a firmware upgrade request, determining a target firmware version to be upgraded; Based on the status information and current load information of the servers in the server cluster, dividing the servers in the server cluster into a first group and a second group; Sending the target firmware corresponding to the target firmware version to the servers in the first group for firmware upgrade; When it is monitored that the firmware upgrade of the servers in the first group is successful, switching the service traffic of the servers in the second group to the first group, and sending the target firmware to the servers in the second group for firmware upgrade.

2. The maintenance method of the server cluster according to claim 1, characterized in that The sending the target firmware corresponding to the target firmware version to the servers in the first group for firmware upgrade includes: Performing integrity verification on the target firmware corresponding to the target firmware version; Upon successful verification, obtaining a key pair to encrypt the target firmware to obtain an encrypted firmware; Sending the encrypted firmware to the servers in the first group for firmware upgrade.

3. The maintenance method of the server cluster according to claim 1, characterized in that The method further includes: When it is monitored that the firmware upgrade of the first server in the first group fails, controlling the first server to execute a preset rollback script to roll back to the firmware version before the upgrade; Obtaining the hardware configuration information of the first server; Updating the compatibility rule library based on the hardware configuration information and the target firmware version.

4. The maintenance method of the server cluster according to claim 1, characterized in that, The method further includes: Upon receiving a firmware release request, obtaining the correspondence between the firmware version of the firmware to be released and the server model, where the firmware release request carries the correspondence and the firmware package of the firmware to be released; Determining candidate servers matching the firmware to be released according to the correspondence; Based on the hardware configuration information and load information of the candidate servers, determining target servers; Sending the firmware package of the firmware to be released to the target servers for firmware upgrade; When it is monitored that there is no abnormality within a preset duration after the target servers perform firmware upgrade, sending the firmware package to the remaining servers in the candidate servers for firmware upgrade.

5. The maintenance method of the server cluster according to claim 1, characterized in that, The method further includes: Obtaining the hardware status data of the second server in the server cluster; Inputting the hardware status data into a pre-trained hardware health diagnosis model for hardware diagnosis, and obtaining a diagnosis result output by the hardware health diagnosis model; When the diagnosis result indicates that there is hardware to be maintained, generating a maintenance work order based on the diagnosis result; Feeding back the maintenance work order to the user.

6. The maintenance method of the server cluster according to claim 5, characterized in that, The method further includes: Upon receiving a hardware maintenance instruction, determining a third server adjacent to the second server according to the distribution information of the servers in the server cluster; After migrating the service traffic of the second server to the third server, starting the maintenance of the hardware to be maintained in the second server.

7. The maintenance method of the server cluster according to claim 6, characterized in that, The method further includes: Obtaining maintenance operation records; Storing the maintenance operation records in a blockchain.

8. A maintenance device for a server cluster, characterized in that Including: A version determination module, configured to determine a target firmware version to be upgraded upon receiving a firmware upgrade request; A grouping module, configured to divide the servers in the server cluster into a first group and a second group based on the status information and current load information of the servers in the server cluster; A sending module, configured to send the target firmware corresponding to the target firmware version to the servers in the first group for firmware upgrade; A service migration module, configured to switch the service traffic of the servers in the second group to the first group when it is detected that the firmware upgrade of the servers in the first group is successful; The sending module is further configured to send the target firmware to the servers in the second group for firmware upgrade.

9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the server cluster maintenance method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the server cluster maintenance method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Cited By

  • Storage cluster system upgrading method and device, electronic equipment, medium and product

    CN120560693A