System upgrading method and device, equipment, storage medium and computer program product
By dividing storage pools and fault domains in a distributed storage cluster and determining node groups and their upgrade sequence, the problem of low upgrade efficiency in distributed storage systems is solved, and fast system upgrades with no data loss are achieved.
Patent Information
- Application Number
- CN202511351487.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the system upgrade efficiency of the distributed storage system is low, and it is difficult to complete the upgrade quickly without affecting the business.
By dividing the distributed storage cluster into storage pools and fault domains, determining the node groups and their upgrade order, and performing system upgrades on the node groups in sequence, we ensure that the nodes meet the upgrade conditions and come from the same fault domain or storage pool, maximizing the number of nodes that can be upgraded simultaneously.
Improves the system upgrade efficiency of distributed storage clusters, reduces upgrade time, ensures data is not lost, and avoids business interruption.
Smart Images

Figure CN120848924A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of distributed storage systems, and more particularly to a system upgrade method, apparatus, device, storage medium, and computer program product. Background Technology
[0002] With the explosive growth of data, traditional centralized storage systems often struggle to meet business needs. Distributed storage systems have been proposed as a related technology, which can better provide data storage services due to their high reliability and scalability.
[0003] However, in practical applications, there is currently no effective solution for efficiently upgrading distributed storage systems. Summary of the Invention
[0004] To address the related technical issues, embodiments of this application provide a system upgrade method, apparatus, electronic device, storage medium, and computer program product.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a system upgrade method applied to a first node in a distributed storage cluster. The distributed storage cluster is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes. The method includes: The system determines first information, second information, and third information. The first information represents the state of the distributed storage cluster, the second information represents one or more second nodes that need to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool partitioning results of the distributed storage cluster. Using the first information, the second information, and the third information, one or more node groups and the upgrade order corresponding to each node group are determined. Each node group contains one or more second nodes. The second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or belong to different storage pools. Following the upgrade order, system upgrades are performed on all second nodes within the node group in sequence.
[0006] In the above scheme, performing a system upgrade on all second nodes included in the node group includes: The isolated node group contains all the second nodes; For each isolated second node, control the second node to perform system upgrade operations.
[0007] In the above scheme, the isolation node group includes all the second nodes, including: For each second node in the node group, when the second node includes a storage node, the second node is instructed to perform the following: determine whether each storage process associated with the second node meets the isolation condition, and obtain the determination result; if the determination result indicates that all storage processes associated with the second node meet the isolation condition, perform a crash operation. And / or, When the second node includes a gateway node, instruct the second node to perform the following: set the load balancing weight corresponding to the second node to 0; perform a liveness check query to obtain the query result; and if the query result indicates that there are no services that maintain connections, perform a shutdown operation.
[0008] The method in the above scheme further includes: Increase the first value, which represents the number of log records to be retained in the placement group associated with the distributed storage cluster; wherein, the distributed storage cluster is configured with one or more placement groups, the placement groups are used to perform distributed management of the data stored in the distributed storage cluster, and the placement group logs are used to record the management operations of the placement groups.
[0009] The method in the above scheme further includes: The second value is reduced to enable the distributed storage cluster to provide storage services during system upgrades. The second value represents the minimum number of replicas associated with the distributed storage cluster.
[0010] In the above scheme, the distributed storage cluster includes at least storage nodes, and each storage node is associated with one or more storage processes. The method further includes: Control each storage node to restart each storage process associated with the storage node in batches after performing a system upgrade operation.
[0011] The method in the above scheme further includes: After performing a system upgrade operation, each second node is controlled to pull a software upgrade package from the third node and use the software upgrade package to perform a software upgrade. The third node is used at least to manage the software source.
[0012] The method in the above scheme further includes: Before performing a system upgrade operation, each second node is controlled to upload the software upgrade package to the third node.
[0013] In the above scheme, the distributed storage cluster includes at least storage nodes, and each storage node is associated with one or more storage processes. The method further includes: After each storage node performs a system upgrade, the pace of data balancing for each storage process is adjusted based on the service bandwidth of that storage process.
[0014] The method in the above scheme further includes: Control each secondary node to back up the system software configuration and / or system logs before performing a system upgrade operation.
[0015] In the above scheme, the backup system software configuration and / or system logs include: The distributed storage cluster is used to back up system software configurations and / or system logs.
[0016] This application embodiment also provides a system upgrade device, set in the first node of a distributed storage cluster, wherein the distributed storage cluster is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes, including: A determining unit is configured to determine first information, second information, and third information. The first information represents the state of the distributed storage cluster. The second information represents one or more second nodes to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool partitioning results of the distributed storage cluster. The unit also uses the first, second, and third information to determine one or more node groups and the upgrade order corresponding to each node group. Each node group contains one or more second nodes, and the second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or different storage pools. The upgrade unit is used to perform system upgrades on all second nodes included in the node group in sequence according to the upgrade order.
[0017] This application also provides an electronic device, including: a processor and a memory for storing a computer program capable of running on the processor. When the processor runs the computer program, it executes the steps of any of the above methods.
[0018] This application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of any of the above methods.
[0019] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above methods.
[0020] The system upgrade method, apparatus, electronic device, storage medium, and computer program product provided in this application embodiment determine first information, second information, and third information for a first node in a distributed storage cluster. The distributed storage cluster is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes. The first information represents the state of the distributed storage cluster, the second information represents one or more second nodes to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool division results of the distributed storage cluster. Using the first, second, and third information, one or more node groups and the upgrade order corresponding to each node group are determined. Each node group contains one or more second nodes, and the second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or belong to different storage pools. According to the upgrade order, system upgrades are performed on all second nodes contained in the node group sequentially. The solution provided in this application embodiment involves a node (i.e., the first node) that initiates a system upgrade of a distributed storage cluster. Based on the status of the distributed storage cluster (i.e., the first information), whether the node needs to be upgraded (i.e., the second information), and the fault domain and storage pool to which the node belongs (i.e., the third information), the nodes that can be upgraded simultaneously are divided into the same node group, and the system upgrade is performed on all nodes in each node group in sequence (i.e., the upgrade order). When partitioning node groups, by considering the node status and system upgrade requirements, it can be ensured that nodes undergoing simultaneous upgrades meet the system upgrade conditions and require system upgrades. Simultaneously, by ensuring that all nodes in each node group come from the same fault domain or different storage pools, the number of nodes in each node group can be maximized (which can also be understood as maximizing the number of nodes undergoing simultaneous system upgrades), thereby reducing the time required for distributed storage cluster upgrades and effectively improving the system upgrade efficiency of the distributed storage cluster. This is because: 1) When storing data, the distributed storage cluster stores multiple blocks corresponding to the data in different fault domains. In this case, even if all nodes in a single fault domain fail simultaneously, the data can still be recovered using data blocks stored in other fault domains, without data loss. Therefore, the distributed storage cluster can tolerate all nodes in the same fault domain undergoing simultaneous system upgrades; 2) When storing data, the distributed storage cluster stores multiple copies of the data in different storage pools. In this case, even if a node in a single storage pool fails, it will not affect the copies stored in other storage pools, without data loss. Therefore, the distributed storage cluster can tolerate nodes in different storage pools undergoing simultaneous system upgrades. Attached Figure Description
[0021] Figure 1This is a schematic diagram of the architecture of the Ceph storage system in related technologies; Figure 2 This is a flowchart illustrating the system upgrade method according to an embodiment of this application; Figure 3 This is a flowchart illustrating a method for replacing the operating system in a distributed storage system, serving as an application example of this application. Figure 4 This is a schematic diagram of the system upgrade device structure according to an embodiment of this application; Figure 5 This is a schematic diagram of the electronic device structure according to an embodiment of this application. Detailed Implementation
[0022] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.
[0023] With the rapid development of information technology, distributed storage systems have been widely used in big data processing, cloud computing, and other fields. A distributed storage system, also known as a distributed storage cluster, typically contains multiple storage devices that work together to provide storage services, offering characteristics such as high performance, high scalability, and high reliability.
[0024] Currently, such as Figure 1 As shown, commonly used distributed storage systems include Ceph storage systems (also known as Ceph storage clusters or Ceph clusters), etc. When providing storage services, Ceph storage systems distribute data evenly across different storage devices according to the Controlled Replication Under Scalable Hashing (CRUSH) pseudo-random algorithm.
[0025] In practical applications, a Ceph storage system can be deployed with one or more nodes. These nodes can include storage nodes for data storage and management, gateway nodes for providing storage service interfaces, etc. Each storage node can run one or more storage processes, each of which manages one or more storage devices and stores data to one or more storage devices. Specifically, storage processes can include object storage processes (OSDs, ObjectStorage Daesons). In this case, the Ceph storage system can also be understood as a large-scale distributed cluster for object storage, capable of providing real-time storage services.
[0026] Currently, operating systems used in distributed storage clusters include OpenEuler, Kylin, and UnionTech OS. When upgrading the operating system of a distributed storage cluster, the following problems typically arise: 1) How to ensure the stability of a massive storage cluster and uninterrupted business operations during operating system upgrades; 2) How to automate the operating system upgrade process and improve the efficiency of operating system upgrades (which can also be understood as migration efficiency or replacement efficiency, etc.). 3) How to promptly identify problems during the operating system upgrade process and automatically restore the operating system.
[0027] Operating system upgrades, also known as system upgrades, can specifically include one or more of the following (one or more can also be understood as at least one): Upgrade the operating system; Replacing the operating system (also known as system replacement); Reinstalling the operating system (also known as system reinstallation).
[0028] In related technologies, system upgrades for storage clusters typically involve the following steps: identifying the nodes in the storage cluster that require operating system upgrades (e.g., storage nodes or gateway nodes); backing up each node (e.g., backing up storage data for storage nodes or configuration information for gateway nodes); performing a downtime (or DOWN) operation on the backed-up nodes (which can also be understood as removing the backed-up nodes from the storage cluster); upgrading the system of the downed nodes (which can also be understood as performing a system upgrade operation, such as running a system upgrade script or reinstalling the operating system); redeploying the relevant service components on the upgraded nodes after the system upgrade; rejoining the nodes with the deployed service components to the storage cluster; performing data balancing (or data recovery) on the rejoined nodes and checking the storage cluster status to verify (or ensure) the performance of the storage cluster; if the verification passes, continuing the system upgrade; or, if the verification fails, performing system recovery and restarting the system upgrade process.
[0029] As can be seen from the above description, system upgrades for a single node (or a single node) take a long time, mainly including the time required for node backup, system upgrade, and data balancing. Therefore, if each single node in the storage cluster is upgraded sequentially, which can be understood as a single node undergoing system upgrade simultaneously, completing the storage cluster upgrade will consume a significant amount of time. Furthermore, each node's system upgrade requires backup, data balancing, and status checks, consuming substantial storage cluster resources and potentially impacting the stability of the storage cluster. This can also be interpreted as causing the storage cluster to become uneven (unable to stably provide storage services).
[0030] For example, assuming the storage cluster has 600 nodes (or can be understood as the storage cluster containing 600 nodes), and the time required to upgrade a single node is 1 hour, then the upgrade of the storage cluster will take a total of 600 hours.
[0031] It is evident that when performing system upgrades on a storage cluster, the more nodes in the cluster that can be upgraded simultaneously, the shorter the total upgrade time and the higher the upgrade efficiency. However, if multiple nodes in the storage cluster are upgraded simultaneously, all nodes undergoing the upgrade will need to be downgraded, which may affect the storage cluster's ability to provide storage services (or, in other words, impact business operations).
[0032] Currently, to ensure that system upgrades of storage clusters do not impact business operations, the maximum number of nodes that the storage cluster can tolerate being down is typically determined based on the erasure code (or erasure code) corresponding to the storage cluster. This, in turn, determines the number of nodes that can be upgraded simultaneously. The role of erasure coding is as follows: when storing data, the storage cluster, based on the erasure code configuration parameters k and m, divides the data into k data blocks, generates m coded blocks, and stores the resulting k+m blocks on k+m nodes. Thus, as long as the number of failed nodes among these k+m nodes is less than or equal to m, the storage cluster can recover the data using the blocks stored on the unaffected nodes. In other words, the storage cluster can tolerate a maximum of m node failures. Here, k is typically an integer greater than 1, and m is typically an integer greater than or equal to 1.
[0033] However, in practical applications, the larger the value of m, the more coding blocks the storage cluster needs to generate when storing data, and the lower the storage efficiency and storage resource utilization. Therefore, in order to ensure the performance of the storage cluster, it is necessary to reasonably set the upper limit of the value of m, such as setting m equal to 1 or m equal to 2.
[0034] Given the limited value of the erasure coding configuration parameter 'm', the number of downtime nodes that a storage cluster determined by erasure coding can tolerate is also relatively small. Therefore, system upgrades for storage clusters still require a long time and are inefficient. In other words, it is difficult to quickly and efficiently carry out system upgrades for ultra-large-scale storage clusters.
[0035] Based on this, in various embodiments of this application, the node that initiates the system upgrade of the distributed storage cluster divides the nodes that can be upgraded at the same time into the same node group according to the status of the distributed storage cluster, whether the node needs to be upgraded, and the fault domain and storage pool to which the node belongs, and performs system upgrades on all nodes in each node group in sequence. When partitioning node groups, by considering the node status and system upgrade requirements, it can be ensured that nodes undergoing simultaneous upgrades meet the system upgrade conditions and require system upgrades. Simultaneously, by ensuring that all nodes in each node group come from the same fault domain (e.g., a single rack) or different storage pools, the number of nodes in each group can be maximized (which can also be understood as maximizing the number of nodes undergoing simultaneous system upgrades). This reduces the time required for distributed storage cluster upgrades and effectively improves the system upgrade efficiency of the distributed storage cluster. The reasons are: 1) When storing data, the distributed storage cluster stores multiple blocks corresponding to the data in different fault domains. In this case, even if all nodes in a single fault domain fail simultaneously, the data can still be recovered using data blocks stored in other fault domains, without data loss. Therefore, the distributed storage cluster can tolerate simultaneous system upgrades for all nodes in the same fault domain. 2) When storing data, the distributed storage cluster stores multiple copies of the data in different storage pools. In this case, even if a node in a single storage pool fails, it will not affect the copies stored in other storage pools, without data loss. Therefore, the distributed storage cluster can tolerate simultaneous system upgrades for nodes in different storage pools.
[0036] This application provides a system upgrade method applied to the first node in a distributed storage cluster. The distributed storage cluster is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes, such as... Figure 2 As shown, the method includes: Step 201: Determine the first information, the second information, and the third information. The first information represents the state of the distributed storage cluster. The second information represents one or more second nodes that need to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool partitioning results of the distributed storage cluster. Step 202: Using the first information, the second information and the third information, determine one or more node groups and the upgrade order corresponding to each node group. The node group contains one or more second nodes. The second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or belong to different storage pools. Step 203: Perform system upgrades on all second nodes in the node group in the order described above.
[0037] In practical applications, the distributed storage cluster can also be understood as a distributed storage system, a storage cluster, or a cluster, etc. Specifically, it may include a Ceph storage cluster, etc. This application embodiment does not limit the name and specific implementation of the distributed storage cluster, as long as its function is implemented.
[0038] The distributed storage cluster includes multiple nodes, and the multiple nodes include at least the first node. The first node can also be understood as the node that initiates the distributed storage cluster system upgrade. That is to say, the first node is at least used to initiate the distributed storage cluster system upgrade process. In this embodiment of the application, the name of the first node is not limited, as long as its function is implemented.
[0039] In practical applications, the distributed storage cluster can be divided into one or more data pools based on preset rules, each data pool can be divided into one or more storage pools, and each storage pool can be divided into one or more fault domains. Each fault domain contains one or more nodes. The distributed storage cluster can contain one or more types of nodes (one or more can also be understood as at least one), such as storage nodes and gateway nodes. The gateway node is at least used to provide a storage service interface, and the storage node is at least used for data storage and management. One or more storage processes (specifically including OSDs) can run on the storage node, and each storage process is at least used to perform data storage and management operations, such as managing one or more storage devices deployed on the storage node.
[0040] Based on the above partitioning results, the distributed storage cluster satisfies the following conditions when storing data: 1) By storing multiple blocks corresponding to the data in different fault domains, even if all nodes in a single fault domain fail simultaneously, the data can still be recovered using the data blocks stored in other fault domains without data loss. Therefore, the distributed storage cluster can tolerate all nodes in the same fault domain undergoing system upgrades simultaneously. The multiple blocks may include k data blocks and m coding blocks obtained based on erasure coding parameters k and m. The values of k and m can be set according to actual needs, and this application embodiment does not limit them. 2) When storing data in a distributed storage cluster, multiple copies of the data are stored in different storage pools. In this case, even if a node in a single storage pool fails, it will not affect the copies stored in other storage pools, and no data loss will occur. Therefore, the distributed storage cluster can tolerate nodes in different storage pools undergoing system upgrades simultaneously. The number of data copies can be set according to actual needs, such as three, etc., and this application embodiment does not limit this.
[0041] In practical applications, users of the distributed storage cluster can determine the need for system upgrades of one or more storage nodes in the cluster based on actual requirements, and trigger the system upgrade process by interacting with the first node. The user can interact with the first node through an interactive interface; this interaction may include inputting system upgrade commands, running system upgrade scripts, or running system upgrade programs, etc., which are not limited in this embodiment.
[0042] Of course, the first node can also automatically execute the system upgrade process according to the pre-configured system upgrade rules, such as periodic execution, etc., and this application embodiment does not limit this.
[0043] In practical applications, the first node can set conditions to trigger the execution of steps 201 to 202 above to upgrade the distributed storage cluster, as needed. This application embodiment does not limit this. Here, the system upgrade refers to one or more operations such as upgrading, replacing, or reinstalling the operating system of the distributed storage cluster.
[0044] In practical applications, before performing a system upgrade operation on the nodes in the distributed storage cluster, in step 201, the first node determines the following information: 1) Which nodes in the distributed storage cluster require system upgrades? 2) Whether the status of the distributed storage cluster meets the system upgrade conditions (or can be understood as being able to perform a system upgrade); 3) Which nodes in the distributed storage cluster can be upgraded simultaneously?
[0045] Regarding 1) above, in practical applications, in step 201, the first node can determine one or more nodes to be upgraded, i.e., the second information, based on one or more of the following: user-input information, pre-configured information, and information sent by relevant devices (such as nodes used for management in the distributed storage cluster). The second information can also be understood as a list of upgrade nodes, etc. This application embodiment does not limit the name of the second information. In this application embodiment, the node to be upgraded is referred to as the second node, which can also be understood as an upgrade node. This application embodiment does not limit the name of the second node.
[0046] Regarding 2) above, in practical applications, in step 201, the first node can initiate a distributed storage cluster status check process to obtain the distributed storage cluster status check result, and use the status check result as the first information. The cluster status may include one or more of the following: overall cluster health status (e.g., the health status of replicas in the cluster), placement group (PG, which can also be understood as a placement group), and storage process status. Here, PG refers to a logical unit used by a distributed storage cluster to distribute and manage a group of blocks (e.g., data blocks and code blocks) and / or a group of replicas corresponding to data across different storage processes.
[0047] For example, when the distributed storage cluster includes a Ceph storage cluster, the first node can run the `ceph -s` command to obtain a status report (or status code) of the Ceph storage cluster, and use this status report as the first information. In practical applications, the first node can be configured to enable cyclic probing. After enabling cyclic probing, the first node periodically runs the `ceph -s` command to obtain a status report. Simultaneously, the first node can use the S3 client tool to probing the basic functions of the Ceph cluster, specifically including operation tests such as listing buckets, uploading objects, and downloading objects, obtaining the probing results for the basic functions. Furthermore, the first node can check the abnormal state and / or error rate of the cluster, obtaining the check results. Then, the first node can determine the first information based on the aforementioned status report, probing results, and check results.
[0048] In practical applications, the first node can also execute an upgrade environment check process, such as checking whether the hardware (e.g., hard drive, network card, central processing unit (CPU)) meets the system upgrade requirements (e.g., whether the hardware drivers meet the requirements), whether the memory meets the system upgrade requirements (specifically, whether the memory capacity meets the requirements), whether the network status of the distributed storage cluster is normal, and whether the software of each component of the distributed storage cluster (e.g., storage processes, PGs) meets the upgrade requirements. Here, the specific implementation of the upgrade environment check process can be understood according to relevant technologies, and this application embodiment does not limit it in this regard.
[0049] As can be seen from the above description, the first node can determine the first information before performing a system upgrade, in order to determine whether the state of the distributed storage cluster meets the system upgrade conditions.
[0050] Regarding point 3) above, the conditions for data storage based on the above distributed storage cluster can be determined that when multiple second nodes in the distributed storage cluster belong to the same fault domain or different storage pools, the multiple second nodes can perform system upgrades simultaneously.
[0051] In practical applications, in step 201, the first node can determine the fault domain and storage pool partitioning result of the distributed storage cluster, i.e., the third information, based on one or more of the information input by the user, the configuration information of the distributed storage cluster, and information sent by relevant devices (such as nodes used for management in the distributed storage cluster). The third information can also be understood as node partitioning information, etc., and the name of the third information is not limited in this embodiment.
[0052] After determining the first information, the second information, and the third information, in step 202, the first node can use the second information and the third information to determine the fault domain and storage pool corresponding to each second node; then, the first node can combine the first information to divide all second nodes that meet the system upgrade conditions into one or more node groups, each node group containing one or more second nodes, and the second nodes contained in different node groups are completely different; at the same time, the first node can set a corresponding upgrade order for each node group.
[0053] In practical applications, the specific implementation of dividing the first node group into one or more node groups may include: selecting one or more second nodes that can be upgraded simultaneously from all second nodes that meet the system upgrade conditions, recording all selected second nodes in the first row of the execution matrix as a node group; then, for the remaining second nodes, re-selecting one or more second nodes that can be upgraded simultaneously, and recording all re-selected second nodes in the next row of the execution matrix, until all second nodes are recorded in the execution matrix. At this point, the order of each row in the execution matrix can be used as the upgrade order of the corresponding node group. The execution matrix can also be called a change execution matrix or an upgrade matrix; this application embodiment does not limit this.
[0054] After determining the one or more node groups, before performing the system upgrade, the first node can also adjust the relevant configurations of the distributed storage cluster to ensure that the storage service can be stably provided during the system upgrade process and that effective data balancing can be performed after the system upgrade.
[0055] In practical applications, the distributed storage cluster can be configured with one or more Group Providers (PGs). Data management operations (such as storage operations) corresponding to each PG can be recorded in the corresponding PG log (PGlog). The PG log can be used for data balancing (also known as data load balancing). The distributed storage cluster can be configured with a first value, which represents the maximum number of log records retained in each PG log associated with the distributed storage cluster.
[0056] When a node in the distributed storage cluster undergoes a system upgrade, it needs to be placed in a downtime state. During this downtime, other nodes in the distributed storage cluster can continue to provide storage services and generate PG logs. Once the node undergoing the upgrade completes its upgrade and rejoins the distributed storage cluster, these nodes can utilize the PG logs for data balancing. In this scenario, by appropriately setting a first value, the PG logs can fully record all data management operations performed during the system upgrade process, ensuring effective data balancing.
[0057] Based on this, in some optional embodiments, the method may further include: Increase the first value, which represents the number of log records to be retained in the placement group associated with the distributed storage cluster; wherein, the distributed storage cluster is configured with one or more placement groups, the placement groups are used to perform distributed management of the data stored in the distributed storage cluster, and the placement group logs are used to record the management operations of the placement groups.
[0058] In practical applications, in order to improve the fault tolerance of the distributed storage cluster, the distributed storage cluster can be set with a second value. The second value represents the minimum number of replicas associated with the distributed storage cluster. That is, when performing a write operation on the data stored in the distributed storage cluster, the number of available replicas must be at least greater than or equal to the second value. When the number of available replicas is less than the second value, the write operation will not succeed.
[0059] However, when nodes in the distributed storage cluster undergo system upgrades, the replicas on the upgraded nodes become unavailable, potentially reducing the number of available replicas. In this situation, setting a reasonable second value can ensure that write operations can continue normally during system upgrades, preventing write failures due to some nodes undergoing system upgrades.
[0060] Based on this, in some optional embodiments, the method may further include: The second value is reduced so that the distributed storage cluster can provide storage services during system upgrades.
[0061] As can be seen from the above description, by setting appropriate first and second values, it is possible to effectively ensure the stable provision of storage services during system upgrades and to achieve effective data balancing after system upgrades.
[0062] After adjusting the first and second values, in step 203, the first node can perform system upgrades on each node group in the order corresponding to the node group, that is, perform system upgrades on all second nodes in the node group at the same time.
[0063] In practical applications, considering that system upgrades of the distributed storage cluster can simultaneously upgrade the software of each component to ensure compatibility with the upgraded system, each second node in the node group can generate its own corresponding software upgrade package before performing the system upgrade.
[0064] In some optional embodiments, the method may further include: Before performing a system upgrade operation, each second node is controlled to upload the software upgrade package to a third node, which is used at least to manage the software source.
[0065] In practical applications, the third node can also be understood as a public software source node or a separate public software source node, etc. The name of the third node is not limited in this application embodiment.
[0066] The third node is capable of receiving software upgrade packages (also known as upgrade packages or software packages, etc.) sent by one or more second nodes, and creating a software yum repository so that the second nodes can obtain the corresponding software upgrade packages from the third node for software upgrades after system upgrades. Based on this, in some optional embodiments, the method may further include: Control each second node to configure the software installation source after the system upgrade and restart, that is, obtain the address information of the software upgrade package from the third node.
[0067] In practical applications, software upgrades to nodes in the distributed storage cluster may result in the loss (or absence) of system software configurations and / or system logs. In such cases, it is advisable to back up the system software configurations and / or system logs for each secondary node beforehand.
[0068] Based on this, in some optional embodiments, the method may further include: Control each secondary node to back up the system software configuration and / or system logs before performing a system upgrade operation.
[0069] In practical applications, the distributed storage cluster can be used to effectively store backup data of system software configurations and / or system logs for all second nodes.
[0070] Based on this, in some optional embodiments, the backup system software configuration and / or system logs include: The distributed storage cluster is used to back up system software configurations and / or system logs.
[0071] After configuring the system software and / or backing up the system logs and uploading the system software upgrade package, when the first node performs a system upgrade for each node group, it can first isolate the second nodes included in the node group to avoid affecting the stability of the distributed storage cluster. Then, it can control each second node to perform the system upgrade separately.
[0072] Based on this, in some optional embodiments, performing a system upgrade on all second nodes included in the node group includes: The isolated node group contains all the second nodes; For each isolated second node, control the second node to perform system upgrade operations.
[0073] In practical applications, the isolation operations may differ depending on the type of the second node (such as a storage node or a gateway node). These will be explained separately below: When the second node includes a storage node, performing a crash operation on the second node requires ensuring that each storage process (such as OSD) meets the isolation conditions (which can also be understood as crash conditions).
[0074] Based on this, in some optional embodiments, the isolated node group includes all second nodes, including: Instruct the second node to perform the following: determine whether each storage process associated with the second node meets the isolation conditions, and obtain the determination result; if the determination result indicates that all storage processes associated with the second node meet the isolation conditions, perform a crash operation. When the second node includes a gateway node, performing a shutdown operation on the second node requires ensuring that there are no services that maintain a connection through the second node.
[0075] Based on this, in some optional embodiments, the isolated node group includes all second nodes, including: Instruct the second node to perform the following: set the load balancing weight corresponding to the second node to 0; perform a liveness check query to obtain the query results; if the query results indicate that there are no services that maintain connections, perform a shutdown operation.
[0076] In practical applications, isolating gateway nodes can also be understood as isolating single-node services on the load balancing side, and this application embodiment does not limit this.
[0077] After isolating each second node, the first node can control that second node to execute a system upgrade script, thereby upgrading the second node to the target system. The system upgrade process may include steps such as software package replacement, software package update, backup and restore of existing configurations, and system startup; this embodiment does not limit the specific steps. Here, the system upgrade can be implemented via online upgrade.
[0078] After the system upgrade is completed, the first node can control the second node to check for faults in the local hardware and network. If the check passes, the second node will obtain the software upgrade package based on the configured software source and perform the software upgrade.
[0079] Based on this, in some optional embodiments, the method may further include: After performing a system upgrade operation, each second node is controlled to pull a software upgrade package from the third node and use the software upgrade package to perform a software upgrade. The third node is used at least to manage the software source.
[0080] After the software upgrade is completed, for the storage node that performed the system upgrade operation, the first node can control the node to restart all associated storage processes (such as OSD) to rejoin the storage processes into the distributed storage cluster.
[0081] In practical applications, to avoid cluster instability caused by a large number of storage processes joining the distributed storage cluster, the first node can control the storage nodes to restart all storage processes in multiple stages.
[0082] Based on this, in some optional embodiments, the method may further include: After performing a system upgrade, each storage node is controlled to restart each storage process associated with it in batches; wherein, the distributed storage cluster contains at least storage nodes, and each storage node is associated with one or more storage processes.
[0083] In practical applications, the number of batches to be restarted and the number of storage processes to be restarted in each batch can be set according to actual needs, such as restarting one storage process each time (i.e., restarting storage processes one by one), etc. This application embodiment does not limit this.
[0084] After each storage process in a storage node restarts and rejoins the distributed storage cluster, data balancing is required to allow the storage process to recover from data inconsistencies caused by not joining the distributed storage cluster for a long time, thereby ensuring the performance of the distributed storage cluster.
[0085] In practical applications, the data balancing step size of each storage process can be dynamically adjusted to ensure that the data balancing process does not affect the data storage bandwidth of the storage process itself.
[0086] Based on this, in some optional embodiments, the distributed storage cluster includes at least storage nodes, each associated with one or more storage processes, and the method may further include: After each storage node performs a system upgrade, the pace of data balancing for each storage process is adjusted based on the service bandwidth of that storage process.
[0087] In practical applications, the step size for data balancing can be understood as the bandwidth used for data balancing.
[0088] The system upgrade method provided in this application embodiment involves a first node in a distributed storage cluster determining first information, second information, and third information. The distributed storage cluster is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes. The first information represents the state of the distributed storage cluster, the second information represents one or more second nodes to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool division results of the distributed storage cluster. Using the first, second, and third information, one or more node groups and an upgrade order corresponding to each node group are determined. Each node group contains one or more second nodes, and the second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or belong to different storage pools. According to the upgrade order, system upgrades are performed on all second nodes contained in the node groups sequentially. The solution provided in this application embodiment involves a node (i.e., the first node) that initiates a system upgrade of a distributed storage cluster. Based on the status of the distributed storage cluster (i.e., the first information), whether the node needs to be upgraded (i.e., the second information), and the fault domain and storage pool to which the node belongs (i.e., the third information), the nodes that can be upgraded simultaneously are divided into the same node group, and the system upgrade is performed on all nodes in each node group in sequence (i.e., the upgrade order). When partitioning node groups, by considering the node status and system upgrade requirements, it can be ensured that nodes undergoing simultaneous upgrades meet the system upgrade conditions and require system upgrades. Simultaneously, by ensuring that all nodes in each node group come from the same fault domain or different storage pools, the number of nodes in each node group can be maximized (which can also be understood as maximizing the number of nodes undergoing simultaneous system upgrades), thereby reducing the time required for distributed storage cluster upgrades and effectively improving the system upgrade efficiency of the distributed storage cluster. This is because: 1) When storing data, the distributed storage cluster stores multiple blocks corresponding to the data in different fault domains. In this case, even if all nodes in a single fault domain fail simultaneously, the data can still be recovered using data blocks stored in other fault domains, without data loss. Therefore, the distributed storage cluster can tolerate all nodes in the same fault domain undergoing simultaneous system upgrades; 2) When storing data, the distributed storage cluster stores multiple copies of the data in different storage pools. In this case, even if a node in a single storage pool fails, it will not affect the copies stored in other storage pools, without data loss. Therefore, the distributed storage cluster can tolerate nodes in different storage pools undergoing simultaneous system upgrades.
[0089] The following section provides a more detailed description of this application with reference to application examples.
[0090] For large-scale distributed storage systems, this application example provides a method for replacing the operating system in a distributed storage system, aiming to achieve more efficient and rapid system upgrades of the entire cluster without affecting the stability of business systems, reducing unpredictable factors caused by long upgrade times, and ensuring the feasibility of the upgrade process and the normal operation of software services after the upgrade. Figure 3 As shown, it includes the following steps: Step 301: Cluster status check; Here, you can use ceph -s to check whether pg is in an active state and does not need to be repaired (can be expressed as active + clean in English), and whether there are any slow operations in the cluster.
[0091] Step 302: Ensure the environment information is updated; In practical applications, we can check whether the hardware meets the upgrade requirements, whether the memory meets the upgrade system space requirements, whether the cluster network status is normal, and whether the software of each component meets the upgrade requirements.
[0092] Step 303: Upload the upgrade system software package; In practical applications, each node in the cluster that needs a system upgrade can upload the upgrade packages for each component to a separate public software source node and create a software yum repository. This enables unified management of upgrade packages and allows nodes to pull the upgrade packages uniformly after the upgrade.
[0093] Step 304: Enable the loop test function; In practical applications, after enabling the cyclic probing function, normal S3 requests to the cluster can be probinged, including basic function probing, cluster abnormal status checks, and cluster error rate checks.
[0094] Step 305: Check if there are any hardware problems in the current node. Hardware includes: CPU, hard drive, network card, etc. In practical applications, steps 301 to 305 can be referred to as the pre-stage.
[0095] Step 306: Generate the execution matrix based on data pool, storage pool, and fault domain; In practical applications, multiple nodes to be upgraded that have all three replicas in the storage cluster and belong to different storage pools or the same fault domain can be placed in the same row of the execution matrix.
[0096] Step 307: Perform the upgrade operation on all nodes contained in each row of the execution matrix in the order of each row.
[0097] In practical application, steps 306 to 307 can be referred to as the upgrade phase. Step 307, which involves performing an upgrade operation for each row, can specifically include the following steps: Step 1: Dynamically adjust the number of PGlogs for the current node to facilitate subsequent reading from the logs to restore balanced data; Step 2: Set the minimum number of replicas in the cluster to ensure that if other replica OSDs encounter problems during the upgrade process after the current node's OSD replica goes down, the cluster can still provide services normally. Step 3: Isolate the node to be upgraded; In practical applications, for storage nodes, an isolation script can be used to proactively determine whether the storage node software service (which can be an OSD service) can be isolated, and if it can be isolated, then isolation can be performed. For gateway nodes, single-node service isolation can be performed on the load balancer side, and services with no known connections can be guaranteed.
[0098] Step 4: Generate a replacement system YUM repository and configure the software installation repository after system upgrade and restart; Step 5: Backup system software configuration and system logs; In practical applications, large-scale data backup is achieved by uploading system software configuration backups and system log backups to object storage distributed storage. Step 6: Execute the system upgrade script to upgrade the operating system online; For example, the operating system can be upgraded from Bcliunx 7.6 to BigCloud Enterprise Linux for Euler.
[0099] Step 7: Check for problems with local hardware and network; Step 8: Restart the OSD service; In practical applications, for single-node instances, each OSD is restarted individually to ensure that the cluster does not become unstable during the service restart process due to a large number of OSDs being added to the cluster.
[0100] Step 9: Dynamically adjust adaptive recovery; Specifically, adjusting the data balance among replicas, without affecting the original actual business bandwidth, and gradually adjusting the balancing step to achieve data balance can accelerate the process of resolving data inconsistencies caused by OSD replicas not being added to the cluster for a long time.
[0101] Step 10: Provide a check script for each node to perform data checks at the PG level, which can avoid data inconsistency issues and check whether each OSD starts up normally; Step 11: Check if the versions of the software packages are consistent and if the operating system has been replaced with the target system; and ensure that the cluster status is normal by performing basic function checks. Step 12: Restore cluster settings parameters; In practical applications, the cluster should be restored to its state before the change to prevent the cluster state from being lost due to the change being paused.
[0102] The solution provided in this application example first achieves maximum concurrency for large-scale cluster operating system replacement and upgrades under conditions of data pools, storage pools, racks, fault domains, and independent components. Secondly, while ensuring upgrade efficiency, it guarantees stability during the upgrade process through multiple monitoring methods. Thirdly, adaptive recovery scripts accelerate the recovery speed after OSDs are replaced and rejoined to the cluster. Finally, the reliability of each service software during the upgrade process perfectly ensures service availability during operating system replacement. Thus, by optimizing upgrade concurrency, accelerating internal cluster balancing, and implementing multi-level health alerts, the efficiency of upgrading ultra-large-scale storage cluster systems is improved.
[0103] To implement the method of the embodiments of this application, the embodiments of this application also provide a system upgrade device, which is set in the first node of a distributed storage cluster. The distributed storage cluster is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes, such as... Figure 4 As shown, the device includes: The determining unit 401 is used to determine first information, second information, and third information. The first information represents the state of the distributed storage cluster, the second information represents one or more second nodes to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool partitioning results of the distributed storage cluster. The unit also uses the first, second, and third information to determine one or more node groups and the upgrade order corresponding to each node group. Each node group contains one or more second nodes. The second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or belong to different storage pools. Upgrade unit 402 is used to perform system upgrades on all second nodes included in the node group in sequence according to the upgrade order.
[0104] In some optional embodiments, the upgrade unit 402 is specifically used for: The isolated node group contains all the second nodes; For each isolated second node, control the second node to perform system upgrade operations.
[0105] In some optional embodiments, the upgrade unit 402 is specifically used for: For each second node in the node group, when the second node includes a storage node, the second node is instructed to perform the following: determine whether each storage process associated with the second node meets the isolation condition, and obtain the determination result; if the determination result indicates that all storage processes associated with the second node meet the isolation condition, perform a crash operation. And / or, When the second node includes a gateway node, instruct the second node to perform the following: set the load balancing weight corresponding to the second node to 0; perform a liveness check query to obtain the query result; and if the query result indicates that there are no services that maintain connections, perform a shutdown operation.
[0106] In some alternative embodiments, the device may further include: The control unit is used to increase a first value, which represents the number of log records to be retained in the placement group associated with the distributed storage cluster; wherein the distributed storage cluster is configured with one or more placement groups, the placement groups are used to perform distributed management of the data stored in the distributed storage cluster, and the placement group logs are used to record the management operations of the placement groups.
[0107] In some optional embodiments, the control unit is further configured to: The second value is reduced to enable the distributed storage cluster to provide storage services during system upgrades. The second value represents the minimum number of replicas associated with the distributed storage cluster.
[0108] In some optional embodiments, the distributed storage cluster includes at least storage nodes, each associated with one or more storage processes, and the control unit is further configured to: Control each storage node to restart each storage process associated with the storage node in batches after performing a system upgrade operation.
[0109] In some optional embodiments, the control unit is further configured to: After performing a system upgrade operation, each second node is controlled to pull a software upgrade package from the third node and use the software upgrade package to perform a software upgrade. The third node is used at least to manage the software source.
[0110] In some optional embodiments, the control unit is further configured to: Before performing a system upgrade operation, each second node is controlled to upload the software upgrade package to the third node.
[0111] In some optional embodiments, the distributed storage cluster includes at least storage nodes, each associated with one or more storage processes, and the control unit is further configured to: After each storage node performs a system upgrade, the pace of data balancing for each storage process is adjusted based on the service bandwidth of that storage process.
[0112] In some optional embodiments, the control unit is further configured to: Control each secondary node to back up the system software configuration and / or system logs before performing a system upgrade operation.
[0113] In practical applications, the determining unit 401, the upgrading unit 402, and the control unit can be implemented by the processor in the system upgrading device combined with the communication interface.
[0114] It should be noted that the system upgrade device provided in the above embodiments is only illustrated by the division of the above-described program units when performing system upgrades. In actual applications, the above processing can be assigned to different program units as needed, that is, the internal structure of the device can be divided into different program units to complete all or part of the processing described above. In addition, the system upgrade device and system upgrade method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0115] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiments of this application, the embodiments of this application also provide an electronic device. The electronic device belongs to a distributed storage cluster, which is divided into one or more storage pools. Each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes, such as... Figure 5 As shown, the electronic device 500 includes: The communication interface 501 is capable of exchanging information with other devices (such as other nodes in the distributed storage cluster); The processor 502 is connected to the communication interface 501 to enable information interaction with other devices and to execute the methods provided by one or more of the above-mentioned technical solutions when running a computer program; The computer program is stored in memory 503.
[0116] Specifically, the processor 502 is used for: In conjunction with the communication interface 501, first information, second information and third information are determined. The first information represents the state of the distributed storage cluster. The second information represents one or more second nodes that need to be upgraded by the system. The one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool partitioning results of the distributed storage cluster. Using the first information, the second information, and the third information, one or more node groups and the upgrade order corresponding to each node group are determined. Each node group contains one or more second nodes. The second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or belong to different storage pools. Following the upgrade order, system upgrades are performed on all second nodes within the node group in sequence.
[0117] In some optional embodiments, the processor 502 is specifically used for: Combined with the communication interface 501, the isolation node group includes all second nodes; For each isolated second node, control the second node to perform system upgrade operations.
[0118] In some alternative embodiments, the processor 502 is specifically used for: In conjunction with the communication interface 501, for each second node included in the node group, when the second node includes a storage node, the second node is instructed to perform: determine whether each storage process associated with the second node meets the isolation conditions, and obtain the determination result; if the determination result indicates that all storage processes associated with the second node meet the isolation conditions, a crash operation is performed. And / or, When the second node includes a gateway node, instruct the second node to perform the following: set the load balancing weight corresponding to the second node to 0; perform a liveness check query to obtain the query result; and if the query result indicates that there are no services that maintain connections, perform a shutdown operation.
[0119] In some optional embodiments, the processor 502 is further configured to: Increase the first value, which represents the number of log records to be retained in the placement group associated with the distributed storage cluster; wherein, the distributed storage cluster is configured with one or more placement groups, the placement groups are used to perform distributed management of the data stored in the distributed storage cluster, and the placement group logs are used to record the management operations of the placement groups.
[0120] In some optional embodiments, the processor 502 is further configured to: The second value is reduced to enable the distributed storage cluster to provide storage services during system upgrades. The second value represents the minimum number of replicas associated with the distributed storage cluster.
[0121] In some optional embodiments, the distributed storage cluster includes at least storage nodes associated with one or more storage processes, and the processor 502 is further configured to: In conjunction with the communication interface 501, each storage node is controlled to restart each storage process associated with the storage node in batches after performing a system upgrade operation.
[0122] In some optional embodiments, the processor 502 is further configured to: After performing a system upgrade operation, each second node is controlled to pull a software upgrade package from the third node and use the software upgrade package to perform a software upgrade. The third node is used at least to manage the software source.
[0123] In some optional embodiments, the processor 502 is further configured to: Before performing a system upgrade operation, each second node is controlled to upload the software upgrade package to the third node.
[0124] In some optional embodiments, the distributed storage cluster includes at least storage nodes associated with one or more storage processes, and the processor 502 is further configured to: After each storage node performs a system upgrade, the pace of data balancing for each storage process is adjusted based on the service bandwidth of that storage process.
[0125] In some optional embodiments, the processor 502 is further configured to: Control each secondary node to back up the system software configuration and / or system logs before performing a system upgrade operation.
[0126] It should be noted that the specific processing procedures of the processor 502 and the communication interface 501 can be understood by referring to the above method.
[0127] Of course, in practical applications, the various components in electronic device 500 are coupled together through bus system 504. It can be understood that bus system 504 is used to realize the connection and communication between these components. In addition to a data bus, bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in... Figure 5 The general designated all buses as Bus System 504.
[0128] The memory 503 in this embodiment is used to store various types of data to support the operation of the electronic device 500. Examples of such data include any computer program used to operate on the electronic device 500.
[0129] The methods disclosed in the embodiments of this application can be applied to the processor 502, or implemented by the processor 502. The processor 502 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 502 or by instructions in the form of software. The processor 502 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 502 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the memory 503. The processor 502 reads the information in the memory 503 and combines its hardware to complete the steps of the aforementioned method.
[0130] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.
[0131] It is understood that the memory (memory 503) in this embodiment of the application can be volatile memory or non-volatile memory, or both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); the magnetic surface memory can be disk storage or magnetic tape storage. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.
[0132] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 503 storing a computer program, which can be executed by the processor 502 of the electronic device 500 to complete the steps described in the aforementioned method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0133] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by the processor 502 of the electronic device 500 to complete the steps described in the aforementioned method.
[0134] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0135] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0136] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. A system upgrade method, characterized in that, The method, applied to the first node in a distributed storage cluster, wherein the distributed storage cluster is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes, includes: The system determines first information, second information, and third information. The first information represents the state of the distributed storage cluster, the second information represents one or more second nodes that need to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool partitioning results of the distributed storage cluster. Using the first information, the second information, and the third information, one or more node groups and the upgrade order corresponding to each node group are determined. The node group contains one or more second nodes. The second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or belong to different storage pools. Following the upgrade order, system upgrades are performed on all second nodes within the node group in sequence.
2. The method according to claim 1, characterized in that, The system upgrade performed on all second nodes included in the node group includes: The isolated node group contains all the second nodes; For each isolated second node, control the second node to perform system upgrade operations.
3. The method according to claim 2, characterized in that, The isolated node group includes all second nodes, including: For each second node in the node group, when the second node includes a storage node, the second node is instructed to perform the following: determine whether each storage process associated with the second node meets the isolation condition, and obtain the determination result; if the determination result indicates that all storage processes associated with the second node meet the isolation condition, perform a crash operation. And / or, When the second node includes a gateway node, instruct the second node to perform the following: set the load balancing weight corresponding to the second node to 0; perform a liveness check query to obtain the query result; and if the query result indicates that there are no services that maintain connections, perform a shutdown operation.
4. The method according to claim 1, characterized in that, The method further includes: Increase the first value, which represents the number of log records to be retained in the placement group associated with the distributed storage cluster; wherein, the distributed storage cluster is configured with one or more placement groups, the placement groups are used to perform distributed management of the data stored in the distributed storage cluster, and the placement group logs are used to record the management operations of the placement groups.
5. The method according to claim 1, characterized in that, The method further includes: The second value is reduced to enable the distributed storage cluster to provide storage services during system upgrades. The second value represents the minimum number of replicas associated with the distributed storage cluster.
6. The method according to claim 1, characterized in that, The distributed storage cluster includes at least storage nodes, each storage node being associated with one or more storage processes; the method further includes: Control each storage node to restart each storage process associated with the storage node in batches after performing a system upgrade operation.
7. The method according to claim 6, characterized in that, The method further includes: After performing a system upgrade operation, each second node is controlled to pull a software upgrade package from the third node and use the software upgrade package to perform a software upgrade. The third node is used at least to manage the software source.
8. The method according to claim 7, characterized in that, The method further includes: Before performing a system upgrade operation, each second node is controlled to upload the software upgrade package to the third node.
9. The method according to claim 1, characterized in that, The distributed storage cluster includes at least storage nodes, each storage node being associated with one or more storage processes; the method further includes: After each storage node performs a system upgrade, the data balancing pace of each storage process is adjusted based on the service bandwidth of that storage process.
10. The method according to claim 1, characterized in that, The method further includes: Control each secondary node to back up the system software configuration and / or system logs before performing a system upgrade operation.
11. The method according to claim 10, characterized in that, The backup system software configuration and / or system logs include: The distributed storage cluster is used to back up system software configurations and / or system logs.
12. A system upgrade device, characterized in that, The first node is set in a distributed storage cluster, which is divided into one or more storage pools, each storage pool is divided into one or more fault domains, and each fault domain contains one or more nodes, including: A determining unit is configured to determine first information, second information, and third information. The first information represents the state of the distributed storage cluster. The second information represents one or more second nodes to be upgraded, and the one or more second nodes belong to the distributed storage cluster. The third information represents the fault domain and storage pool partitioning results of the distributed storage cluster. The unit also uses the first, second, and third information to determine one or more node groups and the upgrade order corresponding to each node group. Each node group contains one or more second nodes, and the second nodes contained in different node groups are completely different. Different nodes contained in the same node group belong to the same fault domain or different storage pools. The upgrade unit is used to perform system upgrades on all second nodes included in the node group in sequence according to the upgrade order.
13. An electronic device, characterized in that, include: The processor and the memory used to store computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 11.
14. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Cluster grouping online upgrading method and system, terminal and storage medium
CN112463195A
Distributed storage grouping online upgrading method and device, terminal and medium
CN115344290A
Storage cluster online upgrading method and device, equipment and medium
CN116431191A
Method, system and equipment for online upgrading of storage system and medium
CN119828980A