Distributed storage cluster disk changing method and electronic equipment

By monitoring the health status and obtaining configuration information when the distributed storage cluster is in a high-water mark state, the capacity overlimit and data consistency issues in the disk replacement operation are resolved, and an efficient and stable hard disk replacement process is achieved.

CN120743633AActive Publication Date: 2025-10-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511232875.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

When the distributed storage cluster is at a high water level, traditional disk replacement operations lead to dynamic capacity overruns, reduced data migration efficiency, and uncontrolled data reconstruction consistency, making it difficult to complete hard drive replacement efficiently, stably, and safely.

Method used

Before replacing the disk, it monitors the cluster health status, obtains object storage device configuration information, outputs operation prompts, and automatically processes misplaced data after the disk replacement is completed, ensuring data consistency through reconstruction operations.

Benefits of technology

This ensures that disk swapping is performed in a safe environment, reduces the impact on cluster performance, improves the efficiency and stability of disk swapping, and reduces the risk of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743633A_ABST
    Figure CN120743633A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed storage cluster disk changing method and electronic equipment, and relates to the technical field of computer storage, the method comprises the following steps: monitoring whether a distributed storage cluster is in a healthy state in response to a disk changing instruction; under the condition that the distributed storage cluster is in the healthy state, object storage device configuration information of a target storage pool is obtained, disk replacement operation prompt information is output, and the target storage pool comprises a to-be-replaced hard disk; when it is detected that the disk changing operation is completed, whether dislocation data exist in the distributed storage cluster or not is checked; if the dislocation data exists, executing a preset recovery operation based on the configuration information of the object storage device so as to eliminate the dislocation data; and executing a preset reconstruction operation after the dislocation data is eliminated, wherein the reconstruction operation comprises writing data stored in the to-be-replaced hard disk into a newly-replaced hard disk. Through the technical scheme provided by the invention, the disk replacement operation of the distributed storage cluster can be efficiently, stably and safely realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer storage technology, and in particular to a distributed storage cluster disk replacement method and electronic equipment. Background Art

[0002] With the development of new-generation information technologies such as cloud computing, big data, and artificial intelligence, distributed storage clusters have gradually become the mainstream solution in the field of data storage due to their significant advantages such as high scalability, high fault tolerance, and elastic resource allocation.

[0003] In a distributed storage cluster, hard drives serve as the core storage medium. Their selection and configuration directly impact the performance of the distributed storage cluster. Hard drives often expire or fail during use, necessitating replacement to extend the lifespan of the distributed storage cluster. However, ensuring efficient, stable, and secure drive replacement in distributed storage clusters, even when the storage pool is at a high level, is a pressing technical challenge. Summary of the Invention

[0004] The present application provides a distributed storage cluster disk replacement method and electronic device, which can efficiently, stably and safely complete the distributed storage cluster disk replacement operation when the storage pool is at a high water level.

[0005] This application provides a method for replacing disks in a distributed storage cluster, the method comprising:

[0006] In response to a disk swap instruction, monitoring is performed to determine whether the distributed storage cluster is in a healthy state; the disk swap instruction includes the name of a target storage pool, and a storage water level of the target storage pool is greater than or equal to a preset water level threshold.

[0007] When the distributed storage cluster is in a healthy state, object storage device configuration information of the target storage pool is obtained and disk replacement operation prompt information is output; the above-mentioned target storage pool includes the hard disk to be replaced, and the object storage device configuration information includes at least one of the following: topology information of the target storage pool, object storage device mapping information, placement group mapping rule information, and placement group distribution information.

[0008] When the disk swap operation is detected to be complete, the distributed storage cluster is checked for misaligned data.

[0009] If there is misplaced data, a preset recovery operation is performed based on the above-mentioned object storage device configuration information to eliminate the misplaced data; and after eliminating the misplaced data, a preset reconstruction operation is performed, which reconstruction operation includes writing the data stored in the hard disk to be replaced to the newly replaced hard disk.

[0010] The present application also provides a distributed storage cluster disk replacement device, comprising:

[0011] The monitoring module is used to monitor whether the distributed storage cluster is in a healthy state in response to a disk swap instruction; the disk swap instruction includes the name of the target storage pool, and the storage water level of the target storage pool is greater than or equal to a preset water level threshold.

[0012] The acquisition module is used to obtain the object storage device configuration information of the target storage pool and output disk replacement operation prompt information when the distributed storage cluster is in a healthy state; the above-mentioned target storage pool includes the hard disk to be replaced, and the object storage device configuration information includes at least one of the following: topology information of the target storage pool, object storage device mapping information, placement group mapping rule information, and placement group distribution information.

[0013] The checking module is used to check whether there is misplaced data in the distributed storage cluster when detecting that the disk swap operation is completed.

[0014] The recovery module is used to perform a preset recovery operation based on the object storage device configuration information to eliminate the misplaced data if there is misplaced data; and to perform a preset reconstruction operation after eliminating the misplaced data, wherein the reconstruction operation includes writing the data stored in the hard disk to be replaced to the newly replaced hard disk.

[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned distributed storage cluster disk replacement methods when executing the computer program.

[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for replacing a disk in a distributed storage cluster are implemented.

[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned distributed storage cluster disk replacement methods when executed by a processor.

[0018] The distributed storage cluster disk replacement method and electronic device provided in the present application first monitor whether the distributed storage cluster is in a healthy state before performing the disk replacement operation. When the distributed storage cluster is in a healthy state, the object storage device configuration information of the target storage pool (the storage water level is greater than or equal to the preset water level threshold) is obtained, and a disk replacement operation prompt information is output, which can ensure that the disk replacement operation is performed in a safe environment. At the same time, by obtaining the object storage device configuration information of the target storage pool, subsequent data recovery can be guaranteed; when the disk replacement operation is completed, if there is misplaced data in the distributed storage cluster, a preset recovery operation can be executed based on the above-mentioned object storage device configuration information to automatically eliminate the misplaced data and ensure data consistency, thereby reducing the impact of the disk replacement operation on the performance of the distributed storage cluster and completing the disk replacement of the distributed storage cluster efficiently, stably and safely. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 A schematic diagram of a process for replacing a disk in a distributed storage cluster provided in an embodiment of the present application;

[0021] Figure 2 This is another flowchart of a method for replacing a disk in a distributed storage cluster provided in an embodiment of the present application;

[0022] Figure 3 This is a structural diagram of a distributed storage cluster disk replacement device provided in an embodiment of the present application;

[0023] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0026] The following is an explanation of some of the data involved in the examples of this application:

[0027] Object Storage Device (OSD): The most basic storage unit in a distributed storage cluster, responsible for data storage, replication, recovery, and load balancing. It is a key component for ensuring high data availability, consistency, and cluster performance.

[0028] Controlled Replication Under Scalable Hashing (CRUSH) algorithm: A controllable, scalable, distributed replica data placement algorithm, primarily used in distributed storage clusters. It calculates the storage location of data in the cluster to achieve efficient data distribution, load balancing, and fault tolerance management.

[0029] A Placement Group (PG) is a core logical unit in a distributed storage cluster used for data distribution, load balancing, and fault tolerance. Its design significantly reduces cluster complexity and improves scalability by grouping objects. As the middle layer between objects and OSDs, PGs perform the following key functions:

[0030] (1) Object aggregation: Grouping and managing massive objects (e.g., millions) in a storage pool avoids maintaining metadata (e.g., location information) for each object, thereby reducing computational overhead. For example, a storage pool containing 1 million objects can be divided into 1,000 PGs, with each PG managing approximately 1,000 objects.

[0031] (2) Data distribution: Map PG to OSD through the CRUSH algorithm to achieve uniform distribution of data in the cluster.

[0032] (3) Load balancing: Dynamically adjust the mapping relationship between PG and OSD to avoid overloading of a single OSD. For example, when the disk usage of an OSD exceeds the threshold, the storage cluster will migrate some PGs to other idle OSDs.

[0033] With the rapid development and deep integration of next-generation information technologies such as cloud computing, big data, and artificial intelligence, the scale of data generated by various business applications is rapidly increasing. In this era of big data, distributed storage clusters, with their significant advantages such as high scalability, high fault tolerance, and flexible resource allocation, have gradually become the mainstream solution for massive data storage. Their application scenarios are becoming increasingly broad, covering a wide range of industries, from internet services to other fields.

[0034] With the continued deployment and deepening application of distributed storage architectures, the amount of data stored in storage clusters continues to break historical records. This explosive growth has led to a sharp increase in the number of hard drive nodes required in distributed storage clusters, resulting in ultra-large-scale storage arrays composed of tens of thousands of hard drives. In this context, hard drives, as the fundamental carrier for persistent data storage, are crucial for ensuring the high availability of the entire distributed storage cluster, with their operational stability, reliability, and durability.

[0035] In the refined operation and maintenance management of distributed storage clusters, a storage pool high water mark (HSM) serves as a key capacity warning indicator (typically defined as the actual capacity utilization of the storage pool reaching or exceeding a preset threshold, such as 95%). When triggered, the storage cluster enters a capacity-sensitive period. However, traditional hard drive replacement strategies fail to fully consider the special constraints of high water mark scenarios during the design phase, resulting in the following multi-dimensional technical issues during actual operation and maintenance:

[0036] Dynamic capacity overrun: When the storage pool is at a high watermark, traditional disk swapping requires a data migration pre-allocation space check. Because remaining available capacity approaches the safety limit, the expansion of intermediate copies or temporary metadata generated during the data migration process can instantly exceed the physical capacity limit, triggering the storage controller's hard capacity isolation mechanism. This situation can directly result in new write requests being denied service or forcing the cluster to automatically initiate a prioritized data eviction policy, leading to a surge in write latency or even a temporary outage of critical business data.

[0037] Reduced data migration efficiency: Traditional disk swapping typically relies on a PG rebalancing algorithm, which triggers a cross-node data redistribution across all PGs within the storage pool. Under high watermark conditions, the storage cluster's input / output (I / O) channels are already operating under heavy load. At this time, PG rebalancing competes with business I / O resources, resulting in a non-linear increase in I / O response time. In extreme cases, this can cause a surge in the I / O request queue depth, resulting in a drastic drop in the overall storage cluster throughput.

[0038] Data reconstruction consistency is out of control: Under high watermark conditions, the distribution density of PG data shards across storage nodes increases significantly. Traditional PG recovery processes after disk swaps lack intelligent distribution verification mechanisms. When a new hard drive is connected, if the PG mapping relationship recovery algorithm fails to accurately match the original data topology, a large number of data shards will be incorrectly placed on non-optimal nodes. This distribution misalignment triggers a secondary reconstruction effect—the storage cluster must perform additional cross-node data correction operations, which not only prolongs data recovery time but also potentially exhausts link bandwidth due to multiple data transfers.

[0039] The above technical issues will systematically threaten the reliability of the storage cluster, specifically manifesting as: high risk of data loss, unstable cluster status and long recovery time, making it difficult to meet the needs of efficient and safe disk replacement.

[0040] Furthermore, if problems arise during the traditional disk swap process, manual adjustments to the PG data distribution and the amount of data being migrated are often required to avoid migrating large amounts of data. However, this approach is not only labor-intensive but also prone to operational errors and poor fault tolerance.

[0041] Based on the above, how to efficiently, stably and safely complete the disk swap operation of distributed storage under high water level conditions is a technical problem that technicians in this field urgently need to solve.

[0042] In response to the above technical problems, an embodiment of the present application provides a method for replacing disks in a distributed storage cluster. Before performing a disk replacement operation, the method first monitors whether the distributed storage cluster is in a healthy state. When the distributed storage cluster is in a healthy state, the object storage device configuration information of the target storage pool (the storage water level is greater than or equal to the preset water level threshold) is obtained to ensure that the disk replacement operation is performed in a safe environment. When the disk replacement operation is completed, if there is misplaced data in the distributed storage cluster, a preset recovery operation can be performed based on the above object storage device configuration information to automatically eliminate the misplaced data, thereby reducing the impact of the disk replacement operation on the performance of the distributed storage cluster, and completing the disk replacement of the distributed storage cluster efficiently, stably and safely.

[0043] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0044] See also Figure 1 , Figure 1 This is a flow chart of a method for replacing a disk in a distributed storage cluster provided in an embodiment of the present application. In some embodiments, the method for replacing a disk in a distributed storage cluster includes:

[0045] S101. In response to a disk swap instruction, monitor whether the distributed storage cluster is in a healthy state. If so, continue to execute S102; if not, exit the current disk swap process.

[0046] Optionally, the disk swap instruction includes the name of a target storage pool, and the storage water level of the target storage pool is greater than or equal to a preset water level threshold (eg, 95%).

[0047] In some implementations, the monitoring interface of the distributed storage cluster can be used to periodically collect status indicators of storage cluster nodes, including but not limited to key parameters such as storage node online status, network link bandwidth utilization, I / O request response latency, and data copy synchronization status.

[0048] By comprehensively analyzing the above status indicators, it can be determined whether the distributed storage cluster is in a healthy state. For example, when all key parameters are within the preset threshold range, it can be determined that the distributed storage cluster is in a maintainable healthy state.

[0049] S102: Obtain object storage device configuration information of a target storage pool, and output disk replacement operation prompt information; wherein the target storage pool includes the hard disk to be replaced.

[0050] In some implementations, after confirming that the distributed storage cluster is in a maintainable healthy state, the OSD configuration information of the target storage pool may be obtained through the object storage management interface.

[0051] Exemplarily, the above-mentioned OSD configuration information may include topology information of the target storage pool, OSD mapping information, PG mapping rule information, PG distribution information, etc., which is not limited in the embodiments of the present application.

[0052] In some embodiments, after obtaining the object storage device configuration information of the target storage pool, a visual disk replacement operation prompt can be output to the operation and maintenance personnel through the management console. The disk replacement operation prompt can include the location of the device to be replaced, the data migration path description, and safety operation precautions. After receiving the disk replacement operation prompt, the operation and maintenance personnel can replace the hard disk to be replaced according to the disk replacement operation prompt.

[0053] S103: When the disk swap operation is detected to be complete, check whether there is misaligned data in the distributed storage cluster. If yes, proceed to S104; if not, proceed to S105.

[0054] In some implementations, the hard drive tray status changes can be monitored in real time through the hardware management interface, combined with the alarm information in the storage log. If it is confirmed that the old hard drive has been physically removed and the new hard drive has completed power-on identification, it can be determined that the above-mentioned disk replacement operation is completed.

[0055] When the disk swap operation is completed, a data consistency check is performed to determine whether there is misaligned data in the distributed storage cluster by comparing the storage pool metadata records with the actual physical layout.

[0056] For example, the following three possible misalignment scenarios can be identified:

[0057] Orphan shard: The metadata record exists but the physical device is missing.

[0058] Ghost shard: The physical device exists but there is no corresponding metadata record.

[0059] Location mismatch: The actual storage location of the shard does not match the metadata record.

[0060] S104: Based on the object storage device configuration information, perform a preset recovery operation to eliminate the misplaced data.

[0061] In some implementations, if misplaced data exists, previously acquired OSD configuration information may be imported and used to restore the original configuration of the target storage pool, thereby eliminating the misplaced data.

[0062] After the recovery operation is performed, a data consistency check may be performed again to determine whether there is misaligned data in the distributed storage cluster. When it is determined that there is no misaligned data in the distributed storage cluster, step S105 is performed.

[0063] S105 : Execute a preset reconstruction operation, where the reconstruction operation includes writing the data stored in the hard disk to be replaced into the new hard disk.

[0064] In some embodiments, when it is determined that there is no misplaced data in the distributed storage cluster, the data reconstruction engine can be started to perform phased, interruptible and efficient data migration based on the current OSD configuration information and the real-time status of the cluster, and write the data stored on the hard disk to be replaced to the newly replaced hard disk.

[0065] The distributed storage cluster disk replacement method provided in the embodiment of the present application first monitors whether the distributed storage cluster is in a healthy state before performing the disk replacement operation. When the distributed storage cluster is in a healthy state, the object storage device configuration information of the target storage pool is obtained, and the disk replacement operation prompt information is output, so as to ensure that the disk replacement operation is performed in a safe environment. At the same time, by obtaining the object storage device configuration information of the target storage pool, it can provide protection for subsequent data recovery; when the disk replacement operation is completed, if there is misplaced data in the distributed storage cluster, the preset recovery operation can be executed based on the above-mentioned object storage device configuration information to automatically eliminate the misplaced data and ensure data consistency, thereby reducing the impact of the disk replacement operation on the performance of the distributed storage cluster and completing the disk replacement of the distributed storage cluster efficiently, stably and safely.

[0066] In some embodiments, monitoring whether the distributed storage cluster is in a healthy state in step S101 includes:

[0067] Obtain the OSD list associated with the target storage pool and check whether the status of each OSD in the OSD list is normal; when the status of each OSD in the OSD list is normal, obtain the operating parameters of the distributed storage cluster; based on the operating parameters, determine whether the distributed storage cluster is in a healthy state.

[0068] If there are OSDs in an abnormal state in the OSD list, it can be determined that the distributed storage cluster is not in a healthy state.

[0069] For example, the management tool of the distributed storage cluster can be used to query the OSD list associated with the target storage pool. The obtained OSD list is traversed and a status check is performed on each OSD. For example, the status information of each OSD is obtained using the command provided by the distributed storage cluster. The OSD status information is checked to determine whether it is normal. If any OSD in an abnormal state is found in the OSD list, the information of the OSD is immediately recorded and the cluster is marked as unhealthy.

[0070] After confirming that all OSDs are in normal state, continue to obtain the overall operating parameters of the cluster, such as cluster health, data distribution, usage, I / O performance indicators, etc. Based on these operating parameters, determine whether the distributed storage cluster is in a healthy state.

[0071] Optionally, the operating parameters include status parameters of the distributed storage cluster and the number of blocked interfaces; and determining whether the distributed storage cluster is in a healthy state based on the operating parameters includes:

[0072] When the state parameter includes a health state field and the number of the blocked interfaces is less than an alarm number threshold, it is determined that the distributed storage cluster is in a healthy state.

[0073] For example, the cluster's health status field (e.g., HEALTH_OK, HEALTH_WARN, HEALTH_ERR) can be checked. If the cluster is in a healthy state and the number of blocked interfaces is less than a warning threshold, the distributed storage cluster can be determined to be healthy. If the cluster's health status is a health warning or a fault, or the number of blocked interfaces is greater than or equal to the warning threshold, the storage cluster is determined to be unhealthy.

[0074] In some implementations, the OSD list associated with the target storage pool can be obtained to verify whether the OSD status is normal, and to determine the health status of the storage cluster and whether there is serious I / O congestion. If the cluster health status check passes, the disk replacement process continues, otherwise the process exits.

[0075] In the embodiment of the present application, by real-time monitoring of the health status of the distributed storage cluster, it can be ensured that the disk replacement operation is performed in a safe environment.

[0076] In some embodiments, obtaining the OSD configuration information of the target storage pool includes:

[0077] Based on the name of the target storage pool, a directory for storing configuration files during the disk swap process is established; the OSD configuration information of the target storage pool is obtained, and the obtained OSD configuration information is saved to the above directory according to the current timestamp.

[0078] In some implementations, the distributed storage cluster may receive user input parameters during script execution, and generate and save OSD configuration information of the target storage pool according to timestamps.

[0079] Optionally, the OSD configuration information includes at least one of the following information:

[0080] The topology information (Crushmap) of the target storage pool, which includes the hierarchical relationship of the hard disks in the target storage pool; OSD mapping information (OSDmap), which is generated based on the above topology information; PG mapping rule information (Upmap); PG distribution information (PG Dump).

[0081] In some implementations, before replacing the hard disk, the OSD configuration information of the target storage pool can be obtained, including generating a Crushmap of the target storage pool to save the hard disk hierarchy relationship, using the OSD mapping information tool (osdmaptool) combined with the obtained OSDmap to generate an Upmap file, and obtaining a PG Dump to solidify the data distribution logic before the disk replacement.

[0082] In some implementations, the topology information of the target storage pool may be generated according to the identification information of the hard disks in the target storage pool.

[0083] Optionally, the hard disk identification information may include the hard disk serial number, device path, slot information, interface type, etc. Based on the identification information of each hard disk in the target storage pool, the connection relationship between the hard disks in the target storage pool is determined, and then the topology information of the target storage pool is generated.

[0084] In some implementations, OSD mapping information of the target storage pool may be generated according to identification information of the OSD in the target storage pool.

[0085] For example, by parsing the unique identifier of each OSD in the storage pool and combining it with its physical / logical affiliation (such as the host, rack, hard disk type, etc.), a structured mapping table is generated to clarify the location and attributes of each OSD in the storage pool.

[0086] In some implementations, placement group mapping rule information of the target storage pool may be generated based on the OSD mapping information.

[0087] For example, by parsing OSD mapping information (such as OSD's tier affiliation, weight, status, etc.), combined with the storage pool's replication strategy and CRUSH algorithm rules, a set of rules is generated to clarify which OSDs each placement group should be mapped to, in order to achieve high availability and balanced distribution of data.

[0088] In some implementations, placement group distribution information of the target storage pool may be generated based on identification information of the placement groups in the target storage pool;

[0089] For example, by parsing the unique identifier of each placement group in the target storage pool, combining the OSD mapping information and CRUSH algorithm rules, a structured distribution table is generated to clarify the OSD combination to which each PG is currently actually mapped and its physical distribution (such as the host, rack, etc.).

[0090] In an embodiment of the present application, an OSD mapping information tool is used to generate Upmap optimization mapping rules, and information such as crushmap and PG Dump are exported to solidify the data distribution logic before disk replacement, providing a reliable data basis for subsequent recovery operations, avoiding complex problems caused by data migration and distribution changes under high water levels, and ensuring the security and consistency of the disk replacement process.

[0091] In some embodiments, the checking whether there is misaligned data in the distributed storage cluster includes:

[0092] Detect whether the distributed storage cluster is in a stable state; when it is determined that the distributed storage cluster is in a stable state, perform a consistency check on the distributed storage cluster to determine whether there is misaligned data in the distributed storage cluster.

[0093] In some implementations, after detecting that the user has replaced the hard disk, a consistency check (peering) may be performed. After the consistency check is completed, it is detected whether the distributed storage cluster is in a stable state. When it is determined that the distributed storage cluster is in a stable state, a consistency check is performed on the distributed storage cluster to determine whether there is misaligned data.

[0094] In the embodiment of the present application, when it is determined that the distributed storage cluster is in a stable state, a consistency check is performed on the distributed storage cluster to ensure the accuracy of the consistency check result.

[0095] In some embodiments, if there is no misaligned data, a cluster reconstruction operation can be initiated; if misaligned data is detected in the storage cluster, a preset recovery operation is performed, including: updating the configuration parameters of the target storage pool based on the object storage device configuration information; and updating the distribution location of the data stored in the distributed storage cluster based on the updated configuration parameters of the target storage pool.

[0096] In some embodiments, if misaligned data is detected in the storage cluster, the OSD configuration information saved before the disk replacement can be re-imported to restore the target storage pool to the original configuration, and wait for the consistency check to complete again. At the same time, wait for the misaligned data to disappear before starting the cluster reconstruction operation.

[0097] In the embodiment of the present application, a recovery operation is performed based on the OSD configuration information, thereby avoiding reconstruction failure or performance degradation caused by improper data distribution under high water level conditions, and improving the recovery efficiency and reliability after disk replacement.

[0098] In some embodiments, if it is detected that there is misplaced data in the storage cluster, the newly generated OSD configuration information after the disk replacement can also be backed up. When the OSD configuration information saved before the disk replacement makes it difficult to make the misplaced data disappear, it can also be restored to the newly generated OSD configuration information after the disk replacement, ensuring the continuity and stability of the storage cluster business.

[0099] In some embodiments, during the execution of the reconstruction operation, the distributed storage cluster is monitored for alarm messages or blocked interfaces; when an alarm message or blocked interface exists in the distributed storage cluster, the reconstruction operation is suspended, and identification information of the alarm message or blocked interface is output; when it is detected that the alarm message is eliminated or the blocked interface returns to normal, the reconstruction operation is continued.

[0100] For example, after the storage cluster starts normal reconstruction, the storage reconstruction progress and cluster health status can be monitored in real time until the reconstruction is completed. If the storage status generates an alarm or I / O congestion during the reconstruction process, the reconstruction is suspended in time and the user is prompted to check. After the alarm is eliminated, the reconstruction operation is continued until the reconstruction is completed.

[0101] In the embodiment of the present application, by automatically detecting and processing misplaced data, reconstruction failure or performance degradation caused by improper data distribution in high watermark conditions is avoided, and recovery efficiency and reliability after disk replacement are improved.

[0102] See also Figure 2 , Figure 2 This is another flowchart of a distributed storage cluster disk replacement method provided in an embodiment of the present application. In some embodiments, the distributed storage cluster disk replacement method includes:

[0103] S201: In response to a disk swap instruction, monitor whether the distributed storage cluster is in a healthy state. If so, proceed to S202.

[0104] It is understandable that since this application is a tool for replacing disks in a distributed storage cluster under high watermark conditions, the script program can be deployed to the distributed storage node and the "chmod+x" command can be used to add executable permissions to the script program.

[0105] In some implementations, the user can trigger the script and enter the target storage pool name and pre-disk swap parameters. The storage cluster is then started to check the entered parameters to ensure their correctness. The OSD list corresponding to the target storage pool is then retrieved and checked for normal status. The OSD list is also monitored to see if the number of cluster I / O-blocked blocks exceeds a threshold. If the cluster is healthy and the number of I / O-blocked blocks is below a warning threshold, the storage cluster is considered healthy.

[0106] S202: Obtain object storage device configuration information of the target storage pool, and output disk replacement operation prompt information.

[0107] In some implementations, a directory for storing various configuration files and logs during the disk swap process can be automatically created based on the target storage pool name, facilitating recording and management of the disk swap process. The current Crushmap, osdmap, Upmap, and PG dump configuration information is then retrieved and saved with the current timestamp for subsequent data consistency checks and recovery operations. Finally, the user is prompted to begin the disk swap operation.

[0108] In some implementations, after manually replacing the physical hard disk according to the prompts, the user can enter the storage pool name and post-disk replacement operation parameters.

[0109] S203: When the disk swap operation is completed, determine whether there is misplaced data.

[0110] In some implementations, after the consistency check (peering) is completed, the storage cluster may perform data misalignment detection and scan misplaced objects. If there is no misplaced data, S204 is directly executed; otherwise, S206 is executed.

[0111] S204: Execute a preset reconstruction operation.

[0112] S205: Monitor the execution status of the reconstruction operation.

[0113] In some implementations, the storage cluster reconstruction progress and storage cluster health status may be monitored in real time until the reconstruction operation is completed.

[0114] After the storage cluster reconstruction operation is started, the reconstruction progress and cluster health status can be monitored in real time until the reconstruction is completed. If the storage status generates an alarm or I / O congestion during the reconstruction process, the reconstruction operation can be paused and the user will be prompted to check. After the alarm is eliminated, the reconstruction operation will continue until the reconstruction operation is completed.

[0115] S206: Import the acquired object storage device configuration information.

[0116] In some implementations, if misaligned data is detected in the storage cluster, the newly generated Crushmap, osdmap, and Upmap configuration information after the disk swap can be backed up first, and then the old object storage device configuration information saved by the information collection module before the disk swap can be automatically imported, and the consistency check can be completed again.

[0117] S207: Check whether the misaligned data disappears. If so, execute S204.

[0118] This application ensures that disk swaps are performed safely by monitoring cluster health and I / O congestion in real time. It also automatically detects and handles data misalignment, preventing reconstruction failures or performance degradation caused by improper data distribution. This significantly improves the stability and reliability of distributed storage clusters under high-water mark conditions.

[0119] In this embodiment of the application, the challenge of replacing hard drives in a distributed storage cluster during high-water-level conditions is addressed by obtaining key configuration information before the disk is replaced, thus ensuring subsequent recovery. After the disk is replaced, data misalignment issues are automatically detected and addressed to ensure data consistency. This embodiment of the application not only simplifies the operation and maintenance process and reduces the risk of manual intervention, but also ensures business continuity and stability through real-time monitoring and dynamic adjustment mechanisms, providing reliable technical support for disk replacement operations in distributed storage clusters and improving the availability and operation and maintenance efficiency of distributed storage clusters.

[0120] It should be noted that the method provided in the embodiments of the present application can be promoted and applied to any industry that relies on large-scale, high-performance storage infrastructure, especially in fields with high requirements for data continuity and security, such as cloud data centers, financial industry storage systems, and telecom operators' storage systems. It can ensure efficient maintenance of large-scale storage infrastructure, reduce service time, and improve service quality.

[0121] Through the description of the above implementation methods, technicians in this field can clearly understand that the distributed storage cluster disk replacement method described in the above embodiment can be implemented with the help of software plus the necessary general hardware platform. Of course, it can also be implemented through hardware, but in many cases the former is a better implementation method.

[0122] The present application also provides a distributed storage cluster disk replacement device. Figure 3 As shown, Figure 3 Schematic diagram of a distributed storage cluster disk replacement device provided in an embodiment of the present application. The distributed storage cluster disk replacement device 30 includes:

[0123] The monitoring module 301 is used to monitor whether the distributed storage cluster is in a healthy state in response to a disk swap instruction; the disk swap instruction includes the name of the target storage pool, and the storage water level of the target storage pool is greater than or equal to a preset water level threshold.

[0124] The acquisition module 302 is used to obtain the object storage device configuration information of the target storage pool and output disk replacement operation prompt information when the distributed storage cluster is in a healthy state; the target storage pool includes the hard disk to be replaced; the object storage device configuration information includes at least one of the following: topology information of the target storage pool, object storage device mapping information, placement group mapping rule information, and placement group distribution information.

[0125] The checking module 303 is configured to check whether there is misplaced data in the distributed storage cluster when it is detected that the disk swap operation is completed.

[0126] The recovery module 304 is used to perform a preset recovery operation based on the above-mentioned object storage device configuration information to eliminate the above-mentioned misplaced data if there is misplaced data; and to perform a preset reconstruction operation after eliminating the misplaced data, and the reconstruction operation includes writing the data stored in the hard disk to be replaced to the newly replaced hard disk.

[0127] In some embodiments, the monitoring module 301 is specifically configured to:

[0128] Obtain a list of object storage devices associated with the target storage pool, and detect whether the status of the object storage devices in the object storage device list is normal; when the status of the object storage devices in the object storage device list is normal, obtain the operating parameters of the distributed storage cluster; based on the operating parameters, determine whether the distributed storage cluster is in a healthy state.

[0129] In some embodiments, the operating parameters include state parameters and the number of blocked interfaces; the monitoring module 301 is specifically configured to:

[0130] When the state parameter includes a health state field and the number of blocked interfaces is less than the alarm number threshold, it is determined that the distributed storage cluster is in a healthy state.

[0131] In some implementations, the acquisition module 302 is further configured to:

[0132] Based on the identification information of the hard disks in the target storage pool, the topology information of the target storage pool is generated, and the topology information includes the hierarchical relationship of the hard disks in the target storage pool; based on the identification information of the object storage devices in the target storage pool, the object storage device mapping information of the target storage pool is generated; based on the object storage device mapping information, the placement group mapping rule information of the target storage pool is generated; based on the identification information of the placement groups in the target storage pool, the placement group distribution information of the target storage pool is generated.

[0133] The acquisition module 302 is specifically used for:

[0134] Based on the name of the target storage pool, a directory is created for storing configuration files during the disk swap process; object storage device configuration information of the target storage pool is obtained, and the obtained object storage device configuration information is saved to the above directory according to the current timestamp.

[0135] In some implementations, the recovery module 304 is specifically configured to:

[0136] Based on the object storage device configuration information, the configuration parameters of the target storage pool are updated; based on the updated configuration parameters of the target storage pool, the distribution location of the data stored in the distributed storage cluster is updated.

[0137] In some implementations, the acquisition module 302 is further configured to:

[0138] When misplaced data is detected in the distributed storage cluster, object storage device configuration information of the target storage pool after the disk swap operation is completed is obtained; and the object storage device configuration information of the target storage pool after the disk swap operation is completed is backed up.

[0139] In some embodiments, the inspection module 303 is configured to:

[0140] Detect whether the distributed storage cluster is in a stable state; when it is determined that the distributed storage cluster is in a stable state, perform a consistency check on the distributed storage cluster to determine whether there is misaligned data in the distributed storage cluster.

[0141] In some implementations, the recovery module 304 is further configured to:

[0142] If there is no misaligned data in the distributed storage cluster, a reconstruction operation is performed.

[0143] In some implementations, the recovery module 304 is further configured to:

[0144] During the reconstruction operation, monitor whether there are alarm messages or blocked interfaces in the distributed storage cluster; when there are alarm messages or blocked interfaces in the distributed storage cluster, suspend the reconstruction operation and output the identification information of the alarm message or blocked interface; when it is detected that the alarm message is eliminated or the blocked interface returns to normal, continue to perform the above reconstruction operation.

[0145] It is understandable that the description of the features in the embodiment corresponding to the above-mentioned distributed storage cluster disk replacement device can refer to the relevant description of the embodiment corresponding to the above-mentioned distributed storage cluster disk replacement method, and will not be repeated here.

[0146] The distributed storage cluster disk replacement device provided in the embodiment of the present application first monitors whether the distributed storage cluster is in a healthy state before performing the disk replacement operation. When the distributed storage cluster is in a healthy state, the object storage device configuration information of the target storage pool is obtained. When the disk replacement operation is completed, if there is misplaced data in the distributed storage cluster, a preset recovery operation can be performed based on the above-mentioned object storage device configuration information to automatically eliminate the misplaced data, thereby reducing the impact of the disk replacement operation on the performance of the distributed storage cluster, and completing the disk replacement of the distributed storage cluster efficiently, stably and safely.

[0147] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4As shown, the electronic device 40 provided in this embodiment includes: at least one processor 401 and a memory 402. Optionally, the electronic device 40 further includes a communication component 403. The processor 401, the memory 402 and the communication component 403 are connected via a bus.

[0148] During the specific implementation process, at least one processor 401 executes the computer-executable instructions stored in the memory 402 , so that the at least one processor 401 executes the above-mentioned embodiment of the method for disk replacement in a distributed storage cluster.

[0149] The specific implementation process of the processor 401 can refer to the above-mentioned embodiment of the distributed storage cluster disk replacement method. The implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0150] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), or application-specific integrated circuits (ASICs). The general-purpose processor may be a microprocessor or any conventional processor. The steps of the distributed storage cluster disk replacement method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0151] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0152] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0153] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned distributed storage cluster disk replacement method embodiments when running.

[0154] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0155] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned distributed storage cluster disk replacement method embodiments are implemented.

[0156] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0157] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the technical solution and core ideas of the present application. It should be pointed out that, for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for replacing a disk in a distributed storage cluster, characterized in that: The method comprises: In response to a disk swap instruction, monitoring whether the distributed storage cluster is in a healthy state; the disk swap instruction includes the name of a target storage pool, and the storage water level of the target storage pool is greater than or equal to a preset water level threshold; When the distributed storage cluster is in a healthy state, object storage device configuration information of the target storage pool is obtained, and disk replacement operation prompt information is output; the target storage pool includes a hard disk to be replaced, and the object storage device configuration information includes at least one of the following: topology information of the target storage pool, object storage device mapping information, placement group mapping rule information, and placement group distribution information; When it is detected that the disk swap operation is completed, checking whether there is misplaced data in the distributed storage cluster; If there is misplaced data, a preset recovery operation is performed based on the object storage device configuration information to eliminate the misplaced data; and after eliminating the misplaced data, a preset reconstruction operation is performed, and the reconstruction operation includes writing the data stored in the hard disk to be replaced to the newly replaced hard disk.

2. The method according to claim 1, characterized in that The monitoring of whether the distributed storage cluster is in a healthy state includes: Obtaining a list of object storage devices associated with the target storage pool, and detecting whether the status of the object storage devices in the list of object storage devices is normal; When the status of the object storage device in the object storage device list is normal, obtaining the operating parameters of the distributed storage cluster; Based on the operating parameters, it is determined whether the distributed storage cluster is in a healthy state.

3. The method according to claim 2, characterized in that The operating parameters include a state parameter and the number of blocked interfaces; and determining whether the distributed storage cluster is in a healthy state based on the operating parameters includes: When the status parameter includes a health status field and the number of blocked interfaces is less than an alarm number threshold, it is determined that the distributed storage cluster is in a healthy state.

4. The method according to claim 1, wherein The method further comprises: generating topology information of the target storage pool according to identification information of the hard disks in the target storage pool, wherein the topology information includes a hierarchical relationship of the hard disks in the target storage pool; generating object storage device mapping information of the target storage pool according to identification information of the object storage device in the target storage pool; Generating placement group mapping rule information of the target storage pool according to the object storage device mapping information; generating placement group distribution information of the target storage pool according to identification information of the placement groups in the target storage pool; The obtaining of the object storage device configuration information of the target storage pool includes: Based on the name of the target storage pool, a directory is established for storing configuration files during the disk swap process; Obtain object storage device configuration information of the target storage pool, and save the obtained object storage device configuration information to the directory according to the current timestamp.

5. The method according to claim 1, wherein The performing of a preset recovery operation based on the object storage device configuration information includes: Based on the object storage device configuration information, updating the configuration parameters of the target storage pool; Based on the updated configuration parameters of the target storage pool, the distribution location of the data stored in the distributed storage cluster is updated.

6. The method according to claim 1 or 5, characterized in that The method further comprises: When misplaced data is detected in the distributed storage cluster, obtaining object storage device configuration information of the target storage pool after the disk swap operation is completed; Backing up the object storage device configuration information of the target storage pool after the disk swap operation is completed.

7. The method according to claim 1, characterized in that The checking whether there is misplaced data in the distributed storage cluster includes: Detecting whether the distributed storage cluster is in a stable state; When it is determined that the distributed storage cluster is in a stable state, a consistency check is performed on the distributed storage cluster to determine whether misaligned data exists in the distributed storage cluster.

8. The method according to claim 1, characterized in that The method further comprises: If there is no misplaced data in the distributed storage cluster, the reconstruction operation is performed.

9. The method according to claim 1 or 8, characterized in that The method further comprises: During the process of performing the reconstruction operation, monitoring the distributed storage cluster for alarm messages or blocked interfaces; When there is an alarm message or a blocked interface in the distributed storage cluster, suspending the reconstruction operation and outputting identification information of the alarm message or the blocked interface; When it is detected that the alarm message is eliminated or the blocked interface returns to normal, the reconstruction operation is continued.

10. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the distributed storage cluster disk replacement method according to any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Disk changing method and device for distributed storage system and medium

    CN116431067A

  • Distributed cluster data reconstruction method and device, and program product

    CN118312103A

  • Data reconstruction method and device, electronic equipment and computer readable storage medium

    CN119597548A

  • Method and apparatus for reconstructing data in object-based storage arrays

    US20060156059A1