Storage system data management method, device and equipment and readable storage medium
By configuring spare hard drives in a distributed storage system, responding to hard drive failures and dynamically adjusting redundancy, the problems of reduced redundancy and storage space utilization caused by hard drive failures are solved, achieving high reliability and efficient utilization of data storage, and improving system performance and security.
Patent Information
- Application Number
- CN202510694368.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-16
AI Technical Summary
In distributed storage systems, hard drive failures lead to reduced redundancy and increase the risk of data loss. At the same time, high redundancy configurations reduce storage space utilization, making it difficult to maximize storage space utilization while ensuring data security.
By configuring a spare hard drive, data is written to the spare hard drive in response to a hard drive failure, and the data is written back to the repair hard drive after the failure is recovered, dynamically adjusting the redundancy configuration to maintain system redundancy and storage space efficiency.
It effectively solves the problem of reduced redundancy caused by hard disk failure, achieves high reliability and efficient utilization of data storage, reduces enterprise storage costs, and improves storage system performance and security.
Smart Images

Figure CN120653195A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of communication technology, and in particular to a storage system data management method, device, equipment, and readable storage medium. Background Art
[0002] Among current distributed storage technologies, Ceph is widely used for its high scalability and performance. Ceph uses erasure coding (EC) as a data protection method, which offers higher storage efficiency than traditional replication. By splitting the original data into k data blocks and generating m parity blocks, erasure coding ensures data security while reducing required storage space.
[0003] When hard drives in a cluster fail due to wear and tear or other reasons, especially when the number of failures reaches or exceeds the set number of parity blocks, m, the redundancy of the entire system decreases significantly, increasing the risk of data loss. For example, in a typical 4+2 configuration (i.e., k = 4, m = 2), if two hard drives fail simultaneously, the system will be unable to recover the original data because there is insufficient parity information to reconstruct the lost data. This situation is unacceptable for enterprise-level applications that require high reliability, especially in large-scale data center environments where hard drive failures are a common problem.
[0004] Furthermore, even under normal operating conditions, users often increase the number of parity blocks to improve system fault tolerance in the event of a hard drive failure. However, this approach directly reduces storage space utilization. Higher redundancy means more storage resources are used to store parity information rather than actual data, which contradicts the goal of efficient storage space utilization. Therefore, how to maximize storage space utilization while ensuring data security has become a pressing issue. Summary of the Invention
[0005] In view of this, the present disclosure provides a storage system data management method, apparatus, device, and readable storage medium to improve the problem of reduced redundancy caused by hard disk failure.
[0006] The specific technical solutions are as follows:
[0007] This specification provides a storage system data management method, which is applied to a management device of a storage system, wherein the storage system includes several hard disks, and the method includes: obtaining the redundancy of the storage system, wherein the redundancy is associated with the number of data hard disks and the number of verification hard disks currently configured in the storage system, and configuring several hard disks as spare hard disks of the storage system according to the redundancy; in response to an event that a data hard disk and / or a verification hard disk fails, calling a corresponding number of spare hard disks to store data that should be written to the failed data hard disk and / or the failed verification hard disk after the failure occurs; in response to an event that a failed data hard disk and / or a verification hard disk recovers from a failure, reading the data stored during the failure from the corresponding spare hard disks and writing the data hard disk and / or the verification hard disk that have recovered from the failure.
[0008] As a technical solution, configuring a number of hard disks as spare hard disks of the storage system according to redundancy includes: the number of configured spare hard disks is greater than or equal to 1 and less than or equal to the number of check hard disks.
[0009] As a technical solution, the event of responding to the failure recovery of a failed data hard disk and / or a verification hard disk includes: responding to the event of the failed data hard disk and / or the verification hard disk being replaced and put back online; or responding to the event of the failed data hard disk and / or the verification hard disk being reloaded and recovered.
[0010] This specification also provides a storage system data management device, which is applied to a management device of a storage system, wherein the storage system includes several hard disks, and the method includes: a first module, which is used to obtain the redundancy of the storage system, wherein the redundancy is associated with the number of data hard disks and the number of verification hard disks currently configured in the storage system, and the several hard disks are configured as spare hard disks of the storage system according to the redundancy; a second module, which is used to call a corresponding number of spare hard disks to store data that should be written to the failed data hard disk and / or the failed verification hard disk after the failure occurs, in response to the event that the failed data hard disk and / or the verification hard disk recovers from the failure, and a third module, which is used to read the data stored during the failure from the corresponding spare hard disk and write it to the recovered data hard disk and / or verification hard disk.
[0011] As a technical solution, configuring a number of hard disks as spare hard disks of the storage system according to redundancy includes: the number of configured spare hard disks is greater than or equal to 1 and less than or equal to the number of check hard disks.
[0012] As a technical solution, the event of responding to the failure recovery of a failed data hard disk and / or a verification hard disk includes: responding to the event of the failed data hard disk and / or the verification hard disk being replaced and put back online; or responding to the event of the failed data hard disk and / or the verification hard disk being reloaded and recovered.
[0013] This specification also provides a storage system, characterized in that it includes: a data hard disk for storing business data blocks; a check hard disk for storing check data blocks; a spare hard disk for storing data that should be written to the faulty data hard disk and / or the faulty check hard disk when the data hard disk and / or the check hard disk fails, and for providing the data stored during the failure period to the recovered data hard disk and / or the check hard disk after the fault of the faulty data hard disk and / or the check hard disk is recovered; the number of the spare hard disks is configured according to redundancy, and the redundancy is associated with the number of data hard disks and the number of check hard disks.
[0014] As a technical solution, the number of configured spare hard disks is greater than or equal to 1 and less than or equal to the number of verification hard disks.
[0015] This specification also provides an electronic device, including a processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the aforementioned storage system data management method.
[0016] This specification also provides a readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the aforementioned storage system data management method.
[0017] The above technical solutions provided in this specification bring at least the following beneficial effects:
[0018] In the event of a hard drive failure, data is stored on a backup hard drive, and after recovery, the data is written back to the repaired hard drive. This effectively addresses the existing issues of reduced redundancy and data loss caused by hard drive failures, achieving a balance between high data storage reliability and efficient storage space utilization, reducing enterprise storage costs and improving overall storage system performance and data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the implementation methods of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the implementation methods of this specification or the description of the prior art. Obviously, the drawings described below are only some implementation methods recorded in this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the implementation methods of this specification.
[0020] Figure 1 is a flow chart of a storage system data management method in one embodiment of this specification;
[0021] Figure 2This is a schematic diagram of a normal state in one embodiment of this specification;
[0022] Figure 3 is a schematic diagram of a fault state in one embodiment of this specification;
[0023] Figure 4 is a schematic diagram of a fault recovery state in one embodiment of this specification;
[0024] Figure 5 is a structural diagram of a storage system data management device in one embodiment of this specification;
[0025] Figure 6 This is a hardware structure diagram of an electronic device in one embodiment of this specification.
[0026] Reference numerals: first module 21 , second module 22 , third module 23 . DETAILED DESCRIPTION
[0027] The terms used in the embodiments of this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a," "the," and "the" used in this specification and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to any or all possible combinations of one or more of the associated listed items.
[0028] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when...", "when...", or "in response to determining."
[0029] In view of this, this specification provides a storage system data management method, device, equipment and readable storage medium to improve the above technical problems.
[0030] The specific technical solution is described below.
[0031] In one embodiment, the present specification provides a storage system data management method, which is applied to a management device of a storage system, wherein the storage system includes several hard disks, and the method includes: obtaining the redundancy of the storage system, wherein the redundancy is associated with the number of data hard disks and the number of verification hard disks currently configured in the storage system, and configuring several hard disks as spare hard disks of the storage system according to the redundancy; in response to an event in which a data hard disk and / or a verification hard disk fails, calling a corresponding number of spare hard disks to store data that should be written to the failed data hard disk and / or the failed verification hard disk after the failure occurs; in response to an event in which a failed data hard disk and / or a verification hard disk recovers from a failure, reading the data stored during the failure from the corresponding spare hard disks and writing the data hard disk and / or the verification hard disk that have recovered from the failure.
[0032] Specifically, if Figure 1 , including the following steps:
[0033] Step S11 , obtaining the redundancy of the storage system, wherein the redundancy is associated with the number of data hard disks and the number of check hard disks currently configured in the storage system, and configuring a number of hard disks as spare hard disks of the storage system according to the redundancy.
[0034] The management device needs to obtain the redundancy of the storage system. Redundancy is an important indicator of the storage system's data redundancy level. It is closely related to the number of data drives and parity drives currently configured in the storage system. Specifically, data drives store actual data information, while parity drives store parity information generated using specific algorithms (such as erasure coding) to enable data recovery in the event of a data drive failure. By understanding redundancy, the management device can gain a clear understanding of the current status of the storage system, providing a basis for subsequent spare drive configuration.
[0035] Based on the achieved redundancy, the management device configures several hard drives as backup drives for the storage system. These backup drives are normally in standby mode and do not participate in regular data storage operations. Their primary function is to quickly deploy them in the event of a failure in another drive, ensuring the normal operation of the storage system and data integrity. The number and method of configuring backup drives should be considered comprehensively, including the scale of the storage system, the importance of the data, and the expected failure tolerance.
[0036] Step S12, in response to a data hard disk and / or a check hard disk failure, calling a corresponding number of spare hard disks to store data that should be written to the failed data hard disk and / or the failed check hard disk after the failure occurs.
[0037] When a data hard drive and / or parity hard drive in a storage system fails, the management device calls on a corresponding number of spare hard drives to store the data that should have been written to the failed data hard drive and / or parity hard drive on these spare hard drives instead. This process requires an efficient fault detection mechanism and a fast spare hard drive activation process to minimize the interruption time of data writing. For example, in a storage system with a 4+2 erasure code configuration, there are normally 4 data hard drives and 2 parity hard drives. If one of the data hard drives fails, the management device will immediately transfer the data write operation mounted on the failed hard drive to a spare hard drive to ensure continuous data storage and system stability.
[0038] Step S13, in response to the event that the failed data hard disk and / or the check hard disk recovers, read the data stored during the failure from the corresponding spare hard disk and write it to the recovered data hard disk and / or the check hard disk.
[0039] After a drive failure is recovered, the data stored during the failure is read from the corresponding spare drive and written to the recovered data drive and / or parity drive. This step is crucial for restoring the storage system's original configuration and ensuring data consistency. For example, when the failed data drive is repaired or replaced and reinserted into the storage system, the management device reads the data stored during the failure from the spare drive and writes it back to the repaired data drive. This restores the storage system's original data layout and redundancy configuration, making it ready for subsequent failures.
[0040] Consider a storage system using an erasure code configuration with k = 4 and m = 2, meaning four data drives and two parity drives. With this configuration, the storage system has a certain level of redundancy and can tolerate a certain number of drive failures. However, to further enhance data security, the management device configures two spare drives based on redundancy.
[0041] During normal operation, data sent by upper-layer services is split into four data blocks and two parity blocks are generated. These blocks and parity blocks are then distributed and stored on four data hard drives and two parity hard drives. At this time, the spare hard drives are in standby mode and do not participate in data storage.
[0042] When a data drive in a storage system fails, a backup drive is activated and the data originally written to the failed drive is stored on it instead. This way, even in the event of a failure, the storage system can still maintain its redundancy because the backup drive takes over the data storage task of the failed drive, avoiding the loss of redundancy due to the failure. This allows the system to adopt a lower redundancy level during initial design, thus saving costs.
[0043] After the failed drive is repaired or replaced and restored, the management device initiates the data recovery process. It reads the data stored during the failure from the spare drive and writes it back to the restored drive. Once this process is complete, the storage system returns to its original 4+2 erasure coding configuration and continues to provide stable and reliable data storage services for upper-layer services.
[0044] If two parity drives in a storage system fail simultaneously, the management device will deploy a corresponding number of spare drives to replace the failed parity drives based on pre-configured policies. Similarly, when the failed drives recover, data is transferred from the spare drives to the restored parity drives, ensuring storage system redundancy and data security.
[0045] In one embodiment, the storage system management device collects storage pool topology information in real time through a redundancy monitoring module deployed on the control node. This module analyzes the distribution of currently active Object Storage Daemon (OSD) nodes based on the CRUSH algorithm mapping table of the Ceph storage cluster and dynamically calculates the effective redundancy parameter R = current number of parity hard disks / (number of data hard disks + number of parity hard disks). Taking a typical 4+2 erasure code configuration as an example, the initial data hard disks D1-D4 correspond to OSD1-OSD4, and the parity hard disks P1-P2 correspond to OSD5-OSD6. At this time, R = 2 / (4+2) = 33.3%. The management device dynamically allocates the number of spare hard disks according to the preset redundancy threshold range: when R ≥ 25%, the spare resources are configured at 50% of the number of parity hard disks. In this case, OSD7-OSD8 are automatically selected as hot spare disks, forming a 4+2+2 physical storage structure. The selection strategy for spare hard drives prioritizes devices that are physically heterogeneous with the existing data / check hard drives. For example, when the primary storage node is deployed on servers S1-S3 in cabinet A, spare hard drives will be allocated to servers S4-S6 in cabinet B to ensure rack-level fault tolerance.
[0046] The storage system management device achieves millisecond-level fault detection through its heartbeat detection subsystem. When an I / O timeout occurs on data drive D3 (OSD3), the fault handling engine executes a multi-stage response mechanism. First, during the primary failure window of 0-15 seconds, the data routing module temporarily caches data blocks that should have been written to D3 in a memory buffer and simultaneously broadcasts the failure event to the cluster. Starting at the 16th second, the backup scheduler activates hot spare drive OSD7, dynamically remaps its logical identifier to D3', and uses CRUSH algorithm weight adjustment to migrate the hash shard range of the original D3 to D3'. At this point, data stripes issued by upper-layer applications are still sharded according to k=4 and m=2, with the third data block being written to D3' instead of the failed D3. This process achieves seamless failover by modifying the object's PG (Placement Group) mapping. Specifically, the physical location of the original PG1.3 corresponding to OSD3 is updated to OSD7, but the PG number remains unchanged, ensuring client-unawareness. At this stage, the system redundancy remains unchanged at R=2 / (4+2)=33.3% because the failed data disk has been completely replaced by the hot spare disk.
[0047] When the failure duration exceeds the 300-second culling threshold set by the cluster, the management device initiates a deep repair protocol. At this point, OSD3 is permanently removed from the storage pool, and D3' (OSD7) officially takes over the storage responsibilities of the original D3. The metadata server (MDS) updates the global namespace, marking D3 as failed and adding D3' to the set of valid storage nodes. It is worth noting that the logic for generating the parity disk is synchronously adjusted during this process: the new parity block calculation will be based on the data shards of D1, D2, D3', and D4, while the parity relationship of the historical data is maintained consistent through background asynchronous recalculation. For example, for the object_A data block stored in D3 before the failure, the reconstruction engine will read the data from D1, D2, D4, and P1-P2, and regenerate the correct data on D3' using the Reed-Solomon decoding matrix. This process preferentially uses temporary data on the hot spare disk as input to reduce I / O pressure.
[0048] After hardware maintenance personnel replaced the failed drive D3, the management device detected the connection of the new drive, D3_new (OSD9), through the hot-swap detection module, triggering the data migration process. At this point, the incremental synchronizer compared the PG distribution maps of D3' (OSD7) and D3_new to identify the object range to be migrated. For data written during the failure window, OSD7's physical sector data was directly cloned to OSD9 via block device-level replication, bypassing the file system layer to improve transfer efficiency. For a 10TB data volume, using RDMA network-accelerated block replication can reduce migration time from 8 hours required by traditional rebuilds to 1.5 hours. After synchronization is complete, the CRUSH mapping table reassociates D3_new with the original PG1.3 and gradually reduces the weight of OSD7 until it returns to hot standby status. Data consistency during this phase is ensured by a two-phase commit protocol: first, write operations to D3' are frozen. After OSD9 completes full data synchronization, the PG mapping relationship is atomically switched, and finally, OSD7 is released back to the standby pool.
[0049] When multiple hard drives experience cascading failures, the system ensures the availability of core data through dynamic priority scheduling. Suppose that while the D3 failure remains unresolved, a media error also occurs on the parity drive P1 (OSD5). At this point, the redundancy monitoring module detects that R = 1 / (4 + 2) = 16.7%, below the safety threshold, and immediately triggers the secondary backup mechanism. The management device selects OSD8 from the backup pool to take over P1's parity role, while OSD9, originally reserved for data drive backup, is temporarily transferred to the parity backup role. The system now forms a 4+2+(2-1) resilient architecture, with OSD7 replacing the data drive and OSD8 replacing the parity drive, while maintaining a total redundancy of R = (2-1+1) / (4+2) = 33.3%. This dynamic role switching capability allows the system to maintain its initial design redundancy level in dual-failure scenarios, whereas traditional solutions would reduce redundancy to zero in this scenario.
[0050] The data write-back optimization module reduces the I / O bottleneck during recovery through an intelligent prefetching strategy. When D3_new completes the physical replacement, the system not only migrates the data written during the failure, but also preloads the frequently accessed data based on the access popularity prediction model. For example, for video surveillance storage scenarios, the system analyzes the temporal locality characteristics of the last 72 hours and prioritizes migrating video blocks written between 20:00 and 22:00. This strategy increases the recovery priority of frequently used data, and actual measurements show that it can shorten the business-perceived recovery time by 40%. At the same time, the background verifier performs full data verification during low-load periods, and uses a differential CRC algorithm to quickly locate silent errors, ensuring that the data consistency error rate between the hot spare disk and the primary disk is less than 10^-15.
[0051] In ultra-large-scale clusters, management devices use a hierarchical backup strategy to achieve resource optimization. For the three cross-regional data centers A, B, and C, a local backup hard disk pool (L1 backup) is deployed in each site, and a global backup pool (L2 backup) across sites is established. When the local backup resources at site A are exhausted, the failover module gives priority to borrowing backup resources from the L1 pool at site B instead of directly using L2 resources, thereby reducing cross-domain bandwidth consumption. This hierarchical model reduced the failover delay from 230ms to 95ms in a 1024-node cluster test, while reducing cross-region data traffic by 45%. The elastic scaling of backup resources is achieved through the Kubernetes containerization platform. When it is detected that the cluster scale has expanded by 20%, the automatic expansion module will proportionally increase the backup hard disk container instance to keep the proportion of backup resources always within the preset range.
[0052] This solution can still maintain data integrity in extreme scenarios. When an 8+4 erasure code configuration is adopted and three spare hard disks are set up, the simultaneous failure of five hard disks (including three data disks and two verification disks) is simulated. The system replaces the failed data disks with three spare disks through dynamic reorganization, and converts one spare disk to the verification role, ultimately maintaining an 8+4 equivalent structure. At this time, although the initial design redundancy is exceeded, the system can still provide degraded read services until new hard disks are added by prioritizing the integrity of the data disk. This elasticity capability that exceeds the theoretical limits of traditional EC reduces the RTO (recovery time objective) in PB-level data loss scenarios from the industry average of 48 hours to 9 hours.
[0053] In one embodiment, configuring a number of hard disks as spare hard disks of the storage system according to redundancy includes: the number of configured spare hard disks is greater than or equal to 1 and less than or equal to the number of check hard disks.
[0054] In one embodiment, the event of responding to the failure recovery of a failed data hard disk and / or verification hard disk includes: responding to the event of a failed data hard disk and / or verification hard disk being replaced and brought back online; or responding to the event of a failed data hard disk and / or verification hard disk being reloaded and recovered.
[0055] In one embodiment, by adding an additional designated spare storage module, the problem of the cluster having no redundancy and being prone to data loss when the number of failed disks with erasure codes reaches a predetermined number of check blocks is solved; by adding an additional designated spare storage module, the redundancy of k+m can still be maintained, so that when subsequent maintenance engineers replace hardware, the cluster has less CPU computing time and more time to copy the data of the spare storage module to the newly replaced hard disk; when the cluster has an additional designated spare storage module, when deploying the cluster, in scenarios where higher data security is considered, the EC mode with higher storage utilization can be selected, and the same data security protection can be achieved.
[0056] like Figure 2 During normal operation, the data from the upper layer business will be divided into k data blocks, and m check data blocks will be generated to form k+m stripes and stored down to k+m physical storage media, such as hard disks.
[0057] like Figure 3 When a hard disk failure occurs in the Ceph cluster, for example, a hard disk failure (it can be a failure of a destination hard disk among k data blocks, or a failure of a destination hard disk for a verification data block), the additional designated spare storage module will replace the original failed hard disk to ensure the k+m erasure code storage mechanism works.
[0058] When a ceph cluster hard disk failure timeout period, the ceph cluster will remove the hard disk that cannot carry storage work from the cluster at the ceph software code level. After removal, the cluster will reselect the latest hard disk for storage of k+m stripe data blocks, such as Figure 4 , then the data of the spare storage module will be copied to the new replacement storage device.
[0059] The number of spare storage modules in the cluster is set to be less than or equal to the number of check data blocks, and users actively choose to use them.
[0060] In the above implementation method, during a cluster hard disk failure, when the data blocks and check blocks refreshed to the hard disk device by the upper-layer business cannot be globally written to the disk, they can be written to the backup storage module; during the initial failure of the cluster hard disk, there is also m redundancy fault tolerance; after the cluster failure is recovered, the backup storage module can copy the temporarily stored data to the new mapping device, and the backup device remains on standby as a backup device.
[0061] In one embodiment, Figure 5This specification also provides a storage system data management device, which is applied to a management device of a storage system, wherein the storage system includes several hard disks, and the method includes: a first module, which is used to obtain the redundancy of the storage system, wherein the redundancy is associated with the number of data hard disks and the number of verification hard disks currently configured in the storage system, and the several hard disks are configured as spare hard disks of the storage system according to the redundancy; a second module, which is used to call a corresponding number of spare hard disks to store data that should be written to the failed data hard disk and / or the failed verification hard disk after the failure occurs, in response to the event that the failed data hard disk and / or the verification hard disk recovers from the failure, and a third module, which is used to read the data stored during the failure from the corresponding spare hard disk and write it to the recovered data hard disk and / or verification hard disk.
[0062] In one embodiment, configuring a number of hard disks as spare hard disks of the storage system according to redundancy includes: the number of configured spare hard disks is greater than or equal to 1 and less than or equal to the number of check hard disks.
[0063] In one embodiment, the event of responding to the failure recovery of a failed data hard disk and / or verification hard disk includes: responding to the event of a failed data hard disk and / or verification hard disk being replaced and brought back online; or responding to the event of a failed data hard disk and / or verification hard disk being reloaded and recovered.
[0064] The device implementation is the same as or similar to the corresponding method implementation, and will not be repeated here.
[0065] In one embodiment, the present specification also provides a storage system, characterized in that it includes: a data hard disk for storing business data blocks; a check hard disk for storing check data blocks; a spare hard disk for storing data that should be written to the faulty data hard disk and / or the faulty check hard disk when the data hard disk and / or the check hard disk fails, and for providing the data stored during the failure to the recovered data hard disk and / or the check hard disk after the fault of the faulty data hard disk and / or the check hard disk is recovered; the number of the spare hard disks is configured according to redundancy, and the redundancy is associated with the number of data hard disks and the number of check hard disks.
[0066] In this storage system, a 4+2 erasure code configuration uses four data blocks and two parity blocks. A certain number of spare drives, typically no more than the number of parity drives, are also configured to ensure rapid response in the event of a failure. Under normal operating conditions, data sent by upper-layer services is divided into k data blocks, and m parity blocks are generated. These blocks are then combined into k+m stripes and stored on the corresponding physical media.
[0067] When a data drive or parity drive failure is detected, such as when a data drive fails due to long-term wear and tear, the system will automatically call on a spare drive to replace the failed drive. During this process, new data blocks and parity blocks will continue to be calculated and stored according to the original erasure coding strategy, but some data will be temporarily stored on the spare drive. This ensures that even during a drive failure, the entire cluster maintains its original k+m redundancy, effectively avoiding the risk of data loss due to reduced redundancy. For example, in a typical 4+2+2 configuration (i.e., k=4, m=2, spare drives=2), if two hard drives fail simultaneously, the system can still maintain the complete erasure code stripe structure by using the spare drive, thereby ensuring data integrity and recoverability.
[0068] When a failed hard drive rejoins the cluster after repair or replacement, the system initiates a data recovery process. During this time, temporary data originally stored on the backup drive will be copied back to the newly repaired or replaced drive, significantly less than the amount of data required for a full copy. This design not only simplifies the hardware replacement process and reduces the operational complexity for maintenance engineers, but also significantly shortens the data recovery window. For example, in the 4+2+2 configuration example above, if one of the data drives fails and is replaced some time later, the system will automatically read the relevant data from the backup drive and write it back to the new drive after confirming that the new drive is fully ready. The backup drive will then return to standby mode, ready to respond to the next possible failure.
[0069] In one embodiment, the number of configured spare hard disks is greater than or equal to 1 and less than or equal to the number of verification hard disks.
[0070] In one embodiment, this specification provides an electronic device, including a processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the aforementioned storage system data management method. From a hardware perspective, the hardware architecture diagram can be found in Figure 6 shown.
[0071] In one embodiment, this specification provides a readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the aforementioned storage system data management method.
[0072] Here, the readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage media, or a combination thereof.
[0073] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.
[0074] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0075] Those skilled in the art will appreciate that embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of a fully hardware implementation, a fully software implementation, or an implementation combining software and hardware. Furthermore, embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0076] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0077] Furthermore, these computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0079] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0080] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A storage system data management method, characterized in that: A management device applied to a storage system, wherein the storage system includes a plurality of hard disks, and the method includes: Obtaining redundancy of the storage system, where the redundancy is related to the number of data hard disks and the number of parity hard disks currently configured in the storage system, and configuring a number of hard disks as spare hard disks of the storage system based on the redundancy; In response to a failure of a data hard disk and / or a check hard disk, calling a corresponding number of spare hard disks to store data that should be written to the failed data hard disk and / or the failed check hard disk after the failure; In response to an event in which a failed data hard disk and / or a parity hard disk recovers, data stored during the failure is read from the corresponding spare hard disk and written to the recovered data hard disk and / or parity hard disk.
2. The method according to claim 1, characterized in that The step of configuring a plurality of hard disks as spare hard disks of the storage system according to redundancy includes: The number of configured spare hard disks must be greater than or equal to 1 and less than or equal to the number of verification hard disks.
3. The method according to claim 1, characterized in that The event of responding to a failure of a data hard disk and / or a verification hard disk and recovering from the failure includes: In response to an event in which a failed data hard drive and / or parity hard drive is replaced and brought back online; Or, in response to an event where a failed data drive and / or parity drive is reloaded and recovered.
4. A storage system data management device, characterized in that: A management device applied to a storage system, wherein the storage system includes a plurality of hard disks, and the device includes: The first module is configured to obtain redundancy of the storage system, where the redundancy is related to the number of data hard disks and the number of check hard disks currently configured in the storage system, and configure a number of hard disks as spare hard disks of the storage system according to the redundancy; The second module is configured to, in response to a failure of a data hard disk and / or a check hard disk, call a corresponding number of spare hard disks to store data that should be written to the failed data hard disk and / or the failed check hard disk after the failure; The third module is used to respond to the event of failure recovery of the failed data hard disk and / or the verification hard disk, read the data stored during the failure from the corresponding spare hard disk and write it to the recovered data hard disk and / or the verification hard disk.
5. The device according to claim 4, characterized in that The step of configuring a plurality of hard disks as spare hard disks of the storage system according to redundancy includes: The number of configured spare hard disks must be greater than or equal to 1 and less than or equal to the number of verification hard disks.
6. The device according to claim 4, characterized in that The event of responding to a failure of a data hard disk and / or a verification hard disk and recovering from the failure includes: In response to an event in which a failed data hard drive and / or parity hard drive is replaced and brought back online; Or, in response to an event where a failed data drive and / or parity drive is reloaded and recovered.
7. A storage system, characterized in that: include: Data hard disk, used to store business data blocks; A check hard disk is used to store check data blocks; The spare hard disk is used to store data that should be written to the failed data hard disk and / or the failed check hard disk when the data hard disk and / or the check hard disk fails, and to provide the data stored during the failure period to the recovered data hard disk and / or the check hard disk after the failed data hard disk and / or the check hard disk recovers; The number of the spare hard disks is configured according to redundancy, which is associated with the number of data hard disks and the number of check hard disks.
8. The storage system according to claim 7, wherein: The number of configured spare hard disks must be greater than or equal to 1 and less than or equal to the number of verification hard disks.
9. An electronic device, characterized in that: include: A processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1 to 3.
10. A readable storage medium, characterized in that: The readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the method according to any one of claims 1 to 3.