A rapid reconstruction method for a storage system based on fine-grained hard disk partitioning
Through fine-grained hard disk partitioning technology, real-time storage system load information is obtained and hard disk partitions are scanned concurrently. The redundant algorithm is used to quickly reconstruct corrupt data, solving the problem of SSD hard disk reconstruction time being too long, and achieving second-level reconstruction and high reliability.
Patent Information
- Application Number
- CN202510600754.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In the prior art, the reconstruction time of SSD hard disks after failure is too long, resulting in a high risk of data loss. Especially in large-capacity hard disk scenarios, the reconstruction time increases by 8-17 times, affecting the normal reading and writing service of the host.
Using a method based on fine-grained hard disk partitioning, all flash physical blocks with the same block ID in the Target of the SSD hard disk are aggregated into one hard disk partition, and the system load information is obtained in real time, concurrently scan multiple hard disk partitions, and reconstruct data on the faulty partition through redundant algorithms to reduce the reconstruction granularity to avoid full disk reconstruction.
The reconstruction time in seconds is realized, which reduces the probability of data loss, improves the utilization rate of hard disk, shortens the reconstruction time, and improves the system reliability and overall performance of hard disk.
Smart Images

Figure CN120123154B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of server heat dissipation, and particularly relates to a method for quickly reconstructing a storage system based on fine-grained hard disk partitioning. Background Art
[0002] With the development of artificial intelligence technology, data centers have gradually changed from previous cost centers to power centers for the development of data economy and application innovation. The amount of data that the storage system needs to process is increasing, and users' requirements for the reliability and performance of the storage system are also getting higher. The capacity of SSD (Solid State Disk) is larger than that of traditional HDD (Hard Disk Drive), and the proportion of the capacity of SSD hard disks in the global market has been increasing rapidly.
[0003] The emergence of flash memory technology has brought about a huge revolution to the storage system, with significant improvements in aspects such as the performance, power consumption, and storage density of the storage system. With the development of flash memory particle technology, the storage density of SSD hard disks has been greatly improved. Currently, the single-disk capacity that can be obtained has reached 61TB. According to the planning of hard disk manufacturers, the single-disk capacity may reach 300TB or even higher within the next two years.
[0004] For SSD hard disks based on the existing block device interface, all flash memory particles of the hard disk are combined into a unified linear space for the application system to use. If a flash memory particle is damaged, the data of the bad block is restored through the bad block replacement method, and the bad block is isolated. In this way, if the bad blocks reach the OP (Over-Provision, hard disk reserved space) ratio of the hard disk, the entire hard disk fails. In addition, the internal hard disk uses the RAID (Redundant Array of Independent Disks) method to ensure data reliability. However, in order to obtain a larger available space, a large proportion of RAID stripe methods are usually adopted. A single RAID stripe will be distributed to all Dies (the wafers in the flash memory chip, which is the smallest unit that independently executes commands and returns status) in the SSD hard disk. If the entire Die goes offline, the RAID will be in a degraded state. If there are more than the redundant Dies offline, the entire hard disk will fail, triggering a full disk reconstruction.
[0005] Meanwhile, as the single-disk capacity of SSD hard drives increases and the storage density improves, the probability of flash memory particle failures is increasing. For example, the endurance of QLC (Quad-Level Cell, four-level storage cell) particles is about 75% of that of TLC (Trinary-Level Cell, three-level storage cell) particles. Correspondingly, the reliability of SSD hard drives based on QLC particles is also lower than that of SSD hard drives based on TLC particles. Comparing the hard drive specifications of the same manufacturer, the DWPD (Drive Writes Per Day) specification of SSD hard drives based on TLC particles is mostly greater than 1, while the DWPD of SSD hard drives based on QLC particles is mostly less than 0.5.
[0006] Existing distributed storage systems generally adopt replication or EC (Erasure Code) implementation methods. When a hard drive fails, hard drive reconstruction will be performed. To reduce the probability of dual-disk failure in a storage pool, it is required that the storage system complete the reconstruction as soon as possible. However, to reduce the performance impact of reconstruction on the normal read and write operations of the host, the storage system needs to balance the reconstruction operation and the normal read and write operations of the host, resulting in a reduction in reconstruction performance. Taking the current mainstream 7.68TB SSD hard drive as an example, according to a reconstruction performance of 100MB / s, it takes 22 hours to complete the reconstruction; if the hard drive capacity reaches 61TB, it will take 177.7 hours, which is 8 times longer; if a larger 128TB hard drive is used, the reconstruction will take 372.8 hours, which becomes 17 times the current reconstruction time. During the reconstruction of a failed SSD hard drive, if another SSD hard drive also fails, it may lead to data loss. The longer the reconstruction time, the greater the risk of data loss. Therefore, in practical applications, too long reconstruction time is the main bottleneck in the use of large-capacity hard drives.
[0007] Therefore, how to improve the hard drive reconstruction process in the existing technology to enhance the hard drive reconstruction performance and reduce the hard drive reconstruction duration is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0008] The purpose of the present invention is to provide a fast reconstruction method for a storage system based on fine-grained hard drive partitioning to improve the hard drive reconstruction process in the existing technology, enhance the hard drive reconstruction performance, and reduce the hard drive reconstruction duration.
[0009] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0010] A fast reconstruction method for a storage system based on fine-grained hard drive partitioning includes the following steps:
[0011] S1: Aggregate the flash physical blocks Block with the same block ID on all Dies within an independent namespace Target of the SSD hard disk into a fine-grained hard disk partition, obtain the load information of the storage system in real time, and calculate the concurrency value of the background scan;
[0012] S2: Obtain the data cold and hot order of the hard disk partition, and combine the scan concurrency value to obtain the scan range of each hard disk;
[0013] S3: Issue scan requests for specified partitions to specified hard disks concurrently, and perform concurrent scans on multiple hard disks and their multiple partitions;
[0014] S4: Judge whether a partition failure is scanned based on the scan result. If so, execute step S5;
[0015] S5: When a hard disk partition failure is scanned, obtain the failure source. The failure source includes a partition failure returned by a service request, query the key value distribution stored on the failed partition, read the data at the corresponding position of the hard disk, and calculate the data of the failed partition through a redundancy algorithm;
[0016] S6: On the hard disk where the failed partition is located, allocate a new partition for reconstruction;
[0017] S7: Re-save the calculated data to other specified partitions of this hard disk.
[0018] Preferably, the specific process of step S1 is as follows:
[0019] S11: Real-time collect the load information of the storage system through the data acquisition module, including the CPU and hard disk where the service is located;
[0020] S12: Calculate the background scan concurrency value. The specific calculation formula is as follows:
[0021] Scan concurrency value = MAX(maximum concurrency value * (1 – avg(∑ service core CPU utilization rate)) * (1 – avg(∑ service hard disk utilization rate)), minimum concurrency value);
[0022] S13: According to the real-time load information of the storage system, judge the change trend of the real-time load of the storage system. When the real-time load of the storage system increases and the utilization rates of the corresponding CPU and hard disk increase, then adjust the concurrency value down by a specified ratio; when the real-time load of the storage system decreases and the utilization rates of the corresponding CPU and hard disk decrease, then adjust the concurrency value up by a specified ratio.
[0023] Preferably, the adjustment range of the concurrency value in step S13 is [minimum concurrency value, maximum concurrency value]. By dynamically adjusting the concurrency value between [minimum concurrency value, maximum concurrency value], a balance is achieved between the scan speed and the impact on the system service.
[0024] Preferably, the specific process of step S3 is as follows:
[0025] S31: By counting the number of requests falling within the partition range during the previous round of scanning, identify the order of data from cold to hot. The larger the number of requests, the higher the partition heat value, and the higher the probability that the hotter data is found to be faulty by business requests;
[0026] S31: According to the order of heat values from cold to hot, and through the concurrency value of the background scan in step S2, prepare several requests to scan the SSD hard disk partitions;
[0027] S32: Clear the heat value and recalculate it in each round of scanning, and update the partition cold and hot values in real time;
[0028] Among them, one round of scanning means scanning all the data on the SSD hard disk once.
[0029] Preferably, in step S5, when a fault in an SSD hard disk partition is scanned, calculate the number of partition faults, compare the calculation result of the number of partition faults with a preset partition fault number threshold. When the calculation result of the number of partition faults is greater than the preset partition fault number threshold, achieve the balance of capacity and life based on a preset weighted jump balancing algorithm.
[0030] Preferably, the specific process of achieving the balance of capacity and life based on the preset weighted jump balancing algorithm is as follows:
[0031] S51: Reduce the weight of the hard disk with too many faulty partitions through the weight setting module;
[0032] S52: Reduce the probability of subsequent new data being stored on the hard disk.
[0033] Preferably, the SSD hard disk includes multiple Packages. A Package is the encapsulation of a NAND chip. Each Package contains multiple independent namespaces Targets. Each independent namespace Target has an independent storage unit, data bus, and chip select signal Chip select. Each of the independent namespace Targets is connected to the Channel of the SSD hard disk through a bus. The Targets connected to the same Channel are selected through Chip enable. Inside a Target, there are multiple Dies. A Die is the smallest unit that can execute commands and return status independently. A Die contains multiple Planes. Each Plane has its own independent Page register and Data Register. A Plane contains multiple Blocks. A Block is the smallest unit for data erasure, that is, the smallest unit for garbage collection. Data can only be written sequentially within a Block.
[0034] Preferably, the fine-grained hard disk partitioning in step S1 is only established on the Blocks with the same block ID on the Dies within one Target, without RAID redundancy.
[0035] The beneficial effects of the present invention include:
[0036] The storage system fast reconstruction method based on fine-grained hard disk partitioning provided by the present invention aggregates the Blocks with the same block ID on all Dies within the Target of the SSD hard disk into a fine-grained hard disk partition; collects the load information of the storage system and calculates the background scan concurrency value; obtains the hot and cold order of the hard disk partitions, and combines the scan concurrency value to obtain the scan range of each hard disk; concurrently issues scan requests for specified partitions to specified hard disks. When a hard disk partition failure is detected, obtain the source of the failure; query the key value distribution stored on the failed partition, read the data at the corresponding position of the hard disk, and calculate the data of the failed partition through a redundancy algorithm; on the hard disk where the failed partition is located, allocate a new partition for reconstruction and save the calculated data to other specified partitions on this hard disk. The fine-grained partition is established on the Blocks with the same block ID on the Dies within one Target. The fine-grained SSD hard disk partition is used as the unit for failure and reconstruction to reconstruct in a timely manner, accurately reconstruct damaged data, improve the hard disk reconstruction performance, reduce the hard disk reconstruction duration, and combine the service load and data hot and cold to scan the bad blocks of multiple hard disks with multiple concurrencies.
[0037] First, by collecting the load information of the real-time acquisition storage system in real time, calculate the concurrency value of the background scan; obtain the hot and cold order of the hard disk partitions, and combine the scan concurrency value to obtain the scan range of each hard disk; issue scan requests for specified partitions to specified hard disks concurrently, and determine whether a partition failure is scanned. When a hard disk partition failure is scanned, obtain the source of the failure. The sources of failure include partition failures returned by service requests; query the key value distribution stored on the failed partition, read the data at the corresponding position of the hard disk, and calculate the data of the failed partition through a redundancy algorithm; on the hard disk where the failed partition is located, allocate a new partition for reconstruction, and re-save the calculated data to other specified partitions on this hard disk. In the data reconstruction process, if a hard disk partition failure is found, the reconstruction control module immediately reconstructs the data of the failed partition. Since the partition granularity is small, the reconstruction time can reach the second level, and the probability of data loss becomes correspondingly smaller. And it will not cause the entire disk to fail because the failure ratio of the flash memory particles exceeds the OP, making the probability of the entire disk failing almost reduced to 0, while improving the utilization rate of the hard disk. It effectively speeds up the data reconstruction speed of large-capacity SSD hard disks in case of failure, further enhances the overall reliability of the system, and realizes the rapid reconstruction of large-capacity hard disks.
[0038] Secondly, when a SSD hard disk partition failure is scanned, calculate the number of partition failures, and compare the calculation result of the number of partition failures with a preset partition failure number threshold. When the calculation result of the number of partition failures is greater than the preset partition failure number threshold, based on the preset weighted jump balancing algorithm, the balance of capacity and life is realized.
[0039] Thirdly, the present invention eliminates the hard disk OP logic, thereby eliminating the concept of disk failure / full disk reconstruction, and completely solves the full disk reconstruction problem. When a bad block occurs, it does not affect the normal use of the remaining hard disk flash memory particles until all hard disk partitions of the hard disk fail, improving the utilization rate of the hard disk particles. At the same time, the performance impact of the reconstruction process on the host read and write operations is also very small; the reconstruction duration of a single partition is at the second level, and the reconstruction duration does not change linearly with the linear increase of the hard disk capacity, but is only proportional to the size of the damaged data capacity, greatly shortening the reconstruction duration. In the scenario of large-capacity disks, compared with the reconstruction time of the scenario where this solution is not implemented, the difference is obvious, improving the reliability of the hard disk, and being able to more quickly and intelligently detect the bad blocks of the hard disk. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the layout diagram of the storage system based on fine-grained hard disk partitions of the present invention.
[0041] Figure 2 It is the schematic flow diagram of the rapid reconstruction of the storage system based on fine-grained hard disk partitions of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The following combines the attachedFigures 1 to 2 The present invention will be further described in detail as follows:
[0043] Embodiment 1
[0044] See the appendix Figure 2 As shown, a method for rapid reconstruction of a storage system based on fine-grained hard disk partitioning includes the following steps:
[0045] S1: Aggregate the flash physical blocks Block with the same block ID on all Dies in an independent namespace Target of the SSD hard disk into a fine-grained hard disk partition, obtain the load information of the storage system in real time, and calculate the concurrency value of the background scan;
[0046] S2: Obtain the cold and hot order of the hard disk partition data, and combine the scan concurrency value to obtain the scan range of each hard disk;
[0047] S3: Send scan requests for specified partitions to specified hard disks concurrently, and perform concurrent scans on multiple hard disks and their multiple partitions;
[0048] S4: Judge whether a partition failure is scanned based on the scan result. If so, execute step S5;
[0049] S5: When a hard disk partition failure is scanned, obtain the failure source. The failure source includes a partition failure returned by a service request, query the key value distribution stored on the failed partition, read the data at the corresponding position of the hard disk, and calculate the data of the failed partition through a redundancy algorithm;
[0050] S6: On the hard disk where the failed partition is located, allocate a new partition for reconstruction;
[0051] S7: Re-save the calculated data to other specified partitions of this hard disk.
[0052] In the distributed storage systems in the prior art, replicas or EC (Erasure Code) implementation methods are generally adopted to ensure data reliability. After a hard disk fails, hard disk reconstruction will be carried out. In order to reduce the probability of double-disk failure in a storage pool, it is required that the storage system complete the reconstruction in the shortest possible time. However, in order to reduce the performance impact of reconstruction on the normal read and write operations of the host, the storage system needs to balance the reconstruction service and the normal read and write operations of the host, resulting in a reduction in the performance of reconstruction. In order to improve the reconstruction performance of the hard disk, through disk control cooperation, when the FTL in the SSD hard disk or the NAND (NOT AND, circuit design of logic gates) particles fail, the hard disk reports the affected area to the upper-layer storage system, and the upper-layer software reconstructs the affected data. There is a reliability drawback: if multiple Dies fail simultaneously or the entire flash memory particle fails, it will cause the SSD single disk to fail and trigger a full-disk reconstruction; the full-disk reconstruction time is long, and there is a risk of data reliability to a certain extent.
[0053] Therefore, the present invention reduces the reconstruction duration by shrinking the reconstruction failure domain to establish a smaller reconstruction granularity. Through a fine-grained hard disk partition design: all Blocks (flash memory physical blocks, which are the smallest units of data erasure) with the same block ID on all Dies within a Target (independent namespace) of the SSD hard disk are aggregated into a hard disk partition, thereby removing the in-disk RAID redundancy and achieving the purpose of a smaller failure domain. The fine-grained hard disk partition is only established on the Blocks with the same block ID on the Dies within a Target, without the need for RAID redundancy, avoiding cross-Die partition establishment. In this way, the failure domain of the hard disk is reduced from the entire disk to the Blocks within the Die, and the reconstruction scope also changes from full-disk reconstruction to reconstruction at the granularity of the Blocks within the Die.
[0054] By collecting the load information of the storage system, calculate the background scan concurrency value; obtain the hot and cold order of the hard disk partitions, and combine the scan concurrency value to obtain the scan range of each hard disk; concurrently issue scan requests for specified partitions to specified hard disks to obtain hard disk partition failures and the sources of failures; query the key value distribution stored on the faulty partition, read the data at the corresponding position of the hard disk, and calculate the data of the faulty partition through a redundancy algorithm; on the hard disk where the faulty partition is located, allocate a new partition for reconstruction, and save the calculated data to other specified partitions on this hard disk. The fine-grained partition is based on the Block with the same block ID on the Die within a Target. The fine-grained SSD hard disk partition is used as the unit of failure and reconstruction to reconstruct in a timely manner, accurately reconstruct damaged data, improve the hard disk reconstruction performance, reduce the hard disk reconstruction duration, combine the service load and data hot and cold, and perform a fast reconstruction process for the occurrence of hard disk bad blocks during multi-hard disk and multi-concurrent scanning. Due to the small partition granularity, the reconstruction time can reach the second level, and the probability of data loss becomes correspondingly smaller. In addition, it will not cause the entire disk to fail because the failure ratio of the flash memory particles exceeds the OP, which almost reduces the probability of the entire disk failure to 0, and at the same time improves the utilization rate of the hard disk. This solution effectively speeds up the data reconstruction speed of large-capacity SSD hard disks in case of failure, further enhances the overall reliability of the system, and realizes the fast reconstruction of large-capacity hard disks.
[0055] Embodiment 2
[0056] Based on Embodiment 1, the specific process of step S1 is as follows:
[0057] S11: Real-time collect the load information of the storage system through the data collection module, including the CPU and hard disks where the services are located;
[0058] S12: Calculate the background scan concurrency value, and the specific calculation formula is as follows:
[0059] Scan concurrency value = MAX(maximum concurrency value * (1 - avg(∑ service core CPU utilization rate)) * (1 - avg(∑ service hard disk utilization rate)), minimum concurrency value);
[0060] S13: According to the real-time load information of the storage system, judge the change trend of the real-time load of the storage system. When the real-time load of the storage system increases and the utilization rates of the corresponding CPU and hard disks increase, then adjust the concurrency value down by a specified ratio; when the real-time load of the storage system decreases and the utilization rates of the corresponding CPU and hard disks decrease, then adjust the concurrency value up by a specified ratio.
[0061] In this embodiment, the adjustment range of the concurrency value in step S13 is [minimum concurrency value, maximum concurrency value]. By dynamically adjusting the concurrency value between [minimum concurrency value, maximum concurrency value], a balance is achieved between the scanning speed and the impact on the system services.
[0062] Embodiment 3
[0063] Based on Embodiment 1 or Embodiment 2, the specific process of step S3 is as follows:
[0064] S31: By counting the number of requests falling within the partition range during the previous round of scanning, identify the order of data from cold to hot. The larger the number of requests, the higher the partition heat value, and the higher the probability that the hotter data is discovered to have a fault by service requests.
[0065] S31: In the order of heat value from cold to hot, and through the concurrency value of the background scanning in step S2, prepare several requests for scanning the SSD hard disk partitions.
[0066] S32: Clear the heat value and recalculate it for each round of scanning, and update the partition cold and hot values in real time.
[0067] Among them, one round of scanning refers to scanning all the data on the SSD hard disk once.
[0068] Embodiment 4
[0069] Based on Embodiment 1 or Embodiment 2 or Embodiment 3, in step S5, when a fault in an SSD hard disk partition is scanned, calculate the number of partition faults, compare the calculation result of the number of partition faults with a preset partition fault number threshold. When the calculation result of the number of partition faults is greater than the preset partition fault number threshold, achieve the balance of capacity and life based on a preset weighted jump balancing algorithm.
[0070] In this embodiment, the specific process of achieving the balance of capacity and life based on the preset weighted jump balancing algorithm is as follows:
[0071] S51: Reduce the weight of the hard disk with too many faulty partitions through the weight setting module;
[0072] S52: Reduce the probability of subsequent new data being stored in the hard disk.
[0073] See Appendix Figure 1As shown in the figure, the SSD hard disk includes multiple Packages. A Package is the encapsulation of a NAND chip. Each Package contains multiple independent namespaces Targets. Each independent namespace Target has an independent storage unit, data bus, and chip select signal Chip select. Each of the independent namespace Targets is connected to the Channel of the SSD hard disk through a bus. The Targets connected to the same Channel are selected through Chipenable. Within a Target, there are multiple Dies. A Die is the smallest unit that can execute commands and return status independently. A Die contains multiple Planes. Each Plane has its own independent Page register and DataRegister. A Plane contains multiple Blocks. A Block is the smallest unit of data erasure, that is, the smallest unit of garbage collection. Data can only be written sequentially within a Block. The fine-grained hard disk partitioning in step S1 is only established on the Blocks with the same block ID on the Dies within one Target, without RAID redundancy.
[0074] The fine-grained hard disk partitioning is only established on the Blocks with the same block ID on the Dies within one Target, without the need for RAID redundancy, avoiding cross-Die partitioning. In this way, the failure domain of the hard disk is reduced from the entire disk to the Blocks within the Die, and the reconstruction scope is also transformed from full-disk reconstruction to reconstruction at the Block granularity within the Die.
[0075] After a hard disk fails, the shorter the time from the failure to the start of reconstruction, the shorter the reconstruction duration can be indirectly shortened. Therefore, it is necessary to detect bad blocks on the hard disk faster and more timely. There are two ways to detect bad blocks. One is when the storage system reports a failure when reading and writing a hard disk partition, and reports the bad block to the storage system. The other is active detection. The storage system initiates different concurrent bad block scans according to its own business load. Such multi-disk multi-concurrent bad block scans at the hard disk partition granularity can greatly improve the scanning speed. At the same time, the concurrency of the scan is dynamically adjusted according to different pressure business loads. When the system is busy, it scans with low concurrency, and when it is idle, it scans with high concurrency. While quickly detecting bad blocks, the impact of the background scan task on the system business is minimized. The storage system maintains the hot and cold order of the data on the disk at the hard disk partition granularity, and the scanning is executed in the order of partitions from hot to cold.
[0076] In summary, the rapid reconstruction method of the storage system based on fine-grained hard disk partitioning provided by the present invention aggregates Blocks with the same block ID on all Dies within the Target of the SSD hard disk into a fine-grained hard disk partition; collects the load information of the storage system, calculates the background scan concurrency value; obtains the cold and hot order of the hard disk partitions, and combines the scan concurrency value to obtain the scan range of each hard disk; concurrently issues scan requests for specified partitions to specified hard disks, and when a hard disk partition failure is detected, obtains the source of the failure; queries the key value distribution stored on the failed partition, reads the data at the corresponding position of the hard disk, and calculates the data of the failed partition through a redundancy algorithm; on the hard disk where the failed partition is located, allocates a new partition for reconstruction, and saves the calculated data to other specified partitions on this hard disk. The fine-grained partition is based on Blocks with the same block ID on Dies within a Target. The fine-grained SSD hard disk partition is used as the unit of failure and reconstruction to perform timely reconstruction, accurately reconstruct damaged data, improve the hard disk reconstruction performance, reduce the hard disk reconstruction duration, and combine the service load and data cold and hot to perform multi-disk and multi-concurrency scanning of hard disk bad blocks.
[0077] By collecting the load information of the storage system in real time, calculating the concurrency value of the background scan; obtaining the cold and hot order of the hard disk partitions, and combining the scan concurrency value to obtain the scan range of each hard disk; concurrently issuing scan requests for specified partitions to specified hard disks, and determining whether a partition failure is detected. When a hard disk partition failure is detected, obtain the source of the failure, and the source of the failure includes a partition failure returned by a service request; query the key value distribution stored on the failed partition, read the data at the corresponding position of the hard disk, and calculate the data of the failed partition through a redundancy algorithm; on the hard disk where the failed partition is located, allocate a new partition for reconstruction, and save the calculated data to other specified partitions on this hard disk for the data reconstruction process. If a hard disk partition failure is found, the reconstruction control module immediately reconstructs the data of the failed partition. Since the partition granularity is small, the reconstruction time can reach the second level, and the probability of data loss becomes correspondingly smaller. And it will not cause the entire disk to fail because the failure ratio of the flash memory particles exceeds the OP, making the probability of the entire disk failure almost reduced to 0, while improving the utilization rate of the hard disk. It effectively speeds up the data reconstruction speed of large-capacity SSD hard disks in case of failure, further enhances the overall reliability of the system, and realizes the rapid reconstruction of large-capacity hard disks.
[0078] When a fault in the SSD hard disk partition is scanned, the number of partition faults is calculated. The calculation result of the number of partition faults is compared with the preset partition fault number threshold. When the calculation result of the number of partition faults is greater than the preset partition fault number threshold, based on the preset weighted jump equalization algorithm, the balance of capacity and lifespan is achieved. The hard disk OP logic is eliminated, thus eliminating the concept of disk failure / full disk reconstruction and completely solving the full disk reconstruction problem. When a bad block occurs, it does not affect the normal use of the remaining hard disk flash particles until all hard disk partitions of the hard disk fail, improving the utilization rate of hard disk particles. At the same time, the performance impact of the reconstruction process on the host read and write operations is also very small. The reconstruction duration of a single partition is at the second level. The reconstruction duration does not change linearly with the linear increase of the hard disk capacity, but is only proportional to the size of the damaged data capacity, greatly shortening the reconstruction duration. In the scenario of large-capacity disks, compared with the reconstruction time of the scenario where this solution is not implemented, the difference is obvious, improving the reliability of the hard disk and being able to more quickly and intelligently detect the bad blocks of the hard disk.
Claims
1. A method for rapid reconstruction of a storage system based on fine-grained hard disk partitioning, characterized in that, It includes the following steps: S1: Aggregate the flash physical blocks (Blocks) with the same block ID on all Dies within an independent namespace Target of the SSD hard disk into a fine-grained hard disk partition, and obtain the load information of the storage system in real time to calculate the concurrency value of the background scan; S2: Obtain the data cold-hot order of the hard disk partition, and combine the scan concurrency value to obtain the scan range of each hard disk; S3: Send scan requests for specified partitions to specified hard disks concurrently to perform concurrent scanning of multiple hard disks and their multiple partitions; S4: Based on the scan results, determine whether a partition failure is scanned. If so, execute step S5; S5: When a hard disk partition failure is scanned, obtain the failure source. The failure source includes a partition failure returned by a service request. Query the key value distribution stored on the failed partition, read the data at the corresponding position of the hard disk, and calculate the data of the failed partition through a redundancy algorithm; S6: On the hard disk where the failed partition is located, allocate a new partition for reconstruction; S7: Re-save the calculated data to other specified partitions of this hard disk; The specific process of step S1 is as follows: S11: Collect the load information of the storage system in real time through a data acquisition module, including the CPU and hard disk where the service is located; S12: Calculate the background scan concurrency value. The specific calculation formula is as follows: Scan concurrency value = MAX(maximum concurrency value * (1 – avg(∑ service core CPU utilization rate)) * (1 – avg(∑ service hard disk utilization rate)), minimum concurrency value); S13: According to the real-time load information of the storage system, judge the change trend of the real-time load of the storage system. When the real-time load of the storage system increases and the utilization rates of the corresponding CPU and hard disk increase, then reduce the concurrency value by a specified proportion; when the real-time load of the storage system decreases and the utilization rates of the corresponding CPU and hard disk decrease, then increase the concurrency value by a specified proportion.
2. A method for quickly reconstructing a storage system based on fine-grained hard disk partitioning according to claim 1, characterized in that The adjustment range of the concurrency value in step S13 is [minimum concurrency value, maximum concurrency value]. By dynamically adjusting the concurrency value between [minimum concurrency value, maximum concurrency value], a balance is achieved between the scan speed and the impact on the system service.
3. A method for rapid reconstruction of a storage system based on fine-grained hard disk partitioning according to claim 1, characterized in that, The specific process of step S3 is as follows: S31: Identify the data cold-to-hot order by counting the number of requests that fell within the partition range during the previous scan. The more requests, the higher the partition heat value, and the higher the probability that the hotter data is found to be faulty by service requests; S31: Prepare several requests to scan the SSD hard disk partitions in the order of heat value from cold to hot and through the background scan concurrency value in step S2; S32: Clear the heat value and recalculate it for each round of scan, and update the partition cold-hot value in real time; Among them, one round of scan refers to scanning all the data of the SSD hard disk once.
4. A rapid reconstruction method for a storage system based on fine-grained hard disk partitioning according to claim 1, characterized in that, In step S5, when a failure of an SSD hard disk partition is scanned, calculate the number of partition failures, and compare the calculation result of the number of partition failures with a preset partition failure number threshold. When the calculation result of the number of partition failures is greater than the preset partition failure number threshold, achieve the balance of capacity and life based on a preset weighted jump-based balancing algorithm; The specific process of realizing the balance between capacity and life based on the preset weighted jump-based balancing algorithm is as follows: S51: Reduce the weight of the partitioned hard disk with a fault through the weight setting module; S52: Reduce the probability of subsequent new data being stored in the hard disk.
5. A method for rapid reconstruction of a storage system based on fine-grained hard disk partitioning according to claim 1, characterized in that The SSD hard disk includes multiple Packages. A Package is the encapsulation of a NAND chip. A Package contains multiple independent namespaces Targets. Each independent namespace Target has an independent storage unit, data bus, and chip select signal Chip select. Each of the independent namespace Targets is connected to the Channel of the SSD hard disk through a bus. The Targets connected to the same Channel are selected through chip enable Chip enable. Within one Target, there are multiple Dies. A Die is the smallest unit that can execute commands and return status independently. A Die contains multiple Planes. Each Plane has its own independent Page register and Data Register. A Plane contains multiple Blocks. A Block is the smallest unit of data erasure, that is, the smallest unit of garbage collection. Data can only be written in sequence within a Block.
6. A method for quickly reconstructing a storage system based on fine-grained hard disk partitioning according to claim 5, characterized in that, The fine-grained hard disk partitioning in step S1 is only established on the Blocks with the same block ID on the Dies within one Target, without RAID redundancy.
Citation Information
Patent Citations
Method and system for quickly storing big data
CN118093584A
Enhanced eMMC / SSD data integrity protection framework
CN119760793A