A storage management method, device, equipment and machine readable storage medium

CN116820337BActive Publication Date: 2026-08-07XINHUASAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XINHUASAN INFORMATION TECH CO LTD
Filing Date
2023-06-28
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]有鉴于此,本公开提供一种存储管理方法、装置及电子设备、机器可读存储介质,以改善上述因修复故障产生大量数据迁移的问题

Benefits of technology

[0018]为缓存盘设定寿命阈值,在其运行寿命数据达到预设阈值时将其所在的存储节点标记为降级状态,限制数据写入避免其使用寿命耗尽损坏的同时阻止其所在存储节点向其他存储节点迁移数据,然后在用于更替的新磁盘加载后直接将缓存的数据迁移至新的磁盘,并相应重新配置缓存盘,从而避免了缓存盘故障造成的大量数据迁移。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116820337B_ABST
    Figure CN116820337B_ABST
Patent Text Reader

Abstract

The present disclosure provides a storage management method, device, equipment and machine readable storage medium, the method comprising: monitoring the cache disk running life data of the storage node, marking the to-be-replaced disk, and configuring the storage node where the to-be-replaced disk is located as a degraded state; migrating the data cached in the to-be-replaced disk to a new disk; configuring the new disk as the cache disk of the storage node where the to-be-replaced disk is located, and deleting the configuration information of the to-be-replaced disk configured as the cache disk. Through the technical solution of the present disclosure, a life threshold is set for the cache disk, when the running life data thereof reaches the preset threshold, the storage node where the cache disk is located is marked as a degraded state, data writing is limited to avoid the use life of the cache disk from being exhausted and damaged, and the storage node where the cache disk is located is prevented from migrating data to other storage nodes, then the cached data is directly migrated to the new disk after the new disk for replacement is loaded, and the cache disk is reconfigured accordingly, thereby avoiding a large amount of data migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of communication technology, and in particular to a storage management method, apparatus, device and machine-readable storage medium. Background Technology

[0002] With the development of technologies such as high-definition video, image processing, and video surveillance, user data volume is increasing rapidly, and users' demands for the read and write performance of stored data are also rising simultaneously. When purchasing storage products, users typically specify performance requirements to meet business needs. In many business scenarios, storage performance is not just a matter of speed; it can even affect whether the business itself can operate normally. For example, in document digitization, archives and libraries digitize paper books and store them on storage servers. As more and more digitized books are stored, the capacity of the storage cluster will increase significantly. Simultaneously, a large number of users will access the storage cluster concurrently. If the read and write performance of the storage cluster is poor and the read and write latency is high, it will reduce user experience and hinder the progress of digitization. Therefore, improving the read and write performance of the storage cluster is crucial.

[0003] With the development of digitalization, distributed storage is being used in more and more fields, and the amount of data on storage systems is also increasing. To improve the read and write performance of distributed storage, a caching disk acceleration solution is usually adopted, that is, using NVMe disks or SSD disks to speed up HDD disks. The typical acceleration solution for NVMe disks (or SSD disks) is to divide an NVMe disk into multiple partitions, with each partition used as a cache for an HDD disk, thereby achieving the acceleration effect.

[0004] Distributed storage typically consists of a cluster of multiple hosts (starting from 3 nodes), with each host using NVMe to accelerate HDD. Because NVMe disks have a limited write lifespan—meaning there's a limit to the total amount of data that can be written to an NVMe disk—once this limit is reached, the disk may become unwritable or even lose data. When an NVMe cache disk reaches 100% lifespan, causing a failure, the host detects the failure and takes the cache disk offline. This also takes the data disk it accelerates offline, requiring data reconstruction on other nodes. When a new cache disk replaces the old one, the data is rebuilt on the new cache disk and its storage nodes using data from other nodes. This results in two full data migrations: data migration out during disk failure and data migration back after recovery, leading to wasted performance resources. Summary of the Invention

[0005] In view of this, the present disclosure provides a storage management method, apparatus, electronic device, and machine-readable storage medium to improve the problem of large-scale data migration caused by fault repair.

[0006] The specific technical solution is as follows:

[0007] This disclosure provides a storage management method applied to a storage cluster, the storage cluster including several storage nodes. The method includes: monitoring the lifetime data of the cache disks of the storage nodes; if the lifetime data of the cache disks of a storage node reaches a preset threshold, marking the cache disk as a disk to be replaced, and configuring the storage node where the disk to be replaced is located in a degraded state, the degraded state including restricting data writing and data migration; in response to an event that a new disk for replacing the disk to be replaced has been loaded on the storage node where the disk to be replaced is located, migrating the cached data in the disk to be replaced to the new disk; configuring the new disk as the cache disk of the storage node where the disk to be replaced is located, and deleting the configuration information of the disk to be replaced that was configured as a cache disk.

[0008] As a technical solution, configuring the new disk as the cache disk of the storage node where the disk to be replaced is located, and deleting the configuration information of the disk to be replaced as a cache disk, includes: in response to the event that the disk to be replaced has been unloaded, canceling the degraded state of the storage node that was configured as a degraded state.

[0009] As a technical solution, the monitoring of the cache disk lifespan data of the storage node, if the cache disk lifespan data of a storage node reaches a preset threshold, marks the cache disk as a disk to be replaced and configures the storage node where the disk to be replaced is located in a degraded state, including: generating alarm information, the alarm information including a replacement prompt for the disk to be replaced and the degraded state information of the storage node where the disk to be replaced is located.

[0010] As a technical solution, the operational lifespan data includes power-on duration and / or data write volume.

[0011] This disclosure also provides a storage management device applied to a storage cluster, the storage cluster including several storage nodes. The device includes: a first module for monitoring the lifetime data of the cache disks of the storage nodes; if the lifetime data of the cache disks of a storage node reaches a preset threshold, the cache disk is marked as a disk to be replaced, and the storage node where the disk to be replaced is located is configured in a degraded state, the degraded state including restricting data writing and data migration; a second module for migrating the cached data in the disk to be replaced to the new disk in response to an event that the new disk used to replace the disk to be replaced has been loaded on the storage node where the disk to be replaced is located; and a third module for configuring the new disk as the cache disk of the storage node where the disk to be replaced is located, and deleting the configuration information of the disk to be replaced that is configured as a cache disk.

[0012] As a technical solution, configuring the new disk as the cache disk of the storage node where the disk to be replaced is located, and deleting the configuration information of the disk to be replaced as a cache disk, includes: in response to the event that the disk to be replaced has been unloaded, canceling the degraded state of the storage node that was configured as a degraded state.

[0013] As a technical solution, the monitoring of the cache disk lifespan data of the storage node, if the cache disk lifespan data of a storage node reaches a preset threshold, marks the cache disk as a disk to be replaced and configures the storage node where the disk to be replaced is located in a degraded state, including: generating alarm information, the alarm information including a replacement prompt for the disk to be replaced and the degraded state information of the storage node where the disk to be replaced is located.

[0014] As a technical solution, the operational lifespan data includes power-on duration and / or data write volume.

[0015] This disclosure also provides an electronic device, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the aforementioned storage management method.

[0016] This disclosure also provides a machine-readable storage medium storing machine-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the aforementioned storage management method.

[0017] The technical solution provided in this disclosure brings at least the following beneficial effects:

[0018] A lifetime threshold is set for the cache disk. When the data of its operation reaches the preset threshold, the storage node where it is located is marked as degraded. Data writing is restricted to prevent it from being damaged due to the end of its lifespan, while preventing the storage node where it is located from migrating data to other storage nodes. Then, after the new disk for replacement is loaded, the cached data is directly migrated to the new disk, and the cache disk is reconfigured accordingly, thereby avoiding a large amount of data migration caused by cache disk failure. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments of this disclosure or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this disclosure.

[0020] Figure 1This is a flowchart of a storage management method according to one embodiment of the present disclosure;

[0021] Figure 2 This is a structural diagram of a storage management device according to one embodiment of the present disclosure;

[0022] Figure 3 This is a hardware structure diagram of an electronic device according to one embodiment of the present disclosure.

[0023] Reference numerals: Module 1 21, Module 22, Module 3 23. Detailed Implementation

[0024] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.

[0025] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."

[0026] This disclosure provides a storage management method, apparatus, electronic device, and machine-readable storage medium to at least improve one of the aforementioned technical problems.

[0027] The specific technical solution is described below.

[0028] In one embodiment, this disclosure provides a storage management method applied to a storage cluster, the storage cluster including several storage nodes. The method includes: monitoring the lifetime data of the cache disks of the storage nodes; if the lifetime data of the cache disks of a storage node reaches a preset threshold, marking the cache disk as a disk to be replaced, and configuring the storage node where the disk to be replaced is located in a degraded state, the degraded state including restricting data writing and data migration; in response to an event that a new disk for replacing the disk to be replaced has been loaded on the storage node where the disk to be replaced is located, migrating the cached data in the disk to be replaced to the new disk; configuring the new disk as the cache disk of the storage node where the disk to be replaced is located, and deleting the configuration information of the disk to be replaced being configured as a cache disk.

[0029] Specifically, such as Figure 1 This includes the following steps:

[0030] Step S11: Monitor the cache disk lifespan data of the storage node. If the cache disk lifespan data of a storage node reaches a preset threshold, mark the cache disk as a disk to be replaced and configure the storage node where the disk to be replaced is located in a degraded state.

[0031] Step S12: In response to the event that the new disk used to replace the disk to be replaced has been loaded on the storage node where the disk to be replaced is located, the cached data in the disk to be replaced is migrated to the new disk.

[0032] Step S13: Configure the new disk as the cache disk of the storage node where the disk to be replaced is located, and delete the configuration information of the disk to be replaced that was configured as a cache disk.

[0033] A lifetime threshold is set for the cache disk. When the data of its operation reaches the preset threshold, the storage node where it is located is marked as degraded. Data writing is restricted to prevent it from being damaged due to the end of its lifespan, while preventing the storage node where it is located from migrating data to other storage nodes. Then, after the new disk for replacement is loaded, the cached data is directly migrated to the new disk, and the cache disk is reconfigured accordingly, thereby avoiding a large amount of data migration caused by cache disk failure.

[0034] In one implementation, configuring the new disk as the cache disk of the storage node where the disk to be replaced is located, and deleting the configuration information of the disk to be replaced as a cache disk, includes: in response to the event that the disk to be replaced has been unloaded, canceling the degraded state of the storage node that was configured as a degraded state.

[0035] In one implementation, the monitoring of the cache disk lifetime data of the storage node, if the cache disk lifetime data of a storage node reaches a preset threshold, marks the cache disk as a disk to be replaced and configures the storage node where the disk to be replaced is located in a degraded state, including: generating alarm information, the alarm information including a replacement prompt for the disk to be replaced and degraded state information of the storage node where the disk to be replaced is located.

[0036] In one implementation, the operational lifespan data includes power-on duration and / or data write volume.

[0037] In one implementation, the storage cluster sets a cache disk (NVMe) lifetime alarm threshold (90%). When the NVMe disk lifetime reaches 90%, an alarm is reported. After receiving the NVMe cache disk lifetime alarm, the operation and maintenance personnel replace the cache disk (NVMe) with a new one. When replacing the cache disk, no data recovery is performed. The storage node is in a degraded and pending recovery state. The data of the original cache disk is read and written to the new cache disk (using dd or system call copying is acceptable). After the cache disk data copying is completed, the faulty disk is replaced and the degraded state is canceled.

[0038] In one embodiment, this disclosure also provides a storage management device applied to a storage cluster, the storage cluster including a plurality of storage nodes. The device includes: a first module, used to monitor the lifetime data of the cache disks of the storage nodes; if the lifetime data of the cache disks of a storage node reaches a preset threshold, the cache disk is marked as a disk to be replaced, and the storage node where the disk to be replaced is located is configured in a degraded state, the degraded state including restricting data writing and data migration; a second module, used to migrate the cached data in the disk to be replaced to the new disk in response to an event that a new disk for replacing the disk to be replaced has been loaded on the storage node where the disk to be replaced is located; and a third module, used to configure the new disk as the cache disk of the storage node where the disk to be replaced is located, and delete the configuration information of the disk to be replaced being configured as a cache disk.

[0039] In one implementation, configuring the new disk as the cache disk of the storage node where the disk to be replaced is located, and deleting the configuration information of the disk to be replaced as a cache disk, includes: in response to the event that the disk to be replaced has been unloaded, canceling the degraded state of the storage node that was configured as a degraded state.

[0040] In one implementation, the monitoring of the cache disk lifetime data of the storage node, if the cache disk lifetime data of a storage node reaches a preset threshold, marks the cache disk as a disk to be replaced and configures the storage node where the disk to be replaced is located in a degraded state, including: generating alarm information, the alarm information including a replacement prompt for the disk to be replaced and degraded state information of the storage node where the disk to be replaced is located.

[0041] In one implementation, the operational lifespan data includes power-on duration and / or data write volume.

[0042] The implementation methods of the apparatus are the same as or similar to the corresponding implementation methods, and will not be described again here.

[0043] In one embodiment, this disclosure provides an electronic device including a processor and a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement the aforementioned storage management method. From a hardware perspective, a hardware architecture diagram can be found... Figure 3 As shown.

[0044] In one embodiment, this disclosure provides a machine-readable storage medium storing machine-executable instructions that, when invoked and executed by a processor, cause the processor to implement the aforementioned storage management method.

[0045] Here, a machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, a machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0046] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0047] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing this disclosure, the functions of each unit can be implemented in one or more software and / or hardware.

[0048] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware implementation, a completely software implementation, or an implementation combining software and hardware aspects. Furthermore, embodiments of this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0049] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0050] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0051] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0052] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware implementation, a completely software implementation, or an implementation combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (which may include, but are not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0053] The above description is merely an embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. A storage management method, characterized in that, Applied to a storage cluster, which includes several storage nodes, the method includes: Monitor the cache disk lifetime data of the storage node. If the cache disk lifetime data of a storage node reaches a preset threshold, mark the cache disk as a disk to be replaced and configure the storage node where the disk to be replaced is located in a degraded state. The degraded state includes restricting data writing and data migration. In response to the event that the new disk used to replace the disk has been loaded on the storage node where the disk to be replaced is located, the cached data in the disk to be replaced is migrated to the new disk; Configure the new disk as the cache disk of the storage node where the disk to be replaced is located, and delete the configuration information of the disk to be replaced that was configured as a cache disk.

2. The method according to claim 1, characterized in that, The configuration of the new disk as the cache disk of the storage node where the disk to be replaced resides involves deleting the configuration information of the disk to be replaced that was configured as a cache disk, including: In response to the event that the disk to be replaced has been unloaded, the degraded state of the storage node that was configured as degraded is cancelled.

3. The method according to claim 1, characterized in that, If the cache disk lifetime data of a monitored storage node reaches a preset threshold, the cache disk is marked as a disk to be replaced, and the storage node containing the disk to be replaced is configured to a degraded state, including: Generate alarm information, which includes a replacement prompt for the disk to be replaced and the degradation status information of the storage node where the disk to be replaced is located.

4. The method according to claim 1, characterized in that, The operational lifespan data includes power-on time and / or data write volume.

5. A storage management device, characterized in that, Applied to a storage cluster, the storage cluster comprising several storage nodes, the device includes: The first module is used to monitor the cache disk lifespan data of storage nodes. If the cache disk lifespan data of a storage node reaches a preset threshold, the cache disk is marked as a disk to be replaced, and the storage node where the disk to be replaced is located is configured to be in a degraded state. The degraded state includes restricting data writing and data migration. The second module is used to migrate the cached data in the disk to be replaced to the new disk in response to the event that the new disk to be replaced has been loaded on the storage node where the disk to be replaced is located. The third module is used to configure the new disk as the cache disk of the storage node where the disk to be replaced is located, and to delete the configuration information of the disk to be replaced that was configured as a cache disk.

6. The apparatus according to claim 5, characterized in that, The configuration of the new disk as the cache disk of the storage node where the disk to be replaced resides involves deleting the configuration information of the disk to be replaced that was configured as a cache disk, including: In response to the event that the disk to be replaced has been unloaded, the degraded state of the storage node that was configured as degraded is cancelled.

7. The apparatus according to claim 5, characterized in that, If the cache disk lifetime data of a monitored storage node reaches a preset threshold, the cache disk is marked as a disk to be replaced, and the storage node containing the disk to be replaced is configured to a degraded state, including: Generate alarm information, which includes a replacement prompt for the disk to be replaced and the degradation status information of the storage node where the disk to be replaced is located.

8. The apparatus according to claim 5, characterized in that, The operational lifespan data includes power-on time and / or data write volume.

9. An electronic device, characterized in that, include: A processor and a machine-readable storage medium storing machine-executable instructions that can be executed by the processor to implement the method of any one of claims 1-4.

10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions that, when invoked and executed by a processor, cause the processor to implement the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Data deletion method and system for distributed storage cluster

    CN106227469A

  • Method for realizing inverse wear equalization

    CN113535082A