Method, apparatus, and computer program product for managing a storage system

CN115220646BActive Publication Date: 2026-08-21EMC IP HLDG CO LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202110432823.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-21
Publication Date
2026-08-21
Estimated Expiration
2041-04-21

AI Technical Summary

Technical Problem

然而,在使用RAID对顺序I/O数据流进行存储时,通常无法充分地利用存储系统中的各个存储设备,从而限制了存储系统对顺序I/O进行存储的性能表现

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115220646B_ABST
    Figure CN115220646B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, an electronic device and a computer program product for managing a storage system. The method comprises determining a plurality of storage units provided by a plurality of storage devices, each of the plurality of storage units having a storage space allocated from a first number of the plurality of storage devices; dividing the plurality of storage units into at least one storage unit group based on a total number of the plurality of storage devices and the first number, each of the at least one storage unit group comprising a second number of the plurality of storage units; and storing to-be-stored data into the at least one storage unit group based on a logical address of the to-be-stored data. Embodiments of the present disclosure can more reasonably allocate storage resources and improve the processing performance of the storage system for sequential data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure generally relate to the field of data storage, and more specifically to methods, apparatus, and computer program products for managing storage systems. Background Technology

[0002] With the development of data storage technology, various data storage devices are now able to provide users with increasingly higher data storage capabilities, and data access speeds have also improved significantly. While improving data storage capabilities, users are also placing increasingly higher demands on the performance of storing sequential data (e.g., input / output (I / O) data streams).

[0003] Redundant Array of Independent Disks (RAID) is commonly used in storage systems for data storage. In RAID, a disk is a logical concept and can be distributed across different storage devices in the storage system. However, when using RAID to store sequential I / O data streams, it is often impossible to fully utilize the individual storage devices in the storage system, thus limiting the performance of the storage system for sequential I / O storage. Summary of the Invention

[0004] Embodiments of this disclosure provide methods, apparatus, and computer program products for managing storage systems.

[0005] In a first aspect of this disclosure, a method for managing storage devices is provided. The method includes determining a plurality of storage cells provided by a plurality of storage devices, each of the plurality of storage cells having storage space allocated from a first number of storage devices; dividing the plurality of storage cells into at least one group of storage cells based on the total number of the plurality of storage devices and the first number, each group of storage cells including a second number of storage cells; and storing data to be stored into the at least one group of storage cells based on the logical address of the data to be stored.

[0006] In a second aspect of this disclosure, an electronic device is provided. The electronic device includes at least one processing unit and at least one memory. The at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform an action including: determining a plurality of storage cells provided by a plurality of storage devices, each of the plurality of storage cells having storage space allocated from a first number of storage devices; dividing the plurality of storage cells into at least one group of storage cells based on the total number of the plurality of storage devices and the first number, each group of storage cells including a second number of storage cells; and storing data to be stored into the at least one group of storage cells based on a logical address of the data to be stored.

[0007] In a third aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a non-transitory computer storage medium and includes machine-executable instructions. When executed by a device, the machine-executable instructions cause the device to perform any step of the method described in the first aspect of this disclosure.

[0008] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or essential features of this disclosure, nor is it intended to limit the scope of this disclosure. Attached Figure Description

[0009] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0010] Figure 1 A schematic diagram of an example system that can be implemented therein according to some embodiments of the present disclosure is shown;

[0011] Figure 2 A schematic block diagram of a conventional scheme for storing sequential data is shown;

[0012] Figure 3 A flowchart illustrating an example method for managing a storage system according to some embodiments of this disclosure is shown;

[0013] Figure 4 A flowchart is shown illustrating an example method for dividing storage cell groups according to some embodiments of the present disclosure;

[0014] Figure 5 A schematic diagram is shown for storing data blocks sequentially into a group of storage cells according to some embodiments of the present disclosure;

[0015] Figure 6 A schematic diagram illustrating updating at least one group of storage cells and migrating stored data according to some embodiments of the present disclosure is shown; and

[0016] Figure 7 A schematic block diagram of an example device that can be used to implement embodiments of the present disclosure is shown.

[0017] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0018] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0019] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0020] In the context of this disclosure, the storage system can be a RAID-based storage system. A RAID-based storage system combines multiple storage devices into a single disk array. By providing redundant storage devices, the reliability of the entire disk group is significantly greater than that of a single storage device. RAID can offer various advantages over a single storage device, such as enhanced data consolidation, enhanced fault tolerance, increased throughput or capacity, and so on. Multiple RAID standards exist, such as RAID-1, RAID-2, RAID-3, RAID-4, RAID-5, RAID-6, RAID-10, RAID-50, and so on.

[0021] Figure 1 A schematic diagram of a storage system 100 in which the methods of this disclosure may be implemented is shown. Figure 1The storage system 100 shown includes multiple storage devices 160-1, 160-2, 160-3, ..., 160-P, where P is an integer greater than 1. For ease of discussion, storage devices 160-1, 160-2, 160-3, ..., 160-P are sometimes collectively referred to or individually as storage device 160. In some embodiments, examples of storage device 160 may include, but are not limited to, digital versatile discs (DVDs), Blu-ray discs (BDs), optical discs (CDs), floppy disks, hard disk drives, magnetic tape drives, optical drives, hard disk drive (HDDs), solid-state storage devices (SSDs), or other hard disk devices.

[0022] Figure 1 The diagram also shows that storage system 100 is divided into multiple disks 110-1, 110-2, ..., 110-N. In the following discussion, disks 110-1, 110-2, ..., 110-N are sometimes collectively referred to as disk 110 or simply as disk 110. For example, for a RAID 6 (6+2) based storage system 100, N can be 8. As another example, for a RAID 4 (4+1) based storage system 100, N can be 5. It should be understood that N can be any suitable integer greater than 1.

[0023] In this article, disk 110 refers to a virtual, logical storage disk. Although Figure 1 The specific correspondence between disk 110 and storage device 160 is not shown, but it should be understood that disk 110 can be distributed across different storage devices 160 (e.g., hard disks) in storage system 100. For example, disk 110 can be distributed across one storage device 160 in storage system 100, or it can be distributed across several storage devices 160 in storage system 100. Disk 110 can be divided into multiple blocks. For example, disk 110-1 can be divided into blocks 120-1, 120-2, 120-3, ..., 120-M (collectively referred to as or individually referred to as block 120), where M is an integer greater than 1. Disk 110-2 can be divided into blocks 130-1, 130-2, 130-3, ..., 130-M (collectively referred to as or individually referred to as block 130), where M is an integer greater than 1. Disk 110-N can be divided into blocks 140-1, 140-2, 140-3, ..., 140-M (collectively referred to as blocks 140 or individually), where M is an integer greater than 1. Each block 120, block 130, ..., block 140 has the same storage space size.

[0024] Each block in the storage system 100 is further divided into multiple storage units 150. For example, storage unit 150-1 may include block 120-1 of disk 110-1, block 130-1 of disk 110-2, ... block 140-1 of disk 110-N. Storage unit 150-2 may include block 120-2 of disk 110-1, block 130-2 of disk 110-2, ... block 140-2 of disk 110-N. Storage unit 150-3 may include block 120-3 of disk 110-1, block 130-3 of disk 110-2, ... block 140-3 of disk 110-N. Storage unit 150-M may include block 120-M of disk 110-1, block 130-M of disk 110-2, ... block 140-M of disk 110-N.

[0025] For each storage unit 150, the individual blocks it comprises correspond to storage space allocated by different storage devices 160. In other words, the storage space of storage unit 150 is provided by N physical storage units 160. Furthermore, storage unit 150 can be viewed as a storage space with contiguous logical addresses, for use by file systems or other storage applications to store data.

[0026] During the use of storage system 100, it is often necessary to store large amounts of sequential data (e.g., sequential I / O data streams). When the amount of sequential data is particularly large, the number N of disks 110 in storage system 100 (i.e., the RAID width) will limit the ability of storage system 100 to process sequential data. Therefore, it is necessary to optimize the sequential data processing capability of storage system 100.

[0027] Figure 2 This illustrates a common scheme for processing sequential data in a storage system. For example... Figure 2 The storage system 200 is divided into disks 210-1, 210-2, ..., 210-N (collectively or individually referred to as "disk 210"). Disk 210 includes multiple blocks. For example, disk 210-1 includes blocks 220-1, 220-2, 220-3, ..., 220-M (collectively or individually referred to as "block 220"). Figure 2 As shown, each block in the storage system 200 is further divided into multiple storage units 250-1, 250-2, ... 250-M (collectively referred to as "storage unit 250" or individually).

[0028] For storage cell 250, each block it comprises corresponds to a storage space allocated to a different storage device. In other words, the storage space of storage cell 250 is provided by N physical storage cells. Storage cell 250 can be viewed as a storage space with contiguous logical addresses.

[0029] When storage system 200 needs to process large amounts of data, refer to Figure 2 As detailed in the diagram, storage cell 250-1 stores sequential data in ascending order of logical address until its storage space is full. For example, the data block 230-1 with the smallest logical address (e.g., a 4.5MB I / O data stripe) is stored first in storage cell 250-1, followed by the data block 230-2 with the second smallest logical address, then data block 230-3, and so on. The storage system 200 continues to store data in ascending order of logical address in a similar manner until storage cell 250-1 is fully occupied.

[0030] In a conventional approach, when the storage space of a certain storage unit 250 is not fully occupied, the storage system 200 will not use other storage units 250 to process data. For example, when the storage space of storage unit 250-1 is not fully occupied, the storage system 200 will not use other storage units 250 to process data. During this process of using only storage unit 250, the N storage devices in the storage system 200 that provide storage space solely for storage unit 250 are used, while other storage devices are not used. When the amount of sequential data is particularly large, during this period, the workload of these N storage devices providing storage space for storage unit 250 is very large, while other storage devices are unnecessarily idle.

[0031] This conventional approach can only utilize a limited number of storage devices simultaneously, thus limiting the maximum amount of sequential data that can be processed. Furthermore, with only a small number of storage devices being used concurrently while others remain idle, it is detrimental to the allocation of storage resources within the storage system. Additionally, this approach can also cause some storage devices to be overloaded compared to others, increasing the risk of device damage.

[0032] Embodiments of this disclosure provide a scheme for managing a storage system to address one or more of the aforementioned problems and other potential problems. In this scheme, multiple storage units are divided into at least one storage unit group based on the total number of storage devices and a first number of storage devices spanned by the storage space of each storage unit. The scheme also includes storing data to be stored into at least one storage unit group, wherein the storage space of each storage unit group is allocated by a number greater than the first number of storage devices. In this manner, sequential data can be processed simultaneously on more storage devices, improving the data processing performance of the storage system. Furthermore, the storage space of the storage system can be allocated more rationally, thereby avoiding damage to storage devices due to overuse.

[0033] The basic principles and several exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings. Figure 3 A flowchart of an example method 300 for managing a storage system 100 according to an embodiment of the present disclosure is shown. Method 300 may be, for example, by... Figure 1 The method is executed using the storage system 100 shown. It should be understood that method 300 may also include additional actions not shown and / or the actions shown may be omitted; the scope of this disclosure is not limited in this respect. The following is in conjunction with… Figure 1 Let me describe method 300 in detail.

[0034] like Figure 3 As shown, at 310, a plurality of storage cells 150 provided by a plurality of storage devices 160 are defined. Each of the plurality of storage cells 150 has a first number (i.e., ...) from the plurality of storage devices 160. Figure 1 The storage space provided by storage device 160 (N in the text).

[0035] For example, storage system 100 may include 16 storage devices 160, and the storage space of the 16 storage devices 160 is divided into blocks of a predetermined size, such as 32GB per block. It should be understood that this is merely illustrative, and the blocks can be set to any suitable predetermined size. M storage units 150 are determined, provided by the plurality of (e.g., 16) storage devices 160. Each storage unit 150 has blocks provided by N (e.g., 8 for RAID 6) storage devices 160 from the plurality of storage devices 160.

[0036] In some embodiments, the total number of storage units 150 can be determined based on a predetermined size of storage space provided by the storage device 160 and a predetermined size of each block. For example, if the total number of storage devices 160 in the storage system 100 is 16, each storage device 160 is predetermined to provide 128GB of storage space, and the predetermined size of each block is 32GB, then for RAID 6 (i.e., N equals 8), the total number M of storage units 150 can be determined to be 8. In some embodiments, the total number of storage units 150 can also be arbitrarily selected, for example, the total number M of storage units 150 can be selected as 24. It should be understood that the 16, 24, 128GB, 32GB, etc. listed above are merely exemplary and do not limit the invention in any way. In some embodiments, other total numbers and storage spaces of storage devices 160 can be selected and allocated in other ways.

[0037] At point 320, based on the total number of storage devices 160 and the aforementioned first number, the multiple storage cells 150 are divided into at least one storage cell group. Each storage cell group in the at least one storage cell group includes a second number of storage cells 150. For example, if the total number of storage devices 160 is 16 and the first number is 8, the second number can be determined to be 2. That is, M (e.g., 8) storage cells 150 are divided into pairs, resulting in a total of 4 storage cell groups.

[0038] In some embodiments, for each of at least one group of storage cells, the group of storage cells has storage space allocated from a third number of storage devices 160 out of a plurality of storage devices 160. The third number is the product of a first number and a second number. For example, in the example described above, the first number is 8, the second number is 2, and the third number is 16. That is, each group of storage cells has storage space allocated from 16 storage devices 160. In this way, the storage space of the group of storage cells can be provided by different storage devices 160 to a greater extent.

[0039] In some embodiments, such as Figure 4 The method 400 shown is used to divide multiple storage cells 150 into at least one group of storage cells. The following will combine... Figure 4 Several embodiments of dividing at least one group of storage cells are described in more detail.

[0040] At position 330, based on the logical address of the data to be stored, the data to be stored is stored into at least one group of storage cells. For example, the data to be stored is stored sequentially into one storage cell of one of the storage cell groups, in ascending order of the logical addresses of the data to be stored.

[0041] In some embodiments, an identifier for a group of storage cells is determined based on the logical address of the data to be stored, the size of the storage cells 150, and a second number. The data to be stored is stored in a target storage cell group with the identifier in at least one storage cell group. For example, the identifier can be determined by the following formula (1):

[0042] PER_GROUP_ID = IO_LBA /

[0043] (PER_SIZE*NUMBER_OF_PERS_IN_THE_GROUP) (1) Where PER_GROUP_ID represents the identifier of the storage unit group, IO_LBA represents the logical address of the data to be stored, PER_SIZE represents the size of storage unit 150, and NUMBER_OF_PERS_IN_THE_GROUP represents the second number.

[0044] For example, in the example described above in conjunction with 320, the size of storage cell 150 is the product of 32GB and a first number N (i.e., 8). The second number is 2. According to equation (1), it can be determined which storage cell group the data to be stored will be stored in. For example, if the identifier PER_GROUP_ID is determined to be 0, the data to be stored can be stored in the storage cell group with identifier 0.

[0045] In some embodiments, multiple sequentially ordered data blocks comprising the data to be stored may be interleaved and stored in a second number of storage cells 150 within a target storage cell group. Reference is made below. Figure 5 Example diagrams are described illustrating the storage of data to be stored into a target group of storage cells in storage system 100 according to some embodiments. For example... Figure 5 As shown, a target storage cell group 510 is illustrated. The target storage cell group 510 includes storage cell 150-1 and storage cell 150-2. It should be understood that, for clarity, Figure 5 Only the target storage unit group 510 is shown in the figure, but it should be understood that the storage system 100 may also include other storage unit groups.

[0046] like Figure 5 As shown, the data block 520-1 with the smallest logical address among multiple data blocks (for example, a 4.5MB data stripe) can be stored in storage unit 150-1. Then, the data block 520-2 with the second smallest logical address is stored in storage unit 150-2. After that, data block 520-3 is stored in storage unit 150-1 in sequence, data block 520-4 in storage unit 150-2, data block 520-5 in storage unit 150-1, data block 520-6 in storage unit 150-2, and so on, until the storage space of storage units 150-1 and 150-2 is full.

[0047] In some embodiments, the storage location of the data block to be stored in at least one group of storage cells can be determined using the following equations (2)-(4) in combination with the previously described equation (1):

[0048] LBA_OFFSET_IN_PER_GROUP=IO_LBA%(PER_SIZE*

[0049] NUMBER_OF_PERS_IN_THE_GROUP) (2)

[0050] PER_ID = PER_GROUP_ID*

[0051] NUMBER_OF_PERS_IN_THE_GROUP+

[0052] LBA_OFFSET_IN_PER_GROUP / LRB%

[0053] NUMBER_OF_PERS_IN_THE_GROUP (3)

[0054] LBA_OFFSET_IN_PER=LBA_OFFSET_IN_PER_GROUP / (LRB

[0055] *NUMBER_OF_PERS_IN_THE_GROUP)*LRB+IO_LBA%

[0056] LRB (4)

[0057] Where LBA_OFFSET_IN_PER_GROUP represents the offset of the data block within the target memory group, IO_LBA represents the logical address of the data block, PER_SIZE represents the size of memory cell 150, NUMBER_OF_PERS_IN_THE_GROUP represents the second number, and PER_GROUP_ID represents the identifier of the memory group. LRB represents the sequence of data blocks, for example, ... Figure 5 As shown, the LRB of data block 520-1 is 1, the LRB of data block 520-2 is 2, and so on. The determined PER_ID represents the identifier of the target storage unit in the target storage unit, and the data block will be stored in the target storage unit with this identifier. The determined LBA_OFFSET_IN_PER represents the address offset of the data block in the target storage unit of the target storage unit group. The data to be stored can be stored in at least one storage unit group of the storage system 100 through the above calculations.

[0058] In this way, sequential data can be stored in groups of storage cells of storage system 100, instead of just in a single storage cell. These groups of storage cells can then be allocated storage space by more storage devices. This increases the number of available storage devices, thereby optimizing the allocation of storage resources. Furthermore, this reduces the workload on certain storage devices, preventing them from failing due to excessive workload. As a result, the performance of the storage system can be improved, especially for processing sequential I / O data streams.

[0059] refer to Figure 4 A method 400 for partitioning at least one group of storage units according to some embodiments is described. Method 400 can be considered as an example implementation of block 320 in method 300. Method 400 can be, for example, by... Figure 1 The method is executed by the storage system 100 shown. It should be understood that method 400 can also be executed by other suitable devices or apparatuses. Method 400 may include additional actions not shown and / or the actions shown may be omitted; the scope of this disclosure is not limited in this respect. For ease of explanation, reference will be made to... Figure 1 To describe process 400.

[0060] At point 410, a threshold number of storage cell groups that the multiple storage devices 160 can provide is determined based on the total number of storage devices 160 and a first number. The second number can be an integer not greater than the threshold number. For example, the threshold number can be determined using the following formula (5):

[0061] T=round_down(NUM_PHY_DISK,N) / N (5)

[0062] Where NUM_PHY_DISK represents the total number of storage devices 160, N represents the first number, round_down represents the round-down operation, and T represents the threshold number. For example, if the total number of storage devices 160 in storage system 100 is 50 and N is 8, then T = round_down(50,8) / 8 = 48 / 8 = 6. That is, the threshold number of the storage unit group is 8.

[0063] By determining a threshold number for a storage cell group and setting a second number as an integer not greater than the threshold number, it can be ensured that each storage cell in the same storage cell group is allocated storage space by a different storage device 160. In other words, there is no situation where a block of one storage cell in the same storage cell group is allocated storage space by the same storage device 160 as a block of another storage cell in the same storage cell group.

[0064] This avoids having multiple blocks allocated to the same storage device 160 within a single storage unit group. This prevents an overloaded storage device from having too many blocks allocated to a single storage unit group. As a result, storage resources are allocated more rationally, preventing storage device failures or damage caused by excessive workload.

[0065] At 420, the second number is determined based on the threshold number and the total number of multiple storage units 150 (i.e., M). For example, the second number can be determined using the following formula (6):

[0066] Q=upper_bound(NUM_PER_IN_VDG_FACTORS,T) (6)

[0067] Where NUM_PER_IN_VDG_FACTORS represents the factor of the total number M of storage cells 150, T represents the threshold number, upper_bound represents the lookup boundary value, and Q represents the second number. For example, if the total number of storage cells 150 is 24, then it has factors 2, 3, 4, 6, 8, and 12. Then the second number can be determined as: Q = upper_bound([2,3,4,6,8,12],6) = 6.

[0068] In this way, the second number can be determined to be such that all storage units 150 can be divided into an integer number of storage unit groups. In this manner, all storage units 150 can be utilized, thereby avoiding waste of storage resources. Furthermore, storage resources can be utilized more fully and rationally, further improving the performance of the storage system.

[0069] Furthermore, by determining the second number as the maximum number that satisfies the threshold and allows the storage cells 150 to be divided into an integer number of storage cell groups, the number of storage cells utilized by each storage cell group can be maximized. In some embodiments, the second number can also be maximized by reasonably selecting the total number M of storage cells 150. For example, compared to a total of 25 storage cells 150, 24 storage cells can be more easily divided into storage cell groups with a larger second number. Therefore, reasonably selecting the total number of storage cells 150 also allows for better partitioning of storage cell groups. This, in turn, enables more efficient use of storage resources, thereby improving storage system performance.

[0070] During use, the number of storage devices 160 may increase. In some embodiments, the second number can be updated in response to an increase in the number of storage devices 160. For example, a reference can be used. Figure 4 The described process is used to update the second number. It should be understood that other methods can also be used to update the second number. If the updated second number does not change, the existing storage unit groups are not modified. Multiple newly added storage units 150 are allocated using only the newly added storage devices 160, and are divided into at least one newly added storage unit group according to the second number.

[0071] If the updated second number is greater than the original second number, then the existing storage cells 150 and the new storage devices 150 allocated by the newly added storage devices 160 are used to determine the updated at least one group of storage cells. Each of the updated at least one group of storage cells has the updated second number of storage cells.

[0072] In some embodiments, determining the updated at least one group of storage cells includes, for each storage cell 150, determining whether the storage cell 150 stores data. If the storage cell 150 already stores data, then the storage cell 150, the other storage cells 150 in the original storage cell group corresponding to the storage cell 150, and a fourth number of storage cells 150 that do not store data are determined as the updated group of storage cells. The fourth number is the difference between the updated fourth number and the original fourth number.

[0073] Figure 6 A schematic diagram illustrating the updating of at least one group of storage cells according to an embodiment of the present disclosure is shown. For example, if the total number of storage devices 160 before the update is 16 and the total number of storage devices after the update is 24, then the second number before the update is 2 and the second number after the update is 3. If the total number of storage cells is 24, then there are 8 groups of storage cells after the update. Figure 6 The original memory cell groups 610-1, 610-2, and 610-3 (collectively or individually referred to as memory cell group 610) are shown. It should be understood that, although... Figure 6 The diagram shows three groups of storage cells 610, but the storage system 100 may also have fewer or more groups of storage cells 610. Figure 6 The updated memory cell groups 630-1 and 630-2 (collectively or individually referred to as memory cell group 630) are also shown. It should be understood that, although... Figure 6 The diagram shows two updated storage cell groups 630, but the storage system 100 may also have fewer or more storage cell groups 630.

[0074] like Figure 6 As shown, the previous storage unit group 610-1 included storage unit 650-1 and storage unit 650-2. Both storage units 650-1 and 650-2 stored data. The updated storage unit group 630-1 consists of storage unit 650-1, storage unit 650-2, and a previously unused storage unit 655-3. Storage unit 655-3 can be a storage unit provided by the newly added storage device 160, or it can be a previously unused storage unit.

[0075] Figure 6 The diagram also illustrates the corresponding storage locations of each data block of data stored before the update within the updated at least one group of storage cells 630. In some embodiments, the corresponding storage locations of each data block of stored data within the updated at least one group of storage cells 630 can be determined based on the logical addresses of the plurality of data blocks comprising the stored data, the size of the storage cells, and the second number of updates. For example, a reference can be used. Figure 5The process described by equations (1)-(4) is used to determine the corresponding storage location of each data block of stored data in at least one updated storage unit group 630.

[0076] Based on the updated storage locations of multiple data blocks, the data blocks are sequentially migrated to at least one updated storage unit group 630 in ascending order of logical address. For example, as... Figure 6 As shown, data blocks 620-1 and 620-2 do not need to be migrated. Data block 620-3 is migrated to storage cell 655-3 of the updated storage cell group 630-1, data block 620-4 is migrated to storage cell 650-1, and so on. Figure 6 The document also shows the respective migrated storage locations of the other data blocks, which will not be detailed here. Although Figure 6 For clarity, the direct correspondence between data blocks in the pre-update memory cell group and their corresponding data blocks in the updated memory cell group is not shown; however, the corresponding labels indicate the storage location of each data block. It should be understood that, for illustrative purposes, Figure 6 Each storage unit in the diagram stores six data blocks, but this is merely illustrative and does not limit the invention in any way. Any suitable number of data blocks can be stored in a storage unit.

[0077] In some embodiments, checkpoints can be set during the migration of stored data. Data blocks before the checkpoint are stored using the updated storage unit group, while data blocks after the checkpoint are stored using the original storage unit group. For example, initially, the checkpoint is set at data block 620-1 in storage unit 650-1 of storage unit group 610-1 (i.e., the storage location where the data block with the smallest logical address is located), and data blocks are migrated sequentially in ascending order of logical address. Each time a data block is migrated, the checkpoint is moved to the next data block. Furthermore, if a data block is invalid (e.g., the client can indicate that the data block is invalid), the data block is skipped and not migrated; instead, the checkpoint is moved directly to the next data block.

[0078] like Figure 6As shown, after all data blocks in storage cell group 610-1 have been migrated, storage cell group 630-1 is not full. The checkpoint moves to data block 625-1 of storage cell 630-1 in the next storage cell group 610-2, and data blocks in storage cell group 610-2 are migrated sequentially. For example, data block 625-1 is migrated to storage cell 650-1. After data block 625-6 of storage cell 650-3 is migrated to storage cell 655-3, storage cell group 630-1 is full. Afterwards, the remaining data blocks in storage cell 110-2 will be migrated to the next updated storage cell group 630-2. In some embodiments, storage cell group 630-2 may consist of storage cell 650-3, storage cell 650-4, and another storage cell that does not store data. Figure 6 (Not shown in the image).

[0079] Alternatively, storage cell group 630-2 can also consist of unused storage cells 655-4, 655-5, and 655-6. The remaining data blocks in storage cell group 610-2 will be migrated to storage cell group 630-2 sequentially. In the example where storage cell group 630-2 consists of unused storage cells, after all the remaining data blocks in storage cell group 610-2 have been migrated to storage cell group 630-2, storage cells 650-3 and 650-4 will no longer contain data blocks. Therefore, during the subsequent data block migration process, storage cells 650-3 and 650-4 can be used as unused storage cells to form a new storage cell group.

[0080] Alternatively or concurrently, the data block migration process described above can be performed in the background during off-peak hours of the storage system 100. Furthermore, if at least one group of storage cells contains only sparse data, the aforementioned data block migration process can be combined with a garbage collection process.

[0081] By performing the data block migration process described above, sequential data can be stored in ascending logical address order within at least one updated group of storage units. After the data migration, all sequential data is stored according to the updated group of at least one storage unit, and the updated group of storage units is allocated storage space by a larger number of storage devices. This allows for better allocation of storage resources and improves the overall performance of the storage system.

[0082] Figure 7 A schematic block diagram of an example device 700 that can be used to implement embodiments of the present disclosure is shown. For example, such as Figure 1 The storage system 100 shown can be implemented by device 700. For example... Figure 7As shown, device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. CPU 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0083] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0084] The various processes and handling described above, such as methods 300 and / or 400, can be executed by processing unit 701. For example, in some embodiments, methods 300 and / or 400 can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of methods 300 and / or 400 described above can be performed.

[0085] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0086] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0087] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0088] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0089] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0090] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0091] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0093] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for managing a storage system, comprising: A plurality of storage cells provided by a plurality of storage devices are determined, each of the plurality of storage cells having storage space allocated from a first number of the plurality of storage devices; Based on the total number of the plurality of storage devices and the first number, the plurality of storage cells are divided into at least one storage cell group, each storage cell group in the at least one storage cell group includes a second number of storage cells, wherein for each storage cell group in the at least one storage cell group, the storage cell group has storage space allocated from a third number of storage devices among the plurality of storage devices, the third number being equal to the first number multiplied by the second number, wherein each storage cell within a single storage cell group has storage space allocated from a different subset of the plurality of storage devices; as well as Based on the logical address of the data to be stored, the data to be stored is stored in the at least one group of storage units.

2. The method according to claim 1, wherein dividing the plurality of storage cells into the at least one group of storage cells comprises: Based on the total number of the plurality of storage devices and the first number, a threshold number of storage unit groups that the plurality of storage devices can provide is determined; The second number is determined based on the threshold number and the total number of the plurality of storage units; as well as Based on the second number, the plurality of storage cells are divided into at least one storage cell group.

3. The method according to claim 1, wherein storing the data to be stored in the at least one group of storage cells based on the logical address of the data to be stored comprises: An identifier for the group of storage cells is determined based on the logical address of the data to be stored, the size of the storage cell, and the second number. as well as The data to be stored is stored in the target storage cell group that has the identifier in the at least one storage cell group.

4. The method of claim 3, wherein storing the data to be stored into a target storage cell group having the identifier in the at least one storage cell group comprises: The data to be stored, comprising multiple sequentially ordered data blocks, is interleaved and stored into a second number of storage cells in the target storage cell group.

5. The method according to claim 1, further comprising: In response to the increase in the total number of the plurality of storage devices, the second number is updated; Determine whether the updated second number is greater than the original second number; as well as If it is determined that the updated second number is greater than the original second number, at least one updated storage cell group is determined, each of the updated at least one storage cell group having the updated second number of storage cells.

6. The method of claim 5, wherein determining the updated at least one group of storage cells comprises: For each of the plurality of storage units, determine whether the storage unit stores data; as well as If it is determined that the storage unit already stores data, the storage unit, each of the other storage units in the storage unit group that included the storage unit before the update, and a fourth number of storage units that do not store data are determined as the updated storage unit group, the fourth number being the difference between the updated second number and the second number before the update.

7. The method of claim 5, further comprising: A group of storage cells containing stored data prior to the update is identified, wherein the stored data includes multiple data blocks; as well as Based on the logical addresses of the multiple data blocks included in the data stored before the update, the size of the storage unit, and the updated second number, the multiple data blocks are migrated sequentially to the updated at least one group of storage units in ascending order of their logical addresses.

8. An electronic device comprising: At least one processor; as well as At least one memory storing computer program instructions, the at least one memory and the computer program instructions being configured, together with the at least one processor, to cause the electronic device to perform actions, the actions including: A plurality of storage cells provided by a plurality of storage devices are determined, each of the plurality of storage cells having storage space allocated from a first number of the plurality of storage devices; Based on the total number of the plurality of storage devices and the first number, the plurality of storage cells are divided into at least one storage cell group. Each storage cell group in the at least one storage cell group includes a second number of storage cells. For each storage cell group in the at least one storage cell group, the storage cell group has storage space allocated from a third number of storage devices among the plurality of storage devices, the third number being equal to the first number multiplied by the second number. Each storage cell within a single storage cell group has storage space allocated from a different subset of the plurality of storage devices. Based on the logical address of the data to be stored, the data to be stored is stored in the at least one group of storage units.

9. The electronic device of claim 8, wherein dividing the plurality of storage cells into the at least one group of storage cells comprises: Based on the total number of the plurality of storage devices and the first number, a threshold number of storage unit groups that the plurality of storage devices can provide is determined; The second number is determined based on the threshold number and the total number of the plurality of storage units; as well as Based on the second number, the plurality of storage cells are divided into at least one storage cell group.

10. The electronic device of claim 8, wherein storing the data to be stored in the at least one group of storage cells based on the logical address of the data to be stored comprises: An identifier for the group of storage cells is determined based on the logical address of the data to be stored, the size of the storage cell, and the second number. as well as The data to be stored is stored in the target storage cell group that has the identifier in the at least one storage cell group.

11. The electronic device of claim 10, wherein storing the data to be stored into a target storage cell group having the identifier in the at least one storage cell group comprises: The data to be stored, comprising multiple sequentially ordered data blocks, is interleaved and stored into a second number of storage cells in the target storage cell group.

12. The electronic device of claim 8, wherein the action further comprises: In response to the increase in the total number of the plurality of storage devices, the second number is updated; Determine whether the updated second number is greater than the original second number; as well as If it is determined that the updated second number is greater than the original second number, at least one updated storage cell group is determined, each of the updated at least one storage cell group having the updated second number of storage cells.

13. The electronic device of claim 12, wherein determining the updated at least one group of storage cells comprises: For each of the plurality of storage units, determine whether the storage unit stores data; as well as If it is determined that the storage unit already stores data, the storage unit, each of the other storage units in the storage unit group that included the storage unit before the update, and a fourth number of storage units that do not store data are determined as the updated storage unit group, the fourth number being the difference between the updated second number and the second number before the update.

14. The electronic device of claim 12, wherein the action further comprises: A group of storage cells containing stored data prior to the update is identified, wherein the stored data includes multiple data blocks; as well as Based on the logical addresses of the multiple data blocks included in the data stored before the update, the size of the storage unit, and the updated second number, the multiple data blocks are migrated sequentially to the updated at least one group of storage units in ascending order of their logical addresses.

15. A computer program product tangibly stored on a non-volatile computer-readable medium and comprising machine-executable instructions that, when executed, cause a device to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method, system, and program for determining a configuration of a logical array including a plurality of storage devices

    US20030074527A1

  • Rebalancing of striped disk data

    US20070118689A1

  • Declustered array of storage devices with chunk groups and support for multiple erasure schemes

    US20180095676A1

  • Resiliency groups

    US20190042407A1

  • Storage system

    US20200210291A1