Data transfer method and apparatus, host, storage system, and storage medium

By performing internal movement of valid data and generating predictive verification data within the logical block group, the write amplification and performance overhead issues when there is a large amount of valid data within the logical block group are resolved, achieving more efficient data movement and reliability.

WO2025236685A1PCT designated stage Publication Date: 2025-11-20HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/142889
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-13
Filing Date
2024-12-26
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

In storage systems, when there is a large amount of valid data within a logical block group, existing technologies result in significant write amplification and performance overhead during system garbage collection.

Method used

Data migration is performed within the logical block group, moving valid data from the first strip to the invalid data address within the logical block group, reducing the overall data migration volume, and generating verification data before migration to ensure data redundancy and avoid data protection gaps.

Benefits of technology

Internal data migration reduces write amplification and performance overhead, improving data reliability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024142889_20112025_PF_FP_ABST
    Figure CN2024142889_20112025_PF_FP_ABST
Patent Text Reader

Abstract

A data transfer method and apparatus, a host, a storage system, and a storage medium, relating to the technical field of storage. In the method, within a first logical block group of a host, valid data in a first stripe of the first logical block group is transferred to a logical address of invalid data outside the first stripe. The valid data in the first stripe in the first logical block group is transferred, so that the entire first stripe can be freed up. Further, when system garbage collection is performed either manually or automatically by the host, the first stripe is completely emptied for subsequent writing of new data. In addition, only part of the valid data in the first logical block group is transferred without needing to transfer all of the valid data in the entire first logical block group; accordingly, compared to the related art in which all valid data in the first logical block group is transferred to another blank logical block group, the amount of valid data being transferred is reduced, thereby helping to reduce write amplification and performance overhead caused by data transfer.
Need to check novelty before this filing date? Find Prior Art

Description

Data migration method and device, host, storage system, and storage medium

[0001] The present application claims priority to the Chinese Patent Application No. 202410608274.2, filed on May 13, 2024, and entitled "Data migration method and device, host, storage system, and storage medium", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the technical field of storage, and in particular to a data migration method and device, a host, a storage system, and a storage medium. BACKGROUND

[0003] The storage system includes a host and a storage, wherein the storage provides a storage space mapped as a plurality of logical block groups, each logical block group includes a plurality of logical blocks, and the plurality of logical blocks in the logical block group are respectively mapped to different storage spaces of the storage. Each logical block group includes a plurality of stripes, each stripe includes a plurality of stripe units, and the plurality of stripe units in the stripe come from different logical blocks in the logical block group. The plurality of stripe units in a stripe include data units and check units, the data units are used to store data in the stripe, and the check units are used to store check data of the data in the stripe, thereby providing data redundancy and improving the reliability of the data.

[0004] In the related art, the host performs data writing in the granularity of a stripe, and when performing system garbage collection, the valid data in the entire logical block group is moved to a blank logical block group in the unit of the entire logical block group to recycle the logical block group. However, in the case of high storage space utilization, that is, in the case of less invalid data in the logical block group, when performing system garbage collection by using the above data migration method, a large amount of valid data in the logical block group is moved, causing large system write amplification and performance overhead. SUMMARY

[0005] Embodiments of the present application provide a data migration method, device, host, storage system, and storage medium, which can reduce write amplification and performance overhead caused by data migration. The technical solution is as follows.

[0006] In a first aspect, a data migration method is provided. The first logical block group of the host includes a plurality of logical blocks, the plurality of logical blocks map the storage space of the storage, the plurality of logical blocks include a plurality of stripes, each stripe in the plurality of stripes is composed of stripe units of different logical blocks, wherein each logical block includes a plurality of stripe units,

[0007] The first logical block group includes valid data, and the method includes: moving the valid data in the first stripe of the first logical block group to a logical address of invalid data outside the first stripe within the first logical block group of the host.

[0008] The data moving method can be applied to a scenario in which the host automatically performs system garbage collection, for example, the host periodically performs system garbage collection, or the host performs system garbage collection when the available capacity of the storage space of the memory is less than the first capacity. The data moving method can also be applied to a scenario in which the host performs system garbage collection manually, for example, the host performs system garbage collection in response to a garbage collection instruction.

[0009] In the method, the valid data in the first stripe of the first logical block group is moved out, so that the first stripe can be completely emptied. When the system garbage collection is performed manually or automatically by the host, the first stripe is completely emptied for subsequent writing of new data. Moreover, the valid data moved out of the first stripe is written to an address outside the first stripe in the first logical block group, that is, the data moving occurs within the first logical block group, and the moving is performed on part (or all) of the valid data in the first logical block group, without moving all the valid data in the entire logical block group. Therefore, compared with moving all the valid data in the first logical block group to another blank logical block group in the related art, the amount of valid data moved is reduced, thereby facilitating reduction of write amplification and performance overhead caused by data moving.

[0010] Optionally, the logical address of the invalid data outside the first stripe is located in a second stripe of the first logical block group, and before the valid data in the first stripe is moved to the logical address of the invalid data outside the first stripe, the method further includes:

[0011] The first check data of the second stripe is generated according to the predicted data in the second stripe after the moving.

[0012] In the method, before the actual moving operation is performed, the data in the second stripe after the moving is predicted, and the new check data of the second stripe is generated based on the predicted data. Compared with generating the new check data after the data moving is performed, there is no data protection gap from the completion of the data moving to the generation of the new check data, the data in the stripe is protected by the redundancy mechanism at all times, and the reliability of the data is improved.

[0013] Optionally, the method further includes:

[0014] If a data error occurs in the process of moving the first valid data in the first stripe to the logical address of the invalid data in the second stripe, data recovery is performed on the second stripe based on the first check data.

[0015] Optionally, after the first valid data in the first stripe is moved to the logical address of the invalid data in the second stripe, the method further comprises:

[0016] deleting the second check data of the second stripe, the second check data being the check data of the original data of the second stripe.

[0017] Optionally, after the valid data in the first stripe is moved to the logical address of the invalid data outside the first stripe, the first stripe does not include valid data.

[0018] Optionally, the method further comprises:

[0019] after the valid data in the first stripe is moved, sending a trim instruction or an unmap instruction to the memory based on the logical address of the invalid data in the first stripe, to instruct the memory that the storage space mapped by the logical address of the invalid data in the first stripe stores invalid data.

[0020] In some embodiments, if the invalid data in the first stripe is distributed in different stripe units of the first stripe, the host sends a trim instruction or an unmap instruction to the memory based on the logical address of the invalid data in the different stripe units. After receiving the trim instruction or the unmap instruction, the memory marks the data stored in the storage space mapped by the logical address carried in the instruction as invalid data, and deletes the mapping relationship between the logical address and the physical address.

[0021] Optionally, the method further comprises:

[0022] performing data writing based on the first stripe.

[0023] In a second aspect, a data moving apparatus is provided, the apparatus comprising at least one function module for performing the data moving method provided in the first aspect or any possible implementation manner of the first aspect.

[0024] In a third aspect, a host is provided, the host comprising a processor and a memory, the memory being configured to store instructions, and the processor being configured to read and execute the instructions in the memory to implement the data moving method provided in the first aspect or any possible implementation manner of the first aspect.

[0025] Optionally, the host further comprises a memory.

[0026] Optionally, the memory is a flash disk or a phase change memory (PCM).

[0027] In a fourth aspect, a storage system is provided, the storage system comprising a memory and a host capable of executing instructions to perform the data migration method as provided in the first aspect or any possible implementation of the first aspect.

[0028] In a fifth aspect, a computer program product comprising instructions, which when executed by a host, cause the host to perform the data migration method as provided in the first aspect or any possible implementation of the first aspect.

[0029] In a sixth aspect, a computer-readable storage medium is provided, the storage medium comprising instructions, which when executed by a host, cause the host to perform the data migration method as provided in the first aspect or any possible implementation of the first aspect.

[0030] On the basis of the implementation manners of the aspects provided by the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF DRAWINGS

[0031] Fig. 1 is a structural schematic diagram of a storage system provided by an embodiment of the present application;

[0032] Fig. 2 is a structural schematic diagram of a host provided by an embodiment of the present application;

[0033] Fig. 3 is a flowchart of a data migration method provided by an embodiment of the present application;

[0034] Fig. 4 is a flowchart of a data migration method provided by an embodiment of the present application;

[0035] Fig. 5 is a flowchart of a data migration method provided by an embodiment of the present application;

[0036] Fig. 6 is a flowchart of a data migration method provided by an embodiment of the present application;

[0037] Fig. 7 is a flowchart of a data migration method provided by an embodiment of the present application;

[0038] Fig. 8 is a structural schematic diagram of a data migration apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION

[0039] To make the purpose, technical scheme and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0040] First, the implementation environment of the present application is introduced.

[0041] Figure 1 is a structural diagram of a storage system according to an embodiment of the present application. The storage system shown in Figure 1 is applicable to both distributed storage system and centralized storage system. The storage system comprises a host 101 and a storage 102. In Figure 1, the storage 102 is exemplified by several hard disks.

[0042] The physical address of the storage space provided by the storage 102 is not directly exposed to the host 101. The storage 102 can be of any type. In the embodiments of the present application, the storage 102 is exemplified by a flash disk. However, the embodiments of the present application are also applicable to phase change memory (PCM) or other types of storage. The storage 102 cannot directly write new data in the location of old data by overwriting. Instead, the old data is removed by garbage collection before new data is written in the original location of the old data. Each storage 102 is divided into several physical chunks. These physical chunks are mapped into logical chunks to form a storage pool. The storage pool is used to provide storage space to the host 101. The storage space actually comes from the storage 102 included in the storage system. Of course, not all storages 102 need to provide storage space to the storage pool. In actual applications, the storage system can include one or more storage pools. One storage pool includes part or all of the storages 102. A plurality of logical chunks from different storages 102 form a logical chunk group (CKG). The logical chunk group is the smallest allocation unit of the storage pool.

[0043] It should be noted that in some embodiments of the present application, the host can not include a storage. In other embodiments, the host can also include a storage. In this case, the physical address of the storage space is also not exposed to the central processing unit (CPU) of the host 101. Therefore, the storage space cannot be directly managed by the software system running on the host CPU.

[0044] The number of logical chunks included in a logical chunk group depends on the mechanism (also referred to as redundancy mode) used to ensure data reliability. Generally, in order to ensure data reliability, the storage system uses an erasure coding (EC) check mechanism to store data. The EC check mechanism refers to dividing the data to be stored into at least two data shards. The check shards of the at least two data shards are calculated according to a certain check algorithm. When one data shard is lost, the data can be recovered by using other data shards and check shards. If the EC check mechanism is used, then a logical chunk group includes at least three logical chunks. Each logical chunk is located on a different storage 102.

[0045] Since the principle of the redundant array of independent disks (RAID) is similar to that of the EC, the RAID is regarded as a kind of EC in the present application. In the EC check mechanism, a plurality of logical blocks from different memories 102 are divided into a data group and a check group according to a set RAID type. The data group includes at least two logical blocks for storing data shards, and the check group includes at least one logical block for storing check shards of the data shards. When the data in the memory is full to a certain size, the host 101 can be divided into a plurality of data shards according to the set RAID type, and the check shards are calculated and obtained, and the data shards and the check shards are sent to a plurality of different memories 101 to be saved in the logical block group. The logical addresses where the data shards and the check shards are located form a stripe. A logical block group can include one or more stripes. The logical addresses where the data shards and the check shards included in the stripe are located can be referred to as a stripe unit (SU), or a strip. For example, taking a capacity of 8 kilobytes (KB) of a stripe as an example, it is assumed that 6 mapped storage spaces from 6 flash disks logical blocks (chunk0-chunk5) form a logical block group (a subset of the storage pool), and the logical block group is grouped based on a set RAID type (taking RAID6 as an example). Among them, chunk0, chunk1, chunk2 and chunk3 are data block groups, and chunk4 and chunk5 are check block groups. When the data stored in the memory of the host 101 reaches 8KBx4=32KB, the host divides the data into 4 data shards (data shard 0, data shard 1, data shard 2, and data shard 3), each of which has a size of 8KB, and then calculates and obtains 2 check shards (P0 and Q0), each of which also has a size of 8KB. The processor of the host 101 sends the data shards and the check shards to the 6 flash disks to store the data in the logical block group. It can be understood that according to the redundancy protection mechanism of RAID6, when any two data shards or check shards fail, the failed shards can be reconstructed according to the remaining data shards or check shards.

[0046] The memory 102 includes a controller and a storage medium, and the controller includes a CPU, a random access memory (RAM), a host interface, a storage medium interface, and a bus. The storage medium includes a plurality of physical blocks. The memory 102 can be an embedded multi media card (eMMC), a universal flash storage (UFS), a serial attached SCSI solid state drive (SAS SSD), a serial advanced technology attachment solid state drive (SATA SSD), a non-volatile memory express (NVMe), a NAND flash, or a disk, etc., and the embodiments of the present application do not limit the same. The bus between the controller and the storage medium in the memory 102 can be an open NAND flash interface (ONFI), a toggle, a common flash interface (CFI), a serial peripheral interface (SPI), or a DDR bus, etc., and the embodiments of the present application do not limit the same. Exemplarily, the memory 102 can receive a data read / write request of the host 101, write data into the storage medium in the memory 102, or read data from the storage medium in the memory 102 and return to the host 101.

[0047] The bus between the host 101 and the plurality of memories 102 can be a peripheral component interconnect express (PCIe), a UFS, an eMMC, an NVMe, a SAS, or a SATA, etc., and the embodiments of the present application do not limit the same.

[0048] It should be noted that the above FIG. 1 is described by taking a distributed storage system or a centralized storage system as an example, and the implementation environment of the embodiments of the present application can also be a host including a memory, and the embodiments of the present application do not limit the same.

[0049] The structure of the host 101 in the implementation environment is described below. FIG. 2 is a structural schematic diagram of a host according to an embodiment of the present application. As shown in FIG. 2, the host 101 includes a bus 1011, a processor 1012, a memory 1013, and a communication interface 1014. The processor 1012, the memory 1013, and the communication interface 1014 communicate with each other through the bus 1011. The host 101 can be a server or a storage array controller. It should be understood that the number of processors and memories in the host 101 is not limited in the present application.

[0050] The bus 1011 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, only one line is shown in FIG. 2, but it does not mean that there is only one bus or only one type of bus. The bus 1011 can include a path for transmitting information between various components (for example, the memory 1013, the processor 1012, and the communication interface 1014) of the host 101.

[0051] The processor 1012 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), or the like.

[0052] The memory 1013 can include a volatile memory, such as a random access memory (RAM). The memory 1013 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a mechanical hard disk drive (HDD), a solid state drive (SSD), a flash drive, or a PCM disk.

[0053] The memory 1013 stores executable program code, and the processor 1012 executes the executable program code to implement the data migration method as shown in FIG. 3. That is, the memory 1013 stores program code for executing the data migration method.

[0054] The communication interface 1014 uses a transceiver module such as, but not limited to, a network interface card, a transceiver, and the like to enable communication between the host 101 and other devices or communication networks.

[0055] The embodiment of the present application provides a data moving method, in which valid data in a first sub-strip in a first logical block group is moved out, so that the first sub-strip can be emptied as a whole. Then, when system garbage collection is manually or automatically performed by a host, the first sub-strip is emptied as a whole, and is used for subsequent writing of new data. Moreover, the valid data moved out from the first sub-strip is written to an address outside the first sub-strip in the first logical block group, that is, data moving occurs in the first logical block group, and the moving is performed on part (or all) of the valid data in the first logical block group, and all valid data in the first logical block group does not need to be moved. Therefore, compared with moving all valid data in the first logical block group to another blank logical block group in the related art, the amount of valid data to be moved is reduced, so that write amplification and performance overhead caused by data moving are reduced.

[0056] In the embodiment of the present application, different logical blocks are mapped to storage spaces of different memories, or are mapped to storage spaces of different servers / disk frames, so that the reliability can be improved. Similarly, different sub-strip units of the same sub-strip are provided with storage spaces of different memories, or are provided with storage spaces of different servers / disk frames, so that the reliability can be improved.

[0057] In the embodiment of the present application, a specific sub-strip can be specified to move valid data. Or, a sub-strip is not specified, for example, valid data is moved from the tail to the head of a logical block group, or is moved from the head to the tail of the logical block group, or is moved from both ends to the middle of the logical block group. In the case that a sub-strip is not specified, the effect of moving valid data in one or more sub-strips can be achieved, and after the moving is completed, the valid data is concentrated in a position in the logical block group.

[0058] The data migration method can be applied to a scenario in which the host performs system garbage collection. For example, the host periodically performs system garbage collection. In a system garbage collection process, the host arranges valid data scattered in the logical block group based on the data migration method. If the data migration method frees at least one stripe in the logical block group, the host sends a trim instruction or an unmap instruction based on a logical address of invalid data in the at least one stripe to the storage. The storage marks data stored in a storage space mapped by the logical address of the invalid data in the at least one stripe as invalid based on the trim instruction or the unmap instruction, so as to complete recycling of the at least one stripe. For another example, the host performs system garbage collection when available capacity of the storage is less than the first capacity. For another example, the host performs system garbage collection in response to a garbage collection instruction. It should be noted that the above examples of application scenarios of the data migration method are only exemplary, and the embodiments of the present application do not limit the application scenarios of the data migration method.

[0059] The following further describes the flow of the data migration method in the scenario in which the host performs system garbage collection. FIG. 3 is a flowchart of a data migration method according to an embodiment of the present application. In the example in which the method is applied to the host, the first logical block group of the host includes a plurality of logical blocks, and the plurality of logical blocks map a storage space of the storage. The plurality of logical blocks includes a plurality of stripes, and each stripe in the plurality of stripes is composed of stripe units of different logical blocks. Each logical block includes a plurality of stripe units. As shown in FIG. 3, the method includes the following steps 301 to 308.

[0060] 301. In response to performing system garbage collection, the host generates a migration plan for the first logical block group. The migration plan is used to indicate that valid data in a first stripe in the first logical block group is migrated to a logical address of invalid data outside the first stripe.

[0061] The timing of the system garbage collection performed by the host can be set according to actual requirements. For example, in some embodiments, the host periodically performs system garbage collection, and accordingly, in response to reaching the time node of the system garbage collection, the host determines a first logical block group from the plurality of logical block groups of the host, the first logical block group being a logical block group to be subjected to garbage collection in the plurality of logical block groups. For another example, in some other embodiments, a user manually instructs the host to perform system garbage collection, and accordingly, the host performs system garbage collection in response to a system garbage collection instruction. For yet another example, in some yet other embodiments, the host performs system garbage collection when the available capacity of the storage space of the memory is less than a first capacity, and accordingly, in response to the available capacity of the storage space of the memory being less than the first capacity, the host determines a first logical block group from the plurality of logical block groups, where the host learns the available capacity of the storage space of the memory in the following manners: the host periodically sends an available capacity acquisition request to the memory, the memory returns the available capacity of the memory to the host upon receiving the request, or the memory reports a capacity shortage message to the host when the available capacity is less than the first capacity to notify the host that the available capacity of the memory is less than the first capacity, or the memory periodically reports the available capacity of the memory to the host, which is not limited in the embodiments of the present application. The value of the first capacity can be set according to actual requirements, which is not limited in the embodiments of the present application.

[0062] For example, taking the first stripe in the first logical block group as an example, the process of generating the migration plan for the first stripe by the host includes: determining a first stripe and a second stripe in the first logical block group, the first stripe being a stripe for migrating out valid data, and the second stripe being a stripe for migrating in valid data; determining a destination migration address of first valid data in the first stripe based on the data amount of the valid data in the first stripe and the data amount of invalid data in the second stripe, the destination migration address of the first valid data being a logical address of the invalid data in the second stripe; and generating the migration plan based on the first valid data and the destination migration address of the first valid data. It should be noted that the host first generates the migration plan for the first stripe, and then generates the migration plan for other stripes based on the migration plan for the first stripe, and further generates the migration plan for the first logical block group. It should be noted that the process of generating the migration plan for other stripes in the first logical block group is the same as that of generating the migration plan for the first stripe, which is not described herein.

[0063] The first stripe (the stripe for moving out valid data) and the second stripe (the stripe for moving in valid data) in the first logical block group are determined based on the moving strategy adopted. For example, in some embodiments, the moving strategy of moving from the tail of a logical block to the head of the logical block is adopted, and correspondingly, the stripe at the tail of the first logical block group is determined as the first stripe, and the stripe at the head of the first logical block group is determined as the second stripe; for another example, in some other embodiments, the moving strategy of moving from the head of a logical block group to the tail of the logical block group is adopted, and correspondingly, the stripe at the head of the first logical block group is determined as the first stripe, and the stripe at the tail of the first logical block group is determined as the second stripe; for yet another example, in some other embodiments, the moving strategy of moving from the head and the tail of a logical block group to the middle of the logical block group is adopted, and correspondingly, the stripe at the head or the tail of the first logical block group is determined as the first stripe, and the stripe in the middle of the first logical block group is determined as the second stripe. It should be noted that the moving strategy described above is a moving strategy based on the moving direction, that is, the first stripe is not specified, and based on the moving strategy described above, the valid data is concentrated in the position set in the first logical block group after the moving is completed. If the first stripe determined based on the moving strategy described above does not include invalid data, the first stripe is excluded, and the first stripe is determined again based on the moving strategy described above, so as to avoid moving the valid data in the stripe that does not include invalid data, and reduce the write amplification caused by the moving.

[0064] In addition to the moving strategy based on the moving direction described above, a moving strategy based on the data moving amount can also be adopted, for example, a plurality of stripes in the first logical block group are sorted according to the data amount of the valid data, the stripe including the least valid data in the first logical block group is determined as the first stripe, and the stripe including the most valid data in the first logical block group is determined as the second stripe. By adopting the moving strategy, the invalid data in the first logical block group can be concentrated in at least one stripe, and at the same time, the data amount of the valid data moved can be reduced, so as to reduce the write amplification caused by the moving. It should be noted that the examples of the corresponding moving strategies described above are only exemplary, and in some embodiments, other moving strategies are also adopted, and the embodiments of the present application do not limit the moving strategy.

[0065] It should be noted that the generation process of the above moving plan is described by taking an example that the data amount of the valid data in the first stripe is less than or equal to the data amount of the invalid data in the second stripe, that is, the second stripe can receive all the valid data in the first stripe. In some embodiments, the data amount of the valid data in the first stripe is greater than the data amount of the invalid data in the second stripe, that is, the second stripe cannot receive all the valid data in the first stripe. Then, the host determines a third stripe in the first logical block group, and the determination manner of the third stripe is the same as that of the second stripe. The host determines the destination moving address of the first valid data in the first stripe and the destination moving address of the second valid data in the first stripe based on the data amount of the valid data in the first stripe, the data amount of the invalid data in the second stripe, and the data amount of the invalid data in the third stripe. The destination moving address of the first valid data is the logical address of the invalid data in the second stripe, and the destination moving address of the second valid data is the logical address of the invalid data in the third stripe. The host generates a moving plan based on the first valid data, the second valid data, the destination moving address of the first valid data, and the destination moving address of the second valid data. In the above embodiment, when the invalid data space of one stripe is insufficient to receive the valid data in the first stripe, the valid data in the first stripe is planned to be moved to the logical address of the invalid data in multiple stripes outside the first stripe, so that the first stripe can be emptied as much as possible after moving, thereby facilitating the recycling of the first stripe.

[0066] It should be noted that step 301 is described by taking an example that the host generates a moving plan. In some embodiments, the host provides the distribution information of the valid data and the invalid data in the first logical block group to a target client in response to performing system garbage collection, and the target client provides the function of generating a moving plan. The target client generates a moving plan for the first logical block group and sends the moving plan to the host. The embodiments of the present application are not limited in comparison.

[0067] The first valid data in the first stripe is moved to the logical address of the invalid data in the second stripe of the first logical block group as an example.

[0068] 302. The host generates first check data of the second stripe of the first logical block group based on the moving plan, and the first check data is the check data of the data in the second stripe after moving according to the moving plan.

[0069] The host generates the first check data by the following process: predicting first data based on the migration plan, the first data being data in the second stripe after migration according to the migration plan; and generating the first check data based on a check algorithm of a set RAID type and the first data. The process is described below with reference to FIG. 4. FIG. 4 is a flowchart of a data migration method according to an embodiment of the present application. As shown in FIG. 4, the first stripe S1 includes six stripe units SU10-SU15, and the second stripe includes six stripe units SU20-SU25. SU10-SU15 and SU20-SU25 are located on six logical blocks chunk0-chunk5 respectively. The storage space mapped by chunk0-chunk5 is from memory0-memory5 respectively. In the first stripe S1, SU10-SU13 store data A (invalid data), B (invalid data), C (valid data and the data amount is less than the capacity of a stripe unit), and D (valid data and the data amount is less than the capacity of a stripe unit). SU14 and SU15 are used to store check data P1 and Q1 of the data in SU10-SU13. In the second stripe S2, SU20-SU23 store data E (valid data), F (valid data), G (invalid data), and H (valid data). SU24 and SU25 are used to store check data P2 and Q2 of the data in SU20-SU23. The migration plan indicates that C in the first stripe S1 is migrated to SU22 in the second stripe S2, and D in the first stripe S1 is migrated to SU22 in the second stripe S2. Based on the migration plan, the host predicts that the data in the second stripe is data E, F, (C, D), and H. Then, the host generates first check data P2' and Q2' of the second stripe based on a check algorithm of a set RAID type and E, F, (C, D), and H. It should be noted that the example shown in FIG. 4 is only illustrative. For example, the data amount of valid data C and D in a first stripe is equal to the data amount of invalid data in a second stripe. The same applies to other cases, and thus is not described herein.

[0070] In step 302, before the actual migration operation, the check data of the data in the second stripe after migration is predicted. Compared with generating new check data after the data migration, there is no data protection gap from the completion of data migration to the generation of new check data. The data in the stripe is protected by the EC check mechanism at all times, and the reliability of the data is improved.

[0071] It should be noted that steps 301 and 302 are an implementation of generating first check data of the second stripe based on the predicted data in the second stripe after migration. In some embodiments, the process is implemented based on other steps, which are not limited in the embodiments of the present application.

[0072] 303、the host stores the first check data of the second stripe into a first stripe unit, the logical block where the first stripe unit is located and the logical block where the second stripe unit is located are mapped to the storage space from the same memory, the second stripe unit is used to store the second check data of the second stripe, and the second check data is the check data of the original data of the second stripe.

[0073] For example, the above step 302 is continued, FIG. 5 is a flow diagram of a data moving method provided by an embodiment of the present application, the check data P2 and Q2 shown in FIG. 5 are the second check data of the second stripe, SU24 and SU25 are the second stripe units, SU24' and SU25' are the first stripe units, the logical block where SU24' is located is chunk4', and the storage space mapped by chunk4' and chunk4 are both from memory 4; the logical block where SU5' is located is chunk5', and the storage space mapped by chunk5' and chunk5 are both from memory 5. The host stores the generated first check data P2' and Q2' into SU24' and SU25' respectively.

[0074] In the above step 303, the predicted moving check data and the check data before moving are stored in different logical blocks provided by the same memory, so as to avoid destroying the RAID organization relationship, and then facilitate reconstructing the second stripe based on the first stripe unit after moving is completed, wherein reconstructing the second stripe based on the first stripe unit is to replace the original second stripe unit in the second stripe with the first stripe unit.

[0075] It should be noted that the above step 303 is described by taking the host storing the first check data into the first stripe unit after the first check data is generated as an example, in some embodiments, the host stores the first check data into the cache after the first check data is generated, and deletes the second check data stored in the second stripe unit after moving is completed, and then stores the first check data into the second stripe unit. In the above embodiment, the host caches the generated first check data first, and then writes it into the second stripe unit after moving is completed, so as to avoid reconstructing the second stripe, and then avoid re-determining the mapping relationship between the logical address of the logical block group and the physical address of the memory. The present embodiment does not limit this.

[0076] 304、the host moves the first valid data in the first stripe to the logical address of the invalid data in the second stripe.

[0077] The host can complete the moving process of the valid data by different instructions. For example, in some embodiments, the host obtains the first valid data from the first logical address in the first stripe, and stores the first valid data to a second logical address in the second stripe, which is a logical address of invalid data in the second stripe. The host sends a read instruction to the memory for the first logical address to obtain the first valid data, and sends a write instruction to the memory for the second logical address to store the first valid data to the second logical address. For another example, in other embodiments, the host sends a data moving instruction to the memory, and the data moving instruction indicates the memory to move the valid data in the physical address corresponding to the first logical address to the physical address corresponding to the second logical address.

[0078] It should be noted that the host can move the first valid data to the second stripe at one time, or move the first valid data to the second stripe in multiple times, and the embodiments of the present application do not limit this.

[0079] It should be noted that the steps 301 to 304 are described by taking moving the valid data in a single stripe as an example, but the moving granularity of the embodiments of the present application is not limited to one stripe. In some embodiments, the moving is performed with multiple stripes as the granularity, and the embodiments of the present application do not limit the moving granularity.

[0080] It should be noted that if the moving plan further indicates to move the third valid data in the fourth stripe of the first logical block group to a third logical address, the logical address of the third valid data is continuous with the logical address of the first valid data, the third logical address is continuous with the second logical address, the host can move the first valid data and the third valid data to the second logical address at one time, that is, in the actual moving process, the host can move the data according to the stripe, first move the data for one stripe, and then move the data for another stripe after the moving of the valid data in the stripe is completed, which is beneficial to ensure the order of the data moving process, thereby reducing the probability of errors occurring in the data moving process. The host can also move the data according to the logical address of the valid data and the destination moving address, and move the valid data with continuous logical addresses and the same destination moving address together, thereby being beneficial to reduce the instruction overhead of the host in the data moving process. The embodiments of the present application do not limit the specific moving process of the host, as long as the moving result is consistent with the moving plan.

[0081] It should be noted that after moving the first valid data in the first stripe to the second stripe, the host marks the original first valid data in the first stripe as invalid data.

[0082] The following continues the example in step 303 above. FIG. 6 is a flow diagram of a data migration method according to an embodiment of the present application. As shown in FIG. 6, the host migrates C in the first stripe S1 to SU22 in the second stripe S2, and migrates D in the first stripe S1 to SU22 in the second stripe S2 according to the migration plan. After the migration, the data in the second stripe S2 is E (valid data), F (valid data), C (valid data), D (valid data), and H (valid data), and the data in the first stripe S1 is A (invalid data), B (invalid data), C (invalid data), and D (invalid data).

[0083] 305. If a data error occurs in the process of migrating the valid data in the first stripe to the logical address of the invalid data in the second stripe, the host performs data recovery on the second stripe based on the first check data.

[0084] The process in which the host determines whether a data error occurs in the second stripe includes: generating third check data of the second stripe based on the data in the second stripe after the migration and a check algorithm of the set RAID type; and comparing the third check data and the first check data. If the comparison is consistent, it indicates that no data error occurs. If the comparison is inconsistent, it indicates that a data error occurs.

[0085] The process in which the host performs data recovery on the second stripe based on the first check data includes: determining the stripe unit in which the data error occurs based on the third check data and the first check data; and performing data recovery on the data in error in the second stripe based on the first check data and the data not in error if the number of the stripe unit in which the data error occurs is less than a first number, the first number being determined based on the error correction capability of the check algorithm of the set RAID type. In some embodiments, if the number of the data in the stripe unit in which the data error occurs is greater than the first number, the host reports a data error message.

[0086] In step 305, if a data error occurs in the data migration process, data recovery is performed on the second stripe based on the predicted check data, thereby improving the reliability of the data in the second stripe.

[0087] It should be noted that step 305 is an optional step, and in some embodiments, step 305 is not performed, which is not limited in the embodiments of the present application.

[0088] 306. The host deletes the second check data of the second stripe.

[0089] In some embodiments, after the host deletes the second check data of the second stripe, the host reconstructs the second stripe based on the first stripe unit, that is, the first stripe unit replaces the original second stripe unit in the second stripe.

[0090] It should be noted that the step 306 is optional, and in some embodiments, the step 306 is not performed, which is not limited in the embodiments of the present application.

[0091] Taking the example in the step 304 as an example, FIG. 7 is a flow diagram of a data moving method provided by the embodiments of the present application. As shown in FIG. 7, after the moving is completed, the host deletes the check data P2 and Q2 in the SU24 and the SU25 of the second stripe S2, and reconstructs the second stripe S2 into S2', which includes the stripe units SU20-SU23, SU24' and SU25'.

[0092] 307. If the valid data in the first stripe is moved to the logical address of the invalid data after the first stripe, the first stripe does not include valid data, and the host sends a trim instruction or an unmap instruction to the storage based on the logical address of the invalid data in the first stripe after the valid data in the first stripe is emptied, to instruct the storage that the storage space mapped by the logical address of the invalid data in the first stripe stores invalid data.

[0093] If the invalid data in the first stripe is distributed in different stripe units of the first stripe, the host sends a trim instruction or an unmap instruction to the storage based on the logical address of the invalid data in the different stripe units. In some embodiments, after the storage receives the trim instruction or the unmap instruction, the storage marks the data stored in the storage space mapped by the logical address as invalid data, and deletes the mapping relationship between the logical address and the physical address.

[0094] It should be noted that the step 307 is optional, and in some embodiments, the step 307 is not performed, which is not limited in the embodiments of the present application.

[0095] It should be noted that the step 307 is described based on the example of sending a trim instruction or an unmap instruction to the storage based on the logical address of the invalid data in a single stripe, that is, the granularity of data cleaning is described based on a single stripe. In some embodiments, the granularity of data cleaning is based on multiple stripes, that is, the host sends a trim instruction and an unmap instruction to the storage based on the logical address of the invalid data in multiple emptied stripes after the valid data in the multiple stripes is emptied, and the granularity of data cleaning is not limited in the embodiments of the present application.

[0096] 308. The host performs data writing based on the first stripe.

[0097] It should be noted that the above step 308 is an optional step, and in some embodiments, the step 308 is not performed, and the embodiments of the present application do not limit this.

[0098] In the above steps 307 and 308, if at least one stripe is emptied by data migration, a trim instruction or an unmap instruction is sent to the memory to recover the emptied stripe, so that data can be written into the emptied stripe, that is, the space inside the logical block group is recovered by performing effective data arrangement inside the logical block group, without the need to migrate all the effective data in the entire logical block group to another logical block group, reducing write amplification and performance overhead, and improving the efficiency of system garbage collection.

[0099] It should be noted that the above steps 302 to 308 are described by taking the first stripe as an example, and other stripes in the first logical block group are the same as the first stripe, and details are not repeated.

[0100] In the method, the valid data in the first stripe in the first logical block group is moved out, so that the first stripe can be emptied as a whole. When the system garbage collection is performed manually or automatically by the host, the first stripe is emptied completely for subsequent writing of new data. Since the valid data moved out from the first stripe is written to an address outside the first stripe in the first logical block group, the data movement occurs within the first logical block group, and only part (or all) of the valid data in the first logical block group is moved, without moving all the valid data in the first logical block group. Therefore, compared with moving all the valid data in the first logical block group to another blank logical block group, the amount of data movement is reduced, so that the write amplification and performance overhead caused by data movement are reduced. Further, before the actual movement operation is performed, the check data of the data in the second stripe after the movement is predicted. Compared with generating new check data after the data movement, there is no data protection gap from the completion of the data movement to the generation of the new check data, the data in the stripe is protected by the EC check mechanism at all times, and the reliability of the data is improved. In addition, if at least one stripe is emptied by the data movement, a trim instruction or an unmap instruction is sent to the storage, so that the emptied stripe is recycled. Therefore, data can be written to the emptied stripe without performing garbage collection by moving all the valid data in the first logical block group to another logical block group, the write amplification and performance overhead are reduced, and the efficiency of the system garbage collection is improved. In the embodiment of the application, the stripe is a logical concept, so the valid data stripe removed can be directly written with data (overwritten). In other embodiments, the valid data stripe removed can be logically cleared of data, and then new data can be written to the data-cleared stripe. The physical data clearing is triggered by the trim instruction or the unmap instruction, and the host (or a server, a storage controller, or another storage management device) sends the trim instruction or the unmap instruction to the storage, so as to inform the storage which data is invalid. After receiving the instruction, the storage physically clears the data by garbage collection at an appropriate time, and recycles the storage space occupied by the invalid data.

[0101] FIG. 8 is a structural schematic diagram of a data movement device provided by an embodiment of the application. The first logical block group of the host includes a plurality of logical blocks, the plurality of logical blocks map the storage space of the storage, the plurality of logical blocks include a plurality of stripes, each stripe in the plurality of stripes is composed of stripe units of different logical blocks, and each logical block includes a plurality of stripe units.

[0102] The device includes a determination module 801 and a movement module 802.

[0103] The determining module 801 is configured to determine valid data.

[0104] The moving module 802 is configured to move the valid data in the first stripe of the first logical block group to a logical address of invalid data outside the first stripe within the first logical block group.

[0105] Optionally, the apparatus further includes:

[0106] The generating module is configured to generate first check data of the second stripe according to the predicted data in the second stripe after moving.

[0107] Optionally, the apparatus further includes:

[0108] The restoring module is configured to, if a data error occurs in the process of moving the first valid data in the first stripe to the logical address of invalid data in the second stripe, perform data restoration on the second stripe based on the first check data.

[0109] Optionally, the apparatus further includes:

[0110] The deleting module is configured to delete second check data of the second stripe, the second check data being check data of original data of the second stripe.

[0111] Optionally, after the valid data in the first stripe is moved to the logical address of invalid data outside the first stripe, the first stripe does not include valid data.

[0112] Optionally, the apparatus further includes:

[0113] The sending module is configured to, after the valid data in the first stripe is moved, send a trim instruction or an unmap instruction to the memory based on the logical address of invalid data in the first stripe, to instruct the memory that a storage space mapped by the logical address of invalid data in the first stripe stores invalid data.

[0114] Optionally, the apparatus further includes:

[0115] The writing module is configured to perform data writing based on the first stripe.

[0116] It should be noted that in other embodiments, the steps responsible for implementation by the above modules can be specified as needed, and the entire function of the above device can be implemented by the above modules respectively implementing different steps in the above data migration method. That is, the data migration device provided in the above embodiments is only used as an example to illustrate the division of the above functional modules when implementing the data migration method. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the corresponding method embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.

[0117] Among them, the determination module 801 and the migration module 802 can be implemented by software or can be implemented by hardware. For example, the implementation of the migration module 802 is introduced below. Similarly, the implementation of the determination module 801 can refer to the implementation of the migration module 802.

[0118] As an example of a software functional unit, the migration module 802 can include code running on a computing instance. Among them, the computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance can be one or more. For example, the migration module 802 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region), or can be distributed in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (availability zone, AZ), or can be distributed in different AZs, and each AZ includes a data center or multiple data centers with similar geographical locations. Among them, usually one region can include multiple AZs.

[0119] Similarly, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (virtual private cloud, VPC), or can be distributed in multiple VPCs. Among them, usually one VPC is set in one region, and communication between two VPCs in the same region and between VPCs in different regions needs to be set in each VPC Communication gateway, and interconnection between VPCs is realized through the communication gateway.

[0120] As an example of a hardware functional unit, the migration module 802 can include at least one computing device, such as a server or the like. Alternatively, the migration module 802 can also be a device implemented with an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0121] The plurality of computing devices included in the migration module 802 can be distributed in the same region or in different regions. The plurality of computing devices included in the migration module 802 can be distributed in the same AZ or in different AZs. Similarly, the plurality of computing devices included in the migration module 802 can be distributed in the same VPC or in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0122] The present application implements a computer program product, which can be a software or a program product containing instructions, capable of running on a computing device, a server, or stored in any available medium. When the computer program product runs on the host in the storage system, the host is caused to perform the data migration method provided in the above method embodiments.

[0123] The present application embodiment provides a computer readable storage medium, which can be any available medium capable of being stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk), etc. The computer readable storage medium includes instructions, when executed by the host in the storage system, the host performs the data migration method provided in the above method embodiments.

[0124] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the data and the like involved in the present application are obtained under full authorization.

[0125] Those skilled in the art can realize that, in combination with the method steps and units described in the embodiments disclosed in the present application, the electronic hardware, computer software or combination of the two can be realized, and in order to clearly show the interchangeability of hardware and software, the steps and components of the embodiments have been described in the above description. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0127] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the unit is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed mutual units can be indirect coupling or communication connection through some interfaces, devices or units, and can also be electrical, mechanical or other forms of connection.

[0128] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiments of the present application.

[0129] In addition, each unit in various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of a software unit.

[0130] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computing device (which can be a personal computer, a server, or a computing device) to execute all or part of the steps of the method in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0131] The terms "first", "second", and the like are used in the present application to distinguish between items or similar items having substantially the same function and action. It should be understood that there is no logical or chronological dependency between "first", "second", and "n", and the quantity and execution order are not limited. It should also be understood that although the following description uses the terms first, second, and the like to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first sub-bar can be called the second sub-bar, and similarly, the second sub-bar can be called the first sub-bar. The first sub-bar and the second sub-bar can both be sub-bars, and in some cases, can be separate and different sub-bars.

[0132] In the present application, the term "at least one" means one or more, and the term "a plurality of" means two or more. The terms "system" and "network" are often used interchangeably.

[0133] It should also be understood that the term "if" can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]".

[0134] The above description is merely illustrative of the application, and not restrictive of the same. Since modifications can be made by those skilled in the art, the scope of the application should be determined by the scope of the claims that follow.

[0135] In the above embodiments, all or part of the steps can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the steps can be implemented in the form of a computer program product. The computer program product includes one or more computer program instructions. When loaded and executed by a computer, the computer program instructions generate the processes or functions in accordance with the embodiments of the present application. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus.

[0136] The computer program instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer program instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired or wireless manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD), or a semiconductor medium (such as a solid state disk) and the like.

[0137] Those skilled in the art can understand that all or part of the steps of the above embodiments can be completed by hardware, or by program instructing relevant hardware, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read only memory, a magnetic disk or an optical disk, etc.

[0138] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit the same. Although the above embodiments of the present application are described in detail, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some technical features can be replaced by equivalent ones. The modifications or replacements do not change the essence of the corresponding technical solutions, and the essence of the corresponding technical solutions is still within the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data migration method, characterized by, The first logical block group of the host includes a plurality of logical blocks, the plurality of logical blocks map storage space of a memory, the plurality of logical blocks include a plurality of stripes, each stripe in the plurality of stripes is composed of stripe units of different logical blocks, wherein each logical block includes a plurality of stripe units, The first logical block group includes valid data, and the method includes: Within the first logical block group, valid data in a first stripe of the first logical block group is moved to a logical address of invalid data outside the first stripe.

2. The method of claim 1, wherein, The logical address of invalid data outside the first stripe is located in a second stripe of the first logical block group, and before the valid data in the first stripe is moved to the logical address of invalid data outside the first stripe, the method further includes: According to the predicted data in the second stripe after the moving, first check data of the second stripe is generated.

3. The method of claim 2, wherein, The method further includes: When a data error occurs in the process of moving the first valid data in the first stripe to the logical address of invalid data in the second stripe, data recovery is performed on the second stripe based on the first check data.

4. The method of claim 2, wherein, After the first valid data in the first stripe is moved to the logical address of invalid data in the second stripe, the method further includes: The second check data of the second stripe is deleted, and the second check data is the check data of the original data of the second stripe.

5. The method according to any one of claims 1 to 4, characterized in that, After the valid data in the first stripe is moved to the logical address of invalid data outside the first stripe, the first stripe does not include valid data.

6. The method of claim 5, wherein, The method further includes: After the valid data in the first stripe is emptied, a trim instruction or an unmap instruction is sent to the memory based on the logical address of invalid data in the first stripe, to instruct the memory that the storage space mapped by the logical address of invalid data in the first stripe stores invalid data.

7. The method of claim 6, wherein, The method further includes: Data is written based on the first stripe.

8. A data migration apparatus, characterized by comprising: The first logical block group of the host includes a plurality of logical blocks, the plurality of logical blocks map storage space of a memory, the plurality of logical blocks include a plurality of stripes, each stripe in the plurality of stripes is composed of stripe units of different logical blocks, wherein each logical block includes a plurality of stripe units, The device includes: A determination module for determining valid data; A moving module for moving valid data in a first stripe of the first logical block group to a logical address of invalid data outside the first stripe within the first logical block group.

9. The apparatus of claim 8, wherein, The logical address of invalid data outside the first stripe is located in a second stripe of the first logical block group, and the device further includes: A generation module for generating first check data of the second stripe according to the predicted data in the second stripe after the moving.

10. The apparatus of claim 9, wherein, The device further includes: A recovery module for performing data recovery on the second stripe based on the first check data when a data error occurs in the process of moving the first valid data in the first stripe to the logical address of invalid data in the second stripe.

11. The apparatus of claim 9, wherein, The device further includes: The deleting module is configured to delete second check data of the second stripe, the second check data being check data of original data of the second stripe.

12. The apparatus of any one of claims 8-11, wherein, After the valid data in the first stripe is moved to the logical address of the invalid data outside the first stripe, the first stripe does not include valid data.

13. The apparatus of claim 12, wherein, The apparatus further includes: The sending module is configured to, after the valid data in the first stripe is emptied, send a trim instruction or an unmap instruction to the memory based on the logical address of the invalid data in the first stripe, to instruct the memory that the storage space mapped by the logical address of the invalid data in the first stripe stores invalid data.

14. The apparatus of claim 13, wherein, The apparatus further includes: The writing module is configured to perform data writing based on the first stripe.

15. A host, characterized by The host further includes a memory.

16. The host of claim 15, wherein, The memory is a flash disk or a phase change memory (PCM).

17. The host of claim 16, wherein, The storage system includes a memory and a host, and the host is capable of executing instructions to perform the data moving method according to any one of claims 1 to 7.

18. A storage system, characterized by The instructions, when executed by the host, cause the host to perform the data moving method according to any one of claims 1 to 7.

19. A computer program product comprising instructions, characterized in that, The instructions, when executed by the host, cause the host to perform the data moving method according to any one of claims 1 to 7.

20. A computer-readable storage medium, characterized in that, ​

Citation Information

Patent Citations

  • Cold and hot data separation method and device, medium and computer program product

    CN113900964A

  • Data processing method and device

    CN115686345A

  • Two-Level Hierarchical Log Structured Array Architecture with Minimized Write Amplification

    US20160179410A1

  • Efficient Implementation of Optimized Host-Based Garbage Collection Strategies Using Xcopy and Arrays of Flash Devices

    US20170242785A1