Migration from a fully allocated volume to a deduplicated volume in a storage system

The method and system for migrating data from fully allocated volumes to deduplicated volumes in storage systems address the inefficiencies of current technologies by utilizing surplus computing power in the drive layer to calculate deduplication hashes, reducing computational and storage requirements.

JP2025519005APending Publication Date: 2025-06-24INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024558259
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-26
Filing Date
2023-05-17
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Current storage systems face challenges in efficiently migrating data from fully allocated volumes to deduplicated volumes, often requiring multiple data copies or reading and writing data to new locations, which increases computational and storage requirements.

Method used

A computer-implemented method and system that migrates data from a fully allocated volume to a deduplicated volume by moving the physical allocation of storage data to a virtual address range within a deduplication domain, setting deduplication metadata as a pass-through, and executing a background deduplication process, thereby calculating deduplication hashes using surplus computing power in the drive layer without reading data outside the physical drive.

Benefits of technology

This approach reduces computational requirements for volume migration by leveraging excess computing power in the drive layer, eliminates the need to read and write data to new locations, and simplifies storage management by avoiding data duplication within the pool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025519005000001_ABST
    Figure 2025519005000001_ABST
Patent Text Reader

Abstract

Method, computer program product, and computer system for migrating from a fully allocated volume to a deduplicated volume in a storage system. The method includes moving a physical allocation of stored data associated with a fully allocated volume to a virtual address range within a deduplication domain and setting deduplication metadata as a pass-through. The method then performs a background deduplication process on the virtual address range into which the physical allocation was placed and performs hashing on the physical drives on which the data is stored using a drive query hash interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a storage system, and more particularly to the migration from a fully allocated volume to a deduplicated volume.

Background Art

[0002] A common feature found in today's storage controllers and appliances is data deduplication. There is an increasing need to store more user data in the same physical capacity, thereby reducing the overall cost of ownership of the storage system. Data deduplication functions by identifying repeated data patterns and creating references to duplicate data stored elsewhere in the system instead of storing the user data. Existing duplicates may be within the same volume, within another volume in the same pool or a different pool within the storage system, or within a volume used by another host.

[0003] A fully allocated volume has all its capacity pre-allocated, and the storage system has minimal interaction with the data except for caching and RAID operations. A customer may have both fully allocated volumes and deduplicated volumes within the same storage pool.

[0004] It is common for a customer to have both deduplicated volumes and fully allocated volumes mixed within a pool. Deduplication is a new technology, and thus the migration path from full allocation to deduplication is both beneficial and needed.

[0005] Current state-of-the-art technologies involve creating multiple copies of the data to perform the migration, or reading all the data and writing it to a new location within the pool.

Summary of the Invention

[0006] According to an embodiment of the present invention, there is provided a computer-implemented method for migrating from a fully allocated volume to a deduplicated volume in a storage system. The method is executed by one or more processors of a computer system and includes moving the physical allocation of storage data associated with the fully allocated volume to a virtual address range within a deduplication domain and setting the deduplication metadata as a pass-through. A background deduplication process is executed on the virtual address range populated with the physical allocation, and hash calculation is performed on the physical drive where the data is stored using a drive query hash interface. The virtual address range becomes a deduplicated volume of the data.

[0007] This provides an improvement in that deduplication hashes can be calculated using the surplus computing power within the drive layer. This eliminates the need to read data outside the physical drive, reducing the computational requirements for volume migration.

[0008] The method may include not moving the data stored on the physical drive during the execution of the method. This provides an improvement in that there is no need to write the data to a new location during volume migration.

[0009] The method may include pausing input / output operations while moving the physical allocation of the storage data associated with the fully allocated volume to a virtual address range within the deduplication domain. This provides an improvement in that conflicts with the underlying data can be avoided during the movement of the physical allocation.

[0010] The method may include reserving, by the deduplication layer, a virtual address range within the deduplication domain that will be the destination of the physical allocation associated with the fully allocated volume.

[0011] Setting the deduplication metadata as a pass-through may include creating a pass-through forward search structure for the assimilation from a fully allocated volume to a deduplicated volume, and creating a stub reverse search structure that is updated with the deduplication metadata when the deduplicated volume is deduplicated.

[0012] Moving the physical allocation of the storage data associated with the fully allocated volume to a virtual address range within the deduplication domain may be performed by the virtualization layer and may include marking the volume as a deduplicated volume and activating the deduplication layer and the pass-through metadata.

[0013] Executing a background deduplication process may include iterating over the grains of the deduplicated volume by the deduplication layer and locking the grains to input / output operations during grain deduplication.

[0014] Iterating over the grains may include the following. If the grain is a deduplication hit, create a referrer metadata that points to the source metadata on the pass-through forward search structure and mark the grain as invalid. If the grain is a deduplication miss, a hash for the grain is stored in the deduplication fingerprint database, and source metadata that points to the grain on the data disk may be created on the pass-through forward search structure. For all grains, the reverse search structure may be updated by invalidation or by the virtual address.

[0015] The method may include performing hash calculation at the FCM on the drive layer without reading data from outside the flash core module (FCM).

[0016] According to another embodiment of the present invention, there is provided a system for migrating from a fully allocated volume to a deduplicated volume in a storage system, including a processor and a memory configured to provide computer program instructions for executing the functions of the components to the processor. The system may include a physical allocation migration component in a virtualization layer for moving the physical allocation of storage data associated with the fully allocated volume to a virtual address range within a deduplication domain. A deduplication metadata setting component is provided for setting the deduplication metadata as a pass-through. A deduplication component of the deduplication layer is provided for executing a background deduplication process on the virtual address range into which the physical allocation is loaded and performing hash calculation on the physical drive where the data is stored using a drive query hash interface. The virtual address range becomes a deduplicated volume of the data.

[0017] Thereby, a system is provided that is improved to calculate deduplication hashes using the excess computing power in the drive layer. Thereby, the computational requirements for volume migration are reduced.

[0018] The physical allocation of the storage data may be a managed disk of a data extent associated with a physical data extent stored across a plurality of physical drives having extra computing power for performing hash calculation of the deduplication process. The plurality of physical drives may be a rush core module (FCM) having hash calculation ability.

[0019] The system may include a suspension component for temporarily suspending input / output operations while moving the physical allocation of the storage data associated with the fully allocated volume to a virtual address range within the deduplication domain.

[0020] The system may include a reservation component within the deduplication layer for reserving a virtual address range within a deduplication domain that is the destination of a physical allocation associated with a fully allocated volume.

[0021] The deduplication metadata setting component may include a forward search component for creating a pass-through forward search structure for the assimilation of a fully allocated volume to a deduplicated volume. A reverse search component for creating a stub reverse search structure that is updated with deduplication metadata when a deduplicated volume is deduplicated may also be included.

[0022] The physical allocation movement component may include marking the volume as a deduplicated volume and activating the deduplication layer and pass-through metadata.

[0023] The deduplication component may include a grain iteration component for iterating over the grains of a deduplication volume by the deduplication layer and locking the grains to input / output operations during grain deduplication.

[0024] The grain iteration component may include a hit component for creating referral metadata that points to source metadata on the pass-through forward search structure and marking the grain as invalid if the grain is a deduplication hit. If the grain is a deduplication miss, a miss component for storing a hash of the grain in a deduplication fingerprint database and creating source metadata that points to the grain on the data disk on the pass-through forward search structure may also be included. A reverse search update component for updating the reverse search structure by invalidation or virtual address may also be included.

[0025] According to another embodiment of the present invention, there is provided a computer program product for migrating from a fully allocated volume to a deduplicated volume in a storage system. The computer program product includes a computer-readable storage medium having program instructions embodied thereon. The program instructions are executable by a processor and cause the processor to move the physical allocation of storage data associated with the fully allocated volume to a virtual address range within a deduplication domain and set the deduplication metadata as a pass-through. A background deduplication process is executed on the virtual address range into which the physical allocation is loaded, and a hash calculation is performed on the physical drive where the data is stored using a drive query hash interface. The virtual address range becomes a deduplicated volume of the data.

[0026] The computer-readable storage medium can be a non-transitory computer-readable storage medium, and the computer-readable program code may be executable by a processing circuit.

[0027] Here, exemplary embodiments of the present invention will be described by way of example only with reference to the accompanying drawings.

Brief Description of the Drawings

[0028]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0029] It should be understood that, for the sake of brevity and clarity of description, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated compared to other elements for clarity. Further, reference numerals may be repeated between figures to indicate corresponding or similar features where appropriate.

[0030] Exemplary embodiments of the present invention are provided that include a method, a system, and a computer program product for volume migration from fully allocated provisioning in a storage pool to a deduplicated volume. However, it should be understood that the scope of the concepts of the present invention is not limited thereto. The disclosed exemplary embodiments are merely examples of the claimed systems, methods, and computer program products. The concepts of the present invention may be embodied in many different forms and should not be construed as limited to only the exemplary embodiments shown herein. Rather, these exemplary embodiments are provided for the sake of completeness of the disclosure and to facilitate understanding by those skilled in the art. In the detailed description, descriptions of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented exemplary embodiments.

[0031] References to "one embodiment", "an embodiment", "exemplary embodiment", etc. in this specification indicate that the described embodiment may include a particular feature, structure, or characteristic, but not necessarily all embodiments include that feature, structure, or characteristic. Also, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly described.

[0032] To avoid obscuring the presentation of exemplary embodiments of the concepts of the present invention, in the following detailed description, some process steps or operations known in the art may be combined for purposes of presentation and explanation, and in some cases may not be described in detail. Further, some process steps or operations known in the art may not be described at all. The following detailed description focuses on the specific features or elements of the concepts of the present invention according to various exemplary embodiments.

[0033] A deduplication migration coordinator component may be provided to handle the in-pool volume migration from a fully allocated volume to a deduplicated volume. The migration is optimized to calculate the deduplication hash of the logical address range of the drive layer itself using the surplus computing power within the drive layer. By calculating the hash value of the deduplication migration on the drive layer, the system computing requirements for the migration are reduced, and there is no need to read data on the disk from outside the drive itself.

[0034] By the described migration method, there is no need to write data to a new location during migration. This migration process requires no additional capacity or "swing space", and the existing fully allocated capacity is quickly assimilated into the deduplication domain, thus simplifying storage management during the migration process.

[0035] Migration to a deduplicated volume is an improvement in the technical field of computer storage systems and provides a more efficient storage solution.

[0036] Physical storage drives are increasingly including extra computing power to perform advanced functions at the drive layer. For example, a flash core module (FCM) has the ability to encrypt or compress data. FCM drives typically do not fully utilize their computing power to provide normal functions. In addition to this, the FCM uses a field programmable gate array (FPGA), which is a very adaptable computing unit that can be utilized to efficiently perform hash calculations (e.g., SHA-1 calculations). The described methods and systems utilize this surplus computing power in a storage environment.

[0037] Deduplication is a technique for determining whether a grain of data already exists on a storage controller, and instead of rewriting the data again, the metadata structure is updated to point to the owner of the existing storage portion and increment the reference count of that entry. The reference is typically achieved by obtaining an incoming IO, hashing that IO using an algorithm such as SHA-1, and comparing the hash fragment to an existing IO database. If there is a deduplication hit, the storage effectively writes a small amount of metadata instead of the data itself, thus resulting in a capacity savings for the user.

[0038] Mixing both deduplicated volumes and fully allocated volumes within a pool is common for storage users. Deduplication is a new technology, and thus a migration path from full allocation to deduplication is both beneficial and needed.

[0039] Referring to FIG. 1, flowchart 100 shows an exemplary embodiment of the described method for migrating from a fully allocated volume in a storage pool to a deduplicated volume.

[0040] An initial stage may be performed by the deduplication layer to reserve the data disk virtual address range of the allocated physical extents currently associated with the fully allocated volume to be migrated (step 101). The reserved data disk virtual address range may be the address range at this stage of the virtual area of the data disk that is not currently associated with any physical disk extent, not currently in use, or both. The migration process may move the allocated physical disk extents associated with the fully allocated volume to this virtual area. Since the reserved virtual area is within the deduplication domain, this reserved virtual area becomes a deduplicated volume.

[0041] The initial stage of the method may set up a pass-through forward search structure and set the deduplication metadata as a pass-through for assimilation of the fully allocated volume (step 102). A desirable property of the forward search structure is that if a read attempt on the structure does not find an entry, the original fully allocated volume needs to be read. This is intended by the "pass-through" where an empty structure results in a read on the fully allocated volume. The pass-through forward search structure may map each virtual address within the fully allocated volume to a virtual address within the reserved virtual address range within the deduplication domain. Also, a stub reverse search structure may be created to update the deduplication metadata (e.g., "on the fly") during the migration process.

[0042] While pausing the I / O to the volume, the physical allocation of the storage data associated with the fully allocated volume may be moved within the reserved data disk virtual address range in the deduplication domain (step 103). This may be executed in the virtualization layer, and the I / O to the volume may be paused while the physical allocation is being moved.

[0043] When the physical allocation extent of the volume is mapped to the reserved virtual area in the deduplication domain, the ownership of this data is passed to the deduplication control, and the data disk can become the owner of the data.

[0044] As part of the movement, the method may mark the volume as a deduplicated volume, and activate the deduplication layer and the pass-through forward search structure for new I / O operations (step 104). During the movement of the physical allocation and while the deduplication metadata is being set, the I / O may be paused so as not to conflict with the underlying data.

[0045] The method may execute a background deduplication process for the grains of the newly introduced virtual address range (step 105).

[0046] The deduplication process may request that the hash calculation be performed by the physical drive storing the data using the drive query hash interface (step 106). The background deduplication may iterate over the grains of the volume while executing the query all the way to the drive layer for the hash calculation.

[0047] The method updates (e.g., on-the-fly) the reverse search structure by the invalidation or virtual address and provides the deduplication metadata (step 107).

[0048] Thus, the method may use data disk virtual reservation, virtual mapping operations, drive layer hash calculation offloading, or combinations thereof. The initial assimilation of a fully allocated volume may be performed via a pass-through forward search structure, and the reverse search structure is updated on-the-fly.

[0049] This method enables the transfer of physical storage allocation from a fully allocated volume to a deduplicated volume without reading the underlying data or writing that data to a new location on the physical storage within the pool. Any data that is still needed remains on the physical drive where that data already exists.

[0050] Referring to FIG. 2, schematic block diagram 200 shows an exemplary embodiment of a storage system 210 in which the described method and system may be implemented. To illustrate the described method, storage system 210 shows a virtualization layer 220, a deduplication layer 250, and a drive layer 280. The arrangement of the components of the system is exemplary and should not be considered as limiting possible embodiments.

[0051] Storage system 210 may provide storage for host applications in one or more host servers having a storage interface through which input / output (IO) operations for reading and writing data to and from storage system 210 may be processed.

[0052] The storage system 210 may include a virtualized storage controller that provides a virtualization layer 220. The virtualization layer 220 may receive host IO operations from a host application for a logical volume of a data disk having a logical address of a data block. The virtualization layer 220 may present a logical address to the host and maintain logical address metadata 221 that provides a mapping between the logical address and the physical address for providing virtualization of the storage system 210. The logical address may be provided as a volume presented to the connected host.

[0053] The described method may include a deduplication migration adjustment component 230 in the virtualization layer 220 that moves a managed disk physical allocation 241 of a backend data extent associated with a fully allocated volume 222 to a reserved data disk virtual address range of a new deduplication volume 224.

[0054] The virtualization layer 220 may maintain a log structured array (LSA) structure used to describe the logical to physical layout of block devices within the storage system. The LSA structure may provide an easy implementation method for many different data reduction techniques and may be used in a storage system regardless of the type of storage backend. The LSA storage system may refer to the physical address of the storage backend using the logical block address designation of the logical block address (LBA) within the virtual domain. The host application may only need to provide the LBA without recognizing the physical backend in some cases.

[0055] The storage system 210 may include a backend storage controller 240. The backend storage controller 240 may include another deduplication migration adjustment component 260 that provides a deduplication layer 250 for deduplication of data stored in the physical drive layer 280. Deduplication may be used to remove duplicates from stored data chunks by indicating the chunks stored using deduplication metadata 251. The deduplication process may start by applying a hash function to create a unique digital fingerprint or signature of a given data chunk. This fingerprint value is stored in an indexed fingerprint database 252 and may be compared with the fingerprint value created for newly arriving data chunks. By comparing the fingerprint values, it may be determined whether the data chunk is unique or a duplicate of a data block already stored.

[0056] The deduplication metadata 251 may provide a mapping to the addresses of stored data blocks. The deduplication metadata 251 may include a forward search structure that describes a virtual-to-physical mapping, such as using a B-tree. A source chunk may be a chunk of data that stores the original copy of the data, and a referrer chunk may be a symbolic link to the data of the source chunk. The source chunk may include a count of the number of referrers that refer back to the source. The source chunk may know how many chunks are referring to it, but may not know which chunks are referring to it.

[0057] The backend storage system controller 240 may also include various functions, such as a garbage collection component for garbage collection of storage extents within the physical drive layer 280.

[0058] The described deduplication layer 250 may include another deduplication migration adjustment component 260 that may include an address range reservation component 261 capable of reserving a range of logical addresses for the new deduplication volume 224. The pass-through forward search structure 262 may be used to assimilate the fully allocated volume 222 by creating source metadata indicating grains on the physical data disk and deduplication referral metadata indicating the source metadata. The reverse search 264 may be used to record the invalidation of the migrated volume or the on-the-fly update of the virtual address or both.

[0059] The managed disk physical allocation 241 may be held within the backend storage controller 240 as a logical unit of physical storage that is not visible to the host. The managed disk physical allocation 241 may be allocated to a storage pool and provide extents that can be used by a volume. The managed disk physical allocation 241 may be presented as a single logical disk that can provide blocks of available physical storage without a one-to-one requirement corresponding to a physical drive.

[0060] The drive layer 280 may include storage drives 291-293 of non-volatile storage capable of providing a physical storage pool. The storage drives 291-293 may be in module form and may each have a processor 297-299. The storage drives 291-293 may each include a deduplication hash component 294-296, and the calculation of the deduplication hash for data on the corresponding drive may be offloaded to these components.

[0061] Storage drives 291-293 may have the ability to perform extra calculations and may be requested to calculate hash values in the data area on the drive. In an exemplary embodiment, storage drives 291-293 may be flash core modules (FCMs). The FCM may be a compression drive and may calculate a hash as part of this mechanism. The FCM may already have the ability to calculate a hash, and depending on the hash scheme, range, and size, the hash may already be calculated and stored.

[0062] The described method may provide a fast and efficient mechanism for migrating a fully allocated volume to a deduplicated volume on a physical flash drive array within a pool. The deduplication hash workload may be offloaded to the drive where the data is located. Deduplication hashing can be performed by the CPU itself by executing a SHA-1 calculation or on a PCI-e connected compute board such as Lewisburg.

[0063] Migration can be performed without unnecessarily reading and writing volume data outside the drive layer. Thus, there is no data duplication within the pool as a result of the migration, and no additional storage is required during the migration. The unique source data may not be moved or changed. Customers may immediately notice a reduction in the consumption of pool capacity for volumes with a high deduplication level.

[0064] Referring to FIGS. 3A and 3B, schematic diagrams 300, 310 illustrate an example of an exemplary embodiment of the described method.

[0065] In FIG. 3A, an initial stage of a fully allocated volume 322 presented to a host by a virtualization layer is shown. The fully allocated volume 322 may have an associated physical allocation provided by a managed disk physical allocation 341 within a backend storage. The managed disk physical allocation 341 may present a logical unit of physical storage that provides a block of physical storage available for use in storage drives 391 - 393 within an allocated storage pool 390.

[0066] The method may reserve a virtual address range 323 within a deduplication domain 320, and data of the fully allocated volume 322 is migrated to that virtual address range 323 and then deduplicated (shown in FIG. 3B) to become a deduplicated volume 324. At the stage shown in FIG. 3A, the volume may still be fully allocated and IO may not be associated with the deduplication domain.

[0067] FIG. 3B shows a situation after the virtualization layer has moved the physical allocation of a volume provided by the managed disk physical allocation 341 to a reserved virtual address range 323 within the deduplication domain 320. During the movement of the physical allocation, IO to the volume may be paused.

[0068] The deduplication domain 320 may be provided with a pass - through forward search structure 362 for deduplication metadata of a new deduplicated volume 324 address indicated from the fully allocated volume 322 address. In the deduplication domain 320, a reverse search structure 364 may be updated to point from the managed disk physical allocation 341 to the deduplicated volume 324.

[0069] Next, a background deduplication process may be performed on the new deduplicated volume 324, and the hash calculation is offloaded to the processing capabilities of the storage drives 391-393 where the physical data is stored. The physical data may not move during the migration to the deduplication volume 324.

[0070] Figures 4A and 4B are flowcharts 400, 420 showing exemplary embodiments of the described method of migrating from a fully allocated volume to a deduplicated volume.

[0071] Referring to Figure 4A, the method may receive a user migration request from a fully allocated volume to a deduplicated volume (step 401).

[0072] The method may reserve the data disk LBA range of the managed disk (MDISK) extent currently associated with the fully allocated volume in the deduplication layer (step 402).

[0073] The method may create a pass-through forward search structure in the deduplication layer, for example in the form of a pass-through forward search tree or table, where each LBA indicates the relative LBA of the reserved LBA range of the data disk (step 403). The method may also create a stub reverse search structure, for example in the form of a reverse search tree or table (step 404). At this stage, the volume is still fully allocated and IO may not interact with the deduplication layer.

[0074] The method may pause IO to the volume (step 405). The method may move the extent of the volume to the reserved data disk area in the virtualization layer (step 406).

[0075] The method may mark the volume as a deduplicated volume, and the deduplication layer and the pass-through forward search are activated for the new IO (407). The method may resume the suspension of the IO for the volume (step 408).

[0076] Referring to FIG. 4B, the method may perform deduplication by iterating through the volume in grain units in the deduplication layer (step 421). The method may execute a query up to the drive layer for hash calculation that can be performed by the modified read IO for each grain (step 422). The method may lock the grain against the IO during grain migration in the deduplication layer to avoid host IO contention (step 423).

[0077] The method may send a hash calculation request to the drive layer, and the drive layer may process the deduplication hash calculation of the grain and return the hash result (step 424). The method may pre-transmit the hashing algorithm and the grain size to the module on the drive layer. Each drive has the ability to calculate the hash of the data stored in each drive itself.

[0078] For each grain or the next grain or both (step 425), it is determined whether the hash of the grain exists in the deduplication fingerprint database (decision 426). If the hash exists, the method may create referral metadata on the forward search structure to indicate the source metadata of the hash (step 427). The method may invalidate the grain in the physical drive and mark it as a candidate for garbage collection (428). Here, the grain can be overwritten with new data.

[0079] If a hash does not exist in the deduplication fingerprint database, the grain may be marked as a deduplication miss. The method may store the hash in the deduplication fingerprint database (step 429). The method may create source metadata on a forward search structure to indicate grains on the data disk (step 430). The data may remain on the drive (431).

[0080] Each time a grain is processed, the method may update a reverse search structure by invalidation or virtual address or both (step 432). It is determined whether there is a next grain (decision 433). If so, the method may loop to process the next grain in step 422. When all grains have been analyzed, the volume migration may be complete (step 434).

[0081] Accordingly, migration can be performed without unnecessarily reading and writing volume data outside the drive layer. As a result of the migration, data does not duplicate within the pool and no additional storage is required during migration. Unique source data may not be moved or changed or both. The customer will immediately notice a reduction in the consumption of pool capacity for volumes with a high deduplication level. Knowledge about hierarchy and heat may be retained for source grains.

[0082] Referring to FIG. 5, the block diagram shows an exemplary embodiment of a computing system 500. The computing system 500 may include a deduplication migration adjustment component 520 and an exemplary drive module 560. The drive module 560 may have extra computing power 561 and may include a deduplication migration interface 562 and a hash calculation component 563 for performing hash calculation for deduplication migration on the drive module 560.

[0083] Computing system 500 may include circuitry for performing the functions of the described components, which may be at least one processor 501, a hardware module, or a software unit executed on at least one processor. A plurality of processors operating parallel processing threads may be provided to enable some or all of the parallel processing of the functions of the components. Memory 502 may be configured to provide computer instructions 503 to at least one processor 501 to perform the functions of the components.

[0084] What is described is shown as a single component on computing system 500. However, the deduplication migration component 520 may be provided across multiple computing devices, such as a virtualization layer and a deduplication layer of a storage system.

[0085] The described deduplication migration component 520 may include a deduplication component 530 that includes a reservation component 535 for reserving a virtual address range within a deduplication domain that is the destination of the physical allocation associated with a fully allocated volume.

[0086] The described deduplication migration component 520 may include a physical allocation migration component 521 within the virtualization layer for moving the physical allocation of stored data associated with a fully allocated volume to a virtual address range within a deduplication domain. The physical allocation migration component 521 may include a suspension component 522 that can suspend input / output operations while moving the physical allocation of stored data associated with a fully allocated volume to a virtual address range within a deduplication domain.

[0087] The described deduplication migration component 520 may include a deduplication metadata setting component 523 that can set the deduplication metadata as a pass-through. The deduplication metadata setting component 523 may include a forward search component 524 for creating a pass-through forward search structure for the assimilation from the fully allocated volume to the deduplicated volume, and a reverse search component 525 for creating a stub reverse search structure that is updated with the deduplication metadata when the deduplicated volume is deduplicated.

[0088] The deduplication component 530 of the deduplication layer may perform a background deduplication process on the virtual address range to which the physical allocation is input, and may perform hash calculation on the physical drive where the data is stored using the drive query hash interface 540. The virtual address range becomes the deduplicated volume of the data.

[0089] The deduplication component 530 may include a grain iteration processing component 531 for iterating the grain of the deduplicated volume by the deduplication layer and locking the grain for input / output operations during grain deduplication. The grain iteration processing component 531 includes a hit component 532 for creating a referrer metadata that indicates the source metadata on the pass-through forward search and marking the grain as invalid when the grain is a deduplication hit, a miss component 533 for storing the hash of the grain in the deduplication fingerprint database and creating the source metadata that indicates the grain on the data disk on the pass-through forward search when the grain is a deduplication miss, and a reverse search update component 534 for updating the reverse search by invalidation or virtual address.

[0090] FIG. 6 shows a block diagram of components of a computing system used in a virtualization layer, a deduplication layer, or a physical layer or a combination thereof, according to an exemplary embodiment of the present invention. It should be understood that FIG. 6 merely provides an illustration of one implementation form and does not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications may be made to the illustrated computing system.

[0091] The computing system can include one or more processors 602, one or more computer-readable RAMs 604, one or more computer-readable ROMs 606, one or more computer-readable storage media 608, a device driver 612, a read / write drive or interface 614, and a network adapter or interface 616, all of which are interconnected via a communication fabric 618. The communication fabric 618 can be implemented using any architecture designed to pass data or control information or both between processors (such as microprocessors, communication and network processors, etc.), system memory, peripheral devices, and any other hardware components within the system.

[0092] One or more operating systems 610 and application programs 611, such as the deduplication migration adjustment components 230, 260, are stored in one or more of one or more computer-readable storage media 608 for execution by one or more processors 602 via one or more of their respective RAMs 604 (typically including cache memory). In the illustrated embodiment, each computer-readable storage media 608 can be a magnetic disk storage device of an internal hard drive, a CD-ROM, a DVD, a memory stick, a magnetic tape, a magnetic disk, an optical disk, a semiconductor storage device such as a RAM, a ROM, an EPROM, a flash memory, or any other computer-readable storage media capable of storing computer programs and digital information, according to embodiments of the present invention.

[0093] The computing system can also include an R / W drive or interface 614 for reading and writing to one or more portable computer-readable storage media 626. Application programs 611 on the computing system can be stored on one or more portable computer-readable storage media 626, read via their respective R / W drives or interfaces 614, and loaded into their respective computer-readable storage media 608.

[0094] The computing system may also include a network adapter or interface 616, such as a TCP / IP adapter card or a wireless communication adapter. An application program 611 on the computing system may be downloaded from an external computer or an external storage device to the computing device via a network (e.g., the Internet, a local area network, or other wide area network or wireless network) and the network adapter or interface 616. The program may be loaded from the network adapter or interface 616 to the computer-readable storage medium 608. The network may include copper wires, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0095] The computing system may also include a display screen 620, a keyboard or keypad 622, and a computer mouse or touchpad 624. The device driver 612 interfaces with the display screen 620 for imaging, the keyboard or keypad 622, the computer mouse or touchpad 624, or the display screen 620 for pressure sensing of alphanumeric input and user selection, or a combination thereof. The device driver 612, the R / W drive or interface 614, and the network adapter or interface 616 may include hardware and software stored in the computer-readable storage medium 608 or the ROM 606 or both.

[0096] The present invention may be a system, method, or computer program product integrated at any possible technical detail level, or a combination thereof. The computer program product may include one or more computer-readable storage media having computer-readable program instructions for causing a processor to implement aspects of the present invention.

[0097] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, punch card, or mechanically encoded devices such as raised structures within grooves in which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transitory signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0098] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network or a combination thereof. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0099] Computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk(R), C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to carry out aspects of the present invention.

[0100] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0101] These computer-readable program instructions, when executed via the processor of a computer or other programmable data processing apparatus, may create means for causing the functions / acts specified in one or more blocks of a flowchart and / or block diagram to be implemented, thereby creating a machine, whether the computer or other programmable data processing apparatus is a processor. These computer-readable program instructions may also be stored in a computer-readable storage medium that includes a manufactured article that includes instructions for implementing the aspects of the functions / acts specified in one or more blocks of a flowchart and / or block diagram, such that the computer-readable medium can be stored and can direct a computer, programmable data processing apparatus, or other device or combination thereof to function in a particular manner.

[0102] The computer-readable program instructions may also be loaded onto a computer, other programmable apparatus, or other device to create a computer-implemented process such that the instructions executed on the computer, other programmable apparatus, or other device cause a series of operational steps to be performed to implement the functions / acts specified in one or more blocks of a flowchart and / or block diagram.

[0103] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions shown in the blocks may be performed in an order different from that shown in the figures. For example, two blocks shown in succession may actually be performed as one step, may be executed simultaneously, substantially simultaneously, in a partially or wholly temporally overlapping manner, or the blocks may sometimes be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagram or flowchart diagram, or both, and combinations of blocks in the block diagram or flowchart diagram, or both, can be implemented by a dedicated hardware-based system that performs the specified function or action, or that performs a combination of dedicated hardware and computer instructions.

[0104] Cloud computing Although this disclosure includes a detailed description of cloud computing, it should be understood that the implementations of the teachings described herein are not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented in combination with any other type of computing environment now known or later developed.

[0105] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or service provider interaction. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0106] The characteristics are as follows.

[0107] On-demand self-service: Cloud consumers can unilaterally provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with the service provider.

[0108] Broad network access: Cloud capabilities are available over the network and can be accessed using standard mechanisms, facilitating use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0109] Resource pooling: Provider computing resources are pooled and provided to multiple users using a multi-tenant model, with various physical and virtual resources dynamically assigned and re-assigned according to demand. Users generally have a sense of location independence in that they are usually unaware of and do not manage the exact location of the provided resources, although at a higher level of abstraction, the location (e.g., country, state, or data center) can be specified.

[0110] Rapid adaptability: The capabilities can be provisioned quickly, flexibly, and in some cases automatically, scale out rapidly, be released quickly, and scale in rapidly. The capabilities available for provisioning often appear to the user as being able to purchase any amount, without limit, at any time.

[0111] Measured services: The cloud system automatically controls and optimizes resource usage by leveraging a metering function at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). The amount of resource usage can be monitored, controlled, and reported, providing transparency to both the provider and the user of the services being utilized.

[0112] The service model is as follows.

[0113] SaaS (Software as a Service): The capabilities provided to the user are the use of the provider's applications running on the cloud infrastructure. Those applications can be accessed from various client devices via a thin-client interface such as a web browser (e.g., web-based email). The user does not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, storage, or individual application functions, except for limited user-specific application configuration settings.

[0114] PaaS (Platform as a Service): The capabilities provided to users are to deploy applications created or obtained by users, which are created using programming languages and tools supported by the provider, onto the cloud infrastructure. Users do not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, or storage, but can control the deployed applications and, in some cases, the configuration of the application hosting environment.

[0115] IaaS (Infrastructure as a Service): The capabilities provided to users are the provisioning of processing, storage, network, and other basic computing resources, and users can deploy and run any software that can include operating systems and applications. Users do not manage or control the underlying cloud infrastructure, but can control the operating systems, storage, and deployed applications, and in some cases, can limitedly control selected network components (such as host firewalls).

[0116] The deployment models are as follows.

[0117] Private cloud: This cloud infrastructure is operated only for an organization. This cloud infrastructure may be managed by the organization or a third party and may exist on-premises or off-premises.

[0118] Community Cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community that shares concerns (e.g., mission, security requirements, policies, and compliance considerations). This cloud infrastructure may be managed by an organization or a third party and may exist on-premises or off-premises.

[0119] Public Cloud: The cloud infrastructure is made available to the general public or a large industry group and is owned by an organization that sells cloud services.

[0120] Hybrid Cloud: The cloud infrastructure remains a unique entity but is a composite of two or more clouds (private, community, or public) joined by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.

[0121] The cloud computing environment is a service-oriented environment that emphasizes statelessness, loose coupling, modularity, and semantic interoperability. At the center of cloud computing is an infrastructure with a network of interconnected nodes.

[0122] Referring now to FIG. 7, an exemplary cloud computing environment 50 is shown. As illustrated, cloud computing environment 50 includes one or more cloud computing nodes 10 that may serve as a communication partner for local computing devices used by cloud consumers such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or an automotive computer system 54N or combinations thereof. Nodes 10 may communicate with each other. These nodes may be physically or virtually grouped within one or more networks such as a private cloud, community cloud, public cloud, or hybrid cloud as described above in this specification or combinations thereof (not shown). Thereby, cloud computing environment 50 can provide an infrastructure, platform, or SaaS, or combinations thereof that a cloud consumer need not maintain resources on a local computing device. The types of computing devices 54A - 54N shown in FIG. 6 are for illustrative purposes only, and it should be understood that cloud computing nodes 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network or network addressable connection or both (e.g., using a web browser).

[0123] Referring now to FIG. 8, a set of functional abstraction layers provided by cloud computing environment 50 (FIG. 7) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 8 are for illustrative purposes only and embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.

[0124] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.

[0125] The virtualization layer 70 provides an abstraction layer that can provide virtual entities such as virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual clients 75.

[0126] In one example, the management layer 80 may provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources utilized to execute tasks within a cloud computing environment. Measurement and pricing 82 provides for cost tracking when resources are utilized within a cloud computing environment, and invoicing or charging or billing for the utilization of those resources. In one example, those resources may include application software licenses. Security provides for cloud user and task identity verification, as well as protection of data and other resources. User portal 83 provides access to the cloud computing environment to users and system administrators. Service level management 84 performs the allocation and management of cloud computing resources so as to meet the required service levels. Planning and fulfillment of service level agreements (SLAs) 85 provides for the pre-arrangement and procurement of cloud computing resources expected to be required in the future in accordance with the SLA.

[0127] The workload layer 90 provides examples of functions that a cloud computing environment can utilize. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and life cycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and storage deduplication processing 96.

[0128] The computer program product of the present invention has one or more computer-readable hardware storage devices in which computer-readable program code is stored, and the program code is executable by one or more processors to perform the method of the present invention.

[0129] The computer system of the present invention includes one or more processors, one or more memories, and one or more computer-readable hardware storage devices, and the one or more hardware storage devices include program code executable by the one or more processors via the one or more memories to execute the method of the present invention.

[0130] The description of various embodiments of the present invention is presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or technical improvements found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

[0131] Modifications may be made to the foregoing without departing from the scope of the invention.

Claims

1. A method for migrating from a fully allocated volume to a deduplicated volume, comprising: moving the physical allocation of the stored data associated with the fully allocated volume to a virtual address range within a deduplication domain and setting the deduplication metadata as a pass-through; executing a background deduplication process on the virtual address range into which the physical allocation is loaded, and performing a hash calculation on the physical drive storing the data using a drive query hash interface, wherein the background deduplication process is executed such that the virtual address range becomes the deduplicated volume of the data.

2. The method according to claim 1, further comprising not moving the data stored on the physical drive during the execution of the method.

3. The method according to claim 1 or 2, wherein the physical allocation of the stored data is a managed disk of a data extent associated with a physical data extent stored across a plurality of physical drives having extra computing power for performing the hash calculation of the deduplication process.

4. The method according to any one of claims 1 to 3, further comprising temporarily suspending input / output operations while moving the physical allocation of the stored data associated with the fully allocated volume to a virtual address range within a deduplication domain.

5. The method according to any one of claims 1 to 4, further comprising reserving, by a deduplication layer, the virtual address range within the deduplication domain that is the destination of the physical allocation associated with the fully allocated volume.

6. Setting the deduplication metadata as a pass-through includes: creating a pass-through forward search structure for assimilation from the fully allocated volume to the deduplicated volume; and creating a stub reverse search structure that is updated with deduplication metadata when the deduplicated volume is deduplicated. The method according to any one of claims 1 to 5.

7. Moving the physical allocation of the memory data associated with the fully allocated volume to a virtual address range within the deduplication domain is performed by a virtualization layer, marking the volume as a deduplicated volume, and activating a deduplication layer and pass-through metadata, the method according to any one of claims 1 to 6.

8. Executing a background deduplication process includes iteratively processing the grains of the deduplicated volume by a deduplication layer and locking the grains against input / output operations during grain deduplication, the method according to any one of claims 1 to 7.

9. Iteratively processing the grains includes when the grain is a deduplication hit, creating referral metadata that points to source metadata on a pass-through forward search structure and marking the grain as invalid; when the grain is a deduplication miss, storing a hash of the grain in a deduplication fingerprint database and creating source metadata that points to the grain on the data disk on the pass-through forward search structure; and for all grains, updating a reverse search structure by invalidation or virtual address, the method according to claim 8.

10. The hash calculation is performed on the FCM on the drive layer without reading data from outside the flash core module (FCM), the method according to any one of claims 1 to 9.

11. A system for migrating from a fully allocated volume to a deduplicated volume in a storage system, including a processor and a memory configured to provide computer program instructions for executing the functions of the components to the processor, a physical allocation movement component within a virtualization layer for moving the physical allocation of the memory data associated with the fully allocated volume to a virtual address range within the deduplication domain; a deduplication metadata setting component for setting deduplication metadata as pass-through; A system comprising a deduplication component of a deduplication layer for performing a background deduplication process on the virtual address range to which the physical allocation has been input, and for performing a hash calculation on a physical drive in which the data is stored using a drive query hash interface, wherein the virtual address range becomes a deduplicated volume of the data.

12. The system according to claim 11, wherein the physical allocation of the stored data is a managed disk of a data extent associated with a physical data extent stored across a plurality of physical drives having extra computing power on which the hash calculation of the deduplication process is performed.

13. The system according to claim 11 or 12, comprising a suspension component for temporarily suspending input / output operations while moving the physical allocation of the stored data associated with the fully allocated volume to a virtual address range within a deduplication domain.

14. The system according to any one of claims 11 to 13, comprising a reservation component within the deduplication layer for reserving the virtual address range within the deduplication domain that is the destination of the physical allocation associated with the fully allocated volume.

15. The deduplication metadata setting component comprises a forward search component for creating a pass-through forward search structure for assimilation from the fully allocated volume to the deduplicated volume, and a reverse search component for creating a stub reverse search structure that is updated with the deduplication metadata when the deduplicated volume is deduplicated, the system according to any one of claims 11 to 14.

16. The system according to any one of claims 11 to 15, wherein the physical allocation movement component marks the volume as a deduplicated volume and activates the deduplication layer and the pass-through metadata.

17. The system according to any one of claims 11 to 16, wherein the duplicate elimination component includes a grain iteration processing component for iteratively processing the grains of the duplicate elimination volume by the duplicate elimination layer and locking the grains against input / output operations during grain duplicate elimination.

18. The grain iteration processing component When a grain is a duplicate elimination hit, a referrer metadata indicating source metadata on a pass-through forward search structure is created, and a hit component for marking the grain as invalid; When a grain is a duplicate elimination miss, a hash for the grain is stored in a duplicate elimination fingerprint database, and a miss component for creating source metadata indicating the grain on the data disk on the pass-through forward search structure; The system according to claim 17, further comprising a reverse search update component for updating the reverse search structure by invalidation or virtual address.

19. The system according to any one of claims 11 to 18, wherein the plurality of physical drives includes a flash core module (FCM).

20. A computer program product for migrating from a fully allocated volume to a deduplicated volume in a storage system, the computer program product including a computer-readable storage medium having program instructions embodied thereon, the program instructions being executable by a processor, and the processor moving the physical allocation of storage data associated with the fully allocated volume to a virtual address range within a deduplication domain and setting deduplication metadata as pass-through; executing a background deduplication process on the virtual address range into which the physical allocation is inserted, and performing a hash calculation on the physical drive storing the data using a drive query hash interface, wherein the virtual address range becomes the deduplicated volume of the data, and executing the background deduplication process.

21. A computer program comprising program code means adapted to perform the method according to any one of claims 1 to 10 when the computer program is executed on a computer.