Log for data cloning operation
By loading the logs into memory in the data deduplication storage system without loading the index and performing cloning operations, the problem of large consumption of cloning operations in the existing technology is solved and the system performance is improved.
Patent Information
- Application Number
- CN202111265821.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-08
- Filing Date
- 2021-10-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-28
AI Technical Summary
When performing cloning operations, existing data deduplication storage systems need to load the entire index into memory, resulting in large consumption of system processing time and bandwidth, affecting performance.
Perform cloning by loading the log into memory without loading the associated index. The log stores only information indicating changes to the data stored in the index, and uses the cloned data structure to accumulate these changes until the event is triggered, and then used to update the index.
It significantly reduces the processing time and bandwidth required for cloning operations, and improves the performance of the data deduplication storage system.
Smart Images

Figure CN115454323B_ABST
Abstract
Description
Background Art
[0001] Data reduction techniques can be applied to reduce the amount of data stored in a storage system. Example data reduction techniques include data deduplication. Data deduplication identifies duplicate data units and seeks to reduce or eliminate the number of instances of duplicate data units stored in the storage system. Brief Description of the Drawings
[0002] Some embodiments are described with reference to the following drawings.
[0003] Figure 1 is a schematic diagram of an example system according to some embodiments.
[0004] Figure 2 is an illustration of an example data structure according to some embodiments.
[0005] Figures 3A to 3C is an illustration of an example data structure according to some embodiments.
[0006] Figure 4 is an illustration of an example process according to some embodiments.
[0007] Figure 5 is an illustration of an example process according to some embodiments.
[0008] Figure 6 is a diagram of an example machine-readable medium storing instructions according to some embodiments.
[0009] Figure 7 is a schematic diagram of an example computing device according to some embodiments.
[0010] In all the drawings, the same reference numerals refer to similar but not necessarily identical elements. The drawings are not necessarily to scale, and the dimensions of some parts may be enlarged to more clearly illustrate the examples shown. Additionally, the drawings provide examples and / or embodiments consistent with the description; however, the description is not limited to the examples and / or embodiments provided in the drawings. Detailed Description
[0011] In this disclosure, unless the context clearly indicates otherwise, the terms "a," "an," or "the" are intended to include the plural forms as well. Similarly, as used in this disclosure, the terms "includes," "including," "comprises," "comprising," or "having" specify the presence of the stated element, but do not preclude the presence or addition of other elements.
[0012] In some examples, a storage system can deduplicate data to reduce the amount of space required to store the data. The storage system can perform a data deduplication process that includes breaking a data stream into discrete data units or "chunks." Further, the storage system can determine an identifier or "fingerprint" of an incoming data unit and can determine which incoming data units are duplicates of previously stored data units. In the case where a data unit is a duplicate, the storage system can store a reference to the previous data unit instead of storing the duplicate incoming data unit.
[0013] As used herein, a "fingerprint" refers to a value obtained by applying a function to the content of a data unit (where "content" can include all or a subset of the content of the data unit). Examples of functions that can be applied include hash functions that produce a hash value based on the incoming data unit. Examples of hash functions include cryptographic hash functions such as the Secure Hash Algorithm 2 (SHA-2) hash functions (e.g., SHA-224, SHA-256, SHA-384, etc.). In other examples, other types of hash functions or other types of fingerprint functions can be employed.
[0014] A "storage system" can include a storage device or an array of storage devices. The storage system can also include one or more storage controllers that manage access to the (multiple) storage devices. A "data unit" can refer to any portion of data that can be individually identified within the storage system. In some cases, a data unit can refer to a chunk, a collection of chunks, or any other portion of data. In some examples, the storage system can store data units in a persistent storage. The persistent storage can be implemented using one or more (multiple) persistent (e.g., non-volatile) storage devices such as (multiple) disk-based storage devices (e.g., (multiple) hard disk drives (HDDs)), (multiple) solid-state devices (SSDs) (such as (multiple) flash storage devices), etc. or a combination thereof.
[0015] "Controller" can refer to a hardware processing circuit, which can include any one or some combination of a microprocessor, a core of a multi-core microprocessor, a microcontroller, a programmable integrated circuit, a programmable gate array, a digital signal processor, or other hardware processing circuits. Alternatively, "controller" can refer to a combination of a hardware processing circuit and machine-readable instructions (software and / or firmware) executable on the hardware processing circuit.
[0016] In some examples, a deduplication storage system can use stored metadata to process an original data stream and reconstruct the original data stream from stored data units. In this way, the deduplication process can avoid storing duplicate copies of duplicate data units, thereby reducing the amount of space required to store the data stream. In some examples, the deduplication metadata can include a data recipe (also referred to herein as a "manifest") that specifies the order of receipt of specific data units (e.g., in a data stream). To retrieve stored data (e.g., in response to a read request or a clone request), the deduplication system can use the manifest to determine the order of receipt of the data units, thereby re-creating the original data stream. The manifest can include a series of records, each record representing a specific set of data units. The records of the manifest can include one or more fields (also referred to herein as "pointer information") that identify an index of storage information including the data units. For example, the storage information can include one or more index fields that specify location information (e.g., container, offset, etc.) of the stored data units, compression and / or encryption characteristics of the stored data units, and the like. In some examples, both the manifest and the index can be read in fixed-size addressable portions (e.g., 4KB portions).
[0017] A "cloning" operation can refer to the process of creating a copy of a specific data stream stored in a deduplication storage system. For example, a cloning operation can include loading a source manifest of a specific data stream into memory and identifying a set of indexes based on the source manifest. Each identified index is loaded into memory as a whole unit, decompressed, and deserialized. The indexes can then be updated to indicate the reference count of the cloned data units (i.e., the number of instances in which the data units appear in the manifest), and written from memory to a persistent storage device. A new manifest (i.e., representing the cloned data) can be formed from components of the source manifest. However, in some examples, when the indexes are loaded into memory, the cloning operation may only affect a relatively small portion of the indexes (e.g., ten records out of ten thousand records in the index). Thus, in such examples, loading the entire index into memory may consume more system processing time and bandwidth than loading only the changed portion into memory.
[0018] According to some embodiments of the present disclosure, a data deduplication storage system can perform a cloning operation by loading a journal into memory without loading the associated index. Each journal can store only information indicating changes to data stored in the corresponding index and can thus be relatively smaller than the corresponding index. Further, the journal can include a data structure (referred to herein as a "clone data structure") dedicated to recording metadata changes associated with the cloning operation. The clone data structure can accumulate these changes until a triggering event (e.g., when the journal becomes full) and can be used to update the corresponding index during a single load into memory. Thus, since each journal is smaller than the corresponding index, performing the cloning operation using the journal can consume relatively less processing time and bandwidth compared to using the corresponding index. Accordingly, the disclosed techniques for cloning operations can significantly improve the performance of a data deduplication storage system.
[0019] 1. Example storage system
[0020] Figure 1 An example of a storage system 100 including a storage controller 110, a memory 115, and a persistent storage device 140 is shown according to some embodiments. As shown, the persistent storage device 140 can include any number of manifests 150, indexes 160, and data containers 170. The persistent storage device 140 can include one or more non-transitory storage media such as hard disk drives (HDDs), solid state drives (SSDs), optical disks, etc., or combinations thereof. The memory 115 can be implemented with a semiconductor memory such as random access memory (RAM).
[0021] In some embodiments, the storage system 100 can perform data deduplication on the stored data. For example, the storage controller 110 can partition an input data stream into data units and can store at least one copy of each data unit in a data container 170 (e.g., by appending the data unit to the end of the container). In some examples, each data container 170 can be partitioned into multiple parts (also referred to herein as "entities").
[0022] In one or more embodiments, the storage controller 110 can generate a fingerprint for each data unit. For example, the fingerprint can include a full or partial hash value based on the data unit. To determine whether an incoming data unit is a replica of a stored data unit, the storage controller 110 can compare the fingerprint generated for the incoming data unit with the fingerprint of the stored data unit. If the comparison results in a match, the storage controller 110 can determine that the storage system 100 has already stored a replica of the incoming data unit.
[0023] As Figure 1As shown, the persistent storage device 140 can store the manifest 150, the index 160, the data container 170, and the log group 120. In some embodiments, the storage controller 110 can generate the manifest 150 to record the order of receipt of data units. Further, the manifest 150 can include pointers or other information indicating the index 160 associated with each data unit. In some embodiments, the associated index 160 can indicate the storage location of the data unit. For example, the associated index 160 can include information specifying that the data unit is stored at a particular offset in an entity and that the entity is stored at a particular offset in the data container 170.
[0024] In some embodiments, the storage controller 110 can receive a read request to access the stored data and, in response, can access the manifest 150 to determine the sequence of data units that make up the original data. The storage controller 110 can then use the pointer data included in the manifest 150 to identify the index 160 associated with the data unit. Further, the storage controller 110 can use the information included in the identified index 160 to determine the storage location of the data unit (e.g., data container 170, entity, offset, etc.) and can then read the data unit from the determined location.
[0025] In some embodiments, a log 130 can be associated with each index 160. The log 130 can include information indicating changes to the data stored in the index 160. For example, when a copy of the index 160 present in the memory 115 is modified to reflect a change to the metadata, the change can also be recorded as an entry in the associated log 130. In some embodiments, multiple logs 130 can be grouped into a log group 120 associated with a single file or object stored in the data deduplication system. For example, multiple logs can correspond to an index storing metadata associated with a single file.
[0026] In some embodiments, the storage controller 110 can receive a request to perform a clone operation on the source manifest 150. The clone operation can include creating a new manifest (also referred to as a "clone manifest") that replicates all or a portion of the source manifest 150 at a particular point in time. Thus, the clone manifest can be used to generate a copy of the sequence of data units at a particular point in time. For example, the clone manifest can be used to provide a backup copy of the sequence of data units as of a particular date.
[0027] In one or more embodiments, during a clone operation of a manifest section, storage system 100 loads log 130 associated with the manifest section into memory 115, but does not load index 160 associated with the manifest section into memory 115. In some embodiments, each log 130 may include a clone data structure (not shown in Figure 1 dedicated to recording metadata changes associated with the clone operation for the manifest section. The clone data structure may accumulate these metadata changes and may subsequently be used to update the corresponding index 160 to reflect the same metadata changes. The use of log 130 during a clone operation is further discussed below with reference to Figures 2 to 7 ).
[0028] 2. Example data structure
[0029] Now referring to Figure 2 , a diagram of an example data structure 200 used in data deduplication according to some embodiments is shown. As shown, data structure 200 may include manifest records 210, container index 220, containers 250, and entities 260. In some examples, container index 220 and containers 250 may generally correspond to example embodiments of index 160 and data containers 170 (as shown in Figure 1 ). In some examples, data structure 200 may be generated and / or managed by storage controller 110 (as shown in Figure 1 ). Further, storage controller 110 may use data structure 200 to obtain stored deduplicated data.
[0030] As shown in Figure 2 , in some examples, manifest records 210 may have multiple fields, including an offset field, a container index field, a length field, and a unit address field. In some embodiments, manifest records 210 may represent a range of data units in a run-length reference format. For example, to represent a range starting from a first data unit and lasting for N data units, the unit address field may indicate the arrival number of the first data unit in the range (i.e., the numerical order in which the first data unit was added to the identified container index), and the length field may indicate the number of data units in the range after the first data unit within the container index.
[0031] In some embodiments, each container index 220 may include any number of data unit records 230 and entity records 240. Each data unit record 230 may include various metadata fields, such as a fingerprint (e.g., a hash of the data unit), a unit address, an entity identifier, a unit offset (i.e., the offset of the data unit within the entity), a count value, and a unit length. Further, each entity record 240 may include various metadata fields, such as an entity identifier, an entity offset (i.e., the offset of the entity within the container), a storage length (i.e., the length of the data unit within the entity), a decompressed length, a checksum value, and compression / encryption information (e.g., compression type, encryption type, etc.). In some embodiments, each container 250 may include any number of entities 260, and each entity 260 may include any number of stored data units.
[0032] In some embodiments, each container index 220 may include a version number 235. The version number 235 may indicate the generation or relative age of the metadata in the container index. For example, the version number 235 may be compared to the version number of an associated log ( Figure 2 not shown). If the version number 235 is greater than the version number of the associated log, it may be determined that the container index 220 includes more up-to-date metadata than the associated log.
[0033] 3A. Example data structure during non-clone operation
[0034] Now referring to Figure 3A , an illustration of the memory 115 during a non-cloning operation (e.g., a read operation) is shown. As shown, during a non-cloning operation, the memory 115 may include a plurality of logs 320 in a log group 310, and may also include a plurality of indexes 330. For example, in response to detecting a non-cloning operation, a particular log 320 and an associated index 330 may be loaded into the memory together. In some examples, the log group 310, the logs 320, and the indexes 330 may generally correspond to example embodiments of the log group 120, the logs 130, and the index 160 (as Figure 1 shown).
[0035] In some embodiments, each log 320 may be associated with a corresponding index 330 and may record changes to the metadata stored in the corresponding index 330. Further, for each log group 120, all corresponding indexes 330 may be associated with a single storage object (e.g., a document, a database table, a data file, etc.). For example, all corresponding indexes 330 may include metadata for data units included in a single file stored in a data deduplication system (e.g., Figure 1 the storage system 100 shown).
[0036] In some embodiments, each log 320 may include or be associated with a version number 325. Further, each index 330 may include or be associated with a version number 335. In some embodiments, during a non-clone operation, the version number 325 may be compared with the version number 335 to determine whether the log 320 or the associated index 330 reflects the most recent version of the metadata. For example, if the version number 325 is greater than the version number 335, it may be determined that the change data included in the log 320 reflects a more recent metadata state than the metadata stored in the index 330. If so, the index 330 may be updated to include the changes recorded in the log 320. However, if the version number 325 is less than the version number 335, it may be determined that the change data included in the log 320 reflects an older metadata state than the metadata stored in the index 330. In this case, the log 320 may be cleared without updating the index 330.
[0037] 3B-3C. Example data structure during clone operation
[0038] Now referring to Figures 3B to 3C , an example of a data structure used during a clone operation is shown. As Figure 3A shown, during a clone operation, the memory 115 may include multiple logs 320 in a log group 310, but does not include the corresponding index 330 (as Figure 3A shown). Further, each log 320 may include a clone data structure 327 dedicated to recording metadata changes during the clone operation.
[0039] In some embodiments, performing a clone operation on a source manifest (or a portion thereof) may include identifying a container index 330 associated with the manifest. A log group 310 including the log 320 corresponding to the identified container index 330 is loaded from a persistent storage device into the memory 115, but the container index 330 is not loaded into the memory 115. In some embodiments, the log group 310 is loaded into the memory as a whole unit. Performing the clone operation includes generating a new clone manifest that references the same data units as the data units referenced by the source manifest, and thus causing the reference count of the cloned data units to increase. In some embodiments, the increment may be recorded in the clone data structure 327 of the log 320 loaded into the memory 115.
[0040] Now referring to Figure 3C, an example implementation of the clone data structure 327 is shown. As shown, the clone data structure 327 can have multiple fields, including a manifest identifier field, a cell address field, a length field, and a reference count field. The clone data structure 327 can include multiple records or rows, each record corresponding to a specific range in the clone manifest. In some examples, the cell address field and the length field of each record can identify the corresponding range in a run-length reference format. For example, the cell address field can indicate the arrival number of the first data unit in the range, and the length field can indicate the number of data units after the first data unit in the range. In some implementations, the reference count field can store or otherwise indicate changes to the reference count that occur during the clone operation. For example, when the clone operation copies a specific range, the reference count field can indicate that the reference count of that range has been incremented by one. In some implementations, the reference count field can store a numerical value representing the total change to the reference count that occurs during one or more clone operations.
[0041] In some implementations, in response to a trigger event, the data stored in the clone data structure 327 of the log 320 can be used to update the associated container index 330 (as Figure 3A shown). This update can be referred to as "folding" the clone data structure 327 into the container index 330. For example, when the log 320 is full (e.g., the data stored in the log 320 exceeds a maximum threshold), the container index 330 can be loaded into the memory 115, and the reference count in the container index 330 can be incremented to include the value in the reference count field of the corresponding record of the clone data structure 327. When each record of the clone data structure 327 has been folded into the container index 330, the stored data of the entire clone data structure 327 is cleared (i.e., all rows are deleted).
[0042] In some implementations, when the clone data structure 327 is folded into the container index 330, the container index 330 can also be updated based on changes in the log 320 that are not included in the clone data structure 327 (i.e., metadata changes that are not related to the clone operation). In some implementations, metadata changes that are not related to the clone operation (also referred to as "non-clone updates") can occur in the manner described above with reference to Figure 3A described. For example, a non-clone update can include comparing the version number 325 of the log 320 with the version number 335 of the container index 330, and performing a non-clone update if the version number 325 is greater than the version number 335.
[0043] In some cases, a failure event (e.g., a power outage) may interrupt the folding of the container index 330 into the clone data structure 327 before completion. After such an interruption, it may be difficult or impossible to determine which records of the clone data structure 327 have been folded into the container index 330, and thus the stored data may be lost or corrupted. In some embodiments, the manifest identifier field can be used to recover the stored data after a failure event. For example, when each record of the clone data structure 327 is folded into the container index 330, the manifest identifier field in that record of the clone data structure 327 can be populated with the identifier of the clone manifest (i.e., the manifest generated by the clone operation). It should be noted that the manifest identifier field of a record remains empty until the folding of that record is complete. After a failure event, the manifest identifier field of each record can be analyzed. The presence of the identifier of the clone manifest in the manifest identifier field can indicate that the current record has been folded into the container index and should not be folded again. Further, in some examples, each record that identifies the clone manifest in the manifest identifier field can be cleared from the log 320.
[0044] 4. Example process
[0045] Now referring to Figure 4 , an example process 400 in accordance with some embodiments is shown. In some examples, process 400 can be performed by a storage controller 110 ( Figure 1 shown) during a clone operation. Process 400 can be implemented in a hardware or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions can be stored in a non-transitory computer-readable medium such as an optical storage device, a semiconductor storage device, or a magnetic storage device. The machine-readable instructions can be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. For illustrative purposes, the details of processes 400 - 404 may be described below with reference to Figures 1 to 3C showing an example in accordance with some embodiments. However, other embodiments are possible.
[0046] Block 410 can include detecting a clone operation for a manifest range. For example, referring to Figures 1 to 3C , the storage controller 110 can receive a request to clone a source manifest 150 and, in response, can initiate the execution of a clone operation to create a clone manifest that replicates a portion of the source manifest 150 at a particular point in time.
[0047] Block 420 can include identifying a container index associated with a manifest portion. Block 430 can include loading log groups into a memory, where the log groups include logs associated with the container index. For example, referring to Figures 1 to 3C, the storage controller 110 may identify the container index 160 associated with the source manifest 150 and, thus, may identify the log 320 associated with the identified container index 160. The storage controller 110 may then load the log group 310 including the identified log 320 from the persistent storage device 140 into the memory 115. However, in some embodiments, the storage controller 110 may not load the identified container index 160 into the memory 115.
[0048] Block 440 may include storing an indication of the metadata change in a clone data structure of the log. For example, referring to Figures 1 to 3C , the storage controller 110 may generate a new clone manifest that references the same data units as the data units referenced by the source manifest and may determine a change in the reference count of the cloned data units. These changes may be recorded in the clone data structure 327 of the log 320 loaded into the memory 115 (e.g., in the reference count field of the record associated with the clone range).
[0049] Decision block 450 may include determining whether the log is full. If it is determined that the log is not full ("No" at block 450), then process 400 may be completed. However, if it is determined that the log is full ("Yes" at block 450), then process 400 may continue at block 460, which may include loading the container index into the memory. Block 470 may include updating the container index based on the clone data structure of the log. Block 480 may include clearing the clone data structure of the log. For example, referring to Figures 1 to 3C , the storage controller 110 may continue the cloning operation until it is completed. However, if it is determined that the log 320 loaded into the memory 115 has reached the maximum threshold of the stored data, the associated container index 160 may be loaded from the persistent storage device 140 into the memory 115. Then the data stored in the log 320 may be copied into the container index 160 in the memory 115, and subsequently the data stored in the log 320 may be deleted. After block 480, process 400 may be completed.
[0050] In some examples, blocks 420 - 440 of process 400 may be repeated multiple times during a single cloning operation. For example, blocks 420 - 440 of process 400 may be executed for a sequence of container indexes associated with the range of the manifest being cloned.
[0051] 5. Example process
[0052] Now referring to Figure 5 , an example process 500 is shown in accordance with some embodiments. In some examples, the storage controller 110 may be used ( Figure 1perform process 500 as shown). Process 500 may be implemented in hardware or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions may be stored in a non-transitory computer-readable medium such as an optical storage device, a semiconductor storage device, or a magnetic storage device. The machine-readable instructions may be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. For illustrative purposes, the following may refer to Figures 1 to 3C describe the details of process 500. However, other embodiments are also possible.
[0053] Block 510 may include detecting a cloning operation for a manifest range by a storage controller of a data deduplication storage system. For example, referring to Figures 1 to 3C , the storage controller 110 may receive a request to clone the source manifest 150 and, in response, may initiate the execution of a cloning operation to create a clone manifest that replicates a portion of the source manifest 150 at a specific point in time.
[0054] Block 520 may include loading a log from a persistent storage device into a memory by the storage controller in response to the detected cloning operation, where the log is for storing changes to a container index associated with the manifest range, and where the container index is not loaded into the memory in response to the detected cloning operation. For example, referring to Figures 1 to 3C , the storage controller 110 may identify the container index 160 associated with the source manifest 150 and may identify the log 320 associated with that container index 160. The storage controller 110 may then identify the log group 310 that includes the identified log 320 and may then load the identified log group 310 as a whole from the persistent storage device 140 into the memory 115. However, in response to the cloning operation, the storage controller 110 may not load the container index 160 into the memory 115.
[0055] Block 530 may include updating the log in the memory by the storage controller to include an indication of changes to the metadata of the container index, where the changes are associated with the detected cloning operation. For example, referring to Figures 1 to 3C , the storage controller 110 may generate a new clone manifest that references the same data units as the data units referenced by the source manifest and may determine changes to the reference counts of the cloned data units. These changes may be recorded in the clone data structure 327 of the log 320 loaded into the memory 115 (e.g., in the reference count field of the record associated with the cloning range).
[0056] 6. Example machine-readable medium
[0057] Figure 6Shown is a machine-readable medium 600 storing instructions 610-630 according to some embodiments. The instructions 610-630 may be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, and the like. The machine-readable medium 600 may be a non-transitory storage medium, such as an optical storage medium, a semiconductor storage medium, or a magnetic storage medium.
[0058] Instructions 610 may be executed to detect a cloning operation for a manifest scope. Instructions 620 may be executed to load a log from a persistent storage device into a memory in response to the detected cloning operation, where the log is for storing changes to a container index associated with the manifest scope, and where the container index is not loaded into the memory in response to the detected cloning operation. Instructions 630 may be executed to update the log in the memory to include an indication of changes to metadata of the container index, where the changes are associated with the detected cloning operation.
[0059] 7. Example computing device
[0060] Figure 7 Shown is a schematic diagram of an example computing device 700. In some examples, the computing device 700 may generally correspond to some or all of the storage system 100 (as Figure 1 shown). As shown, the computing device 700 may include a hardware processor 702 and a machine-readable storage device 705 containing instructions 710-730. The machine-readable storage device 705 may be a non-transitory medium. The instructions 710-730 may be executed by the hardware processor 702 or by a processing engine included in the hardware processor 702.
[0061] Instructions 710 may be executed to detect a cloning operation for a manifest scope. Instructions 720 may be executed to load a log from a persistent storage device into a memory in response to the detected cloning operation, where the log is for storing changes to a container index associated with the manifest scope, and where the container index is not loaded into the memory in response to the detected cloning operation. Instructions 730 may be executed to update the log in the memory to include an indication of changes to metadata of the container index, where the changes are associated with the detected cloning operation.
[0062] According to the embodiments described herein, a data deduplication storage system may perform a cloning operation by loading a log into a memory rather than loading an associated index into the memory. Each log may include a cloning data structure dedicated to recording or otherwise indicating metadata changes associated with the cloning operation. The cloning data structure may accumulate these changes until a triggering event, and may be used to update a corresponding index during a single load into the memory. In some examples, performing a cloning operation using a log may consume relatively less processing time and bandwidth compared to that required to load the associated index into the memory. Thus, the techniques disclosed for cloning operations may significantly improve the performance of a data deduplication storage system.
[0063] Note that although Figures 1 to 7 various examples are shown, embodiments are not limited in this regard. For example, referring to Figure 1 , it is contemplated that storage system 100 may include additional devices and / or components, fewer components, different components, different arrangements, etc. In another example, it is contemplated that the functionality of storage controller 110 described above may be included in any other engine or software of storage system 100. Other combinations and / or variations are possible.
[0064] Data and instructions are stored in respective storage devices implemented as one or more computer-readable or machine-readable storage media. The storage media include different forms of non-transitory memory, including: semiconductor memory devices such as dynamic random access memory or static random access memory (DRAM or SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory; magnetic disks such as fixed floppy disks and removable disks; other magnetic media including magnetic tape; optical media such as compact discs (CDs) or digital video discs (DVDs); or other types of storage devices.
[0065] Note that the instructions discussed above may be provided on one computer-readable or machine-readable storage media, or alternatively, may be provided on multiple computer-readable or machine-readable storage media distributed in a large system having potentially multiple nodes. Such one or more computer-readable or machine-readable storage media are considered to be part of an article (or article of manufacture). An article or article of manufacture may refer to any single manufactured component or multiple components. One or more storage media may be located within a machine running machine-readable instructions, or at a remote site from which the machine-readable instructions may be downloaded over a network for execution.
[0066] In the foregoing description, numerous specific details are set forth to provide a thorough understanding of the subject matter disclosed herein. However, implementations may be practiced without some of these specific details. Other implementations may include modifications and variations of the details discussed above. The appended claims are intended to cover such modifications and variations.
Claims
1. A computer-implemented method, comprising: detecting, by a storage controller of a deduplication storage system, a clone operation for a manifest range; loading a log from a persistent storage device into a memory in response to the detected clone operation, wherein the log is for storing changes to a container index associated with the manifest range, and wherein the container index is not loaded into the memory in response to the detected clone operation; and updating the log in the memory to include an indication of changes to metadata of a container index not loaded into the memory, wherein the changes to the metadata are associated with the detected clone operation.
2. The computer-implemented method according to claim 1, comprising: detecting a non-clone operation for a second manifest range; loading, in response to the detected non-clone operation, a second log and a second container index together from the persistent storage device into the memory; and updating the second log in the memory to store second changes associated with the second container index, wherein the second changes are associated with the non-clone operation.
3. The computer-implemented method according to claim 1, wherein loading the log into the memory comprises: identifying the container index associated with the manifest range; identifying the log based on the identified container index; identifying a log group including the identified log; and loading the identified log group as a whole into the memory, wherein the identified log group is a data structure for including a plurality of logs, and wherein each log in the plurality of logs in the identified log group is for recording changes associated with a different container index.
4. The computer-implemented method according to claim 1, wherein updating the log to include the indication of changes to the metadata comprises: adding a record to a clone data structure included in the log, wherein the clone data structure is for exclusively recording indications of metadata changes during a clone operation.
5. The computer-implemented method according to claim 4, comprising: updating the record to include a unit address and a length value, wherein the unit address and the length value identify the manifest range in a run-length reference format.
6. The computer-implemented method according to claim 4, comprising: updating the record to include an indication of a reference count of the manifest range.
7. The computer-implemented method according to claim 4, comprising: in response to determining that the log in the memory has reached a maximum threshold of stored data: loading the container index from the persistent storage device into the memory; updating the loaded container index in the memory based on records of the clone data structure; and after updating the loaded container index in the memory based on records of the clone data structure, clearing the records of the clone data structure.
8. The computer-implemented method according to claim 7, comprising: for each record of the clone data structure: updating the loaded container index in the memory based on the record; and After updating the loaded container index based on the record in the memory, update the record to include an identifier of the manifest generated by the cloning operation.
9. A non-transitory machine-readable medium storing instructions that, when executed, cause a processor to perform the following operations: Detect a cloning operation for a manifest range; In response to the detected cloning operation, load a log from a persistent storage device into the memory, wherein, the log is used to store changes to a container index associated with the manifest range, and wherein the container index is not loaded into the memory in response to the detected cloning operation; and Update the log in the memory to include an indication of changes to metadata of a container index not loaded into the memory, wherein the changes to the metadata are associated with the detected cloning operation.
10. The non-transitory machine-readable medium of claim 9, comprising instructions that, when executed, cause the processor to perform the following operations: Identify the container index associated with the manifest range; Identify the log based on the identified container index; Identify a log group including the identified log; and Load the identified log group as a whole into the memory, wherein, the log group is a data structure for including a plurality of logs, and wherein each log in the plurality of logs in the log group is used to record changes associated with a different container index.
11. The non-transitory machine-readable medium of claim 9, comprising instructions that, when executed, cause the processor to perform the following operations: Add a record to a cloning data structure included in the log, wherein, the cloning data structure is used to exclusively record an indication of metadata changes during a cloning operation.
12. The non-transitory machine-readable medium of claim 11, comprising instructions that, when executed, cause the processor to perform the following operations: Update the record to include a unit address and a length value, wherein, the unit address and the length value identify the manifest range in a run-length reference format.
13. The non-transitory machine-readable medium of claim 11, comprising instructions that, when executed, cause the processor to perform the following operations: Detect a non-cloning operation for a second manifest range; In response to the detected non-cloning operation, load a second log and a second container index together from the persistent storage device into the memory; and Update the second log in the memory to store second changes associated with the second container index, wherein, the second changes are associated with the detected non-cloning operation.
14. The non-transitory machine-readable medium of claim 11, comprising instructions that, when executed, cause the processor to perform the following operations: In response to determining that the log in the memory has reached a maximum threshold of stored data: Load the container index from the persistent storage device into the memory; Update the loaded container index in the memory based on the record of the cloning data structure; and After updating the loaded container index in the memory based on the records of the clone data structure, clear the records of the clone data structure.
15. A storage system, comprising: a processor including a plurality of processing engines; and a machine-readable storage device storing instructions that can be executed by the processor to perform the following operations: detect a clone operation for a manifest range; in response to the detected clone operation, load a log from a persistent storage device into the memory, where the log is used to store changes to a container index associated with the manifest range, and where the container index is not loaded into the memory in response to the detected clone operation; and update the log in the memory to include an indication of changes to metadata of a container index that is not loaded into the memory, where the changes to the metadata are associated with the detected clone operation.
16. The storage system of claim 15, comprising instructions that can be executed by the processor to perform the following operations: identify the container index associated with the manifest range; identify the log based on the identified container index; identify a log group including the identified log; and load the identified log group as a whole into the memory, where the log group is a data structure for including a plurality of logs, and where each of the plurality of logs in the log group is used to record changes associated with a different container index.
17. The storage system of claim 15, comprising instructions that can be executed by the processor to perform the following operations: add a record to a clone data structure included in the log, where the clone data structure is used to exclusively record an indication of metadata changes during a clone operation.
18. The storage system of claim 17, comprising instructions that can be executed by the processor to perform the following operations: update the record to include a unit address and a length value, where the unit address and the length value identify the manifest range in a run-length reference format.
19. The storage system of claim 17, comprising instructions that can be executed by the processor to perform the following operations: update the record to include an indication of a reference count of the manifest range.
20. The storage system of claim 15, comprising instructions that can be executed by the processor to perform the following operations: in response to determining that the log in the memory has reached a maximum threshold of stored data: load the container index from the persistent storage device into the memory; update the loaded container index in the memory based on the records of the clone data structure; and after updating the loaded container index in the memory based on the records of the clone data structure, clear the records of the clone data structure.
Citation Information
Patent Citations
Partition level operation with concurrent activities
US20150261807A1
System and method for generating backups of a protected system from a recovery system
US20170083540A1