Repair string for deduplication storage system

By selecting the comparison window and matching score in the deduplication storage system, identifying and replacing the corrupted string, the problem of corrupted data string irreparable is solved, and the system's data integrity and performance is improved.

CN120353640APending Publication Date: 2025-07-22HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410878253.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2024-07-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In existing deduplication storage systems, corrupted data strings cannot be effectively repaired in snapshots, resulting in data loss and system performance degradation.

Method used

By selecting the comparison window in the storage controller, identifying the corrupted string, and comparing the matching score with the candidate list, selecting the uncorrupted string as the repair string, replacing the reference of the corrupted string, and realizing the repair of the corrupted string.

Benefits of technology

In reducing data loss, the performance of the storage system is improved and data integrity and reliability are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353640A_ABST
    Figure CN120353640A_ABST
Patent Text Reader

Abstract

A repair string for a data deduplication storage system is provided. Example implementations relate to data deduplication operations in a storage system. Examples include selecting a comparison window in a first manifest of a deduplication storage system, where the comparison window includes a plurality of data units and includes a corrupted string. The example also includes identifying a plurality of manifests; determining a match score for the manifest based on a match to the comparison window; and identifying a plurality of repair strings based on the match scores. The example also includes recording the corrupted string and the plurality of repair strings in a first entry of the repair string data structure, and repairing the corrupted string in the first manifest using at least one repair string of the plurality of repair strings.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Data reduction techniques can be applied to reduce the amount of data stored in a storage system. Example data reduction techniques include deduplication. Deduplication identifies duplicate data units and seeks to reduce or eliminate the number of instances of duplicate data units stored in the storage system. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Some embodiments are described with reference to the following drawings.

[0003] Figure 1 is a schematic diagram of an example storage system according to some embodiments.

[0004] Figures 2A to 2E is an illustration of an example data structure according to some embodiments.

[0005] Figure 3 is an illustration of an example process according to some embodiments.

[0006] Figures 4A to 4N is an illustration of an example operation according to some embodiments.

[0007] Figure 5 is an illustration of an example process according to some embodiments.

[0008] Figures 6A to 6E is an illustration of an example operation according to some embodiments.

[0009] Figure 7 is an illustration of an example process according to some embodiments.

[0010] Figure 8 is a diagram of an example machine-readable medium storing instructions according to some embodiments.

[0011] Figure 9 is a schematic diagram of an example computing device according to some embodiments.

[0012] In all the drawings, the same reference numerals refer to similar but not necessarily identical elements. The drawings are not necessarily to scale, and the dimensions of some parts may be enlarged to more clearly illustrate the examples shown. Additionally, the drawings provide examples and / or embodiments consistent with the description; however, the description is not limited to the examples and / or embodiments provided in the drawings. DETAILED DESCRIPTION

[0013] In the present disclosure, unless the context clearly indicates otherwise, the use of the terms "a", "an", or "the" is intended to include the plural forms as well. Similarly, when used in the present disclosure, the terms "includes", "including", "comprises", or "comprising" or "have", "having" specify the presence of the element but do not preclude the presence or addition of other elements.

[0014] In some examples, a storage system may back up a data collection (referred to herein as a "stream" or "data stream" of data) in a deduplicated form, thereby reducing the amount of storage space required to store the data stream. The storage system may create "backup items" to represent the data stream in a deduplicated form. The storage system may perform a deduplication process that includes breaking the data stream into discrete data units (or "chunks") and determining "fingerprints" (as described below) of these incoming data units. Further, the storage system may compare the fingerprints of the incoming data units with the fingerprints of the stored data units and may thereby determine which incoming data units are duplicates of previously stored data units (e.g., when the comparison indicates matching fingerprints). In the case where a data unit is a duplicate, the storage system may store a reference to the previously stored data unit rather than storing the duplicate incoming data unit. The process for receiving an inbound data stream and deduplicating it may be referred to herein as the "data ingestion" process of the storage system.

[0015] As used herein, a "fingerprint" refers to a value obtained by applying a function to the content of a data unit (where the "content" may include all or a subset of the content of the data unit). Examples of functions that may be applied include hash functions that produce a hash value based on the content of the incoming data unit. Examples of hash functions include cryptographic hash functions such as the Secure Hash Algorithm 2 (SHA-2) hash functions (e.g., SHA-224, SHA-256, SHA-384, etc.). In other examples, other types of hash functions or other types of fingerprint functions may be employed.

[0016] "Storage system" may include a storage device or an array of storage devices. The storage system may also include (a) storage controller(s) that manages access to the (multiple) storage devices. "Data unit" may refer to any portion of data that can be individually identified within the storage system. In some cases, a data unit may refer to a chunk, a collection of chunks, or any other portion of data. In some examples, the storage system may store data units in a persistent storage device. The persistent storage device may be implemented using one or more (multiple) persistent (e.g., non-volatile) storage devices such as (multiple) disk-based storage devices (e.g., (multiple) hard disk drives (HDD)), (multiple) solid state devices (SSD) (such as (multiple) flash storage devices), etc. or a combination thereof. "Controller" may refer to a hardware processing circuit, which may include any one or some combination of a microprocessor, a core of a multi-core microprocessor, a microcontroller, a programmable integrated circuit, a programmable gate array, a digital signal processor, or other hardware processing circuits. Alternatively, "controller" may refer to a combination of a hardware processing circuit and machine-readable instructions (software and / or firmware) executable on the hardware processing circuit.

[0017] In some examples, a deduplication storage system may use metadata to process an inbound data stream (e.g., a backup item). For example, such metadata may include a data recipe (also referred to herein as a "manifest") that specifies the order of receipt of specific data units for each backup item. Further, such metadata may include item metadata to represent each received backup item in a deduplicated form. The item metadata may include an identifier of a set of manifests and may indicate the order of the set of manifests. Processing each backup item may be referred to herein as a "backup process". Subsequently, in response to a read request, the deduplication system may use the item metadata and the set of manifests to determine the order of receipt of the data units, such that the original data stream of the backup item can be recreated. Thus, the set of manifests may be a representation of the original backup item. A manifest may include a series of records, each record representing a specific set of (multiple) data units. Records in the manifest may include one or more fields that identify container indices that index the data units (e.g., including storage information of the data units). For example, a container index may include one or more fields that specify location information (e.g., container, offset, etc.) of the stored data units, compression and / or encryption characteristics of the stored data units, etc. Further, the container index may include reference counts that indicate the number of manifests that reference each data unit.

[0018] In some examples, when a data unit (e.g., in a data stream) is received, it can be matched with one or more container indexes to determine whether the same chunk has already been stored in a container of the deduplication storage system. For example, the deduplication storage system can compare the fingerprint of the received data unit with the fingerprints in one or more container indexes. If no matching fingerprint is found in the searched container indexes, the received data unit can be added to the container, and an entry for the received data unit can be added to the container index corresponding to that container. However, if a matching fingerprint is found in the searched container indexes, it can be determined that a data unit identical to the received data unit has already been stored in the container. In response to that determination, the reference count of the corresponding entry is incremented, and the received data unit is not stored in the container (since it already exists in one of the containers), thus avoiding storing duplicate data units in the deduplication storage system. As used herein, the term "matching operation" can refer to an operation for comparing the fingerprints of a set of multiple data units (e.g., from a specific backup data stream) with the fingerprints stored in a container index.

[0019] In some examples, a deduplication storage system can process an inbound data stream to store deduplicated copies (also referred to herein as "snapshots") of all data blocks in a source data collection (also referred to herein as "source items") at a particular point in time. Further, in some examples, at least some of the data blocks in the source items may change over time. For example, a sales database (i.e., the source item) can be replicated in a snapshot at a first point in time, and the sales database can subsequently be updated to include new records of additional sales transactions. Accordingly, the deduplication storage system can generate and store a sequence of snapshots to capture the changing state of the source item at different points in time. In some examples, some or all of the data units in the source item may be lost or become unavailable (e.g., due to a malware attack, system failure, etc.). In such examples, the stored snapshots can be used to recover at least some of the lost data of the source item. For example, if the sales database is lost due to a system failure, the most recent snapshot can be replicated to regenerate the sales database as it existed at the time the snapshot was created. However, this replication of the most recent snapshot does not recover the changes made to the sales database after the creation of that most recent snapshot.

[0020] In some examples, one or more contiguous data units (also referred to herein as "strings") included in a snapshot may be corrupted. As used herein, the term "corrupted" may refer to a (plural) data unit that does not include correct data content. For example, a string of (plural) data units in a source item may be corrupted by a malware attack, and the corrupted string may be copied into each subsequent snapshot of the source item. In this way, the (plural) most recent snapshots that include the corrupted string may not be usable to repair the source item. Therefore, it may be necessary to use an older snapshot that does not include the corrupted string to perform a repair of the source item. Further, because the older snapshot is less similar to the source item than the most recent snapshot (e.g., due to changes to the source item accumulating over time), using the older snapshot to restore the source item is likely to result in greater data loss.

[0021] According to some embodiments of the present disclosure, a controller of a deduplication storage system may repair a snapshot that includes a corrupted string. The controller may identify the corrupted string included in a snapshot of a source item. For example, the controller may identify the corrupted string by matching a specific fingerprint sequence corresponding to the data units included in the string. The controller may identify a first manifest that references a portion of the snapshot that includes the corrupted string. The controller may then select a comparison window (e.g., a set of data units that includes the corrupted string) from the first manifest, and may compare the comparison window with a set of candidate manifests associated with a previous snapshot of the source item. The controller may assign a matching score to each candidate manifest based on the similarity of each candidate manifest to the comparison window. If the candidate manifest with the highest matching score also references the corrupted string, the controller may select a new comparison window from that candidate manifest, and may then compare the remaining candidate manifests with the new comparison window. When determining that the candidate manifest with the highest score does not reference the corrupted string, the controller may identify an uncorrupted string from that candidate manifest, and may determine that the uncorrupted string may be used as a repair for the corrupted string (i.e., the uncorrupted string may be used in the snapshot instead of the corrupted string). Further, the controller may replace the reference to the corrupted string with a reference to the uncorrupted string (in the first manifest). In this way, the controller may repair the affected manifest with reduced (or no) data loss, and thereby may improve the performance of the storage system. Aspects of the disclosed repair process are further discussed below with reference to Figures 1 to 9 Further aspects of the disclosed repair process are discussed.

[0022] Figure 1 -Example storage system

[0023] Figure 1FIG. 0 shows an example of a storage system 100 including a storage controller 110, a memory 115, and a persistent storage device 140, according to some embodiments. The persistent storage device 140 may include one or more non-transitory storage media, such as a hard disk drive (HDD), a solid state drive (SSD), an optical disk, etc., or a combination thereof. The memory 115 may be implemented with a semiconductor memory such as random access memory (RAM). In some examples, the storage controller 110 may be implemented via hardware (e.g., electronic circuitry) or a combination of hardware and programming (e.g., including at least one processor and instructions executable by the at least one processor and stored on at least one machine-readable storage medium).

[0024] As Figure 1 shown, the memory 115 and the persistent storage device 140 may store various data structures, which at least include item metadata 130, a manifest 150, a container index 160, data containers 170, and repair data 180. In some examples, copies of the item metadata 130, the manifest 150, the container index 160, the data containers 170, and the repair data 180 may be transferred between the memory 115 and the persistent storage device 140 (e.g., via read-write input / output (I / O) operations).

[0025] In some embodiments, the storage system 100 may perform a data ingestion operation to deduplicate received data. For example, the storage controller 110 may receive an inbound data stream including a plurality of data units, and may store at least one copy of each data unit in the data containers 170 (e.g., by appending the data unit to the end of the data containers 170). In some examples, each data container 170 may be partitioned into entities, where each entity includes a plurality of stored data units. Further, in some examples, the inbound stream may be deduplicated and stored as a backup item.

[0026] In one or more embodiments, the storage controller 110 may generate a fingerprint for each received data unit. For example, the fingerprint may include a full or partial hash value based on the data unit. To determine whether an incoming data unit is a duplicate of a stored data unit, the storage controller 110 may perform a matching operation to compare the fingerprint generated for the incoming data unit with the fingerprints in at least one container index 160. If a match is identified, the storage controller 110 may determine that the storage system 100 has stored a duplicate of the incoming data unit. The storage controller 110 may then store a reference to the previous data unit instead of storing the duplicate incoming data unit.

[0027] In some embodiments, the storage controller 110 may generate item metadata 130 to represent each backup item in a deduplicated form. In some examples, a set of backup items (represented by a set of item metadata 130) may be a sequence of snapshots that capture the changing data content of a source item (e.g., a storage volume) at multiple points in time. Each item metadata 130 may include an identifier for a set of manifests 150 and may indicate the order of the set of manifests 150. The manifests 150 record the order in which data units are received.

[0028] In some embodiments, the manifest 150 may include a pointer or other information indicating a container index 160 that indexes each data unit. In some embodiments, the container index 160 may include a fingerprint (e.g., a hash) of the stored data unit for use in a matching process of the deduplication process. Further, the container index 160 may indicate the location where the data unit is stored. For example, the container index 160 may include information specifying that the data unit is stored at a particular offset in an entity and the entity is stored at a particular offset in a data container 170. The container index 160 may also include a reference count indicating the number of manifests 150 that reference each data unit.

[0029] In some embodiments, the storage controller 110 may receive a read request to access the stored data and, in response, may access the item metadata 130 and the manifests 150 to determine the sequence of data units that make up the original data. The storage controller 110 may then use the pointer data included in the manifests 150 to identify the container index 160 that indexes the data units. Further, the storage controller 110 may use the information included in the identified container index 160 (and the information included in the manifests 150) to determine the location where the data unit is stored (e.g., data container 170, entity, offset, etc.) and may then read the data unit from the determined location.

[0030] In some embodiments, the storage controller 110 may identify a corrupted string included in a first snapshot of a source item and may identify the manifests 150 that reference the corrupted string. For example, the storage controller 110 may receive (e.g., from a user or a program) a message or command identifying a data range (e.g., based on a scan or analysis of the snapshot) that includes a string of consecutively corrupted data units in the snapshot. In some examples, the location of the corrupted string may be specified as an offset and size within a given backup item (e.g., a particular item metadata 130). In such an example, the storage controller 110 may determine the particular manifest 150 that represents the offset and size in the item metadata 130.

[0031] In some embodiments, the storage controller 110 may select a set of data unit references (also referred to herein as a “comparison window”) in the manifest 150, where the set of data unit references includes a subset of references to the data units that make up the corrupted string. For example, the storage controller 110 may select a comparison window in the manifest 150 that includes a subset of data unit references, where the reference to the corrupted string is located at the center of the comparison window (also referred to herein as the “center position”). Reference is made below to Figure 4B Describe an example implementation of the comparison window.

[0032] In some embodiments, the storage controller 110 may identify a set of candidate manifests 150 that reference data units included in a previous snapshot of the source item (e.g., multiple snapshots generated prior to the snapshot that includes the identified corrupted string). The storage controller 110 may compare the comparison window with each candidate manifest 150 and may assign a matching score to each candidate manifest based on the similarity of each candidate manifest to the comparison window. As used herein, the term “matching score” may refer to a numerical value that measures the similarity between the data units referenced in the comparison window and the set of data units referenced in the candidate manifest 150. For example, the matching score may indicate how many data units are referenced in both the comparison window and a portion of the candidate manifest 150. Reference is made below to Figures 4D to 4G Describe an example calculation of the matching score.

[0033] In some embodiments, the storage controller 110 may identify a candidate manifest 150 having the highest matching score (e.g., most similar to the comparison window). Further, if the identified candidate manifest 150 also references a corrupted string, the storage controller 110 may select a new comparison window from the identified candidate manifest 150 and then may compare the remaining candidate manifests 150 with the new comparison window. If necessary, the storage controller 110 may repeat this process (e.g., determine that the candidate manifest 150 having the highest matching score references a corrupted string and then select a new comparison window) within multiple iterations until it is determined that the candidate manifest 150 having the highest matching score does not reference a corrupted string. Further, upon determining that the candidate manifest 150 having the highest matching score does not reference a corrupted string, the storage controller 110 may select the uncorrupted string referenced in that candidate manifest 150. In some embodiments, the uncorrupted string may be determined to include the correct data units that should be present in the snapshot, as opposed to the corrupted string. Thus, the storage controller 110 may use the uncorrupted string to repair the manifest 150 that references the corrupted string (e.g., by replacing the reference to the corrupted string with a reference to the uncorrupted string). The uncorrupted string that can repair the manifest 150 may be referred to herein as a "repair string". In this manner, the storage controller 110 may repair the manifest 150 with reduced (or no) data loss and thereby may improve the performance of the storage system 100. The following references Figure 3 and Figures 4A to 4N describe an example process for repairing manifests included in a snapshot.

[0034] In some embodiments, the storage controller 110 may identify multiple repair strings that can be used to attempt to repair a given corrupted string. For example, after comparing the candidate manifests 150 with the comparison window, the storage controller 110 may identify multiple candidate manifests 150 that do not reference the corrupted string and may identify different repair strings from each of the multiple candidate manifests 150. In some embodiments, the storage controller 110 may store information in an entry of the repair data 180 to record the multiple repair strings that can be used to attempt to repair the corrupted string. Subsequently, the storage controller 110 may compare the received data string with the repair string entries of the repair data 180. If the received data string matches the corrupted string recorded in an entry of the repair data 180, the storage controller 110 may attempt one or more repairs of the corrupted string using the repair string listed in that entry. In some embodiments, the repair data 180 may also include historical data regarding repairs that have been attempted using the repair string. For example, the historical data may record the corrupted string, the repair string, and the location of each attempted repair. Such historical information may be used to roll back an attempted repair (e.g., if the attempted repair is determined to be invalid). The following references Figure 2C describe an example implementation of the repair data 180.

[0035] Figures 2A to 2E - Example data structure

[0036] Figure 2A Illustrated is a diagram of an example data structure 200 used in deduplication according to some embodiments. As shown, the data structure 200 may include item metadata 202, a manifest 203, a container index 220, and a data container 250. In some examples, the item metadata 202, the manifest 203, the container index 220, and the data container 250 may generally correspond to example embodiments of the item metadata 130, the manifest 150, the container index 160, and the data container 170 (as Figure 1 shown). In some examples, the data structure 200 may be generated and / or managed by a storage controller 110 (as Figure 1 shown).

[0037] In some embodiments, the item metadata 202 may include a plurality of manifest identifiers 205. Each manifest identifier 205 may identify a different manifest 203. In some embodiments, the manifest identifiers 205 may be arranged in a stream order (i.e., based on the order of receipt of data units represented by the identified manifest 203). Further, the item metadata 202 may include a container list 204 associated with each manifest identifier 205. In some embodiments, the container list 204 may include identifiers of a set of container indexes 220 that index data units included in the associated manifest 203 (i.e., the manifest 203 identified by the associated manifest identifier 205).

[0038] Although only one of each data structure is shown for simplicity Figure 2A in the illustration, the data structure 200 may include multiple instances of the item metadata 202, each instance including or pointing to one or more manifests 203. In such an example, the data structure 200 may include multiple manifests 203. The manifest 203 may reference multiple container indexes 220, each container index corresponding to one of the multiple data containers 250. Each container index 220 may include one or more data unit records 230 and one or more entity records 240.

[0039] As Figure 2AAs shown, in some examples, each manifest 203 may include one or more manifest records 210. Each manifest record 210 may include various fields such as an offset, a length, a container index, and a cell address. In some embodiments, each container index 220 may include any number of data cell records 230 and entity records 240. Each data cell record 230 may include various fields such as a fingerprint (e.g., a hash of the data cell), a cell address, an entity identifier, a cell offset (i.e., the offset of the data cell within the entity), a reference count value, and a cell length. In some examples, the reference count value may indicate the number of manifest records 210 that reference the data cell record 230. Further, each entity record 240 may include various fields such as an entity identifier, an entity offset (i.e., the offset of the entity within the container), a storage length (i.e., the length of the data cell within the entity), a decompressed length, a checksum value, and compression / encryption information (e.g., a compression type, an encryption type, etc.). In some embodiments, each data container 250 may include any number of entities 260, and each entity 260 may include any number of stored data cells.

[0040] In some embodiments, the cell address (included in the manifest record 210 and the data cell record 230) may be an identifier that deterministically identifies a particular data cell within a given container index 220. In some examples, the cell address may be a numerical value (referred to as an “arrival number”) that indicates (e.g., when receiving an inbound data stream and deduplicating it) the arrival order (also referred to as the “ingestion order”) of the data cell indexed in the given container index 220. For example, an arrival number “1” may be assigned to the first data cell indexed in the container index 220 (e.g., by creating a new data cell record 230 for the first data cell), an arrival number “2” may be assigned to the second data cell, an arrival number “3” may be assigned to the third data cell, and so on. However, other embodiments are possible.

[0041] In some embodiments, the manifest record 210 may use a run-length reference format to represent a range of contiguous data units (e.g., a portion of a data stream) that are indexed within a single container index 220. The run-length reference may be recorded in the cell address field and the length field of the manifest record 210. For example, the cell address field may indicate the arrival number of the first data unit in the represented range of data units, and the length field may indicate the number N (where "N" is an integer) of data units that are after the data unit specified by the arrival number in the cell address field within the range of data units. The data units within the range of data units may have contiguous arrival numbers (e.g., because they are contiguous in the ingested data stream). Thus, the range of data units may be represented by the arrival number of the first data unit in the range of data units (e.g., specified in the cell address field of the manifest record 210) and the number N of additional data units in the range of data units (e.g., specified in the length field of the manifest record 210). The additional data units that are after the first data unit within the range of data units may be deterministically obtained by calculating N arrival numbers that are sequentially after the specified arrival number of the first data unit, where these N arrival numbers identify the additional data units within the range of data units. In such an example, the manifest record 210 may include the arrival number "X" in the cell address field and the number N in the length field to indicate a range of data units that includes the data unit specified by arrival number X and the data units specified by arrival number X+i, where i = 0 to i = N (including 0 and N) (where "i" is an integer). In this way, the manifest record 210 may be used to identify all of the data units within the range of data units.

[0042] In one or more embodiments, the data structure 200 may be used to retrieve the stored deduplicated data. For example, a read request may specify an offset and a length of data within a given file. These request parameters may be matched against the offset and length fields of a particular manifest record 210. The container index and cell address of the particular manifest record 210 may then be matched against a particular data unit record 230 included in the container index 220. Further, the entity identifier of the particular data unit record 230 may be matched against the entity identifier of a particular entity record 240. Additionally, one or more other fields of the particular entity record 240 (e.g., entity offset, storage length, checksum, etc.) may be used to identify the container 250 and the entity 260, and then the data units may be read from the identified container 250 and entity 260.

[0043] In some embodiments, each container index 220 may include an inventory list 222. The inventory list 222 may be a data structure for storing a set of entries, where each entry stores information about a different inventory 203 (or inventory record 210) indexed by the container index 220 (including the inventory list 222). For example, each time the container index 220 is generated or updated to include information about a specific inventory record 210, the inventory list 222 in the container index 220 is updated to store the identifier of the inventory record 210. Further, the entries of the inventory list 222 may be arranged in the respective order in which they enter the inventory list 222 (also referred to herein as the "arrival order" of the entries). In some examples, when the container index 220 is no longer associated with an inventory record 210, the identifier of the inventory record 210 is removed from the inventory list 222.

[0044] Now referring to Figure 2B , an example embodiment of the inventory list 222 is shown. As Figure 2A shown, in some embodiments, each entry of the inventory list 222 may store only the inventory identifier 205. However, other examples are possible. For example, each entry of the inventory list 222 may instead include the inventory identifier 205 and at least one data unit range (e.g., a set of one or more data units included in the inventory 203 and indexed by the container index 220).

[0045] Now referring to Figure 2C , an example embodiment of the repair data 280 is shown. The repair data 280 may generally correspond to an example embodiment of the repair data 180 (as Figure 1 shown). As Figure 2C shown, the repair data 280 may include repair string data 270 and a repair history 275. The repair string data 270 and the repair history 275 may be stored in separate data structures (e.g., tables, databases, formatted text files, etc.).

[0046] In some embodiments, the repair string data 270 may include multiple entries. Each entry of the repair string data 270 may store an identifier of a corrupted string and may also store an identifier of a corresponding set of repair strings (i.e., uncorrupted strings that have been determined to be possible repair alternatives for the corrupted string). In some embodiments, the string identifier (i.e., the unique identifier of the corrupted string or the repair string) may include a fingerprint corresponding to each data unit in the string. For example, Figure 2D illustrates an example string 290 including a single data unit. Thus, as Figure 2D shown, the string identifier 292 is a single fingerprint corresponding to the single data unit in the string 290. In another example, Figure 2EIllustrated is an example string 295 that includes a contiguous sequence of four data units. Thus, as Figure 2E shown, the string identifier 297 includes a sequence of four fingerprints that correspond to the data units in string 295.

[0047] Referring again to Figure 2C , when one or more repair strings for repairing a particular corrupted string are identified, entries can be added to the repair string data 270. Subsequently, when a new data string to be stored is received, the received string can be compared to the entries in the repair string data 270. If the received data string matches a corrupted string recorded in an entry of the repair string data 270, the received data string can be replaced with the repair string listed in the matching entry. Example processes for performing repairs using the repair string data 270 are described below with reference to Figure 5 and Figures 6A to 6D .

[0048] In some embodiments, the repair history 275 can include multiple entries. Each entry in the repair history 275 can store information about different repair operations that have been attempted to repair a corrupted string. For example, an entry in the repair history 275 can record the location where the repair was attempted (e.g., offset, container identifier, address, etc.), the identifier of the corrupted string, the identifier of the repair string, and a timestamp of the repair operation. In some embodiments, the repair history 275 can be used to roll back repairs recorded in an entry. For example, a repair recorded in an entry can be rolled back (i.e., undone) by overwriting the repair string with the corrupted string at the recorded location.

[0049] Figure 3 and Figures 4A to 4N - Example process for repairing a snapshot

[0050] Figure 3 Illustrated is an example process 300 for repairing a snapshot in a deduplication storage system according to some embodiments. In some examples, process 300 can be performed using a storage controller 110 as Figure 1 shown. Process 300 can be implemented in a hardware or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions can be stored in a non-transitory computer-readable medium such as an optical storage device, a semiconductor storage device, or a magnetic storage device. The machine-readable instructions can be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. In some embodiments, process 300 can be executed by a single processing thread. In other embodiments, process 300 can be executed in parallel by multiple processing threads (e.g., using a work graph and performing multiple housekeeping jobs simultaneously).

[0051] The box 310 may include a corruption string included in the first snapshot that identifies the backup item. The box 320 may include a first manifest that references the corruption string. In some embodiments, the manifest may reference a total of T (where "T" is an integer) data units, and may record the reception order of the T data units in the manifest. For example, referring to Figure 1 and Figure 4A , the storage controller 110 receives an indication (e.g., a user command, a software message, etc.) that identifies a data range (e.g., an offset and a length in the backup item) of the corruption string 412 (labeled "CS") included in the snapshot. The corruption string 412 may include N (where "N" is an integer) corrupted data units. Note that for illustrative purposes, the corruption string 412 is illustrated as a single corrupted data unit in Figure 4A . However, in other examples, the corruption string 412 may include multiple consecutive corrupted data units (i.e., N>1). In some examples, the presence of the corruption string 412 may be detected by malware scanning, integrity testing, etc. The storage controller 110 reads the item metadata 130 representing the snapshot, identifies the manifest 410 that references the corruption string 412, and loads the manifest 410 into the memory 115.

[0052] Referring again to Figure 3 , the box 330 may include identifying a comparison window in the first manifest. In some embodiments, the comparison window may have a width W (where "W" is an integer), which indicates the number of data units referenced in the comparison window, where the width W is less than the total number "T" of data units referenced in the manifest. Further, in some embodiments, the width W may be greater than the number N of data units in the corruption string. For example, referring to Figure 1 and Figure 4B , the storage controller selects a comparison window 414 that includes the corruption string 412 in the manifest 410. In some embodiments, the corrupted data unit 412 is referenced at the center position of the comparison window 414. However, other sizes and / or arrangements of the comparison window 414 are possible.

[0053] Referring again to Figure 3 , the box 340 may include a first container index that indexes the corruption string based on the first manifest. The box 350 may include identifying a set of candidate manifests based on the list of manifests included in the first container index. For example, referring to Figures 1 to 2A and Figure 4C, the storage controller 110 reads the manifest 150 to identify the container index 160 that indexes the corrupted string 412, and then loads the container index 160 into the memory 115. Further, the storage controller 110 reads the manifest list 222 (stored in the container index 160) to identify the set of manifests 416 indexed by the container index 160 (e.g., candidate manifests 420, 430, 440, 450). In some embodiments, the set 416 is arranged in the order of arrival (e.g., in the order in which they enter the manifest list 222). The storage controller 110 loads each candidate manifest in the set 416 into the memory 115 (e.g., one at a time or as a group). In some examples, each candidate manifest in the set 416 may reference data units in different snapshots of a particular source item. In such examples, each candidate manifest may record the order of data units in a particular portion of the source item at different points in time (i.e., corresponding to each snapshot).

[0054] Referring again to Figure 3 , block 360 may include determining a match score based on a sliding comparison of the set of candidate manifests with the comparison window. Block 370 may include selecting the candidate manifest with the highest match score. As used herein, "sliding comparison" may refer to an operation that includes: sliding a selection window (also referred to as a "scan window") across the candidate manifests to select different subsets of data units referenced in the candidate manifests, and comparing each subset of data units with the data units referenced in the current comparison window. In some embodiments, the width of the scan window used in the sliding comparison may be equal to the width W of the comparison window. For example, referring to Figure 4D , the controller (in the set 416) selects the candidate manifest 420 that is closest to the manifest 410 (identified at block 320) before. The controller initiates the sliding comparison of the candidate manifest 420 by placing the scan window 422 in the first position, thereby selecting the leftmost W data units (e.g., the earliest W data units in the order of arrival) referenced in the candidate manifest 420, where W is the width of the comparison window 414. The controller then determines the similarity value between the comparison window 414 and the scan window 422 in the first position. In some embodiments, the similarity value may be calculated as the total number of uncorrupted data units (i.e., data units not included in the corrupted string 412) in the comparison window 414 that are also included in the current scan window 422. For example, as Figure 4D shown, two uncorrupted data units 423 (i.e., data units "32" and "42") are referenced in both the comparison window 414 and the scan window 422. Thus, the similarity value of the scan window 422 in the first position is equal to two.

[0055] Now referring to Figure 4E, the controller continues the sliding comparison of candidate list 420 by sliding the scan window 422 one data unit reference to the right to reach the second position. As Figure 4E shown, the controller then identifies three undamaged data units 425 (i.e., data units "32", "42", and "76") that are referenced in both the comparison window 414 and the scan window 422 at the second position. Thus, the similarity value of the scan window 422 at the second position is equal to three.

[0056] Now referring to Figure 4F , the controller completes the sliding comparison of candidate list 420 by sliding the scan window 422 one data unit reference to the right to reach the third position. As Figure 4F shown, the controller then identifies two undamaged data units 427 (i.e., data units "42" and "76") that are referenced in both the comparison window 414 and the scan window 422 at the third position. Thus, the similarity value of the scan window 422 at the third position is equal to two.

[0057] Now referring to Figure 4G , the controller determines that the highest similarity score (i.e., three) in the sliding comparison of candidate list 420 is obtained at the second position of the scan window 422 (as Figure 4E shown). Thus, the controller identifies the second position of the scan window 422 as the matching window 424 of candidate list 420. As used herein, "matching window" may refer to a subset of data unit references in the candidate list that have the highest similarity with respect to the data unit references in the current comparison window (i.e., the comparison window being used for the sliding comparison). Further, as Figure 4G shown, the controller sets the matching score of candidate list 420 to three ("S = 3"), corresponding to the similarity score of the matching window 424.

[0058] Now referring to Figure 4H , the controller performs a sliding comparison of the sets 416 (i.e., lists 420, 430, 440, and 450) with the comparison window 414 and determines the matching scores of the candidate lists. Further, the controller selects the candidate list 420 in the set 416 that has the best matching score (i.e., three).

[0059] Referring again to Figure 3, the decision block 375 may include determining whether the selected manifest references a corrupt string (identified at block 310). If it is determined that the selected manifest references a corrupt string ("yes"), then process 300 may continue at block 380, including removing the selected manifest from the set of candidate manifests. Block 385 may include identifying a new comparison window in the selected manifest. After block 385, process 300 may return to block 360 (i.e., determining a new match score based on a sliding comparison of the remaining candidate manifests with the new comparison window). In some examples, blocks 360, 370, 375, 380, and 385 may be repeated in one or more loops until a negative determination is reached at decision block 375 (i.e., when it is determined that the selected candidate manifest does not reference a corrupt string). For example, referring to Figure 4H , the controller determines that in the candidate manifest 420 (selected at block 370), the match window 424 includes a reference to the corrupt string 412 (which is also referenced in manifest 410). In response to this determination, as Figure 4I shown, the controller removes the candidate manifest 420 from the set 416 to obtain a reduced set 417, and selects the match window 424 in the candidate manifest 420 as the second comparison window 428.

[0060] Now referring to Figure 4J , the controller performs a sliding comparison of the reduced set 417 of candidate manifests (i.e., manifests 430, 440, 450) with the second comparison window 428, and determines the match scores of these candidate manifests. The controller then determines that the candidate manifest 430 has the best match score (i.e., three), corresponding to the similarity score of the second match window 434 in the candidate manifest 430. Further, the controller determines that the second match window 434 includes a reference to the corrupt string 412. In response to this determination, as Figure 4K shown, the controller removes the candidate manifest 430 from the reduced set 417 to obtain a second reduced set 418, and selects the second match window 434 in the candidate manifest 430 as the third comparison window 438.

[0061] Referring again to Figure 3 , if it is determined at decision block 375 that the selected manifest does not reference a corrupt string ("no"), then process 300 may continue at block 390, including identifying the repair string(s) referenced in the selected manifest. Block 395 may include replacing the reference(s) to the corrupt string(s) with reference(s) to the identified repair string(s) in the first manifest. After block 395, process 300 may be completed. For example, referring to Figure 4L, the controller performs a sliding comparison of the second reduced set 418 (i.e., listings 440, 450) with the third comparison window 438 and determines the matching scores of these candidate listings. The controller then determines that candidate listing 440 has the best matching score (i.e., three), corresponding to the similarity score of the third matching window 444 in candidate listing 440. Further, the controller determines that the third matching window 444 does not include a reference to the corrupted string 412 but rather refers to an uncorrupted string 445 (labeled "NS"). In response to this determination, as Figure 4M shown, the controller uses the uncorrupted string 445 to perform a repair 460 (i.e., by replacing the reference(s) to the corrupted string 412 with a reference(s) to the uncorrupted string 445) in each of the listings 410, 420, 430 (i.e., the listings previously determined to include corrupted data units at decision block 375 as Figure 3 shown). In this way, the controller can repair the affected listings with reduced (or no) data loss and thereby improve the performance of the system including the stored data.

[0062] In some embodiments, the controller may determine whether the repaired listings 410, 420, 430 are valid (i.e., contain uncorrupted data). For example, after performing the repair 460 (i.e., using the uncorrupted string 445), the controller may prompt a human user to check or otherwise evaluate the repaired data and confirm that the repaired data is valid. In another example, after performing the repair 460, the controller may perform an automated test to determine whether the repaired data is valid. In some embodiments, if it is determined that the repair 460 is invalid, the controller may roll back (i.e., undo) the repair 460 to restore the listings 410, 420, 430 to their previous condition (i.e., including the corrupted string 412). Further, the controller may use a candidate listing with the next highest matching score that does not include the corrupted string 412 to perform another repair. For example, referring to Figure 4N , the controller determines that candidate listing 450 with the next best matching score (i.e., S = 2) includes an alternative string 447 (i.e., instead of the corrupted string 412). Thus, the controller may attempt a second repair 462 using the alternative string 447.

[0063] In some embodiments, the controller may create an entry in the stored data structure (e.g., Figure 2C shown repair string data 270) to record what has been determined to be available for repairing a particular corrupted string (e.g., using Figure 3One or more repair strings for the process 300 shown. For example, an entry may record the identifier of a given corrupted string (e.g., corrupted string 412), and may also record the identifiers of multiple repair strings (e.g., uncorrupted string 445 and alternative string 447). In some embodiments, the entry may indicate a preference or ranking order of the repair strings (e.g., the first ranking or preference is the uncorrupted string 445, and the second ranking or preference is the alternative string 447). Subsequently, if the received data string matches the corrupted string 412 recorded in the entry, the controller may use the entry to perform one or more attempts to repair using the repair strings recorded in the entry. The following refers to Figure 5 Describe an example process for using the stored repair data.

[0064] In some embodiments, the matching score may be determined based on a sliding comparison (e.g., in Figure 3 The box 360 shown) may be performed using comparison windows of different sizes. For example, the controller may perform an initial set of sliding comparisons using a first width of the comparison window and may not be able to determine the best matching window (in the candidate list) based on the matching score (e.g., if all matching scores are equal, if all matching scores are below a minimum threshold of a useful matching score, etc.). In this example, the controller may initiate a second set of sliding comparisons using a larger width of the comparison window and may again attempt to determine the best matching window based on the matching score. This process may be repeated (with the width of the comparison window increasing continuously) until a valid result is obtained from the sliding comparison (i.e., the best matching window is determined in the candidate list).

[0065] Note that the embodiments are not limited to Figures 4A to 4N The example shown. For example, it is conceivable that any or all of the comparison windows 414 may not be centered within the list 410. Similarly, the matching window 424 may not be centered within the candidate list 420, and the second matching window 434 may not be centered within the candidate list 430. Further, although Figure 4H Shows an example where the selected candidate list 420 with the best matching score is directly adjacent to the first list 410 (including the comparison window 414), it is conceivable that in other examples, the selected candidate list may not be adjacent to the list including the comparison window. For example, if alternatively the candidate list 420 has a lower matching score compared to the candidate list 430, the candidate list 430 may be selected as having the highest matching score in the set 416. In this example, the candidate list 420 may not refer to the same data unit stream (e.g., in a particular source item) as the other candidate lists in the set 416.

[0066] Figure 5 and Figures 6A to 6E - Example process for using an alias list

[0067] Figure 5Illustrates an example process 500 for using stored repair data in a deduplication storage system. In some examples, the process 500 may be performed using a storage controller 110 as Figure 1 shown. The process 500 may be implemented in hardware or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions may be stored in a non-transitory computer-readable medium such as an optical storage device, a semiconductor storage device, or a magnetic storage device. The machine-readable instructions may be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. In some embodiments, the process 500 may be executed by a single processing thread. In other embodiments, the process 500 may be executed in parallel by multiple processing threads (e.g., using a working map and performing multiple housekeeping jobs simultaneously).

[0068] Block 510 may include identifying a set of repair strings for repairing a damaged string. Block 520 may include storing the damaged string and the set of repair strings in an entry of a repair string data structure. For example, referring to Figures 4A to 4N , the controller identifies a damaged string 412 in a snapshot, identifies a manifest 410 that references the damaged string 412, and selects a comparison window 414 in the manifest 410. The controller identifies a container index that indexes the damaged string 412, and identifies a set of candidate manifests 416 that are indexed by the container index (e.g., using the manifest list 222 shown in Figure 2B ). Further, the controller determines a matching score based on a sliding comparison of the set of candidate manifests 416 with the comparison window 414, selects a matching window 424 with the highest matching score in the candidate manifest 420, and determines whether the matching window 424 includes a reference to the damaged data unit 412. If so, the controller performs one or more iterations of the following operations: selecting a comparison window, performing a sliding comparison of the remaining candidate manifests, and selecting a matching window from the remaining candidate manifests. Further, the controller determines that repair strings 445, 447 may be used to repair the damaged string 412, and records identifiers (e.g., fingerprints) of the damaged string 412 and the repair strings 445, 447 in an entry 470 of a stored data structure (e.g., the repair string data 270 shown in Figure 2C ).

[0069] Referring again to Figure 5, the box 530 may include comparing the received data string with the entries of the repair string data structure. The decision box 540 may include determining whether the received input string matches any of the corrupted strings recorded in the entries of the repair string data structure. If there is no match ("No"), then the process 500 may be completed. In other cases, if it is determined that the received input string matches the corrupted string recorded in an entry ("Yes"), then the process 500 may continue at box 550, including performing an attempt to repair using the highest-ranked repair string recorded in that matching entry.

[0070] In some embodiments, a match with an entry of the repair string data structure may be identified if both the received input string and the corrupted string in the entry include the same sequence of data units. For example, referring to Figure 6A , the controller compares the received input string with the index field (e.g., the first field in each entry) of the stored data structure 600. The controller determines that the string "corruption 1" stored in the index field of entry 610 matches the received input string, and thus determines that the received input string was previously identified as a corrupted string. Further, referring to Figure 6B , the controller reads entry 610 to determine the highest-ranked repair string "repair 1" recorded for that corrupted string. The controller then performs a first attempt at repair by replacing the corrupted string with the highest-ranked repair string from entry 610. The stored data structure 600 may generally correspond to Figure 2C an example embodiment of the repair string data 270 shown.

[0071] In some embodiments, a match with an entry of the repair string data structure may be identified if the received input string includes a modified version of the corrupted string stored in the entry. The modified version of the corrupted string may include a first set of data units and a second set of data units, where the first set of data units is the same as the corrupted string recorded in the entry of the repair string data structure, and where the second set of data units includes a number M of additional data units not included in the corrupted string recorded in the first entry, where M is a positive integer. For example, now referring to Figure 6C , the received input string includes the same sequence of data units as the string "corruption 1" stored in the index field of entry 610 (in the stored data structure 600), and also includes two additional data units 620 (i.e., the units "X" and "Y") not included in the string "corruption 1". Further, in the example shown in Figure 6C , the maximum number M of additional data units is equal to 3. Thus, because the number of additional data units 620 (i.e., two) is less than M, the controller identifies a match between the received input string and the string "corruption 1" stored in the index field of entry 610. Additionally, referring to Figure 6D, the controller reads entry 610 to determine the highest-ranked repair string "Repair 1" recorded for the damaged string. The controller then performs a first attempt at repair by replacing the matching portion of the input string (i.e., the data units also included in the string "Damaged 1") with the matching portion of the highest-ranked repair string from entry 610, while leaving the non-matching portions (i.e., data units 620 "X" and "Y") in their original positions within the input string. Thus, the output string includes data units 620 "X" and "Y" in the same positions they occupied in the input string.

[0072] Referring again to Figure 5 , block 555 may include recording the attempted repair in a repair history data structure. For example, referring to Figure 2C , the controller may record information about the attempted repair in an entry of the repair history 275. The recorded information may include the location of the repair, the damaged string that was replaced in the repair, the repair string that was used as a replacement in the repair, and the timestamp of the repair.

[0073] Referring again to Figure 5 , decision block 560 may include determining whether the attempted repair is effective. If it is effective ("Yes"), then process 500 may be completed. In an alternative case, if it is determined that the attempted repair is ineffective ("No"), then process 500 may continue at block 570, including undoing the repair using the repair history data structure. For example, referring to Figure 2C , the controller determines that the attempted repair is ineffective and, in response, obtains information about the attempted repair from the repair history 275. Based on the obtained repair information, the controller undoes the attempted repair by replacing the repair string with the damaged string at the repair location.

[0074] Referring again to Figure 5 , decision block 580 may include determining whether another repair string is effective. If it is ineffective ("No"), then process 500 may be completed. In an alternative case, if it is determined that another repair string is available ("Yes"), then process 500 may continue at block 590, including attempting another repair using the next-highest-ranked string in the matching entry. After block 590, process 500 may return to block 555 (i.e., record the attempted repair in the repair history data structure) and then may proceed to decision block 560 (i.e., determine again whether the attempted repair is effective). For example, referring to Figure 6E, the controller reads the entry 610 of the stored data structure 600 to determine the second-ranked repair string "Repair 2" corresponding to the corrupted string "Corruption 1". The controller then performs a second attempt at repair by replacing the corrupted string with the second-ranked repair string from entry 610. In some embodiments, process 500 may repeat multiple cycles of blocks 555, 560, 570, 580, 590 until it is determined that the repair is effective (at decision block 560), or it is determined that there are no more repair strings available (at decision block 580). When an effective repair is performed, the corrupted string is not stored as part of the deduplicated backup items (representing the received stream). In this way, some embodiments can reduce the amount of corrupted data stored in storage system 100. Further, some embodiments can reduce the data loss and performance impact associated with recovering from each instance of the corrupted string.

[0075] In some embodiments, determining whether an attempted repair is effective (e.g., at decision block 560) can be based on input from a human user. For example, the controller can generate a request or alert (e.g., via a user interface) to ask the user to examine or evaluate the repaired data and confirm that the repaired data is valid (e.g., not corrupted). Further, the request can ask the user whether to attempt another repair operation (e.g., using a different repair string stored in entry 610). In other embodiments, the controller can perform an automated test process to evaluate the repaired data and can automatically perform another repair operation when it is determined that the previous repair was ineffective. For example, such an automated test can include a file system consistency check, an application format verification test, etc.

[0076] In some embodiments, performing an effective repair using process 500 can prevent the corrupted string from being stored as part of the deduplicated backup items (representing the received stream). In this way, some embodiments can reduce the amount of corrupted data stored in storage system 100. Further, some embodiments can reduce the data loss and performance impact associated with recovering from each instance of the corrupted string.

[0077] Figure 7 - Example process for snapshot repair

[0078] Figure 7 An example process 700 for snapshot repair according to some embodiments is shown. In some examples, storage controller 110 can be used (as Figure 1Process 700 is performed as shown. Process 700 can be implemented in hardware or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions can be stored in a non-transitory computer-readable medium such as an optical storage device, a semiconductor storage device, or a magnetic storage device. The machine-readable instructions can be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc.

[0079] Block 710 can include a storage controller selecting a comparison window in a first manifest of a deduplication storage system, where the comparison window includes a plurality of data units and where the comparison window includes a corruption string. Block 720 can include a storage controller identifying a plurality of manifests including the first manifest, where each manifest in the plurality of manifests is included in a different snapshot of a backup item. For example, referring to Figures 4A to 4C , the controller identifies corruption string 412 in a snapshot, identifies manifest 410 that references corruption string 412, and selects comparison window 414 in manifest 410. The controller identifies a container index that indexes corruption string 412 and identifies a set of candidate manifests 416 indexed by the container index (e.g., using Figure 2B the manifest list 222 shown).

[0080] Referring again to Figure 7 , block 730 can include a storage controller determining a match score for the plurality of manifests based on a match with the comparison window. For example, referring to Figures 4D to 4K , the controller determines a match score based on a sliding comparison of the set of candidate manifests 416 with comparison window 414, selects a match window 424 with the highest match score in candidate manifest 420, and determines whether match window 424 includes a reference to corrupted data unit 412. If so, the controller performs one or more iterations of the following operations: select a comparison window, perform a sliding comparison of the remaining candidate manifests, and select a match window from the remaining candidate manifests.

[0081] Referring again to Figure 7 , block 740 can include a storage controller identifying a plurality of repair strings based on the match scores of the plurality of manifests. Block 750 can include a storage controller recording the corruption string and the plurality of repair strings in a first entry of a repair string data structure. For example, referring to Figures 4L to 4M , the controller determines that repair strings 445, 447 can be used to repair corruption string 412 and records identifiers (e.g., fingerprints) of corruption string 412 and repair strings 445, 447 in entry 470 of a stored data structure (e.g., Figure 6A the stored data structure 600 shown).

[0082] Referring again to Figure 7, the frame 760 may include using at least one repair string among the identified multiple repair strings by the storage controller to repair the corrupted string in the first manifest. After the frame 760, the process 700 may be completed. For example, referring to Figure 6A , the controller compares the received input string with the index field of the stored data structure 600. The controller determines that the string "corruption 1" stored in the index field of the entry 610 matches the received input string, and thus determines that the received input string was previously identified as a corrupted string. Further, referring to Figure 6B , the controller reads the entry 610 to determine the highest-ranked repair string "repair 1" recorded for the corrupted string. The controller performs a first attempt at repair by replacing the corrupted string with the highest-ranked repair string from the entry 610. Further, the controller may determine whether the first attempt at repair is effective (e.g., based on the input of a human user or an automated testing process). If it is determined that the first attempt at repair is ineffective, the controller undoes the first attempt at repair (e.g., using the Figure 2C shown repair history 275). Further, referring to Figure 6E , the controller reads the entry 610 to determine the second-highest-ranked repair string "repair 2" corresponding to the corrupted string "corruption 1", and then performs a second attempt at repair by replacing the corrupted string with the second-highest-ranked repair string from the entry 610. In some examples, multiple attempts at repair may be made until it is determined that the repair is effective or it is determined that no more repair strings are available.

[0083] Figure 8 - Exemplary machine-readable medium

[0084] Figure 8 shows a machine-readable medium 800 storing instructions 810 to 850 according to some embodiments. The instructions 810 to 850 may be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. The machine-readable medium 800 may be a non-transitory storage medium, such as an optical storage medium, a semiconductor storage medium, or a magnetic storage medium. The instructions 810 to 850 may generally correspond to the examples described above with reference to frames 710 to 750 (as Figure 7 shown).

[0085] The instruction 810 may be executed to select a comparison window in the first manifest of the deduplication storage system, where the comparison window includes a plurality of data units, and where the comparison window includes a corrupted string. The instruction 820 may be executed to identify a plurality of manifests including the first manifest, where each manifest in the plurality of manifests is included in a different snapshot of a backup item.

[0086] Instruction 830 can be executed to determine a matching score for multiple manifests based on a match with a comparison window. Instruction 840 can be executed to identify multiple repair strings based on the matching scores of the multiple manifests. Instruction 850 can be executed to record the corrupted string and the multiple repair strings in the first entry of a repair string data structure. Instruction 860 can be executed to repair the corrupted string in the first manifest using at least one of the identified multiple repair strings.

[0087] Figure 9 -Example computing device

[0088] Figure 9 FIG. shows a schematic diagram of an example computing device 900. In some examples, the computing device 900 may generally correspond to some or all of the storage system 100 (as Figure 1 shown). As shown, the computing device 900 may include a hardware processor 902, a memory 904, and a machine-readable storage device 905 including instructions 910-960. The machine-readable storage device 905 may be a non-transitory medium. The instructions 910 to 960 may be executed by the hardware processor 902 or by a processing engine included in the hardware processor 902. The instructions 910 to 950 may generally correspond to the examples described above with reference to blocks 710 to 750 (as Figure 7 shown).

[0089] Instruction 910 can be executed to select a comparison window in a first manifest of a deduplication storage system, where the comparison window includes multiple data units and where the comparison window includes a corrupted string. Instruction 920 can be executed to identify multiple manifests including the first manifest, where each manifest in the multiple manifests is included in a different snapshot of a backup item.

[0090] Instruction 930 can be executed to determine a matching score for multiple manifests based on a match with a comparison window. Instruction 940 can be executed to identify multiple repair strings based on the matching scores of the multiple manifests. Instruction 950 can be executed to record the corrupted string and the multiple repair strings in the first entry of a repair string data structure. Instruction 960 can be executed to repair the corrupted string in the first manifest using at least one of the identified multiple repair strings.

[0091] According to some embodiments of the present disclosure, a controller of a deduplication storage system can repair a snapshot that includes a corrupted string. The controller can identify the corrupted string included in the snapshot of the source item. The controller can identify a first manifest that references a portion of the snapshot that includes the corrupted string. The controller can then select a comparison window from the first manifest and can compare the comparison window with a set of candidate manifests associated with a previous snapshot of the source item. The controller can assign a matching score to each candidate manifest based on the similarity of each candidate manifest to the comparison window. Further, the controller can identify a set of repair strings based on the matching scores and can record the corrupted string and the set of repair strings in an entry of the stored data structure. Subsequently, the controller can compare the received input string with the stored data structure and can thereby determine whether the input string matches the corrupted string recorded in the entry of the stored data structure. If there is a match, the controller can read the entry to identify the repair string for the corrupted string. The controller can use the repair string to perform a repair attempt and can determine whether the repair attempt is effective (e.g., based on input from a human user or an automated testing process). If it is determined that the first repair attempt is not effective, the controller can use the stored repair history data structure to undo the first repair attempt, read the entry to identify a second repair string, and use the second repair string to perform another repair attempt. In this way, multiple repair attempts can be made until it is determined that the repair is effective or that there are no more repair strings available. Accordingly, some embodiments can reduce the amount of corrupted data stored in the storage system. Further, some embodiments can reduce the data loss and performance impact associated with recovering from each instance of the corrupted string.

[0092] Note that although Figures 1 to 9 various examples are shown, embodiments are not limited in this regard. For example, referring to Figure 1 , it can be envisioned that storage system 100 can include additional devices and / or components, fewer components, different components, different arrangements, etc. In another example, it can be envisioned that the functionality of the storage controller 110 described above can be included in any other engine or software of the storage system 100. Other combinations and / or variations are also possible.

[0093] Data and instructions are stored in respective storage devices implemented as one or more computer-readable or machine-readable storage media. The storage media include different forms of non-transitory memory, including: semiconductor memory devices such as dynamic random access memory or static random access memory (DRAM or SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory; magnetic disks such as fixed floppy disks and removable disks; other magnetic media including magnetic tape; optical media such as compact discs (CDs) or digital video discs (DVDs); or other types of storage devices.

[0094] Note that the instructions discussed above can be provided on a single computer-readable or machine-readable storage medium or, alternatively, can be provided on multiple computer-readable or machine-readable storage media distributed in a large system having potentially multiple nodes. Such one or more computer-readable or machine-readable storage media are considered to be part of an article (or article of manufacture). An article or article of manufacture can refer to any single manufactured component or multiple components. The one or more storage media can be located within the machine that executes the machine-readable instructions or, alternatively, can be located at a remote site from which the machine-readable instructions can be downloaded over a network for execution.

[0095] In the foregoing description, numerous details are set forth in order to provide an understanding of the subject matter disclosed herein. However, embodiments may be practiced without some of these details. Other embodiments may include modifications and variations of the details discussed above. The appended claims are intended to cover such modifications and variations.

Claims

1. A computing device, comprising: A processor; A memory; And A machine-readable storage device storing instructions executable by the processor to perform the following operations: Select a comparison window in a first manifest of a deduplication storage system, wherein the comparison window includes a plurality of data units, and wherein the comparison window includes a corrupted string; Identify a plurality of manifests including the first manifest, wherein each manifest in the plurality of manifests is included in a different snapshot of a backup item; Determine a matching score of the plurality of manifests based on a match with the comparison window; Identify a plurality of repair strings based on the matching scores of the plurality of manifests; Record the corrupted string and the plurality of repair strings in a first entry of a repair string data structure; and Use at least one of the plurality of repair strings to repair the corrupted string in the first manifest.

2. The computing device according to claim 1, including instructions executable by the processor to perform the following operations after repairing the corrupted string in the first manifest: Receive a new data string to be stored in the deduplication storage system; and In response to determining that the new data string matches the corrupted string recorded in the first entry of the repair string data structure, use one or more of the plurality of repair strings recorded in the first entry of the repair string data structure to repair the new data string.

3. The computing device according to claim 2, including instructions executable by the processor to perform the following operations: Select a first repair string with the highest ranking from the plurality of repair strings recorded in the first entry; Perform a first attempt to repair the new data string using the first repair string; And Record information about the first attempt to repair in a repair history data structure.

4. The computing device according to claim 3, including instructions executable by the processor to perform the following operations: Determine whether the first attempt to repair is effective; In response to determining that the first attempt to repair is ineffective: Obtain the information about the first attempt to repair from the repair history data structure; and Use the obtained information about the first attempt to repair to undo the first attempt to repair.

5. The computing device according to claim 4, including instructions executable by the processor to perform the following operations in response to determining that the first attempt to repair is ineffective: Select a second repair string with the second highest ranking from the plurality of repair strings recorded in the first entry; and Perform a second attempt to repair the new data string using the second repair string.

6. The computing device according to claim 3, wherein, The information about the first attempt to repair includes: The location of the first attempt to repair; The corrupted string repaired in the first attempt to repair; The first repair string used in the first attempt to repair; and The timestamp of the first attempt to repair.

7. The computing device according to claim 2, including instructions executable by the processor to perform the following operations: Compare the new data string with multiple entries in the repair string data structure; Determine that the new data string includes a first set of data units and a second set of data units, wherein, The first set of data units matches the corrupted string recorded in the first entry of the repair string data structure, and wherein the second set of data units includes a count of data units not included in the corrupted string recorded in the first entry; and In response to determining that the count of data units in the second set does not exceed a maximum threshold of additional data units, determine that the new data string matches the corrupted string recorded in the first entry of the repair string data structure.

8. The computing device according to claim 1, wherein, Identify each repair string among the multiple repair strings in different manifests of the multiple manifests.

9. A method, comprising: Select, by a storage controller, a comparison window in a first manifest of a deduplication storage system, wherein the comparison window includes multiple data units, and wherein the comparison window includes a corrupted string; Identify, by the storage controller, multiple manifests including the first manifest, wherein each of the multiple manifests is included in a different snapshot of a backup item; Determine, by the storage controller, a match score of the multiple manifests based on a match with the comparison window; Identify, by the storage controller, multiple repair strings based on the match score of the multiple manifests; Record, by the storage controller, the corrupted string and the multiple repair strings in a first entry of a repair string data structure; and Repair, by the storage controller, the corrupted string in the first manifest using at least one repair string among the multiple repair strings.

10. The method of claim 9, comprising, after repairing the corrupted string in the first manifest: Receive a new data string to be stored in the deduplication storage system; Determine whether the new data string matches the corrupted string recorded in the first entry of the repair string data structure; And In response to determining that the new data string matches the corrupted string recorded in the first entry of the repair string data structure, repair the new data string using one or more repair strings among the multiple repair strings recorded in the first entry of the repair string data structure.

11. The method of claim 10, comprising: Select a first repair string with the highest ranking from the multiple repair strings recorded in the first entry; Perform a first attempt to repair the new data string using the first repair string; And Record information about the first attempt to repair in a repair history data structure.

12. The method of claim 11, comprising: Determine whether the first attempt to repair is effective; In response to determining that the first attempt to repair is ineffective: Obtain the information about the first attempt to repair from the repair history data structure; Use the obtained information about the first attempt to repair to undo the first attempt to repair.

13. The method according to claim 12, Comprising, in response to determining that the first attempt to repair is ineffective: Select a second repair string with the second highest ranking from the multiple repair strings recorded in the first entry; And Perform a second attempt to repair the new data string using the second repair string.

14. The method according to claim 10, comprising: comparing the new data string with a plurality of entries in the repair string data structure; determining that the new data string includes a first set of data units and a second set of data units, wherein the first set of data units matches the damaged string recorded in the first entry of the repair string data structure, and wherein the second set of data units includes a count of data units not included in the damaged string recorded in the first entry; and responsive to determining that the count of data units in the second set does not exceed a maximum threshold of additional data units, determining that the new data string matches the damaged string recorded in the first entry of the repair string data structure.

15. A non-transitory machine-readable medium storing instructions that, when executed, cause a processor to perform the following operations: Select a comparison window in the first manifest of the deduplication storage system, where, The comparison window includes a plurality of data units, and wherein the comparison window includes a damaged string; identifying a plurality of manifests including the first manifest, wherein each of the plurality of manifests is included in a different snapshot of a backup item; determining a match score of the plurality of manifests based on a match with the comparison window; identifying a plurality of repair strings based on the match scores of the plurality of manifests; recording the damaged string and the plurality of repair strings in a first entry of a repair string data structure; and using at least one of the plurality of repair strings to repair the damaged string in the first manifest.

16. The non-transitory machine-readable medium according to claim 15, comprising instructions that, when executed, cause the processor to perform the following operations after repairing the damaged string in the first manifest: receiving a new data string to be stored in the deduplication storage system; and responsive to determining that the new data string matches the damaged string recorded in the first entry of the repair string data structure, using one or more of the plurality of repair strings recorded in the first entry of the repair string data structure to repair the new data string.

17. The non-transitory machine-readable medium according to claim 16, comprising instructions that, when executed, cause the processor to perform the following operations: selecting a first repair string with the highest ranking from the plurality of repair strings recorded in the first entry; performing a first attempt to repair the new data string using the first repair string; and recording information about the first attempt to repair in a repair history data structure.

18. The non-transitory machine-readable medium according to claim 17, comprising instructions that, when executed, cause the processor to perform the following operations: determining whether the first attempt to repair is effective; responsive to determining that the first attempt to repair is ineffective: obtaining the information about the first attempt to repair from the repair history data structure; using the obtained information about the first attempt to repair to undo the first attempt to repair.

19. The non-transitory machine-readable medium according to claim 18, comprising instructions that, when executed, cause the processor to perform the following operations in response to determining that the first attempt to repair is ineffective: Select a second repair string having a second highest rank from the plurality of repair strings recorded in the first entry; and Perform a second attempt to repair the new data string using the second repair string.

20. The non-transitory machine-readable medium of claim 16, comprising instructions that, when executed, cause the processor to perform the following operations: Compare the new data string with a plurality of entries in the repair string data structure; Determine that the new data string includes a first set of data units and a second set of data units, where, The first set of data units matches the corrupted string recorded in the first entry of the repair string data structure, and wherein the second set of data units includes a count of data units not included in the corrupted string recorded in the first entry; and In response to determining that the count of data units in the second set does not exceed a maximum threshold of additional data units, determine that the new data string matches the corrupted string recorded in the first entry of the repair string data structure.