Method and device for controlling magnetic tape storage

By creating data segment tables and optimizing read sequences in tape storage, the problem of long search time in tape storage is solved, and efficient data reading and deduplication are achieved.

CN120390957APending Publication Date: 2025-07-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380088013.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, tape storage takes a long search time when performing deduplication, resulting in low data reading efficiency, and existing methods require a large amount of preprocessing and reduce the deduplication rate.

Method used

By creating a data segment table, including the identifier, fingerprint and position offset of the data segment, generate a read sequence to optimize the read order of the data segment, reduce complexity with anchor points and adjacent position offsets, and select a read sequence with the minimum read time.

Benefits of technology

It effectively shortens the seek time of tape storage, improves data reading performance, reduces complexity and improves reading efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390957A_ABST
    Figure CN120390957A_ABST
Patent Text Reader

Abstract

To control tape storage, a data segment table is created that includes, in each entry, an identifier of a data segment stored on a tape and one or more positional offsets on the tape that store the data segment or a copy of the data segment. In response to receiving a recovery request identifying a user data object to be recovered, related data segments to be read from the tape are determined, a set of read option lists is created to read all of the related data segments, and a recovery process is performed by selecting a position offset from each read option list and determining a read data order with a minimum read time. A read sequence is generated to read the relevant data segments and / or copies of the relevant data segments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of data storage, and more particularly, to methods and apparatuses for controlling tape storage having a plurality of data segments stored on a tape composed of a plurality of wraps. Background Art

[0002] Generally, various storage devices such as hard disks, pen drives, memory cards, etc. are used to store digital data. Additionally, the demand for storing digital data is increasing, such that tape technology is being used for implementing auxiliary (slower) data storage due to its relatively low cost and long retention time. However, since the seek time for reading data from traditional tapes is relatively long, tapes are used without any optimization other than built-in compression.

[0003] Traditionally, certain attempts have been made to address the long seek time issue when using tapes, such as writing data to the tape in its original (i.e., without deduplication) form. Another existing method for addressing the excessive seek time issue when storing deduplicated data on a tape involves using closed data groups. In this case, data groups (e.g., including a full backup and a series of incremental backups) are created and sent to a tape where deduplication is performed only within the group (i.e., a closed unit formed together with other parts of the group, such as similarity locality (SILO)), such that the data group can be read from the entire tape. In some scenarios, to mitigate the long seek time issue, deduplicated data can also be stored on the tape in a special way to shorten the recovery time, e.g., by writing common data to two files such that the common data can be read sequentially from the tape. However, such attempts require significant preprocessing of the common data and greatly reduce the deduplication rate. Thus, there is a technical problem of how to read data (data segments scattered across tape storage) from a tape with a shorter seek time (the time required to retrieve data stored in a data segment), while allowing deduplication to optimize future reads and improve recovery performance.

[0004] Therefore, in light of the foregoing discussion, there is a need to overcome the above-mentioned drawbacks associated with traditional techniques for writing data to and / or reading data from a tape storing deduplicated data. Summary of the Invention

[0005] The present invention provides a method and apparatus for controlling tape storage, the tape storage having a plurality of data segments stored on a tape composed of a plurality of reels. The present invention provides a solution to an existing problem, namely how to implement data deduplication on the tape while shortening the seek time when reading data, that is, how to improve the data reading efficiency of the data on which data deduplication has been performed. The object of the present invention is to provide a solution that at least partially overcomes the problems encountered in the prior art and provides an improved method for controlling tape storage (the tape storage having a plurality of data segments stored on a tape composed of a plurality of reels), for example, for restoring data from an archive with a shorter seek time.

[0006] One or more objects of the present invention are achieved by the solutions provided in the appended independent claims. Advantageous implementations of the present invention are further defined in the dependent claims.

[0007] In one aspect, the present invention provides a method for controlling tape storage, the tape storage having a plurality of data segments stored on a tape composed of a plurality of reels. The method includes creating a data segment table, the data segment table including, in each entry, an identifier (ID) of a data segment stored on the tape. Each entry is associated with unique data segment metadata, the data segment metadata including a fingerprint of the data segment, a size of the data segment, and one or more position offsets on the tape where the data segment or a copy of the data segment is stored. The method further includes receiving a recovery request identifying a user data object to be restored from the tape storage, and determining relevant data segments to be read from the tape within the recovery of the user data object. The method further includes creating a set of read option lists for reading all the relevant data segments from the tape. Based on the data segment table, each read option list includes position offsets of different relevant data segments among the relevant data segments, the position offsets of the different relevant data segments being supplemented by position offsets of all copies of the relevant data segments, and generating a read sequence for reading the relevant data segments and / or copies of the relevant data segments from the tape by selecting one position offset from each read option list and determining an order for reading data from the selected position offset with the minimum read time.

[0008] The method is used to effectively control the tape storage and shorten the seek time for restoring a data object from the tape. The method is used to improve the reading performance of the tape, for example, by processing the repeated positions of the data segments in the tape and generating the read sequence to provide the best order for reading data segments / copies of data segments, so as to read the relevant data more effectively. The method is used to read data from tape storage, reducing complexity and improving performance.

[0009] In one implementation, the method further includes setting one or more lists of read options consisting of a single position offset as an anchor. Additionally, generating the read sequence to read the relevant data segment and / or a copy of the relevant data segment from the tape includes: selecting the anchor and the position offset of the data segment adjacent to the anchor from the list of read options, and determining a read order to form one or more consecutive read subsequences within the read sequence.

[0010] In this implementation, setting an anchor that must be read from a single position in any case and selecting the position offset adjacent to the anchor significantly reduces the complexity of determining the optimal read sequence for recovering the relevant data. Additionally, determining one or more consecutive read subsequences in the read sequence is used to reduce the seek time required to read the relevant data from the tape and shorten the total processing time.

[0011] In another implementation, determining the relevant data segment to be read from the tape is based on a data object table that identifies one or more tapes, the one or more tapes including data segments of each user data object stored in the tape storage and the ID and / or position offset of the data segments on the one or more tapes.

[0012] In this implementation, using the data object table to determine the relevant data segment to be read enables the optimized reading of relevant data from one or more tapes.

[0013] In another implementation, selecting a read sequence from the read sequences to read the relevant data segment and / or a copy of the relevant data segment from the tape includes: considering the current head of the position offset and the tape position that make up the read sequence, and selecting the read sequence with the minimum read time.

[0014] In this implementation, selecting the read sequence based on the minimum read time is used to improve the reading performance and shorten the total processing time required to read the relevant data segment from the tape.

[0015] In another aspect, the present invention provides a controller for tape storage, the tape storage having a plurality of data segments stored on a tape composed of a plurality of reels. The controller is configured to create a data segment table, the data segment table including, in each entry, an identifier ID of the data segment stored on the tape. Each entry is associated with unique data segment metadata, the data segment metadata including a fingerprint of the data segment, a size of the data segment, and one or more position offsets on the tape where the data segment or a copy of the data segment is stored. The controller is further configured to receive a recovery request identifying a user data object to be recovered from the tape storage, and determine relevant data segments to be read from the tape within the recovery of the user data object. The controller is further configured to create a set of read option lists for reading all the relevant data segments from the tape. Based on the data segment table, each read option list includes position offsets of different relevant data segments among the relevant data segments, the position offsets of the different relevant data segments being supplemented by position offsets of all copies of the relevant data segments, and a read sequence is generated for reading the relevant data segments and / or copies of the relevant data segments from the tape by selecting one position offset from each read option list and determining an order for reading data from the selected position offset with a minimum read time.

[0016] The controller is configured to effectively control the tape storage and reduce the seek time for recovering a data object from the tape. The controller is configured to improve the read performance of the tape, for example, by processing the repeated positions of the data segments in the tape and generating the read sequence to provide an optimal order for reading the data segments / copies of the data segments, thereby reading the relevant data more effectively. The controller is configured to read data from the tape storage, reducing complexity and improving performance.

[0017] In one implementation, the controller is configured to set one or more read option lists consisting of a single position offset as an anchor point. Further, generating the read sequence for reading the relevant data segments and / or copies of the relevant data segments from the tape includes: selecting the anchor point and position offsets of the data segments adjacent to the anchor point from the read option lists, and determining a read order to form one or more consecutive read subsequences within the read sequence.

[0018] In this implementation, setting the anchor point that must be read from a single position in any case and selecting the position offsets adjacent to the anchor point significantly reduces the complexity of determining the optimal read sequence for recovering the relevant data. Further, determining one or more consecutive read subsequences in the read sequence is used to reduce the seek time required for reading the relevant data from the tape and shorten the total processing time.

[0019] In another implementation, determining the relevant data segment to be read from the tape is based on a data object table that identifies one or more tapes. The one or more tapes include data segments of each user data object stored in the tape storage, as well as the IDs and / or position offsets of the data segments on the one or more tapes.

[0020] In this implementation, using the data object table to determine the relevant data segment to be read enables optimized reading of relevant data from one or more tapes.

[0021] In another implementation, selecting one of the read sequences in the read sequence to read the relevant data segment and / or a copy of the relevant data segment from the tape includes: considering the current head of the position offsets and tape positions that make up the read sequence, and selecting the read sequence with the minimum read time.

[0022] In this implementation, selecting the read sequence based on the minimum read time is used to improve the read performance and shorten the total processing time required to read the relevant data segment from the tape.

[0023] It should be understood that all of the above implementations can be combined. It should be noted that all devices, elements, circuit systems, units, and modules described in this application can be implemented in software or hardware elements or any type of combination thereof. All steps performed by the various entities described in this application and the functions to be performed by the various entities described are intended to mean that each entity is adapted or used to perform the corresponding steps and functions. Although in the description of the following specific embodiments, the specific functions or steps to be performed by external entities are not reflected in the description of the specific detailed elements of the entities performing the specific steps or functions, those skilled in the art should be clear that these methods and functions can be implemented in the corresponding software or hardware elements or any type of combination thereof. It should be understood that the features of the present invention are easily combinable in various combinations without departing from the scope of the present invention defined by the appended claims.

[0024] Other aspects, advantages, features, and purposes of the present invention will become apparent from the drawings and the detailed description of the illustrative implementations explained in conjunction with the following appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above invention content and the detailed description of the following illustrative embodiments are better understood when read in conjunction with the drawings. For the purpose of illustrating the present invention, an exemplary structure of the present invention is shown in the drawings. However, the present invention is not limited to the specific methods and tools disclosed herein. In addition, those skilled in the art will understand that the drawings are not drawn to scale. Where possible, the same elements are denoted by the same reference numerals.

[0026] Embodiments of the present invention will now be described by way of example only with reference to the following drawings, in which:

[0027] Figure 1 is a flowchart of a method for controlling tape storage according to an embodiment of the present invention;

[0028] Figure 2 is a diagram depicting the mapping of multiple data segments to a data segment table according to an embodiment of the present invention;

[0029] Figure 3 is a diagram of a system for performing a method of controlling tape storage according to an embodiment of the present invention.

[0030] In the drawings, underlined reference numerals are used to denote the item in which the underlined reference numeral is located or the item adjacent to the underlined reference numeral. Non-underlined reference numerals are related to the items identified by the lines associating the non-underlined reference numerals with the items. When a reference numeral is non-underlined and accompanied by an associated arrow, the non-underlined reference numeral is used to identify the general item to which the arrow points. Detailed Description of the Invention

[0031] The following detailed description illustrates embodiments of the present invention and the ways in which these embodiments can be implemented. Although some modes of implementing the present invention have been disclosed, those skilled in the art will recognize that other embodiments for implementing or practicing the present invention are also possible.

[0032] Figure 1 is a flowchart of a method for controlling tape storage according to an embodiment of the present invention, the tape storage having multiple data segments stored on a tape composed of multiple reels of tape. Referring to Figure 1 , a flowchart of method 100 including steps 102 to 110 is shown.

[0033] The present invention provides a method 100 for controlling tape storage, which includes a plurality of data segments stored on a tape composed of multiple reels. In an implementation, the tape storage includes a plurality of tapes, and each tape is composed of tracks on which data is written and read. Additionally, the length of each track from the beginning to the end of the tape is one length of the tape, where a group of adjacent tracks is called a reel. The plurality of data segments are created by using a segmentation process. In an implementation, a large buffer is divided into a plurality of data segments of the same size (e.g., 16 kilobytes (KB), 32 KB, etc.), and this type of segmentation is called fixed-size segmentation. However, fixed-size segmentation is difficult to use because inserting or deleting a few bytes in small segments changes all data segments, which further increases the complexity of processing the tape. In another implementation, the large buffer is divided (e.g., according to the window size of the data) into a plurality of data segments. In this case, a rolling hash function is applied on a rolling window (e.g., with a size range from 64 bytes to 255 bytes). Additionally, when the rolling hash function meets certain conditions, a plurality of data segments are created, and the conditions are selected in such a way as to obtain the desired size (such as the minimum, maximum, or average size) of the plurality of data segments. Further, the tape is divided into a number of bands such that the tape drive head covers the width of a single band, and each band is composed of an even number of reels of a plurality of reels. The head of the tape drive is used for reading and writing and includes a plurality of reading and writing elements such that multiple adjacent tracks can be read / written. In an example, such a group of adjacent tracks is called a reel (the tape includes a plurality of reels). For example, a linear tape open ninth generation (LTO 9) tape includes four (4) bands, each band having fifty-two (52) tape reels (including twenty-six (26) tape reels from the beginning to the end and twenty-six (26) tape reels from the end to the beginning), and each tape reel has thirty-two (32) tracks. Additionally, for example, in LTO 9, the length of the tape (i.e., the length of each tape reel) is slightly greater than one (1) kilometre (km). The current invention is also applicable to LTO tapes of the tenth, eleventh, twelfth, or thirteenth generation, i.e., LTO 10, LTO 11, LTO 12, or LTO 13.

[0034] At step 102, method 100 includes creating a data segment table that includes, in each entry, an identifier ID of a data segment stored on a tape. In other words, the ID of the data segment is included in each entry of the data segment table, and each ID is a unique identifier of the data segment stored on the tape. Additionally, each entry is associated with unique data segment metadata, where the data segment metadata includes a fingerprint of the data segment, a size of the data segment, and one or more position offsets on the tape where the data segment or a copy of the data segment is stored. The fingerprint of the data segment refers to a representation of the data segment stored on the tape. For example, the fingerprint of the data segment can be extracted by feeding the data segment sampled from the data into a strong hash (e.g., SHA1 that outputs a 20-byte cryptographic hash or a newer version SHA that outputs a longer hash value). The fingerprint of the data segment is used in data-specific deduplication mechanisms, such as for comparing the fingerprints of data buffers instead of comparing the data buffers byte by byte. In an example, the probability of a false positive comparison can be selected. For example, when using a strong hash, the probability of a false positive is very low. For example, for a one (1) petabyte (PB) storage with a segment size of approximately 16 kilobytes (KB), obtaining 2 36 segments, where the SHA1 signature is 20 bytes and equal to 2 160 , i.e., the false positive probability is 1:2 124 .

[0035] In addition, the data stored in the tape is represented as a full index stored in the fast storage layer. For example, if the number of entries stored in each full tape = 45TB / segment size (45KB) = ~1G entries, then in this case, without optimization, each entry in the tape storage includes an ID (e.g., 8 bytes), an offset (e.g., 8 bytes), a size (e.g., 4 bytes), a hash (e.g., 20 bytes, such as SHA1), and thus, the total size of the tape storage (or full cartridge) is up to 40GB. In addition, the full index includes a list of all fingerprints of multiple data segments. However, using the full index to compare the fingerprints of multiple data segments is not efficient due to the long processing time. Therefore, a sparse index is used to compare the fingerprints of multiple data segments. The sparse index is a subset of the fingerprints from the full index (i.e., a smaller index), for example, a subset of the fingerprints of the original data of a determined length (e.g., there are two entries for every 4 megabytes (MB) of data instead of one entry for every 16KB segment). In an implementation, the sparse index of the tape storage is established at a ratio of N hashes per gigabyte (GB) of data, and each hash is correspondingly selected, for example, selected according to the minimum or maximum value. In addition, the ratio of N hashes can be adjusted according to the size of the tape storage. In an example, the number of entries in the sparse index is 45TB x N / 1GB = 45K x N. In addition, each entry includes a weak hash (e.g., 8 to 16 bytes) and a block ID (e.g., 4 bytes), and thus, the total size of the sparse index of each full tape is less than 1MB x N. In addition, the sparse index resides in the global sparse index database (DB) as an array as a separate table in the tape storage. In addition, the global sparse index corresponds to a second sparse index, and the second sparse index includes a lower resolution for each tape storage, for example, there are M entries for every 10GB, for example, 10GB / M >> 1GB / N. First, the sparse index uses the weak hash to point to the area with the relevant part of the original data. After that, the relevant part corresponding to the full index is loaded to compare the fingerprints of multiple data segments. In addition, the weak hash is a part of the strong hash (e.g., bytes 10 to 17 of the 20-byte hash value are the weak hash). In addition, the weak hash is used to detect areas in the old data that have a relatively high probability of including similar data. However, because the chance of false positives is high, the weak hash is not used to determine the actual data. In addition, the sparse index is managed in a key-value database (KVDB). In an example, the KVDB stores all entries of all tapes in the storage system. For example, the storage system can use one thousand (1000) tapes, and all entries are processed by the KVDB residing in the memory (e.g., in the tape).

[0036] According to an embodiment, method 100 further includes: when a new data segment is written to a tape, adding the ID and position offset of the new data segment to the data segment table. Additionally, if the new data segment is not a copy written to the tape, the ID and position offset are added as a new entry to the data segment table; or, if the new data segment is a copy of a data segment that has been written to the tape, the ID and position offset are added as part of an existing entry to the data segment table by adding the ID and position offset to the existing entry. In an example, the ID of the new data segment is added to an entry in the data segment table, and the position offset is added to the data segment metadata associated with that entry. In an implementation, if the new data segment is not a copy written to the tape, the ID and position offset are added as a new entry to the data segment table. In another implementation, if a copy of the new data segment is written to the tape, the ID and position offset are added to an existing entry in the data segment table that includes the ID of the copy. Thus, the method enables adding information about a new copy to an existing entry and keeping the data segment table up-to-date when new data is written to the tape.

[0037] Additionally, at step 104, method 100 includes receiving a recovery request that identifies a user data object to be recovered from tape storage. In other words, the recovery request corresponds to a request sent from a user device or a production site for identifying a user data object that needs to be recovered from tape storage. At step 106, method 100 further includes determining the relevant data segments to be read from the tape within the recovery of the user data object. According to an embodiment, determining the relevant data segments to be read from the tape is based on a data object table that identifies one or more tapes, the one or more tapes including the data segments of each user data object stored in the tape storage and the ID and / or position offset of the data segments on the one or more tapes. In an implementation, the data object table identifies one or more tapes, the one or more tapes including the data segments of each user data object stored in the tape storage and the ID of the data segments. In another implementation, the data object table identifies one or more tapes, the one or more tapes including the data segments of each user data object stored in the tape storage, and the position offset of the data segments on the corresponding one or more tapes. In yet another implementation, the data object table identifies one or more tapes, the one or more tapes including the data segments of each user data object stored in the tape storage, and the ID and position offset of the data segments on the corresponding one or more tapes.

[0038] At step 108, method 100 further includes creating a set of read option lists to read all relevant data segments from the tape. Based on the data segment table, each read option list in the set of read option lists includes the position offsets of different relevant data segments in the relevant data segments, and the position offsets of all copies of the relevant data segments supplement the position offsets of the different relevant data segments. The set of read option lists enables determination of all possible read options that can be used to read all relevant data segments from the tape, i.e., the positions storing the same data segments. In other words, to create each read option list, the position offset of one relevant data segment in the relevant data segments is supplemented by the position offsets of all copies of the relevant data segment obtained from the data segment table.

[0039] At step 110, method 100 generates a read sequence to read the relevant data segments and / or copies of the relevant data segments from the tape by selecting a position offset from each read option list and determining the order to read data from the selected position offset with the minimum read time. In different implementations, according to the selected position offset and the determined order to read data with the minimum read time, the read sequence is used to read only the relevant data segments from the tape, or only the copies of the relevant data segments from the tape, or most likely to read the relevant data segments and the copies of the relevant data segments from the tape to achieve the minimum read time.

[0040] According to an embodiment, method 100 further includes setting one or more read option lists consisting of a single position offset as anchors. In addition, in this embodiment, generating a read sequence to read the relevant data segments and / or copies of the relevant data segments from the tape includes: selecting the position offsets of the anchors and the data segments adjacent to the anchors from the read option lists, and determining the read order to form one or more continuous read subsequences within the read sequence. Therefore, determining one or more continuous read subsequences shortens the seek time required to read the relevant data from the tape and shortens the total processing time.

[0041] According to an embodiment, selecting a read sequence from the read sequences to read the relevant data segments and / or copies of the relevant data segments from the tape includes: considering the position offsets constituting the read sequence and the current head of the tape position, and selecting the read sequence with the minimum read time. Therefore, by selecting the read sequence based on the position offsets of the relevant data to be read from the tape and considering the current head of the tape position (i.e., the current position of the head of the tape drive), to achieve the minimum read time, method 100 is used to improve the read performance and shorten the total processing time required to read the relevant data segments from the tape.

[0042] Method 100 is for effectively controlling tape storage and shortening the seek time. Method 100 is for improving the read performance of a tape, such as by processing the repeated positions of the required data segments and generating a read sequence to provide an optimal order for more effectively reading the data. Accordingly, Method 100 is for reading data from a tape with a shortened seek time, thereby reducing complexity and improving read performance.

[0043] Steps 102 to 110 are merely illustrative and other alternatives may also be provided, where one or more steps are added, one or more steps are deleted, or one or more steps are provided in a different order, without departing from the scope of the claims herein.

[0044] Figure 2 is a diagram depicting the mapping of multiple data segments to a data segment table according to an embodiment of the present invention. Figure 2 Combined Figure 1 the elements in are described. Referring to Figure 2 , a mapping of multiple data segment IDs 202 to a data segment table 204 is shown.

[0045] A conventional or traditional data segment table includes the IDs, offsets, sizes, and related data of multiple data segments stored on a tape, such as in the exemplary table given in Table 1 below:

[0046] Table 1

[0047]

[0048]

[0049] As shown in Table 1 above, a conventional data segment table may include multiple entries for multiple data segments, regardless of whether there are any copies of the data segments stored on the tape. According to the present invention, the conventional data segment table (such as Table 1) is replaced or supplemented by a data segment table 204 to eliminate the entries with copies, as Figure 2As shown. In addition, when a new data segment is written to the tape, the data segment table 204 supplements the ID and position offset of the new data segment in the form of a new entry, or supplements the ID and position offset of the new data segment by adding this information to an existing entry based on whether the tape has any copies of the new data segment stored on the tape. In addition, the data segment table 204 is searched by the ID of the data segment from among multiple data segments to locate and read the data segment stored on the tape. In an implementation, if there is no copy of the new data segment written to the tape, the ID and position offset of the new data segment are added as a new entry to the data segment table. For example, since there is no copy stored on the tape, the data segment 202A (i.e., ID 9) is added as a new entry to the data segment table 204. In another implementation, if the new data segment to be written to the tape has a copy that has been written to the tape, in this case, the ID and position offset of the new data segment are added as part of an existing entry to the data segment table 204 (by adding the duplicate data segment position offset to the existing entry). For example, since the tape already has the same data segment 202C (i.e., ID 1) stored on it, the data segment 202B (i.e., ID 13) is added to the existing entry. Therefore, the duplicate ID and position offset are input into the existing entry to prevent adding duplicate data as a new entry to the data segment table 204. Creating and further maintaining the enhanced data segment table 204 provides a basis for implementing the method 100 as described above.

[0050] Figure 3 is an exemplary schematic diagram of a system for performing a method of controlling tape storage according to an embodiment of the present invention, the tape storage having a plurality of data segments stored on the tape. Refer to Figure 3 , an exemplary schematic diagram 300 is shown depicting a system 302 including a fast storage layer 304 and a tape storage 306.

[0051] In an implementation, system 302 is used to perform all the necessary calculations for segmenting, hashing, and maintaining multiple data segments. In an exemplary implementation, system 302 includes at least a fast storage layer 304, such as a first storage device 304A, a second storage device 304B, up to an nth storage device 304N. In an example, each of the first storage device 304A, the second storage device 304B, up to the nth storage device 304N corresponds to fast data storage such as a solid state drive (SSD). System 302 also includes a tape storage 306. In addition, examples of the tape storage 306 may include, but are not limited to, a tape library, an internal tape drive, etc. In an example, the tape storage 306 is configured according to LTO-9 or includes LTO-9 tapes, which have a volume of 18 terabytes (TB) and an average built-in compression of 1:2.5. Additionally, the data stored on the tape storage 306 is written to separate tapes or cartridges in a tape library 308, which includes a first tape 308A, a second tape 308B, up to an nth tape 308N. Furthermore, the tape storage 306 includes a fast storage portion 306A (a smaller fast storage layer compared to the fast storage layer 304), which is used to store deduplication-related metadata, such as using multiple solid state drives (SSDs) for storage, as Figure 3 shown. Specifically, the data stored on each of the tapes 308A to 308N is represented in corresponding full indexes in a full index database 310, which includes a first full index 310A, a second full index 310B, up to an nth full index 310N stored in the fast storage portion 306A, as Figure 3As shown in. Each full index in the full index databases 310A to 310N of the full index database 310 includes a list of all the fingerprints of multiple data segments stored on the corresponding tapes among the tapes 308A to 308N. In addition, the fast storage section 306A includes a sparse index database 312, which includes a first sparse index 312A, a second sparse index 312B, up to an nth sparse index 312N, which is used to accelerate data processing, for example, in order to initially compare data buffers. Each sparse index in the sparse indexes 312A to 312N is a subset of the fingerprints from the corresponding full index (310A to 310N), for example, a subset of the fingerprints of the original data of a determined length. In addition, each sparse index in the sparse indexes 312A to 312N resides in the fast storage section 306A of the tape storage 306 as an array or a separate table in a global sparse index 314 database (DB), and this global sparse index 314 database includes information (such as weak hashes) about the data stored on all the tapes 308A to 308N of the tape storage 306. In addition, the global sparse index 314 corresponds to a second sparse index that includes a lower-resolution subset of the data segment fingerprints stored in the tape storage 306.

[0052] In operation, to implement the control of the tape storage 306 according to the present invention, the system 302 is used to create a data segment table as described above (i.e., Figure 2 the data segment table 202). After that, the system 302 is used to receive a recovery request that identifies a user data object to be recovered from the tape storage 306, for example, by executing Figure 1 the method 100. Then, the system 302 is used to determine the relevant data segments to be read from one of the tapes 308A to 308N during the recovery of the user data object, and use the data segment table to create a set of read option lists to read all the relevant data segments from the tape. Finally, the system 302 is used to generate a read sequence to read the relevant data segments and / or copies of the relevant data segments from the target tape by selecting a position offset from each read option list and determining the order to read the data from the selected position offset with the minimum read time. In an example, the read sequence is generated based on the nearest neighbor algorithm so that the tape head can quickly move through multiple reel lengths of the target tape. Therefore, both the seek time required to read the relevant data and the total processing time of the system 302 to recover the request are shortened.

[0053] The embodiments of the present invention described above may be modified without departing from the scope of the present invention as defined by the appended claims. Expressions used to describe and claim the present invention such as "comprising", "including", "incorporating", "having", "being" are intended to be construed in a non-exclusive manner, i.e., items, components or elements not explicitly described may also be present. References to the singular should also be construed as referring to the plural. The term "exemplary" as used herein means "serving as an example, instance, or illustration". Any embodiment described as "exemplary" is not necessarily to be construed as more preferred or advantageous than other embodiments, and / or to exclude combinations of features of other embodiments. The term "optionally" as used herein is used to mean "provided in some embodiments and not provided in other embodiments". It should be understood that certain features of the present invention that are described in the context of separate embodiments for clarity may also be provided in combination in a single embodiment. Conversely, various features of the present invention that are described in the context of a single embodiment for clarity may also be provided separately or in any suitable combination or as appropriate in any other described embodiment of the present invention.

Claims

1. A method (100) for controlling a tape storage (306), the tape storage (306) having a plurality of data segments stored on a tape consisting of a plurality of reels, the method (100) comprising: Creating a data segment table (204), the data segment table (204) including in each entry an identifier ID of a data segment stored on the tape, wherein each entry is associated with unique data segment metadata, the data segment metadata including a fingerprint of the data segment, a size of the data segment, and one or more position offsets on the tape where the data segment or a copy of the data segment is stored, Receiving a recovery request identifying a user data object to be recovered from the tape storage (306), Determining the relevant data segments to be read from the tape within the user data object recovery, Based on the data segment table (204), creating a set of read option lists to read all the relevant data segments from the tape, wherein each read option list includes position offsets of different relevant data segments among the relevant data segments, the position offsets of the different relevant data segments being supplemented by position offsets of all copies of the relevant data segments, and Generating a read sequence to read the relevant data segments and / or copies of the relevant data segments from the tape by selecting one position offset from each read option list and determining an order to read data from the selected position offsets with a minimum read time.

2. The method (100) according to claim 1, further comprising: Setting one or more read option lists consisting of a single position offset as an anchor point, wherein generating the read sequence to read the relevant data segments and / or copies of the relevant data segments from the tape includes: selecting the anchor point and position offsets of data segments adjacent to the anchor point from the read option lists, and determining a read order to form one or more consecutive read subsequences within the read sequence.

3. The method (100) according to claim 1 or 2, wherein, Determining the relevant data segments to be read from the tape is based on a data object table that identifies one or more tapes, the one or more tapes including data segments of each user data object stored in the tape storage (306) and IDs and / or position offsets of the data segments on the one or more tapes.

4. The method (100) according to any one of claims 1 to 3, wherein Selecting one of the read sequences in the read sequence to read the relevant data segments and / or copies of the relevant data segments from the tape includes: considering the current head of the tape in terms of the position offsets and tape positions constituting the read sequence, and selecting the read sequence with the minimum read time.

5. The method (100) according to any one of claims 1 to 4, further comprising: When a new data segment is written to the tape, the ID and position offset of the new data segment are added to the data segment table (204), where if the new data segment is not written to a copy of the tape, the ID and position offset are added as a new entry to the data segment table (204), or if the new data segment is a copy of a data segment that has been written to the tape, the ID and position offset are added as part of the existing entry by adding the ID and position offset to the existing entry in the data segment table (204).

6. A controller for tape storage (306), the tape storage (306) having a plurality of data segments stored on a tape consisting of a plurality of reels, the controller for: Create a data segment table (204), where each entry in the data segment table (204) includes an identifier ID of a data segment stored on the tape, where Each entry is associated with unique data segment metadata, the data segment metadata including a fingerprint of the data segment, a size of the data segment, and one or more position offsets on the tape where the data segment or a copy of the data segment is stored, Receiving a recovery request identifying a user data object to be recovered from the tape storage (306), Determining relevant data segments to be read from the tape within the recovery of the user data object, Based on the data segment table (204), creating a set of read option lists to read all the relevant data segments from the tape, where each read option list includes position offsets of different relevant data segments among the relevant data segments, the position offsets of the different relevant data segments being supplemented by position offsets of all copies of the relevant data segments, and Generating a read sequence to read the relevant data segments and / or copies of the relevant data segments from the tape by selecting one position offset from each read option list and determining an order to read data from the selected position offset with the minimum read time.

7. The controller according to claim 6, further for: Set a list of one or more read options consisting of a single position offset as an anchor, where, Generating the read sequence to read the relevant data segments and / or copies of the relevant data segments from the tape includes: selecting the anchor point and position offsets of the data segments adjacent to the anchor point from the read option list, and determining a read order to form one or more consecutive read subsequences within the read sequence.

8. The controller according to claim 6 or 7, wherein, Determining the relevant data segments to be read from the tape is based on a data object table that identifies one or more tapes, the one or more tapes including data segments of each user data object stored in the tape storage (306) and IDs and / or position offsets of the data segments on the one or more tapes.

9. The controller according to any one of claims 6 to 8, wherein, Selecting one read sequence from the read sequences to read the relevant data segments and / or copies of the relevant data segments from the tape includes: considering the current head of the position offsets and tape positions constituting the read sequence, and selecting the read sequence with the minimum read time.

10. The controller according to any one of claims 6 to 9, further for: When a new data segment is written to the tape, the ID and position offset of the new data segment are added to the data segment table (204), where, If the new data segment is not written to the copy of the tape, add the ID and the position offset as a new entry to the data segment table (204), or, if the new data segment is a copy of a data segment that has been written to the tape, add the ID and the position offset as part of the existing entry to the data segment table (204) by adding the ID and the position offset to the existing entry.