Blood relationship-based tape deduplication

Through the combination of strong hashing and weak hashing based on historical usage patterns, the deduplication method of magnetic storage tapes is quickly identified, which solves the problems of low deduplication rate and long reading time of magnetic storage tapes, and achieves more efficient deduplication and cost-reducing effects.

CN120344949APending Publication Date: 2025-07-18HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084810.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When deduplication of existing magnetic storage tapes, there is a problem of low deduplication rate and long reading time, especially when using magnetic storage tapes, existing methods fail to effectively improve efficiency.

Method used

By identifying magnetic storage tapes for deduplication based on the historical usage pattern of incoming data, using the combination of strong hash and weak hash, quickly identify candidate tapes and update identifier mappings, optimizing the deduplication process.

Benefits of technology

It improves the deduplication rate and running time of magnetic storage tapes, reduces the seek time of reading magnetic storage tapes, and reduces the cost of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120344949A_ABST
    Figure CN120344949A_ABST
Patent Text Reader

Abstract

Systems and methods for tape data deduplication are described that include at least one range of data to be written to at least one tape, a list of identifiers for data objects in the data, segmenting the data and outputting a strong hash for each segment, determining a plurality of search representatives from the strong hash, the at least one processor outputs a weak hash for each selected strong hash, searches for identifiers in a mapping of identifiers to tapes to determine candidate tapes and searches for the weak hash in a sparse index, selects a tape having a maximum weak hash match as a result of the search, and outputs a weak hash for each selected strong hash. The strong hash of each segment is compared with a region of the one tape to which the match points, and the mapping is updated such that the identifier now points to at least one tape.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical Field and Background Art

[0002] In some embodiments, the present invention relates to systems and methods for improving the deduplication rate and runtime on magnetic storage tapes. More specifically, but not exclusively, the present invention relates to systems and methods for improving the deduplication rate and runtime by quickly identifying magnetic storage tapes on which to perform deduplication to obtain a higher deduplication rate.

[0003] The demand for data storage is increasing so much that the capacity of existing storage disks cannot meet this demand. Due to the relatively low price and long retention time of magnetic storage tapes, the focus of secondary storage and archiving has gradually shifted from storage disks with deduplication capabilities to magnetic storage tape technology. Since the seek time in magnetic storage tapes can be relatively long, for example, dozens of seconds, many applications do not perform any optimization when using tapes, except for the built-in compression function. Introducing a deduplication function in such applications may make the use of magnetic storage tapes less costly, but may result in a longer seek time for reading magnetic storage tapes.

[0004] One widely used method is not to use the deduplication function on magnetic storage tapes. A second method that has been applied is to maintain closed units ("silos") that include all deduplicated data. Each time data in a silo is needed, even if only a part of the deduplicated data is needed (e.g., to be restored), the entire silo is read. Summary of the Invention

[0005] The embodiments described herein are directed to improving the deduplication rate and runtime on magnetic storage tapes by quickly identifying magnetic storage tapes on which to perform deduplication to obtain a higher deduplication rate. By identifying magnetic storage tapes for deduplicating incoming data based on historical usage patterns of past incoming data, the performance of deduplication can be enhanced. The incoming data to be written to the tape is typically part of one or more logical entities (e.g., data of files, objects, backups, etc.). In earlier versions, such entities may have been written to at least one magnetic storage tape. To increase the deduplication rate, ideally, the incoming data is compared and deduplicated with tapes containing past data of the same entity.

[0006] According to one aspect of some embodiments of the present invention, there is provided at least one data range in a buffer to be written to at least one magnetic storage tape; a list of identifiers of data objects in the identified data; at least one processor for segmenting the data in the buffer and outputting a strong hash for each segment of the data; a plurality of search representatives in the determined strong hashes; a weak hash output by the at least one processor for each selected strong hash; a search engine for performing a search for each identifier in a mapping from identifiers to magnetic storage tapes to determine which magnetic storage tapes are candidate tapes and accordingly searching for the calculated weak hash in a sparse index of each candidate tape; a magnetic storage tape selected as a result of the search; the one magnetic storage tape having the largest number of weak hash matches; the at least one processor for comparing the strong hash of each segment of the data and the region of the one magnetic storage tape pointed to by the weak hash match; updating the mapping from identifiers to magnetic storage tapes such that the identifiers can point to the plurality of magnetic storage tapes.

[0007] According to one aspect of some embodiments of the present invention, there is provided a method including: identifying at least one data range in a buffer to be written to at least one magnetic storage tape; generating a list of identifiers of data objects in the identified data; segmenting the data in the buffer; outputting a strong hash for each segment of the data; selecting a plurality of search representatives from the determined strong hashes; outputting a weak hash for each selected strong hash; searching for each identifier in a mapping from identifiers to magnetic storage tapes to determine which magnetic storage tapes are candidate tapes; searching for the calculated weak hash in a sparse index of each candidate tape; selecting a magnetic storage tape having the largest number of weak hash matches; comparing the strong hash of each segment of the data and the region of the one magnetic storage tape pointed to by the weak hash match; updating the mapping from identifiers to magnetic storage tapes such that the identifiers now point to the at least one magnetic storage tape.

[0008] According to some embodiments of the present invention, if no identifier in the identifier mapping is determined in the search, a global index of data objects is queried to identify a list of magnetic storage tapes to be searched.

[0009] According to some embodiments of the present invention, the at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

[0010] According to some embodiments of the present invention, if the sub-range is less than a predetermined threshold, the sub-range is ignored.

[0011] According to some embodiments of the present invention, if the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, the search results are ignored and the data is searched among all stored magnetic storage tapes.

[0012] According to some embodiments of the present invention, the mapping of the identifier to the magnetic storage tape is stored on at least one of a hard disk drive or a solid state drive.

[0013] According to one aspect of some embodiments of the present invention, there is provided at least one data range to be written to at least one magnetic storage tape in a buffer; a list of identifiers of data objects in the identified data; at least one processor for segmenting the data in the buffer and outputting a strong hash for each segment of the data; a plurality of search representatives in the determined strong hashes; a weak hash output by the at least one processor for each selected strong hash; a search engine for performing a search for each identifier in a global index of data objects to identify a list of magnetic storage tapes to be searched and outputting to the processor one of the results of searching the global index or the results of mapping the identifiers generated based on the results of the search, wherein if the results of searching the global index provide a high deduplication rate, the results of searching the global index are output, otherwise, the results of the identifier mapping are output, whereby the processor locates a magnetic storage tape based on the output and updates the mapping of the identifier to the magnetic storage tape to use the one located magnetic storage tape such that the identifier now points to the at least one magnetic storage tape.

[0014] According to one aspect of some embodiments of the present invention, there is provided a method including: identifying at least one data range to be written to at least one magnetic storage tape in a buffer; generating a list of identifiers of data objects in the identified data; segmenting the data in the buffer; calculating a strong hash for each segment of the data; selecting a plurality of search representatives from the determined strong hashes; calculating a weak hash for each selected strong hash; searching for each identifier in a mapping of the identifier to the magnetic storage tape to determine which magnetic storage tapes are candidate tapes; searching a global index of data objects to identify a list of magnetic storage tapes to be searched; selecting one of the results of searching the global index or the results of mapping the identifiers generated based on the results of the search, wherein if the results of searching the global index provide a high deduplication rate, the results of searching the global index are selected, otherwise, the results of the identifier mapping are selected; locating a magnetic storage tape based on the selected results; and updating the mapping of the identifier to the magnetic storage tape to use the one located magnetic storage tape such that the identifier now points to the at least one magnetic storage tape.

[0015] According to some embodiments of the present invention, if an identifier in the identifier mapping is not determined in the search, a global index of data objects is queried to identify a list of magnetic storage tapes to be searched.

[0016] According to some embodiments of the present invention, the at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

[0017] According to some embodiments of the present invention, if the sub-range is less than a predetermined threshold, the sub-range is ignored.

[0018] According to some embodiments of the present invention, if the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, the search results are ignored and the data is searched in all stored magnetic storage tapes.

[0019] According to some embodiments of the present invention, the mapping of the identifier to the magnetic storage tape is stored on at least one of a hard disk drive or a solid state drive.

[0020] According to one aspect of some embodiments of the present invention, there is provided at least one data range to be written to at least one magnetic storage tape in a buffer; a list of identifiers of data objects in the identified data; at least one processor for segmenting the data in the buffer and outputting a strong hash for each segment of the data; a plurality of search representatives in the determined strong hashes; a weak hash output by the at least one processor for each selected strong hash; a search engine for performing a search for the weak hash in a global index of data objects to identify a list of candidate magnetic storage tapes to be searched, and if the search for the weak hash does not return a result, updating the global index with the weak hash; one magnetic storage tape as a result of the search selection, the one magnetic storage tape having the largest number of weak hash matches.

[0021] In accordance with one aspect of some embodiments of the present invention, a method is provided, comprising: identifying at least one data range in a buffer to be written to at least one magnetic storage tape; generating a list of identifiers of data objects in the identified data; segmenting the data in the buffer; calculating a strong hash for each segment of the data; selecting a plurality of search representatives from the determined strong hashes; calculating a weak hash for each selected strong hash; searching for each identifier in a mapping of identifiers to magnetic storage tapes to determine which magnetic storage tapes are candidate tapes; searching a global index of data objects for the weak hashes to identify a list of candidate magnetic storage tapes to be searched; selecting one magnetic storage tape selected as a result of the search of the identifier mapping, the one magnetic storage tape having the largest number of weak hash matches in the global index; and if searching the global index does not return a result, updating the global index with the weak hashes.

[0022] In accordance with some embodiments of the present invention, if an identifier in the identifier mapping is not determined in the search, a global index of data objects is queried to identify a list of magnetic storage tapes to be searched.

[0023] In accordance with some embodiments of the present invention, the at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

[0024] In accordance with some embodiments of the present invention, if the sub-range is less than a predetermined threshold, the sub-range is ignored.

[0025] In accordance with some embodiments of the present invention, if the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, the search result is ignored and the data is searched in all stored magnetic storage tapes.

[0026] In accordance with some embodiments of the present invention, the mapping of identifiers to magnetic storage tapes is stored on at least one of a hard disk drive or a solid state drive.

[0027] In accordance with one aspect of some embodiments of the present invention, there is provided a computer program for bloodline-based tape deduplication, the computer program comprising program instructions that, when executed by at least one processor, cause the at least one processor to: identify at least one data range in a buffer to be written to at least one magnetic storage tape; generate a list of identifiers of data objects in the identified data; segment the data in the buffer; output a strong hash for each segment of the data; select a plurality of search representatives from the determined strong hashes; output a weak hash for each selected strong hash; search for each identifier in the mapping from identifiers to magnetic storage tapes to determine which magnetic storage tapes are candidate tapes; search for the computed weak hash in the sparse index of each candidate tape; select one magnetic storage tape having the largest number of weak hash matches; compare the strong hash of each segment of the data with the region of the one magnetic storage tape pointed to by the weak hash match; update the mapping from identifiers to magnetic storage tapes such that the identifier now points to at least one magnetic storage tape.

[0028] In accordance with one aspect of some embodiments of the present invention, there is provided a computer program for bloodline-based tape deduplication, the computer program comprising program instructions that, when executed by at least one processor, cause the at least one processor to: identify at least one data range in a buffer to be written to at least one magnetic storage tape; generate a list of identifiers of data objects in the identified data; segment the data in the buffer; compute a strong hash for each segment of the data; select a plurality of search representatives from the determined strong hashes; compute a weak hash for each selected strong hash; search for each identifier in the mapping from identifiers to magnetic storage tapes to determine which magnetic storage tapes are candidate tapes; search a global index of data objects to identify a list of magnetic storage tapes to search; select one of the result of searching the global index or the result of searching the identifier mapping generated based on the result of the search, wherein if the result of searching the global index provides a high deduplication rate, then select the result of searching the global index, otherwise, select the result of the identifier mapping; locate one magnetic storage tape based on the selected result; update the mapping from identifiers to magnetic storage tapes to use the one located magnetic storage tape such that the identifier now points to at least one magnetic storage tape.

[0029] According to one aspect of some embodiments of the present invention, there is provided a computer program for tape deduplication based on blood relationship. The computer program includes program instructions that, when executed by at least one processor, cause the at least one processor to: identify at least one data range in a buffer to be written to at least one magnetic storage tape; generate a list of identifiers of data objects in the identified data; segment the data in the buffer; calculate a strong hash for each segment of the data; select a plurality of search representatives from the determined strong hashes; calculate a weak hash for each selected strong hash; search for each identifier in a mapping of identifiers to magnetic storage tapes to determine which magnetic storage tapes are candidate tapes; search a global index of data objects for the weak hashes to identify a list of candidate magnetic storage tapes to be searched; select one magnetic storage tape selected as a result of the search of the identifier mapping, the one magnetic storage tape having the largest number of weak hash matches in the global index; and update the global index with the weak hashes if the search of the global index returns no results.

[0030] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present invention, only exemplary methods and / or materials are described below. In case of conflict, the present patent specification (including definitions) shall prevail. In addition, these materials, methods, and examples are illustrative only and not necessarily limiting.

[0031] Implementations of the methods and / or systems provided by embodiments of the present invention may involve performing or completing selected tasks manually, automatically, or in a combination of both ways. In addition, according to the actual instruments and devices of the method and / or system embodiments of the present invention, multiple selected tasks may be completed by using hardware, software, firmware, or a combination of the three through an operating system.

[0032] For example, the hardware for performing selected tasks according to embodiments of the present invention may be implemented as a chip or a circuit. For software, performing selected tasks according to embodiments of the present invention may be a computer executing multiple software instructions through any suitable operating system. In an exemplary embodiment of the present invention, one or more of the tasks in the exemplary embodiments of the methods and / or systems described herein are performed by a data processor, such as a computing platform for executing multiple instructions. Optionally, the data processor includes volatile memory and / or non-volatile memory for storing instructions and / or data, such as a hard disk and / or a removable medium for storing instructions and / or data. Optionally, a network connection is also provided. Optionally, a display and / or a user input device such as a keyboard or a mouse are also provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Some embodiments of the present invention are described herein by way of example only in conjunction with the accompanying drawings. Specifically, with reference to the accompanying drawings in detail, it should be emphasized that the details shown are by way of example and for illustrative discussion of the embodiments of the present invention. In this regard, it will be apparent to those skilled in the art how to practice the embodiments of the present invention in accordance with the description of the accompanying drawings.

[0034] In the drawings:

[0035] Figure 1 A first embodiment of a system for bloodline-based tape deduplication described herein is shown;

[0036] Figure 2 is a simplified flowchart showing a first operating method of a system for bloodline-based tape deduplication described herein;

[0037] Figure 3A and Figure 3B A second embodiment of a system for bloodline-based tape deduplication described herein is shown;

[0038] Figure 4 is a simplified flowchart showing a second operating method of a system for bloodline-based tape deduplication described herein;

[0039] Figure 5 A third embodiment of a system for bloodline-based tape deduplication described herein is shown; and

[0040] Figure 6 is a simplified flowchart showing a third operating method of a system for bloodline-based tape deduplication described herein. Detailed Description of the Invention

[0041] Before explaining in detail at least one embodiment of the present invention, it should be understood that the present invention is not necessarily limited to the details of the structure and arrangement of components and / or methods set forth in the following description and / or illustrated in the drawings and / or examples. The present invention may have other embodiments or may be practiced or carried out in various ways.

[0042] Reference is now made to the accompanying drawings, Figure 1Shows a first embodiment of the blood-relationship-based tape deduplication system described herein. The figure depicts that buffer 101 includes at least one data range 106. The data objects in the data range have data object identifiers (IDs), such as 1234, 1235, 2367... 2379, etc. This data range is to be written to magnetic storage tapes, described herein as mst01, mst02, mst03, and mst04. Accordingly, processor 110 for writing the data range to the magnetic storage tape can compile a list of IDs.

[0043] The processor can segment the data in the buffer and output a strong hash for each segment of the data. For example, processor 110 can divide the data object with ID 2367 into multiple segments, such as segment 2367A, segment 2367B, segment 2367C. Segments such as segment 2367A, segment 2367B, segment 2367C generate strong hashes output by processor 110, such as strong hash (SH) SH: 7391, SH: 8259, SH9461. The processor can execute hash functions such as SHA1, SHA256, MD5, Tiger, and Blake3 to generate strong hashes.

[0044] Then, multiple search representatives can be selected from the strong hashes. For example, among the various segments output by the processor, for generating strong hashes, in Figure 1 the example, SH: 7391 and SH9461 can be selected as multiple search representatives. It should be understood that selecting SH: 7391 and SH9461 as search representatives in this example is only for illustration, and other more widely distributed strong hashes can also be selected in the actual implementation of the present invention.

[0045] Then, at least one processor 110 outputs a weak hash for each selected strong hash. For example, strong hash SH: 7391 can generate weak hash WH: 791, while strong hash SH9461 can generate weak hash WH: 941. The weak hash can be a part of the strong hash. For example, if the length of the strong hash is 20 bytes, the weak hash may be 10 bytes to 17 bytes out of the entire 20 bytes.

[0046] Search engine 120 can be executed by processor 110 or by a different processor similar to processor 110 on a different computer. The search engine can search for a given data object identifier, such as identifier 2367, in the mapping 130 from the identifier to magnetic storage tapes such as mst01, mst02, mst03, mst04. The search engine can determine which magnetic storage tapes are candidate tapes for further search. For example, as Figure 1As shown, when the search engine 130 searches for the identifier 2367, mst01 and mst03 are generated as candidate tapes for further search.

[0047] Then, for each candidate tape, a sparse index 140 can be searched based on one of the weak hashes associated with the identifier (detailed below). For example, in Figure 1 , the search engine has identified candidate tapes mst01 and mst03. Then, the search engine searches for WH: 941 in the sparse index 140. The search result is mst03. Then, the magnetic storage tape mst03 is selected and can then be loaded into the tape reading device. The relevant data objects are searched on mst03. If there are multiple matches, the search engine can indicate which magnetic storage tape has the largest number of weak hash matches and select that tape.

[0048] There may be a full index that includes mappings of all indexes to their associated strong hashes. However, to achieve efficient and fast processing, the full index is usually too large. Therefore, a sparse index 140 can be used. The sparse index 140 is smaller than the full index and typically includes a subset of the weak hashes associated with the strong hashes. For example, the full index can include entries for each 16KB data segment in buffer 101. In contrast, the sparse index 140 can include two entries for every four megabytes of data in the buffer. The sparse index 140 typically maps weak hashes (such as WH: 791 and WH: 941) to magnetic storage tapes.

[0049] For ease of description, the sparse index 140 is shown here to have a single entry relevant to the current example. It should be understood that this is only an example and not a limitation.

[0050] Figure 1 There may also be a global index (not shown) in the system shown, which includes a second sparse index where the resolution for each magnetic storage tape is lower than that of the magnetic storage tape 140. If the identifier in the identifier mapping is not determined in the search of the search engine 120, the global index can be queried. For example, the global index can set a certain number of entries (e.g., M entries) for every 10GB of data stored on the magnetic storage tape, while the number of entries per 1GB in the sparse index is larger (e.g., N entries, where 10GB / M >> 1GB / N). Each entry in the global index may include a mapping of the weak hash to the identification of a magnetic storage tape. The global index can include a key-value database where the weak hash is used as the search key. The global index (see the discussion below about the global index 250, reference Figure 2) may include an entry for each tape in the magnetic storage tapes in the system. Similar to the sparse index, since data may be stored on multiple magnetic storage tapes, searching for a single weak hash (i.e., key) may return multiple magnetic storage tapes.

[0051] Processor 110 may divide the data in the data range into data sub-ranges for each data object identifier. In this case, the processor performs segmentation so that no segment includes data across multiple data object identifiers. Additionally, if a data sub-range includes a data volume less than a predetermined threshold, the search engine may ignore the data in that sub-range.

[0052] If the number of magnetic storage tapes including the minimum index returned by the search engine is greater than a given number (which may be a predetermined number), the search result is ignored and the data is searched in all stored magnetic storage tapes. This is because, if there are many possible tapes to search in a search, after a certain point, there is little or no time saved by only looking at the index and not at the tapes themselves.

[0053] The mapping 130 of identifiers to magnetic storage tapes such as mst01, mst02, mst03, mst04 may be stored on a fast drive such as a hard disk drive or a solid state drive, rather than on a tape drive including magnetic storage tapes.

[0054] Now refer to Figure 2 , Figure 2 is a simplified flowchart showing a first operating method of the system for lineage-based tape deduplication described herein. At step 210, at least one data range to be written to at least one magnetic storage tape in the buffer is identified. At step 215, a list of identifiers of data objects in the identified data is generated. At step 220, the data in the buffer is segmented. At step 225, processor 110 outputs a strong hash for each segment of the data. At step 230, a plurality of search representatives are selected from the determined strong hashes. At step 235, processor 110 outputs a weak hash for each selected strong hash. At step 240, each identifier is searched in the mapping of identifiers to magnetic storage tapes to determine which tapes are candidate tapes. At step 245, the sparse index is searched for each candidate tape for the calculated weak hash. At step 250, one magnetic storage tape with the largest number of weak hash matches is selected. At step 255, the strong hash of each segment of the data is compared with the region of one magnetic storage tape pointed to by the weak hash match. At step 260, the mapping of identifiers to magnetic storage tapes is updated such that the identifier now points to at least one magnetic storage tape.

[0055] Now refer to Figure 3A and Figure 3B , Figure 3A andFigure 3B Shows a second embodiment of the bloodline-based tape deduplication system described herein. Figure 3A The data range 106, the processor 110, the strong hashes (e.g., SH: 7391, SH: 8259, and SH: 9461), and the weak hashes (WH: 791 and WH: 941) in the buffer 101 of Figure 1 the buffer 101 of

[0056] Now turning to Figure 3B , the search engine 220 can be the same as or similar to Figure 1 the search engine 120 of Figure 1 and, like Figure 1 the search engine 120 of

[0057] the search engine 120, can be executed by the processor 110 or a different processor similar to the processor 110 on a different computer. The search engine 220 can perform a search in the global index (as described above, referring to

[0058] As with the sparse index, since data may be stored on multiple magnetic storage tapes, searching for a single weak hash (i.e., key) may return multiple magnetic storage tapes. The search engine 220 can then determine whether to utilize the results or search the global index 250, or the data object map 230 and the sparse index 240. Such a decision can be based on the search path that returns fewer results.

[0059] It should be understood that the global index 250 generally allows for deduplication of data across various magnetic storage tapes. If the deduplication rate returned by the search results based on the data object map 230 and the sparse index 240 is lower than the search results based on the global index 250, then the search results based on the data object map 230 and the sparse index 240 are generally preferred. On the other hand, if the deduplication rate returned by the search results based on the global index 250 is lower than the search results based on the data object map 230 and the sparse index 240, then the search results based on the global index 250 are generally preferred. In the case where the search results of the global index are preferred, the data object map 230 can be updated to reflect these results. The appropriate magnetic storage tape (in this case mst03) can be loaded into the tape drive to read the desired data object from the magnetic storage tape.

[0060] Now refer to Figure 4 , Figure 4 is a simplified flowchart showing a second operational method of the system for lineage-based tape deduplication described herein. At step 410, at least one data range to be written to at least one magnetic storage tape in the buffer is identified. At step 415, a list of identifiers of data objects in the identified data is generated. At step 420, the data in the buffer is segmented. At step 425, a strong hash is calculated for each segment of the data. At step 430, a plurality of search representatives are selected from the determined strong hashes. At step 435, a weak hash is calculated for each selected strong hash. At step 440, each identifier is searched in the mapping of identifiers to magnetic storage tapes to determine which tapes are candidate tapes. At step 445, the global index of data objects is searched to identify a list of magnetic storage tapes to be searched. At step 450, one of the results of searching the global index or the results of searching the identifier map generated based on the search results is selected, where if the search results of the global index provide a high deduplication rate, the results of searching the global index are selected, otherwise, the results of the identifier map are selected. At step 455, a magnetic storage tape is located based on the selected result. At step 460, the mapping of identifiers to magnetic storage tapes is updated to use the located magnetic storage tape such that the identifier now points to at least one magnetic storage tape.

[0061] Now refer to Figure 5 , Figure 5Shows a third embodiment of the blood-relationship-based tape deduplication system described herein. Figure 5 The data range 106, the processor 110, the strong hashes (e.g., SH: 7391, SH: 8259, and SH: 9461) and the weak hashes (WH: 791 and WH: 941) in the buffer 101 of Figure 1 and Figure 3A The data range 106, the processor 110, the strong hashes and the weak hashes, and the magnetic storage tapes mst01, mst02, mst03, and mst04 in the buffer 101 of Figure 3B can be the same as or similar to Figure 1 and Figure 3B The global index 250 can be the same as or similar to

[0062] The global index 250. The search engine 320 can be the same as or similar to

[0063] the search engines 120 and 220 of Figure 6 , Figure 6FIG. 0 shows a simplified flowchart of a second method of operation of a blood-based tape deduplication system described herein. At step 610, at least one data range to be written to at least one magnetic storage tape is identified in a buffer. At step 615, a list of identifiers of data objects in the identified data is generated. At step 620, the data in the buffer is segmented. At step 625, a strong hash is computed for each segment of the data. At step 630, a plurality of search representatives are selected from the determined strong hashes. At step 635, a weak hash is computed for each selected strong hash. At step 640, each identifier is searched in a mapping of identifiers to magnetic storage tapes to determine which tapes are candidate tapes. At step 645, a global index 250 of data objects is searched to identify a list of magnetic storage tapes to be searched. At step 650, one magnetic storage tape selected as a result of the search identifier mapping is selected, i.e., one magnetic storage tape having the largest number of weak hash matches in the global index 250. At step 655, if the search of the global index does not return a result, the global index 250 is updated with the weak hashes.

[0064] It is expected that during the life of the patent for this application, many relevant methods and systems for magnetic storage tape deduplication will be developed, and the scope of the term magnetic storage tape deduplication is intended to a priori include all such new technologies.

[0065] The terms "comprising", "having", and variations thereof mean "including but not limited to".

[0066] The term "comprising" means "including and limited to".

[0067] The term "consisting essentially of" means that the composition, method, or structure may include additional ingredients, steps, and / or parts, provided that the additional ingredients, steps, and / or parts do not materially alter the basic and novel characteristics of the claimed composition, method, or structure.

[0068] Unless the context clearly dictates otherwise, the singular forms "a" and "the" as used herein include plural referents. For example, the term "a complex" or "at least one complex" may include multiple complexes, including mixtures thereof.

[0069] It should be understood that certain features of the invention that are described in the context of separate embodiments for clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention that are described in the context of a single embodiment for brevity may also be provided separately or in any suitable sub-combination or as any suitable other embodiment of the invention. Some features described in the context of the various embodiments are not considered to be essential features of those embodiments unless the embodiment is inoperable without those elements.

[0070] Although the present invention has been described in connection with specific embodiments thereof, it will be apparent to those skilled in the art that many alternatives, modifications and variations will be obvious. Therefore, all such alternatives, modifications and variations are intended to be covered within the spirit and broad scope of the appended claims.

[0071] It is the intention of the applicant that all publications, patents and patent applications mentioned in this specification be incorporated herein by reference in their entirety as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference when mentioned herein. Further, the citation or identification of any reference to this application is not to be construed as allowing such reference to take precedence over the present invention in the prior art. In respect of the use of section headings, the section headings should not be construed as necessarily limiting. Additionally, any one or more priority documents of this application are incorporated herein by reference in their entirety.

Claims

1. A system, characterized in that, Comprising: At least one data range in a buffer to be written to at least one magnetic storage tape; A list of identifiers of data objects in the identified data; At least one processor for segmenting the data in the buffer and outputting a strong hash for each segment of the data; Multiple search representatives in the determined strong hashes; Weak hashes output by the at least one processor for each selected strong hash; A search engine for performing a search for each identifier in a mapping of identifiers to magnetic storage tapes to determine which tapes are candidate tapes and searching for the calculated weak hashes in the sparse index of each candidate tape accordingly; One magnetic storage tape selected as a result of the search, the one magnetic storage tape having the largest number of weak hash matches; and The at least one processor for comparing the strong hash of each segment of the data and the area of the one magnetic storage tape pointed to by the weak hash match and updating the mapping of identifiers to magnetic storage tapes such that the identifier now points to at least one magnetic storage tape.

2. The system according to claim 1, wherein If an identifier in the identifier mapping is not determined in the search, query a global index of data objects to identify a list of magnetic storage tapes to search.

3. The system according to claim 1 or 2, characterized in that, The at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

4. The system according to claim 3, characterized in that, If the sub-range is less than a predetermined threshold, ignore the sub-range.

5. The system according to any one of claims 1 to 4, characterized in that If the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, ignore the search result and search for the data in all stored magnetic storage tapes.

6. The system according to any one of claims 1 to 5, characterized in that, The mapping of identifiers to magnetic storage tapes is stored on at least one of a hard disk drive or a solid state drive.

7. A method, characterized in that, Comprising: Identify at least one data range in a buffer to be written to at least one magnetic storage tape; Generate a list of identifiers of data objects in the identified data; Segment the data in the buffer; Output a strong hash for each segment of the data; Select multiple search representatives from the determined strong hashes; Output a weak hash for each selected strong hash; Search for each identifier in a mapping of identifiers to magnetic storage tapes to determine which tapes are candidate tapes; Search for the calculated weak hashes in the sparse index of each candidate tape; Select one magnetic storage tape having the largest number of weak hash matches; Compare the strong hash of each segment of the data and the area of the one magnetic storage tape pointed to by the weak hash match; And Update the mapping of identifiers to magnetic storage tapes such that the identifier now points to the one magnetic storage tape.

8. The method according to claim 7, wherein Further comprising: If an identifier in the identifier mapping is not determined in the search, search a global index of data objects to identify a list of magnetic storage tapes.

9. The method according to claim 7 or 8, characterized in that, Further comprising: The at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

10. The method according to claim 9, wherein If the sub-range is less than a predetermined threshold, ignore the sub-range.

11. The method according to any one of claims 7 to 10, characterized in that, If the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, the search results are ignored, and the data is searched among all stored magnetic storage tapes.

12. The method according to any one of claims 7 to 11, characterized in that The mapping of the identifier to the magnetic storage tape is stored on at least one of a hard disk drive or a solid state drive.

13. A system, characterized in that, Comprising: At least one data range in the buffer to be written to at least one magnetic storage tape; A list of identifiers of data objects in the identified data; At least one processor for segmenting the data in the buffer and outputting a strong hash for each segment of the data; A plurality of search representatives in the determined strong hashes; Weak hashes output by the at least one processor for each selected strong hash; And A search engine for performing a search for each identifier in a global index of data objects to identify a list of magnetic storage tapes to be searched, and outputting one of the results of searching the global index or the results of mapping the identifiers generated based on the results of the search to a processor, wherein if the results of searching the global index provide a high deduplication rate, the results of searching the global index are output, otherwise, the results of mapping the identifiers are output, whereby the processor locates a magnetic storage tape based on the output and updates the mapping of the identifier to the magnetic storage tape using the located magnetic storage tape such that the identifier now points to at least one magnetic storage tape.

14. The system according to claim 13, wherein If the identifier in the identifier mapping is not determined in the search, the global index of data objects is queried to identify a list of magnetic storage tapes to be searched.

15. The system according to claim 13 or 14, characterized in that, The at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

16. The system according to claim 15, wherein If the sub-range is less than a predetermined threshold, the sub-range is ignored.

17. The system according to any one of claims 13 to 16, characterized in that, If the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, the search results are ignored, and the data is searched among all stored magnetic storage tapes.

18. The system according to any one of claims 13 to 17, characterized in that, The mapping of the identifier to the magnetic storage tape is stored on at least one of a hard disk drive or a solid state drive.

19. A method, characterized in that, Comprising: Identifying at least one data range in the buffer to be written to at least one magnetic storage tape; Generating a list of identifiers of data objects in the identified data; Segmenting the data in the buffer; Calculating a strong hash for each segment of the data; Selecting a plurality of search representatives from the determined strong hashes; Calculating a weak hash for each selected strong hash; Searching for each identifier in the mapping of the identifier to the magnetic storage tape to determine which tapes are candidate tapes; Searching the global index of data objects to identify a list of magnetic storage tapes to be searched; Selecting one of the results of searching the global index or the results of mapping the identifiers generated based on the results of the search, wherein if the results of searching the global index provide a high deduplication rate, the results of searching the global index are selected, otherwise, the results of mapping the identifiers are selected; Locate a magnetic storage tape based on the selected result; and Update the mapping of the identifier to the magnetic storage tape to use the located magnetic storage tape such that the identifier now points to at least one magnetic storage tape.

20. The method according to claim 19, wherein Further comprising: If the identifier in the identifier mapping is not determined in the search, search the global index of data objects to identify a list of magnetic storage tapes.

21. The method according to claim 19 or 20, characterized in that, Further comprising: The at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

22. The method according to claim 21, characterized in that, If the sub-range is less than a predetermined threshold, ignore the sub-range.

23. The method according to any one of claims 19 to 22, characterized in that, If the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, ignore the search result and search for the data among all stored magnetic storage tapes.

24. The method according to any one of claims 19 to 23, characterized in that The mapping of the identifier to the magnetic storage tape is stored on at least one of a hard disk drive or a solid state drive.

25. A system, characterized in that, Comprising: At least one data range in a buffer to be written to at least one magnetic storage tape; A list of identifiers of data objects in the identified data; At least one processor for segmenting the data in the buffer and outputting a strong hash for each segment of the data; Multiple search representatives in the determined strong hashes; A weak hash output by the at least one processor for each selected strong hash; A search engine for performing a search for the weak hash in the global index of data objects to identify a list of candidate magnetic storage tapes to be searched, and if the search for the weak hash does not return a result, update the global index with the weak hash; and One magnetic storage tape selected as a result of the search, the one magnetic storage tape having the largest number of weak hash matches.

26. The system according to claim 25, wherein If the identifier in the identifier mapping is not determined in the search, query the global index of data objects to identify a list of magnetic storage tapes to be searched.

27. The system according to claim 25 or 26, characterized in that, The at least one processor segments the data into sub-ranges for each identifier such that no segment includes data across multiple identifiers.

28. The system according to claim 25, wherein If the sub-range is less than a predetermined threshold, ignore the sub-range.

29. The system according to any one of claims 25 to 28, characterized in that, If the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, ignore the search result and search for the data among all stored magnetic storage tapes.

30. The system according to any one of claims 25 to 29, characterized in that, The mapping of the identifier to the magnetic storage tape is stored on at least one of a hard disk drive or a solid state drive.

31. A method, characterized in that, Comprising: Identify at least one data range in a buffer to be written to at least one magnetic storage tape; Generate a list of identifiers of data objects in the identified data; Segment the data in the buffer; Calculate a strong hash for each segment of the data; Select multiple search representatives from the determined strong hashes; Calculate a weak hash for each selected strong hash; Search for each identifier in the mapping of the identifier to the magnetic storage tape to determine which tapes are candidate tapes; Search the global index of data objects for the weak hash to identify a list of candidate magnetic storage tapes to be searched; Select one magnetic storage tape selected as a result of searching for the identifier mapping, the one magnetic storage tape having the largest number of weak hash matches in the global index; and If searching the global index returns no results, update the global index with the weak hashes.

32. The method according to claim 31, wherein Further includes: If the identifier in the identifier mapping is not determined in the search, search the global index of data objects to identify a list of magnetic storage tapes.

33. The method according to claim 31 or 32, characterized in that, Further includes: The at least one processor segments the data into sub - ranges for each identifier such that no segment includes data across multiple identifiers.

34. The method according to claim 33, wherein If the sub - range is less than a predetermined threshold, ignore the sub - range.

35. The method according to any one of claims 31 to 34, characterized in that, If the number of magnetic storage tapes returned in the search is greater than a predetermined threshold, ignore the search results and search for the data among all stored magnetic storage tapes.

36. The method according to any one of claims 31 to 35, characterized in that, The mapping of the identifier to the magnetic storage tape is stored on at least one of a hard disk drive or a solid - state drive.

37. A computer program for blood-based tape deduplication, characterized in that, The computer program includes program instructions that, when executed by at least one processor, cause the at least one processor to: Identify at least one data range in a buffer to be written to at least one magnetic storage tape; Generate a list of identifiers of data objects in the identified data; Segment the data in the buffer; Output a strong hash for each segment of the data; Select a plurality of search representatives from the determined strong hashes; Output a weak hash for each selected strong hash; Search for each identifier in the mapping of the identifier to the magnetic storage tape to determine which tapes are candidate tapes; Search for the calculated weak hashes in the sparse index of each candidate tape; Select one magnetic storage tape having the largest number of weak hash matches; Compare the strong hash of each segment of the data with the region of the one magnetic storage tape pointed to by the weak hash match; And Update the mapping of the identifier to the magnetic storage tape such that the identifier now points to at least one magnetic storage tape.

38. A computer program for tape deduplication based on blood relationship, characterized in that, The computer program includes program instructions that, when executed by at least one processor, cause the at least one processor to: Identify at least one data range in a buffer to be written to at least one magnetic storage tape; Generate a list of identifiers of data objects in the identified data; Segment the data in the buffer; Calculate a strong hash for each segment of the data; Select a plurality of search representatives from the determined strong hashes; Calculate a weak hash for each selected strong hash; Search for each identifier in the mapping of the identifier to the magnetic storage tape to determine which tapes are candidate tapes; Search the global index of data objects to identify a list of magnetic storage tapes to be searched; Select one of the result of searching the global index or the result of searching the identifier mapping generated based on the result of the search, wherein, if the result of searching the global index provides a high deduplication rate, select the result of searching the global index, otherwise, select the result of the identifier mapping; Locate a magnetic storage tape based on the selected result; and Update the mapping of the identifier to the magnetic storage tape to use the located magnetic storage tape such that the identifier now points to at least one magnetic storage tape.

39. A computer program for tape deduplication based on consanguinity, characterized in that, The computer program includes program instructions that, when executed by at least one processor, cause the at least one processor to: Identify at least one data range in a buffer to be written to at least one magnetic storage tape; Generate a list of identifiers of data objects in the identified data; Segment the data in the buffer; Compute a strong hash for each segment of the data; Select a plurality of search representatives from the determined strong hashes; Compute a weak hash for each selected strong hash; Search for each identifier in the mapping of identifiers to magnetic storage tapes to determine which tapes are candidate tapes; Search a global index of data objects for the weak hashes to identify a list of candidate magnetic storage tapes to be searched; Select one magnetic storage tape selected as a result of the search of the identifier mapping, the one magnetic storage tape having the largest number of weak hash matches in the global index; and If searching the global index returns no results, update the global index with the weak hashes.