Method for managing data blocks, electronic device, and computer program product

By employing multiple hash algorithms to verify data block duplicates, the method reduces system performance overhead by avoiding unnecessary data retrieval and compression, thus improving efficiency in data deduplication.

CN114691011BActive Publication Date: 2025-07-15EMC IP HLDG CO LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011562733.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-25
Publication Date
2025-07-15
Estimated Expiration
2040-12-25

AI Technical Summary

Technical Problem

In the prior art, when identifying duplicate data blocks, misjudgment caused by hash algorithm collisions requires execution of operations with high overhead, resulting in reduced system performance.

Method used

Two different hashing algorithms are used to generate fingerprints of data blocks, and duplicate data blocks are identified by comparing fingerprints generated by different hashing algorithms, avoiding bit-by-bit comparison and reading data blocks from storage devices.

Benefits of technology

It effectively reduces the overhead of identifying duplicate data blocks and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691011B_ABST
    Figure CN114691011B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, an electronic device, and a computer program product for managing data blocks. A method for managing data blocks includes generating, based on a first hashing algorithm, a first fingerprint for a first data block to be stored in a storage device. The method further includes determining whether there is a third fingerprint generated for the second data block based on a second hashing algorithm in a fingerprint database if it is determined that there is a second fingerprint generated for the second data block based on the first hashing algorithm in the fingerprint database that matches the first fingerprint, where the fingerprint database records fingerprints of data blocks stored in the storage device. The method further includes generating, based on the second hashing algorithm, a fourth fingerprint for the first data block if it is determined that there is a third fingerprint in the fingerprint database; and determining whether the first data block and the second data block are duplicates by comparing the third fingerprint and the fourth fingerprint. Embodiments of the present disclosure can effectively reduce the overhead of identifying duplicate data blocks in data deduplication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of data storage, and more particularly, to methods, electronic devices, and computer program products for managing data blocks. Background Art

[0002] Generally, before storing data blocks into a storage device, a deduplication operation may be performed to avoid storing duplicate data blocks into the storage device. The deduplication operation generally proceeds as follows. First, the fingerprint (e.g., hash value) of the data block to be stored is determined. Then, the determined fingerprint is compared with the fingerprints of the data blocks already stored in the storage device. If the determined fingerprint does not match any of the fingerprints of the already stored data blocks, it indicates that the data block to be stored is not a duplicate data block. If the determined fingerprint matches the fingerprint of an already stored data block, to avoid misjudgment due to hash algorithm collisions, the already stored data block may be read from the storage device and decompressed. By comparing the decompressed data block bit by bit with the data block to be stored, it is determined whether the two are duplicate data blocks. If it is determined that the data block to be stored is not a duplicate data block, the data block to be stored is compressed and then stored into the storage device. Summary of the Invention

[0003] Embodiments of the present disclosure provide methods, electronic devices, and computer program products for managing data blocks.

[0004] In a first aspect of the present disclosure, a method for managing data blocks is provided. The method includes: generating a first fingerprint for a first data block to be stored into a storage device based on a first hash algorithm; if it is determined that there is a second fingerprint generated for a second data block based on the first hash algorithm in a fingerprint database that matches the first fingerprint, determining whether there is a third fingerprint generated for the second data block based on a second hash algorithm in the fingerprint database, where the fingerprint database records the fingerprints of the data blocks stored in the storage device; if it is determined that there is a third fingerprint in the fingerprint database, generating a fourth fingerprint for the first data block based on the second hash algorithm; and determining whether the first data block and the second data block are duplicates by comparing the third fingerprint and the fourth fingerprint.

[0005] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit and at least one memory. The at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform actions, including: generating, based on a first hashing algorithm, a first fingerprint for a first data block to be stored in a storage device; determining, if it is determined that there is a second fingerprint generated based on the first hashing algorithm for a second data block in a fingerprint database that matches the first fingerprint, whether there is a third fingerprint generated based on a second hashing algorithm for the second data block in the fingerprint database, where the fingerprint database records fingerprints of data blocks stored in the storage device; if it is determined that there is a third fingerprint in the fingerprint database, generating a fourth fingerprint for the first data block based on the second hashing algorithm; and determining whether the first data block and the second data block are duplicates by comparing the third fingerprint and the fourth fingerprint.

[0006] In a third aspect of the present disclosure, a computer-readable storage medium is provided, on which machine-executable instructions are stored. When executed by a device, the machine-executable instructions cause the device to perform any of the steps of the method described in the first aspect above.

[0007] In a fourth aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a non-transitory computer storage medium and includes machine-executable instructions. When executed by a device, the machine-executable instructions cause the device to perform any of the steps of the method described in the first aspect of the present disclosure.

[0008] The Summary of the Invention is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention is not intended to identify key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. Brief Description of the Drawings

[0009] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, where, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.

[0010] Figure 1 A schematic diagram of an example system in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A flowchart of an example method for managing data blocks according to an embodiment of the present disclosure is shown; and

[0012] Figure 3 A schematic block diagram of an example device that can be used to implement embodiments of the present disclosure is shown.

[0013] In each of the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners

[0014] Preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure will be more thorough and complete, and can fully convey the scope of the present disclosure to those skilled in the art.

[0015] The term "including" and its variations used herein mean open inclusion, that is, "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions below.

[0016] As described above, before storing data into a storage device, a deduplication operation can generally be performed to avoid storing duplicate data blocks into the storage device. The deduplication operation generally proceeds as follows. First, the fingerprint (e.g., hash value) of the data block to be stored is determined, and then the determined fingerprint is compared with the fingerprints of the data blocks that have been stored in the storage device. If the determined fingerprint does not match the fingerprints of the stored data blocks, it indicates that the data block to be stored is not a duplicate data block. If the determined fingerprint matches the fingerprint of a stored data block, to avoid misjudgment due to hash algorithm collision, the stored data block can be read from the storage device and decompressed. By comparing the decompressed data block bit by bit with the data block to be stored, it is determined whether the two are duplicate data blocks. If it is determined that the data block to be stored is not a duplicate data block, the data block to be stored is compressed, and then the compressed data block is stored in the storage device.

[0017] In the above traditional solution, when the determined fingerprint matches the fingerprint of a stored data block, to avoid misjudgment due to hash algorithm collision, a series of operations with relatively large overhead are required to identify whether the data block to be stored is a duplicate data block, resulting in a reduction in system performance.

[0018] Embodiments of the present disclosure propose a solution for managing data blocks to address one or more of the above problems and other potential problems. In this solution, based on a first hashing algorithm, a first fingerprint is generated for a first data block to be stored in a storage device. If it is determined that there is a second fingerprint generated for a second data block based on the first hashing algorithm in a fingerprint database that matches the first fingerprint, it is determined whether there is a third fingerprint generated for the second data block based on a second hashing algorithm in the fingerprint database, where the fingerprint database records fingerprints of data blocks stored in the storage device. If it is determined that there is a third fingerprint in the fingerprint database, a fourth fingerprint is generated for the first data block based on the second hashing algorithm. Then, by comparing the third fingerprint and the fourth fingerprint, it is determined whether the first data block and the second data block are duplicates. In this way, embodiments of the present disclosure can effectively reduce the overhead of identifying duplicate data blocks in data deduplication.

[0019] Figure 1 FIG. shows a block diagram of an example system 100 in which embodiments of the present disclosure can be implemented. As Figure 1 shown, the system 100 includes a host 110, a storage manager 120, and a storage device 130. It should be understood that the structure and function of the system 100 are described only for illustrative purposes and do not imply any limitation on the scope of the present disclosure. For example, embodiments of the present disclosure can also be applied to systems different from the system 100.

[0020] In the system 100, the host 110 can be, for example, any physical computer, virtual machine, server, etc. that runs a user application. The host 110 can send input / output (I / O) requests to the storage manager 120, such as requests to read data from and / or write data to the storage device 130. In response to receiving a read request from the host 110, the storage manager 120 can read data from the storage device 130 and return the read data to the host 110. In response to receiving a write request from the host 110, the storage manager 120 can write data to the storage device 130. The storage device 130 can be any currently known or future-developed non-volatile storage medium, such as a disk, a solid-state drive (SSD), or a disk array, etc.

[0021] As Figure 1As shown, for example, the storage manager 120 may include a data cache 121 and a fingerprint database 122. The data cache 121 may cache data to be written to the storage device 130. For example, the data in the data cache 121 may be organized in units of pages (also referred to as "data blocks" herein), and the pages may be flushed to the storage device 130 to achieve persistent storage of the data. The pages to be flushed to the storage device 130 are also referred to as "dirty pages". The fingerprint database 122 may record the fingerprints of the pages that have been flushed to the storage device 130, thereby preventing duplicate pages from being written to the storage device 130. The fingerprint may be a hash value of a page calculated based on a predetermined hash algorithm. For example, for the same page, the fingerprint database 122 may record one or more fingerprints generated based on one or more hash algorithms.

[0022] It should be understood that the data cache 121 and / or the fingerprint database 122 may be implemented using any currently known or future-developed volatile storage medium, non-volatile storage medium, or a combination of both. In addition, the data cache 121 and the fingerprint database 122 may be implemented using the same or different storage media. The scope of the present disclosure is not limited in this regard.

[0023] Figure 2 A flowchart of an example method 200 for storing data according to an embodiment of the present disclosure is shown. The method 200 may be executed, for example, by the storage manager 120 as Figure 1 shown. It should be understood that the method 200 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard. The following will describe the method 200 in detail in conjunction with Figure 1 this.

[0024] As Figure 2 shown, at block 205, the storage manager 120 generates a first fingerprint for a first data block to be stored in the storage device 130 based on a first hash algorithm. In some embodiments, the first hash algorithm may be, for example, the Murmur3 hash algorithm, or any other suitable hash algorithm.

[0025] At block 210, the storage manager 120 queries the fingerprint database 122 to determine whether there is a second fingerprint generated for a second data block based on the first hash algorithm that matches the first fingerprint.

[0026] If it is determined that there is no fingerprint in the fingerprint database 122 that matches the first fingerprint, the storage manager 120 may determine that the first data block does not duplicate any of the data blocks already stored in the storage device 130. In this case, method 200 proceeds to block 235, where the storage manager 120 stores the first data block in the storage device 130 and stores the first fingerprint of the first data block (the fingerprint generated based on the first hash algorithm) in the fingerprint database 122. In some embodiments, the storage manager 120 may compress the first data block and then write the compressed first data block to the storage device 130. In other embodiments, the compression operation may be omitted.

[0027] If it is determined that there is a second fingerprint in the fingerprint database 122 that matches the first fingerprint of the second data block, at block 215, the storage manager 120 determines whether there is a third fingerprint in the fingerprint database 122 that is generated for the second data block based on a second hash algorithm. The second hash algorithm may be any suitable hash algorithm with a lower collision probability than the first hash algorithm. As used herein, "collision" refers to the situation where two data blocks are different but their hash values are the same. In some embodiments, for example, the second hash algorithm is the SHA-1 hash algorithm, whose computational overhead is approximately three times that of the Murmur3 algorithm, but the collision probability is much lower than that of the Murmur 3 algorithm (e.g., the collision probability of a 160-bit SHA-1 hash value is 0.00000000000019%).

[0028] If it is determined that there is a third fingerprint in the fingerprint database 122 that is generated for the second data block based on the second hash algorithm, at block 220, the storage manager 120 generates a fourth fingerprint of the first data block based on the second hash algorithm. At block 225, the storage manager 120 compares the third fingerprint and the fourth fingerprint to determine whether they match.

[0029] If the third fingerprint and the fourth fingerprint match, due to the extremely low collision probability of the SHA-1 algorithm, the storage manager 120 may determine that the first data block and the second data block are duplicate data blocks. In this case, the storage manager 120 will not store the first data block in the storage device 130 to avoid duplicate storage. In addition, since the fingerprint database 122 already stores the fingerprint of the second data block, the storage manager 120 will not write the fingerprint of the first data block to the fingerprint database 122 either. Compared with the traditional solution, identifying duplicate data blocks by comparing fingerprints generated based on different hash algorithms can avoid the operation of reading data blocks from the storage device and performing bit-by-bit comparison, and thus can effectively reduce the overhead of identifying duplicate data blocks.

[0030] If the third fingerprint does not match the fourth fingerprint, the storage manager 120 may determine that the first data block and the second data block are not duplicates. In this case, at block 230, the storage manager 120 may store the first data block in the storage device 130 and store the first fingerprint of the first data block (i.e., the fingerprint generated based on the first hashing algorithm) and the fourth fingerprint (i.e., the fingerprint generated based on the second hashing algorithm) in the fingerprint database 122 for subsequent queries. In some embodiments, the storage manager 120 may compress the first data block and then write the compressed first data block to the storage device 130. In other embodiments, the compression operation may be omitted.

[0031] If it is determined at block 215 that there is no third fingerprint in the fingerprint database 122 generated for the second data block based on the second hashing algorithm, the storage manager 120 will further determine whether the first data block and the second data block are duplicates according to a conventional scheme.

[0032] Specifically, at block 240, the storage manager 120 obtains the second data block from the storage device 130. If the obtained second data block is a compressed data block, the storage manager 120 may decompress it. At block 245, the storage manager 120 compares the first data block and the second data block bit by bit to determine whether they are duplicates.

[0033] If it is determined that the first data block and the second data block are duplicates, then at block 250, the storage manager 120 may generate a third fingerprint for the second data block based on the second hashing algorithm, and then at block 255, store the generated third fingerprint in the fingerprint database 122 for subsequent queries. By delaying the calculation of the third fingerprint (e.g., based on the SHA-1 algorithm) until the time point when it is determined that the first data block and the second data block are duplicate data blocks, the overhead of calculating multiple fingerprints for duplicate data blocks can be avoided.

[0034] If it is determined that the first data block and the second data block are not duplicates, the method 200 proceeds to block 235, where the storage manager 120 stores the first data block in the storage device 130 and stores the first fingerprint of the first data block (the fingerprint generated based on the first hashing algorithm) in the fingerprint database 122. In some embodiments, the storage manager 120 may compress the first data block and then write the compressed first data block to the storage device 130. In other embodiments, the compression operation may be omitted.

[0035] Although Figure 2It only describes generating fingerprints of data blocks based on at most two hash algorithms. It should be understood that embodiments of the present disclosure can also be extended to identify duplicate data blocks by comparing more fingerprints generated based on more than two hash algorithms. In addition, to resist malicious collision attacks, a private data header and / or data tail can be added to the data block before generating the fingerprint of the data block, and then the corresponding fingerprint is generated for the data block with the added data header and / or data tail based on the hash algorithm.

[0036] As can be seen from the above description, embodiments of the present disclosure propose a solution for managing data blocks to solve one or more of the above problems and other potential problems. In this solution, based on a first hash algorithm, a first fingerprint is generated for a first data block to be stored in a storage device. If it is determined that there is a second fingerprint generated for a second data block based on the first hash algorithm in the fingerprint database that matches the first fingerprint, it is determined whether there is a third fingerprint generated for the second data block based on a second hash algorithm in the fingerprint database, where the fingerprint database records fingerprints of data blocks stored in the storage device. If it is determined that there is a third fingerprint in the fingerprint database, a fourth fingerprint is generated for the first data block based on the second hash algorithm. Then, by comparing the third fingerprint and the fourth fingerprint, it is determined whether the first data block and the second data block are duplicates. Compared with the traditional solution, embodiments of the present disclosure can avoid the operation of reading data blocks from the storage device and performing bit-by-bit comparison by comparing fingerprints generated based on different hash algorithms, and thus can effectively reduce the overhead of identifying duplicate data blocks.

[0037] Figure 3 The schematic block diagram of an example device 300 that can be used to implement embodiments of the present disclosure is shown. For example, as Figure 1 shown, the storage manager 120 can be implemented by the device 300. As Figure 3 shown, the device 300 includes a central processing unit (CPU) 301, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 302 or computer program instructions loaded from a storage unit 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 can also be stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0038] Multiple components in device 300 are connected to I / O interface 305, including: input unit 306, such as a keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as a disk, optical disc, etc.; and communication unit 309, such as a network card, modem, wireless communication transceiver, etc. Communication unit 309 allows device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0039] Each of the processes and treatments described above, such as method 200, may be executed by processing unit 301. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by CPU 301, one or more actions of method 200 described above may be performed.

[0040] The present disclosure may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.

[0041] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0042] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0043] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0044] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0045] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, result in an apparatus that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0046] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, whereby the instructions that execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0047] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.

[0048] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for managing data blocks, comprising: Generating a first fingerprint for a first data block to be stored in a storage device based on a first hashing algorithm; If it is determined that there is a second fingerprint generated based on the first hashing algorithm for a second data block in a fingerprint database that matches the first fingerprint, determining whether there is a third fingerprint generated based on a second hashing algorithm for the second data block in the fingerprint database, wherein the fingerprint database records fingerprints of data blocks stored in the storage device; If it is determined that there is the third fingerprint in the fingerprint database, generating a fourth fingerprint for the first data block based on the second hashing algorithm; And Determining whether the first data block and the second data block are duplicates by comparing the third fingerprint and the fourth fingerprint; And The method further comprises: Before generating the first fingerprint, generating the third fingerprint based on the second hashing algorithm and storing the third fingerprint in the fingerprint database, and Wherein determining whether there is the third fingerprint includes querying for the third fingerprint in the fingerprint database, wherein it is determined that there is the third fingerprint in the database.

2. The method according to claim 1, further comprising: If it is determined that there is no fingerprint in the fingerprint database that matches the first fingerprint, Storing the first data block in the storage device; And Storing the first fingerprint in the fingerprint database.

3. The method according to claim 1, further comprising: If it is determined that there is no the third fingerprint in the fingerprint database, retrieving the second data block from the storage device; And Determining whether the first data block and the second data block are duplicates by comparing the first data block and the second data block.

4. The method according to claim 3, further comprising: If it is determined that the first data block and the second data block are duplicates, Generating the third fingerprint for the second data block based on the second hashing algorithm; And Storing the third fingerprint in the fingerprint database.

5. The method according to claim 3, further comprising: If it is determined that the first data block and the second data block are not duplicates, Storing the first data block in the storage device; And Storing the first fingerprint in the fingerprint database.

6. The method according to claim 1, wherein determining whether the first data block and the second data block are duplicates includes: If the third fingerprint and the fourth fingerprint match, determining that the first data block and the second data block are duplicates; And If the third fingerprint and the fourth fingerprint do not match, determining that the first data block and the second data block are not duplicates.

7. The method according to claim 6, further comprising: If it is determined that the first data block and the second data block are not duplicates, Storing the first data block in the storage device; And Storing the first fingerprint and the fourth fingerprint in the fingerprint database.

8. The method according to claim 1, wherein the collision probability of the second hashing algorithm is lower than the collision probability of the first hashing algorithm.

9. The method according to claim 1, wherein the first hashing algorithm is the Murmur3 hashing algorithm, and the second hashing algorithm is the SHA-1 hashing algorithm.

10. An electronic device, comprising: at least one processing unit; at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform actions, the actions including: generating, based on a first hashing algorithm, a first fingerprint for a first data block to be stored in a storage device; if it is determined that there exists in a fingerprint database a second fingerprint generated for a second data block based on the first hashing algorithm that matches the first fingerprint, determining whether there exists in the fingerprint database a third fingerprint generated for the second data block based on a second hashing algorithm, wherein the fingerprint database records fingerprints of data blocks stored in the storage device; if it is determined that there exists the third fingerprint in the fingerprint database, generating a fourth fingerprint for the first data block based on the second hashing algorithm; and determining whether the first data block and the second data block are duplicates by comparing the third fingerprint and the fourth fingerprint; and the actions further include: before generating the first fingerprint, generating the third fingerprint based on the second hashing algorithm and storing the third fingerprint in the fingerprint database, and wherein determining whether there exists the third fingerprint includes querying for the third fingerprint in the fingerprint database, and determining that there exists the third fingerprint in the database.

11. The electronic device according to claim 10, wherein the actions further include: if it is determined that there is no fingerprint in the fingerprint database that matches the first fingerprint, storing the first data block in the storage device; and storing the first fingerprint in the fingerprint database.

12. The electronic device according to claim 10, wherein the actions further include: if it is determined that there is no third fingerprint in the fingerprint database, obtaining the second data block from the storage device; and determining whether the first data block and the second data block are duplicates by comparing the first data block and the second data block.

13. The electronic device according to claim 12, wherein the actions further include: if it is determined that the first data block and the second data block are duplicates, generating the third fingerprint for the second data block based on the second hashing algorithm; and storing the third fingerprint in the fingerprint database.

14. The electronic device according to claim 12, wherein the actions further include: if it is determined that the first data block and the second data block are not duplicates, storing the first data block in the storage device; and storing the first fingerprint in the fingerprint database.

15. The electronic device according to claim 10, wherein determining whether the first data block and the second data block are duplicates includes: If the third fingerprint matches the fourth fingerprint, determine that the first data block and the second data block are duplicates; And If the third fingerprint does not match the fourth fingerprint, determine that the first data block and the second data block are not duplicates.

16. The electronic device according to claim 15, wherein the action further comprises: If it is determined that the first data block and the second data block are not duplicates, Store the first data block in the storage device; And Store the first fingerprint and the fourth fingerprint in the fingerprint database.

17. The electronic device according to claim 10, wherein the collision probability of the second hash algorithm is lower than the collision probability of the first hash algorithm.

18. The electronic device according to claim 10, wherein the first hash algorithm is the Murmur3 hash algorithm, and the second hash algorithm is the SHA-1 hash algorithm.

19. A computer program product, the computer program product being tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Reducing hash collisions in large scale data deduplication

    US10762051B1