Device, system and method for deduplication optimization

By dividing the hash metadata table into multiple parts in the global deduplication server and uploading only relevant parts to memory, the performance degradation caused by hash storage is solved, and the response time and memory management efficiency is improved.

CN113227993BActive Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980086508.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-29
Publication Date
2025-06-27
Estimated Expiration
2039-11-29

AI Technical Summary

Technical Problem

In Global Deduplication Server (GDS), memory management is inefficient due to the need to store the hash values ​​of multiple storage servers, resulting in poor performance, especially when replying to requests to store or delete hash values.

Method used

By dividing the hash metadata table into multiple parts and associating each part with a different hash range, only the relevant hash portions are uploaded into memory, thereby reducing cache misses and improving response time.

Benefits of technology

By reducing cache missing, the response time of the global server is improved, the I/O operation is reduced, and the memory management efficiency is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113227993B_ABST
    Figure CN113227993B_ABST
Patent Text Reader

Abstract

An advanced deduplication method is disclosed, particularly having an additional deduplication layer. Specifically, the present invention proposes a global server for deduplicating multiple storage servers. The global server is used to maintain information related to a set of hash values, where each hash value is associated with a data block of data stored in the global server and / or the storage servers, and notify the storage servers of a first range of the hash values. The global server is further used to receive requests from one or more of the storage servers to modify the information for one or more hash values falling within the first range of the hash values. The global server is further used to modify the information for one or more hash values falling within the first range of the hash values based on the requests received from the one or more storage servers. The present invention also proposes a storage server for deduplication in a global server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data storage deduplication method, and more particularly to an optimized memory management method in a global deduplication server (abbreviated as GDS, also simply referred to as the global server). The present invention solves the performance degradation by optimizing the memory management in the GDS. Background Art

[0002] Data deduplication (also known as data optimization) refers to reducing the physical number of bytes of data that need to be stored on disk or transmitted over a network without compromising the fidelity or integrity of the original data, i.e., the reduction of bytes is lossless and the original data can be fully restored. By reducing the storage resources for storing and / or transmitting data, data deduplication saves hardware costs (for storage and network transmission) and data management costs (such as backups). As the amount of digital storage data grows, these cost savings become very important.

[0003] Data deduplication typically uses a combination of techniques to eliminate redundancy within and between files in persistent storage. One technique is used to identify identical data regions in one or more files, and physically store only a single unique region (called a block), while maintaining pointers to the blocks associated with that file. Another technique is to mix data deduplication and compression, for example by storing compressed blocks.

[0004] Many organizations use dedicated servers to store data (i.e., storage servers). The data stored on different servers often duplicates, resulting in space wastage. One solution to this problem is deduplication, including identifying duplicate data by using hashing and storing only unique data that reaches a certain granularity. However, deduplication is performed at the granularity of a specific single storage server.

[0005] To prevent data duplication between multiple storage servers, the concept of deduplication of deduplication (nested deduplication) is introduced, including an additional deduplication layer (performing deduplication of multiple deduplication servers). In particular, a GDS is proposed, which will store highly duplicate data as a solution to the problem of performing deduplication between multiple storage servers.

[0006] The GDS storage stores the hashes sent by the storage servers (i.e., the storage server cluster) that conform to nested deduplication to determine whether the hash values appear in a sufficient number of storage servers, ensuring that the GDS obtains ownership of the data represented by the hash values. The hash values can be used to uniquely identify the corresponding data blocks. Since the GDS will store all the hash values of the storage servers in the storage server cluster (regardless of whether they also store their data), this results in a large storage space being required to store all these hash values. Therefore, it is impossible to save all the hashes in the memory of the GDS, which further affects the performance of the GDS, especially when responding to requests to store or delete hash values.

[0007] The standard solution to this problem is to cache data in memory. Due to the way the GDS architecture is constructed (multiple storage servers form a storage server cluster that communicates with one or more GDSs), without appropriate caching policies and methods, it may lead to many "cache misses", causing forced disk access every time.

[0008] The present invention aims to solve this performance degradation by optimizing the memory management efficiency of the GDS when hash values are stored in the memory of the GDS. Summary of the Invention

[0009] In view of the above problems, embodiments of the present invention aim to provide a data management solution to optimize the response time of the global server by reducing cache misses. The purpose is to enable more hash values to be read from memory and fewer from disk. Another purpose is to reduce I / O operations.

[0010] The above object is achieved by the embodiments provided in the appended independent claims. Advantageous implementations of the embodiments of the present invention are further defined in the dependent claims.

[0011] A first aspect of the present invention provides a global server for deduplicating multiple storage servers. The global server is configured to: notify the storage servers of a first range of a set of hash values, where information related to the set of hash values is maintained, and each hash value is associated with a data block of data stored in the global server and / or the storage servers; receive requests from one or more of the storage servers to modify the information for one or more hash values that fall within the first range of the hash values; and modify the information for the one or more hash values that fall within the first range of the hash values based on the requests received from the one or more storage servers.

[0012] The present invention provides a global server that stores highly repetitive data in a nested deduplication system. The nested deduplication system may include a global server and multiple storage servers. In particular, the present invention proposes a communication method between the global server and the multiple storage servers, where the communication is initiated by the global server. That is, a specific range of hash values will be provided to the storage servers, and only requests for hash values falling within this range can be reported to the global server. In this way, the global server can control the requested hash values. Therefore, cache misses in the global server are reduced, thereby increasing the response time of the global server.

[0013] The term "global server" is an abbreviation of "global deduplication server", which refers to a server used to process highly repetitive data in a storage system containing multiple deduplication servers. In implementation, the GDS can be implemented as a centralized device (e.g., a server), or deployed on one of the multiple storage servers, or implemented in a distributed manner (e.g., multiple servers form a "virtual" global deduplication server).

[0014] In one implementation of the first aspect, information related to the set of hash values can be maintained in the global server or a separate storage device accessible to the global server. This improves the implementation diversity of the deduplication system.

[0015] In one implementation of the first aspect, the global server is used to send a broadcast message carrying the first range of the hash values to the storage servers.

[0016] In particular, the global server can notify the storage servers of the range of hash values it is willing to accept through a broadcast message carrying this range. This improves the efficiency of the deduplication system.

[0017] In one implementation of the first aspect, the global server includes a disk memory; and the information includes a hash metadata table containing the set of hash values, and the hash metadata table is stored in the disk memory.

[0018] A table storing hash values and information related to each hash value is stored in the local disk of the global server, that is, the hash metadata table.

[0019] In one implementation of the first aspect, the global server is used to: divide the hash metadata table into N parts, where N is a positive integer not less than 2, and each part of the hash metadata table is associated with a different range of hash values.

[0020] It should be noted that the hash metadata table is a sorted table, and all hash values are stored in order. For example, the hash values can be stored in the table in ascending or descending order. When the hash metadata table is divided into parts, the hash values are correspondingly divided into different ranges.

[0021] In one implementation of the first aspect, the global server further includes a memory; and the global server is further configured to: upload a first part of the hash metadata table to the memory, where the first part of the hash metadata table is associated with a first range of the hash values; and modify the first part of the hash metadata table based on the request received from the one or more storage servers.

[0022] Since the hash metadata table is divided into N parts, the range of the hash values is divided into N parts. The global server traverses all parts in a cyclic manner, for example. The global server requests from the storage servers to provide the hash values stored in the storage servers, and these hash values are included in the corresponding parts. It should be noted that each part is the uploaded part currently in the memory. Therefore, the global server can read more from the memory and less from the disk.

[0023] In one implementation of the first aspect, the hash metadata table includes information of each hash value in the set of hash values and information related to one or more storage servers registered for the hash value.

[0024] The embodiment of the present invention is based on the following fact: when a storage server requests the global server to add a hash value, the global server registers the storage server for the hash value.

[0025] In one implementation of the first aspect, the global server is configured to:

[0026] - In response to the request received from the storage server including a request to add a first hash value:

[0027] - In response to the first hash value not being included in the first part of the hash metadata table, add the first hash value to the first part of the hash metadata table, create a first watermark associated with the first hash value, and register the storage server that sent the request related to the first hash value, where the first watermark indicates whether the data block having the first hash value is highly repeated among the storage servers; or

[0028] -In response to the first hash value being included in the first part of the hash metadata table, increasing the value of the first watermark associated with the first hash value, and registering the storage server that sent the request related to the first hash value; and / or

[0029] -In response to a request received from a storage server including a request to delete a second hash value:

[0030] -Decreasing the value of the second watermark associated with the second hash value, and deregistering the storage server that sent the request related to the second hash value, where the second watermark indicates whether the data block having the second hash value is highly replicated among the storage servers; and

[0031] -In response to the value of the second watermark being equal to 0, deleting the second hash value from the first part of the hash metadata table.

[0032] Generally, when the global server receives a request to add a hash value, it creates or raises the watermark associated with the hash value, and registers the storage server that sent the request for the hash value. Correspondingly, when receiving a request to delete data from a storage server, the global server decreases the value of the watermark associated with the hash value of the data, and deregisters the storage server for the hash value. It should be noted that according to an embodiment of the present invention, the requests received from the storage server within a period of time are related to the hash values that fall within a specific range of the memory of the global server within the same period of time.

[0033] In an implementation manner of the first aspect, the global server is configured to: persist the updated first part of the hash metadata table to the disk memory.

[0034] After updating the corresponding part of the hash metadata table currently located in the memory of the global server, the part of the hash metadata table will be persisted back to the disk memory of the global server. That is, the updated part overwrites the old data of the same part.

[0035] In an implementation manner of the first aspect, the global server is configured to: modify the information for one or more hash values that fall within the first range of the hash values based on all the requests received from the one or more storage servers within a predetermined period of time.

[0036] In particular, after uploading a part of the hash metadata table to the memory, all requests received from the storage server within a specific time period shall be processed by the global server. That is, before persisting the updated part to the disk memory, the global server needs to process every request related to adding / deleting hash values and take actions according to each request (modifying the uploaded part of the hash metadata table).

[0037] A second aspect of the present invention provides a storage server for deduplication in a global server. The storage server is configured to: in response to a request to add or delete a first data block, record a request associated with a first hash value of the first data block, where the first hash value is included in a set of hash values, maintain the set of hash values, and each hash value is a hash value of a stored data block; receive a notification of a first range of hash values from the global server; and if the first hash value falls within the first range of hash values, send the recorded request to the global server to modify information in the global server for the first hash value.

[0038] In this topology, multiple storage servers can be connected to the global server. Each storage server can operate in a similar manner. It is worth noting that the storage server communicates with the global server by means of communication initiated by the global server. In particular, the global server notifies the storage server of the way to send requests. For example, the global server can indicate when the storage server can send a request and which request can be sent. Obviously, before sending a request to the global server, the storage server needs to store all requests received from the user. The storage server supports the global server in reducing cache misses, thereby optimizing the response time of the global server. In particular, the latency of the global server can be reduced.

[0039] In an implementation manner of the second aspect, the storage server is configured to: delete the recorded request after sending the request to the global server.

[0040] It is worth noting that after sending a request related to a hash value falling within a specific range provided by the global server to the global server, assuming that the global server will process the request and take corresponding actions. To prevent redundant requests from being sent to the global server again, the storage server will delete those requests that have been sent to the global server from the record.

[0041] In an implementation of the second aspect, the storage server is configured to: after receiving a request to delete the record, maintain information on the storage location of the first data block associated with the maintained first hash value, where the storage location of the first data block is the storage server and / or the global server.

[0042] In particular, the storage server stores the request to delete the record (which is held until processed by the global server), and the storage server maintains a record of the hash value. For example, the storage server maintains information on whether the data associated with the hash value is locally stored in the storage server and / or the global server, and a reference count associated with the hash value, which indicates how frequently the users of the storage server require the hash value.

[0043] In an implementation of the second aspect, the storage server is configured to: receive from the global server a broadcast message carrying a first range of the hash values.

[0044] In an implementation of the second aspect, the storage server is configured to: compare the first range of the hash values with the first hash value; and determine whether the first hash value falls within the first range of the hash values.

[0045] In particular, for each request associated with a respective hash value of a record, the storage server determines whether the request should be sent to the global server.

[0046] In an implementation of the second aspect, if a user requests to add the first data block, the request includes a request for the first hash value of the first data block to be added; or

[0047] if a user requests to delete the first data block, the request includes a request for the first hash value of the first data block to be deleted.

[0048] A third aspect of the present invention provides a method executed by a global server. The method includes: notifying a plurality of storage servers of a first range of a set of hash values, where information related to the set of hash values is maintained, and each hash value is associated with a data block of data stored in the global server and / or the storage server; receiving from one or more of the storage servers requests to modify the information for one or more hash values falling within the first range of the hash values; and modifying the information for the one or more hash values falling within the first range of the hash values based on the requests received from the one or more storage servers.

[0049] The method of the third aspect and its implementation manners provides the same advantages and effects as the global server of the first aspect and its respective implementation manners described above.

[0050] A fourth aspect of the present invention provides a method executed by a storage server. The method includes: in response to a request for adding or deleting a first data block, recording a request associated with a first hash value of the first data block, where the first hash value is included in a set of hash values, and each hash value is a hash value of a stored data block; receiving a notification of a first range of hash values from the global server; and if the first hash value falls within the first range of hash values, sending the recorded request to the global server to modify information in the global server for the first hash value.

[0051] The method of the fourth aspect and its implementation manners provides the same advantages and effects as the storage server of the second aspect and its respective implementation manners described above.

[0052] A fifth aspect of the present invention provides a computer program product, including computer-readable code instructions. When running on a computer, the computer-readable code instructions will cause the computer to execute the method described in the third or fourth aspect and its implementation manners.

[0053] A sixth aspect of the present invention provides a computer-readable storage medium, including computer-executable computer program code instructions. When running on a computer, the computer program code instructions will execute the method described in the third or fourth aspect and its implementation manners. The computer-readable storage medium includes one or more of the following groups: Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), flash memory, Electrically EPROM (EEPROM), and hard disk drive.

[0054] A seventh aspect of the present invention provides a global server for deduplicating multiple storage servers, including a processor and a memory. The memory stores instructions that cause the processor to execute the method described in the third aspect and its implementation manners of the present invention.

[0055] An eighth aspect of the present invention provides a storage server for deduplicating in a global server, including a processor and a memory. The memory stores instructions that cause the processor to execute the method described in the fourth aspect and its implementation manners of the present invention.

[0056] It should be noted that all the devices, components, units, and methods described in this application can be implemented in software or hardware components or any combination thereof. All the steps performed by various entities described in this application and the functions described to be performed by various entities are intended to indicate that each entity is suitable or used to perform its respective steps and functions. Although in the description of the following specific embodiments, the specific functions or steps performed by external entities are not reflected in the description of the specific components of the entity performing the specific steps or functions, those skilled in the art should clearly understand that these methods and functions can be implemented in their respective hardware or software components or any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In conjunction with the accompanying drawings, the following description of specific embodiments will elaborate on the various aspects of the present invention and their implementation forms, where:

[0058] Figure 1 shows the global server provided by an embodiment of the present invention;

[0059] Figure 2 shows the topology provided by an embodiment of the present invention;

[0060] Figure 3 shows the data memory in the global server provided by an embodiment of the present invention;

[0061] Figure 4 shows the storage server provided by an embodiment of the present invention;

[0062] Figure 5 shows the flowchart of the method provided by an embodiment of the present invention; and

[0063] Figure 6 shows the flowchart of another method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] Schematic embodiments of methods, apparatuses, and program products for efficient message transmission in a communication system are described in conjunction with the accompanying drawings. Although this description provides detailed examples of possible implementation manners, it should be noted that the details are for illustration purposes only and in no way limit the scope of this application.

[0065] In addition, one embodiment / example can refer to other embodiments / examples. For example, any description including but not limited to terms, elements, processes, explanations, and / or technical advantages mentioned in one embodiment / example is applicable to other embodiments / examples.

[0066] Figure 1FIG. 0 shows the global server 100 provided by an embodiment of the present invention. The global server 100 may include a processing circuit (not shown) for performing, implementing, or initiating various operations of the global server 100 described herein. The processing circuit may include hardware and software. The hardware may include analog circuits or digital circuits, or both analog and digital circuits. The digital circuit may include components such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a multi-functional processor. In one embodiment, the processing circuit includes one or more processors and a non-transitory memory connected to the one or more processors. The non-transitory memory may carry executable program code that, when executed by the one or more processors, causes the global server 100 to perform, implement, or initiate the operations or methods described herein.

[0067] The global server 100 is used to perform deduplication on multiple storage servers 110 (one of which is shown in the figure). The global server 100 may be used to maintain information 101 related to a set of hash values, where each hash value is associated with a data block of data stored in the global server and / or the storage server 110. The global server 100 is also used to notify the storage server 110 of a first range 1011 of the hash values. Then, the global server 100 is used to receive requests 111 from one or more of the storage servers 110 to modify the information 101 for one or more hash values that fall within the first range 1011 of the hash values. Further, the global server 100 is also used to modify the information 101 for one or more hash values that fall within the first range 1011 of the hash values based on the requests 111 received from the one or more storage servers 110.

[0068] For those skilled in the art of storage, information related to a set of hash values may also be maintained in a separate device (e.g., a storage server) accessible to the global server 100. The above description of the global server should not be regarded as a limitation on the implementation of the global server 100.

[0069] The embodiments of the present invention are applied to a nested deduplication topology. Figure 2Shows the topology of the global server 100 and storage servers A, B, and C 110 provided according to an embodiment of the present invention. It should be noted that the actual number of storage servers 110 in the nested deduplication topology implementation is not limited here. The global server 100 provides an additional deduplication layer. That is, it performs deduplication on multiple deduplication servers (i.e., storage servers) 110. Generally, several application servers can access their respective storage servers. Users can perform data read / write on the storage servers through the application servers.

[0070] By performing a hash function or hash algorithm on a data block, the hash value of the data block can be obtained. The hash value can be used to uniquely identify the corresponding data block. The present invention does not limit the types of hash and chunking techniques used in the storage server, as long as they are the same in all servers. When a user writes data to the storage server 110, the storage server 110 can chunk and hash the data to obtain the hash value of the data block. Since the data stored in multiple deduplication servers or storage servers is often repetitive, in order to avoid space wastage, the storage server can request to store some data in the GDS.

[0071] In particular, Figure 2 The GDS shown is the global server 100 provided according to an embodiment of the present invention as Figure 1 shown. Figure 2 The storage servers A, B, and C shown are all the storage servers 110 provided according to an embodiment of the present invention as Figure 1 shown. The global server 100 is intended to store highly repetitive data of the storage server 110. The global server 100 generally determines highly repetitive data according to some configurable thresholds. The storage server 110 communicates with the global server 100 and sends a request to store data. The global server 100 can accept or reject the request according to the configured threshold. The highly repetitive data can be stored in the global server 100. Accordingly, the storage server 110 can delete these highly repetitive data from its local storage and reply to the global server 100. In some scenarios, the storage server can also save a copy of the highly repetitive data.

[0072] The present invention proposes to optimize the memory management of the hash values stored in the GDS. In this solution, the GDS will initiate a request to the storage server to provide the hash values to the GDS.

[0073] It should be noted that the global server 100 can include a disk memory 102, such as Figure 3As shown. In particular, the set of hash values is included in a table, namely, the hash metadata table. This hash metadata table is stored in the disk memory 102 of the global server 100.

[0074] Optionally, the global server 100 can be used to divide the hash metadata table into N parts, where N is a positive integer not less than 2. Each part of the hash metadata table is associated with a different range of the hash values. In particular, the hash values are arranged in a specific order in the hash metadata table, for example, in ascending or descending order. That is to say, if the hash values are arranged in ascending order, the hash values in the Nth part of the hash metadata table will have a larger value than the hash values in the (N - 1)th part of the hash metadata table. The ranges of the hash values in the respective parts of the hash metadata table do not overlap with each other.

[0075] It should be noted that the global server 100 may also include a memory 103, as Figure 3 shown.

[0076] Optionally, the global server 100 can also be used to upload the first part of the hash metadata table to the memory 103. In particular, the first part of the hash metadata table is associated with the first range 1011 of the hash values. Correspondingly, the global server 100 can be used to modify the first part of the hash metadata table based on the request 111 received from the one or more storage servers 110.

[0077] Figure 3 Illustrates the data storage management in the global server 100 provided by an embodiment of the present invention. As Figure 3 shown, the first part of the hash metadata table is uploaded from the disk memory 102 to the memory 103. The first part of the uploaded hash metadata table corresponds to the first range 1011 of the hash values. Optionally, the global server 100 can send a broadcast message to all the storage servers 110 connected to it, and this message carries a specific range, that is, the first range 1011 provided in this embodiment. This is to indicate to the storage servers 110 the range of hash values that the global server 100 is willing to accept. According to an embodiment of the present invention, the specific range corresponds to a part of the hash metadata table currently uploaded to the memory 103.

[0078] In particular, the hash metadata table includes information of each hash value in the set of hash values and information related to one or more storage servers 110 registered for the hash value. For example, for each hash value, the hash metadata table may include the data block having the hash value, the water level line associated with the hash value, and information of the storage server 110 that requests to add the hash value.

[0079] Typically, each storage server 110 will record new hash values and deleted hash values between notifications from the global server 100. Once a notification with a specific range of hash values is received, the storage server 110 sends a request to record to add or delete the hash values within that specific range.

[0080] Possibly, in one embodiment of the present invention, the request received from the storage server 110 may include a request to add a first hash value. When the first hash value is not included in the first part of the hash metadata table, the global server 100 can be used to add the first hash value to the first part of the hash metadata table, create a first watermark associated with the first hash value, and register the storage server 110 that sent the request related to the first hash value. The first watermark indicates whether the data block with the first hash value is highly repeated among the storage servers 110. For example, if the value of the first watermark is 1, it means that one storage server 110 requests to add the first hash value. It should be noted that when the first hash value is not included in the first part of the hash metadata table, it means that although the first hash value falls within the first range 1011, it is not currently stored in the global server 100. When the first hash value is included in the first part of the hash metadata table, the global server 100 can be used to increase the value of the first watermark associated with the first hash value and register the storage server 110 that sent the request regarding the first hash value.

[0081] Possibly, in another embodiment of the present invention, the request received from the storage server 110 may include a request to delete a second hash value. The global server 100 can be used to decrease the value of the second watermark associated with the second hash value and unregister the storage server 110 that sent the request related to the second hash value. Similarly, the second watermark indicates whether the data block with the second hash value is highly repeated among the storage servers 110. Further, when the value of the second watermark is equal to 0, the global server 100 can be used to delete the second hash value from the first part of the hash metadata table. It should be noted that when the value of the second watermark is equal to 0, it means that currently no storage server 110 still requests to add the second hash value, so it can be deleted from the hash metadata table.

[0082] Possibly, the global server 100 can receive two requests simultaneously, namely a request to add the first hash value and a request to delete the second hash value. Possibly, the first hash value and the second hash value can even be the same hash value. For example, storage server A can request to add a hash value, while storage server B can request to delete the same hash value.

[0083] It should be noted that when processing the requests, the global server 100 inserts the hash value into the table part currently in the memory 103, or deletes the hash value from the table part currently in the memory 103. After inserting the hash value, if the water level line associated with the hash value reaches the high water level line, the global server 100 can request data with that hash value from the storage server 110. After deleting the hash value, if the water level line associated with the hash value reaches the low water level line, the global server 100 can delete the data with that hash value. That is, based on whether the water level line associated with the corresponding hash value is higher / lower than certain thresholds, the global server 100 will request data with that hash value from a certain storage server 110, or will decide to remove the data with that hash value and will notify all relevant storage servers 110 to take back the ownership of the data.

[0084] After updating the first part of the hash metadata table currently in the memory 103 of the global server 100, the updated part will be persisted back to the disk storage 102 of the global server 100. That is, the updated data correspondingly overwrites the old data in the same part.

[0085] The global server 100 will traverse each part one by one in a cyclic manner. That is, after updating the first part of the hash metadata table in the disk storage 102, the global server 100 will continue the process of updating the second part of the hash metadata table. Correspondingly, the global server 100 can be used to notify the storage server 110 of the second range of hash values. The global server 100 can also be used to receive requests from one or more of the storage servers 110 to modify the information 101 for one or more hash values falling within the second range of hash values. Then, the global server 100 can be used to modify the information 101 for one or more hash values falling within the second range of hash values based on the requests received from the one or more storage servers 110.

[0086] In particular, the global server 100 may upload the second part of the hash metadata table from the disk memory 102 to the memory 103. The second part of the hash metadata table is associated with the second range of the hash values.

[0087] It should be understood that, according to an embodiment of the present invention, such a process may be executed once every X minutes / hours / days, where X is a positive number that can be dynamically configured or changed. In this way, only requests related to hash values falling within a predefined finite range can be sent to the global server 100 for further processing. The global server 100 is capable of controlling the requested hash values. Therefore, the global server 100 does not need to continuously access the disk memory 102 to retrieve and update hash values, but can only process a part of the hash values in the memory 103. That is to say, when processing one of the parts, the global server 100 only needs to access the disk memory 102 once or twice. In addition, this can better control the amount of network traffic generated within a predefined time period.

[0088] In particular, after uploading a part of the hash metadata table to the memory 103, all requests received from the storage server 110 within a specific time period should be processed by the global server 100. That is to say, before persisting the updated part to the disk memory 102, the global server 100 needs to process each request for adding / deleting hash values and take actions according to each request (modifying the uploaded part of the hash metadata table).

[0089] Accordingly, the global server 100 can be used to modify the information 101 for one or more hash values falling within the first range 1011 of the hash values based on all requests received from the one or more storage servers 110 within a predetermined time period.

[0090] By enabling the global server 100 to initiate communication and control the range of hash values processed at any point in time, the storage server 110 only reports a limited range of hash values to the global server 100 within a specific time period. This ensures that the hash values within this range are in the memory 103 of the global server 100. This further avoids cache misses. In this way, the memory management of the global server 100 is optimized.

[0091] Figure 4FIG. 0 shows a storage server 110 provided by an embodiment of the present invention. The storage server 110 may include a processing circuit (not shown) for performing, implementing, or initiating various operations of the storage server 110 described herein. The processing circuit may include hardware and software. The hardware may include analog circuits or digital circuits, or both analog and digital circuits. The digital circuits may include components such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a multi-functional processor. In one embodiment, the processing circuit includes one or more processors and a non-transitory memory connected to the one or more processors. The non-transitory memory may carry executable program code, which, when executed by the one or more processors, causes the storage server 110 to perform, implement, or initiate the operations or methods described herein.

[0092] The storage server 110 is used for deduplication on the global server. In particular, Figure 4 the storage server 110 shown is Figure 1 or Figure 2 the same as the storage server 110 shown. And Figure 4 the global server 100 shown is Figures 1 - 3 the same as the global server 100 shown. According to an embodiment of the present invention, the storage server 110 is used to maintain a set of hash values, each hash value being the hash value of a stored data block. When a user requests to add or delete a first data block, the storage server 110 is used to record a request 111 associated with the first hash value of the first data block. Further, the storage server 110 is used to receive a notification of a first range 1011 of hash values from the global server 100. Accordingly, if the first hash value falls within the first range 1011 of hash values, the storage server 110 is used to send the recorded request 111 to the global server 100 to modify the information in the global server 100 for the first hash value.

[0093] Since the data stored in multiple deduplication servers or storage servers is usually duplicate, to avoid space wastage, the storage server 110 may request to store some data in the global server 100. According to an embodiment of the present invention, the global server 100 initiates communication between the global server 100 and the storage server 110 and controls which ranges of hash values can be provided to it. Therefore, according to an embodiment of the present invention, the storage server 110 does not randomly send requests to add or delete data to the global server 100, but records the requests and sends these requests (which can be sent individually) according to notifications or instructions from the global server 100.

[0094] Optionally, after sending the request 111 to the global server 100, the storage server 110 may be used to delete the request 111 for this record.

[0095] It is worth noting that after sending the request 111 related to the hash values falling within a specific range provided by the global server to the global server 100, it is assumed that the global server 100 will process the request 111 and take corresponding actions. To prevent redundant requests from being sent to the global server again, the storage server 110 will delete those requests that have been sent to the global server 100 from the record.

[0096] However, it should be noted that although the storage server 110 deletes the request 111 for this record (which will be saved until it is processed by the global server 100), the storage server 110 still retains the record regarding this hash value. For example, the storage server 110 will still store information regarding whether the data associated with this hash value is locally saved in the storage server 110 and / or the global server 100, as well as the local count associated with this hash value, which indicates the frequency at which the users of the storage server need this hash value.

[0097] Correspondingly, after deleting the request 111 for the record, the storage server 110 may be used to maintain information on the storage location of the first data block associated with the maintained first hash value, where the storage location of the first data block is the storage server 110 and / or the global server 100.

[0098] Optionally, the storage server 110 may be used to receive a broadcast message from the global server 100 carrying the first range 1011 of the hash values.

[0099] Optionally, the storage server 110 may also be used to compare the first range 1011 of the hash values with the first hash value. Further, the storage server 110 is used to determine whether the first hash value falls within the first range 1011 of the hash values.

[0100] Specifically, for each request for a record associated with its respective hash value, the storage server 110 will determine whether the request should be sent to the global server 100. That is, according to this embodiment, the request that can be sent must be related to a hash value that falls within a given range, i.e., the first range 1011.

[0101] It is worth noting that if the user requests to add the first data block, the request 111 may include a request for the first hash value of adding the first data block. Possibly, if the user requests to delete the first data block, the request may include a request for the first hash value of deleting the first data block.

[0102] Optionally, for a request for a record, even if it is allowed to send the request to the global server 100, the storage server 110 may also decide whether to send it. For example, for frequently accessed data, the storage server 110 may decide not to shunt it to the global server 100. Therefore, such data can be retained in the local storage server to allow for a lower read latency. Further, the storage server 110 may also decide not to shunt some data, such as some private data, or for security reasons, to the global server 100.

[0103] It is worth noting that after the storage server 110 sends the request 111 for the record to the global server 100 for processing, new incoming requests from the user will continue to be recorded in the storage server 110. The global server 100 may request the storage server 110 to send more requests for records, especially when the global server 100 has successfully processed all the received requests.

[0104] It should be understood that according to an embodiment of the present invention, when a user requests to add or delete a second data block, the storage server 110 may also be used to record a request associated with the second hash value of the second data block. Then, the storage server 110 can be used to receive a notification of a second range of hash values from the global server 100. Accordingly, if the second hash value falls within the second range of hash values, the storage server 110 can be used to send the recorded request to the global server 100 to modify the information in the global server 100 for the second hash value. Similarly, as in the foregoing embodiment, the storage server 110 can be used to delete the recorded request after successfully sending the request to the global server 100.

[0105] In the present invention, the storage server 110 sends requests to the global server 100 only during a specific time period (when instructed) and only sends requests that meet the conditions provided by the global server 100. In particular, the requests must be related to hash values that fall within a limited range notified by the global server 100. This can better control the amount of network traffic generated within a predefined time period.

[0106] Figure 5 The method 500 for deduplicating multiple storage servers 110 executed by the global server 100 provided by an embodiment of the present invention is shown. In particular, the global server 100 corresponds to Figure 1 the global server 100 described above. The method 500 includes: Step 501: maintaining information 101 related to a set of hash values, where each hash value is associated with a data block of data stored in the global server 100 and / or the storage server 110; Step 502: notifying the storage server 110 of a first range 1011 of the hash values; Step 503: receiving requests 111 from one or more of the storage servers 110 to modify the information 101 for one or more hash values that fall within the first range 1011 of the hash values; and Step 504: modifying the information 101 for one or more hash values that fall within the first range 1011 of the hash values based on the requests 111 received from the one or more storage servers 110. In particular, the storage server 110 is Figure 1 the storage device 110 described above. For those skilled in the art, Step 501 may be optional in the implementation of the method 500.

[0107] Figure 6 The method 600 for deduplicating in the global server 100 executed by the storage server 110 provided by an embodiment of the present invention is shown. In particular, the global server 100 is Figure 4the global server 100, the storage server 110 is Figure 4 the storage server 110. The method 600 includes: Step 601: maintaining a set of hash values, where each hash value is the hash value of a stored data block; Step 602: when a user requests to add or delete a first data block, recording a request 111 associated with the first hash value of the first data block; Step 603: receiving a notification of a first range 1011 of hash values from the global server 100; Step 604: if the first hash value falls within the first range 1011 of hash values, sending the recorded request to the global server 100 to modify information in the global server 100 for the first hash value. For those skilled in the art, Step 601 may be optional in the implementation of the method 600.

[0108] The present invention has been described in connection with different embodiments and implementation manners taken as examples. However, those skilled in the art can understand and obtain other variants by practicing the claimed invention, studying the drawings, the present disclosure, and the independent claims. In the claims and the description, the term "comprising" does not exclude other elements or steps, and "a" does not exclude a plurality of possibilities. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not mean that a combination of these measures cannot be used in an advantageous implementation manner.

[0109] In addition, any method according to an embodiment of the present invention can be implemented in a computer program having an encoding manner, which, when run by processing measures, can cause the processing measures to execute the method steps. The computer program is included in a computer-readable medium of a computer program product. The computer-readable medium can basically include any memory, such as ROM (Read Only Memory), PROM (Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), flash memory, EEPROM (Electrically Erasable Programmable Read Only Memory), and hard disk drive.

[0110] In addition, those skilled in the art will recognize that the user equipment 100 and the access node 110 include necessary communication capabilities in the form of, for example, functions, devices, units, elements, etc. for implementing the solutions of the present invention. Examples of other such methods, units, elements, and functions include: processors, memories, buffers, logic controls, encoders, decoders, rate matchers, de-rate matchers, mapping units, multipliers, decision units, selection units, switches, interleavers, de-interleavers, modulators, demodulators, inputs, outputs, antennas, amplifiers, receiving units, transmitting units, DSPs, trellis-coded modulation (TCM) encoders, TCM decoders, power units, power feeds, communication interfaces, and communication protocols, etc., which are reasonably arranged together to implement the said solutions.

[0111] In particular, processors 100 and 103 may include, for example, one or more instances of a central processing unit (CPU), a processing unit, a processing circuit, a processor, an application specific integrated circuit (ASIC), a microprocessor, or other processing logic that can interpret and execute instructions. The term "processor" may thus represent a processing circuit that includes multiple processing circuits, and the multiple processing circuit instances are any, some, or all of the above-listed items. The processing circuit may further perform data processing functions, input, output, and process data, and the functions include data buffering and device control functions, such as call processing control, user interface control, etc.

Claims

1. A global server (100) for deduplicating multiple storage servers (110), characterized in that, The global server includes a processor and a memory. The memory is used to store executable program code. When the processor executes the executable program code, the following is implemented: Notify the storage server (110) of a first range (1011) of the hash value set, where information (101) related to the hash value set is maintained, and each hash value is associated with a data block of data stored in the global server (100) and / or the storage server (110); Receive a request (111) from one or more of the storage servers (110) to modify the information (101) for one or more hash values that fall within the first range (1011) of the hash values; and Based on the request (111) received from the one or more storage servers (110), modify the information (101) for the one or more hash values that fall within the first range (1011) of the hash values; In response to the request received from the storage server (110) including a request to add a first hash value: In response to the first hash value not being included in the first part of the hash metadata table, add the first hash value to the first part of the hash metadata table, create a first watermark associated with the first hash value, and register the storage server that sent the request related to the first hash value, where the first watermark indicates whether the data block with the first hash value is highly replicated among the storage servers (110). The first part of the hash metadata table is associated with the first range (1011) of the hash values, and the hash metadata table contains the hash value set.

2. The global server (100) according to claim 1, characterized in that, For: Send a broadcast message carrying the first range (1011) of the hash values to the storage server (110).

3. The global server (100) according to claim 1 or 2, characterized in that The global server (100) includes a disk memory (102); and The information (101) includes a hash metadata table containing the hash value set, and the hash metadata table is stored in the disk memory (102).

4. The global server (100) according to claim 3, characterized in that, For: Divide the hash metadata table into N parts, where N is a positive integer not less than 2, and each part of the hash metadata table is associated with a different range of the hash values.

5. The global server (100) according to claim 4, characterized in that The global server (100) further includes a memory (103); and The global server (100) is further used for: Upload the first part of the hash metadata table to the memory (103); and Based on the request (111) received from the one or more storage servers (110), modify the first part of the hash metadata table.

6. The global server (100) according to any one of claims 3 to 5, characterized in that The hash metadata table includes information on each hash value in the hash value set and information related to one or more storage servers (110) registered for the hash value.

7. The global server (100) according to claim 6, characterized in that, For: In response to the first hash value being included in the first part of the hash metadata table, increasing the value of the first watermark associated with the first hash value and registering the storage server that sent the request related to the first hash value; and / or In response to a request received from a storage server (110) including a request to delete a second hash value: Decreasing the value of the second watermark associated with the second hash value and deregistering the storage server that sent the request related to the second hash value, where the second watermark indicates whether data blocks with the second hash value are highly replicated among the storage servers (110); And In response to the value of the second watermark being equal to 0, deleting the second hash value from the first part of the hash metadata table.

8. The global server (100) according to any one of claims 5 to 7, characterized in that, For: Persisting the updated first part of the hash metadata table to the disk memory (102).

9. The global server (100) according to any one of claims 1 to 8, characterized in that, For: Based on all the requests (111) received from the one or more storage servers (110) within a predetermined time period, modifying the information (101) for one or more hash values falling within the first range of the hash values.

10. A storage server (110) for deduplication in a global server (100), characterized in that, The storage server includes a processor and a memory, the memory is used to store executable program code, and when the processor executes the executable program code, it realizes: In response to a request to add or delete a first data block, recording the request (111) associated with the first hash value of the first data block, where the first hash value is included in the hash value set, maintaining the hash value set, and each hash value is the hash value of the stored data block; Receiving a notification of the first range (1011) of hash values from the global server (100); And If the first hash value falls within the first range (1011) of the hash values, sending the recorded request (111) to the global server (100) to modify the information in the global server (100) for the first hash value; Wherein, the recorded request (111) includes a request to add a first hash value, the request to add the first hash value is used to indicate adding the first hash value to the first part of the hash metadata table, creating a first watermark associated with the first hash value, and registering the storage server that sent the request related to the first hash value, the first watermark indicates whether data blocks with the first hash value are highly replicated among the storage servers (110), the first part of the hash metadata table is associated with the first range (1011) of the hash values, and the hash metadata table contains the hash value set.

11. The storage server (110) according to claim 10, wherein, For: Deleting the recorded request after sending the request to the global server (100).

12. The storage server (110) according to claim 11, characterized in that, For: After the request to delete the record, maintain information on the storage location of the first data block associated with the maintained first hash value, where the storage location of the first data block is the storage server and / or the global server (100).

13. The storage server (110) according to any one of claims 10 to 12, characterized in that, For: Receive (502) from the global server (100) a broadcast message carrying a first range of the hash values.

14. The storage server (110) according to any one of claims 10 to 13, characterized in that, For: Compare the first range of the hash values with the first hash value; and Determine whether the first hash value falls within the first range of the hash values.

15. The storage server (110) according to any one of claims 10 to 14, characterized in that If the user requests to add the first data block, the request includes a request to add the first hash value of the first data block; or If the user requests to delete the first data block, the request includes a request to delete the first hash value of the first data block.

16. A method (500) for deduplicating multiple storage servers (110), characterized in that, The method includes: Notify (502) the plurality of storage servers (110) of a first range (1011) of a set of hash values, where information (101) related to the set of hash values is maintained, and each hash value is associated with a data block of data stored in the storage server (110); Receive (503) from one or more of the storage servers a request (111) to modify the information for one or more hash values that fall within the first range (1011) of the hash values; and Based on the request (111) received from the one or more storage servers (110), modify (504) the information (101) for the one or more hash values that fall within the first range of the hash values; In response to the request received from the storage server (110) including a request to add a first hash value: In response to the first hash value not being included in the first part of the hash metadata table, add the first hash value to the first part of the hash metadata table, create a first watermark associated with the first hash value, and register the storage server that sent the request related to the first hash value, where the first watermark indicates whether the data block having the first hash value is highly replicated among the storage servers (110), the first part of the hash metadata table is associated with the first range (1011) of the hash values, and the hash metadata table contains the set of hash values.

17. A method (600) for deduplication in a global server, characterized in that, The method includes: In response to a request to add or delete a first data block, record (602) the request (111) associated with the first hash value of the first data block, where the first hash value is included in the set of hash values, and each hash value is the hash value of a data block of stored data; Receive (603) from the global server (100) a notification of a first range (1011) of hash values; and If the first hash value falls within a first range (1011) of the hash values, a request (111) for the record is sent (604) to the global server (100) to modify information in the global server (100) for the first hash value; Wherein, the recorded request (111) includes a request for adding the first hash value, the request for adding the first hash value is used to indicate adding the first hash value to a first part of the hash metadata table, creating a first watermark associated with the first hash value, and registering a storage server that has sent a request related to the first hash value, the first watermark indicating whether data blocks having the first hash value are highly duplicated among the storage servers (110), the first part of the hash metadata table is associated with the first range (1011) of the hash values, and the hash metadata table contains the set of hash values.

18. A computer program product, characterized in that, Comprising computer-readable code instructions which, when run on a computer, will cause the computer to perform the method according to claim 16 or 17.

19. A computer-readable storage medium, characterized in that, Comprising computer-executable computer program code instructions which, when run on a computer, execute the method according to claim 16 or 17.

Citation Information

Patent Citations

  • Using index partitioning and reconciliation for data deduplication

    CN102591946A