Method, device and storage medium for deduplicating data in object storage system
By determining the target granularity in the object storage system, filtering duplicate data and adding soft links, the problem that the object storage system does not support duplicate data deletion is solved, and automated deletion, performance optimization and user experience improvement are achieved.
Patent Information
- Application Number
- CN202410841039.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-06-26
AI Technical Summary
Existing object storage systems do not support deduplication, resulting in waste of storage space and excessive performance consumption.
By determining the target granularity of deduplication of the object storage system, filtering out duplicate data, and adding soft links to the baseline data to the metadata, thereby reclaiming the storage space of duplicate data.
It realizes automated deduplication, reduces performance overhead, improves system stability and security, and does not require additional complex data management and operation and maintenance operations, improving user experience.
Smart Images

Figure CN118656025B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the technical field of computers and network communications, and in particular to a method, device, and storage medium for deleting duplicate data in an object storage system. Background Art
[0002] Data deduplication technology is an effective method to save data storage space. Currently, data deduplication technology has been widely used in storage systems in data centers, mainly in backup and archive storage systems, primary storage systems, such as all-flash arrays, etc.
[0003] Object storage is a technology that stores and manages data in an unstructured format (called an object). Currently, object storage cloud services have been widely used in various industries and serve hundreds of millions of customers. However, currently, large-scale object storage systems in the industry do not support the deduplication function. Summary of the invention
[0004] The embodiments of the present disclosure provide a method, device and storage medium for deduplicating data in an object storage system, so as to realize automatic deduplication of data in the object storage system.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for deduplicating data in an object storage system, comprising:
[0006] Determine a target granularity for deduplication of the object storage system, where the target granularity is an object granularity or a slice granularity, wherein the object granularity takes the entire actual data of the object as a data processing unit, and the slice granularity takes the slice data corresponding to the actual data of the object as a data processing unit;
[0007] Performing duplicate data screening on each data of target granularity in the object storage system;
[0008] One data in any set of duplicate data is used as the reference data, and soft links are added to the metadata of other data in any set of duplicate data except the reference data to point to the reference data, and storage space of the other data is reclaimed.
[0009] In a second aspect, an embodiment of the present disclosure provides a data deduplication device for an object storage system, including:
[0010] A policy setting unit, used to determine a target granularity of deduplication of the object storage system, wherein the target granularity is an object granularity or a slice granularity, wherein the object granularity is to use the actual data of the object as a whole as a data processing unit, and the slice granularity is to use the slice data corresponding to the actual data of the object as a data processing unit;
[0011] A screening unit, used for screening duplicate data of each data of target granularity in the object storage system;
[0012] The deleting unit is used to use one data in any set of duplicate data as the benchmark data, add a soft link in the metadata of other data in any set of duplicate data except the benchmark data to point to the benchmark data, and reclaim the storage space of the other data.
[0013] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;
[0014] The memory stores computer-executable instructions;
[0015] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method for deduplication of the object storage system as described in the first aspect and various possible designs of the first aspect.
[0016] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method for deduplicating data in the object storage system as described in the first aspect and various possible designs of the first aspect is implemented.
[0017] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including computer executable instructions. When a processor executes the computer executable instructions, it implements the object storage system deduplication method described in the first aspect and various possible designs of the first aspect.
[0018] The object storage system deduplication method, device and storage medium provided by the embodiment of the present disclosure determine the target granularity of deduplication of the object storage system, the target granularity is the object granularity or the slice granularity, wherein the object granularity is to use the actual data of the object as a whole as the data processing unit, and the slice granularity is to use the slice data corresponding to the actual data of the object as the data processing unit; duplicate data screening is performed on each data of the target granularity in the object storage system; one data in any group of duplicate data is used as the reference data, and a soft link is added to the metadata of other data in any group of duplicate data except the reference data to point to the reference data, and the storage space of the other data is reclaimed. By determining the target granularity of duplicate data to maximize the benefits of deduplication, and by screening duplicate data, and pointing to the reference data through soft links in the metadata of duplicate data and reclaiming the storage space of duplicate data, the performance overhead of deduplication in the object storage system is reduced, and the stability and security are improved; and there is no need to add complex data management and operation and maintenance operations to customers, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0020] Figure 1a and Figure 1b An example diagram of a method for deduplication of an object storage system in the related art;
[0021] Figure 2 A schematic diagram of a method for deduplicating data in an object storage system according to an embodiment of the present disclosure;
[0022] Figure 3 A schematic diagram of a method for deduplicating data in an object storage system provided by another embodiment of the present disclosure;
[0023] Figure 4 A schematic diagram of a method for deduplicating data in an object storage system provided by another embodiment of the present disclosure;
[0024] Figure 5 A schematic diagram of a method for deduplicating data in an object storage system provided by another embodiment of the present disclosure;
[0025] Figure 6 A structural block diagram of a data deduplication device for an object storage system provided by an embodiment of the present disclosure;
[0026] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0028] Data deduplication technology is an effective method to save data storage space. Currently, data deduplication technology has been widely used in storage systems in data centers, mainly in backup and archive storage systems, primary storage systems, such as all-flash arrays, etc.
[0029] Object storage is a technology that stores and manages data in an unstructured format (called an object). Currently, object storage cloud services have been widely used in various industries and serve hundreds of millions of customers. The industry's object storage cloud services generally have the following characteristics: 1) large-scale, support for massive storage, multi-tenancy; 2) elasticity, strong scalability; 3) persistence, high reliability, and low cost.
[0030] However, due to factors such as technical complexity, large-scale object storage systems in the industry currently do not support the deduplication function. For example, the performance consumption of deduplication in object storage systems is too high, affecting the operation of the object storage system, or it is easy to have inaccurate duplicate data searches and large data granularity, resulting in little final benefit and possible security risks.
[0031] However, since object storage services are used by a large number of customers, with a wide range of business types and rich customer data styles, for some users, such as those using object storage for archiving and backup, the data backed up multiple times usually has a lot of duplicate data. Deduplication is very valuable and can greatly reduce the storage costs of tenants. For such scenarios with high data duplication, the industry's commonly used solutions are mainly the following two:
[0032] Solution 1: Add a deduplication system middle layer between users and the object storage system:
[0033] For example Figure 1aAs shown, the user first transmits the original data to the middle-layer deduplication system or deduplication software, which may be deployed in a local data center or a public cloud; after the deduplication system deduplicates the original data, it calls the API (Application Programming Interface) interface provided by the object storage system to store the deduplicated data in the object storage system.
[0034] Solution 2: Based on the analysis and reporting capabilities provided by the object storage system, users manage duplicate data themselves:
[0035] For example Figure 1b As shown, users write raw data in the form of objects into the public cloud object storage system through the API interface provided by the object storage system; the object storage system provides multi-dimensional analysis capabilities based on the object metadata list, including analysis reports on duplicate objects; users deploy data management software on the public cloud, rely on duplicate object analysis reports as input, make some decisions and object data management, and may call the object storage system to delete duplicate objects to save storage costs on the cloud.
[0036] However, both Solution 1 and Solution 2 essentially build duplicate data management capabilities outside the object storage system, so both have the problem of adding complex data management and operation and maintenance operations to customers.
[0037] Among them, the deduplication system in Solution 1, users can deploy mature commercial application software, but there are several problems: 1) It needs to hijack the protocol, and the data will be fragmented, and additional metadata management is required, which requires online processing, resulting in increased write latency; 2) It increases some operation and maintenance burdens. The backend object storage is elastic and supports unlimited expansion. The deduplication system needs to have sufficient elasticity and scalability to match; 3) This solution is coupled with the deduplication system, which increases the management difficulty for users and requires high IT operation and maintenance capabilities of customers. The customer base of cloud storage is huge, and for some customers, it is difficult to meet the above requirements.
[0038] Solution 2 essentially exposes the complex deduplication process to customers. Users need to develop data management software and configure duplicate object processing strategies based on their own business needs. Data management software needs to consider complex concurrency with user operations and requires a unified mechanism to control concurrency, otherwise serious problems such as object data loss may occur. In addition, this solution is currently only applicable to object-level deduplication. In some scenarios, such as database backup, there are a large number of similar files and a small number of identical files, so the deduplication effect is poor.
[0039] In order to solve at least one of the above technical problems, an embodiment of the present disclosure provides a method for deduplicating data in an object storage system, by determining the target granularity of deduplicating data in the object storage system, the target granularity is the object granularity or the slice granularity, wherein the object granularity is to use the actual data of the object as a whole as a data processing unit, and the slice granularity is to use the slice data corresponding to the actual data of the object as a data processing unit; duplicate data screening is performed on each data of the target granularity in the object storage system; one data in any group of duplicate data is used as the reference data, and a soft link is added to the metadata of other data in any group of duplicate data except the reference data to point to the reference data, and the storage space of the other data is reclaimed. By determining the target granularity of duplicate data to maximize the benefits of deduplication, and by screening duplicate data, and pointing to the reference data through soft links in the metadata of duplicate data and reclaiming the storage space of duplicate data, the performance overhead of deduplication in the object storage system is reduced, the stability and security are improved, and the purpose of saving storage costs can be effectively achieved; and there is no need to add complex data management and operation and maintenance operations to customers, thereby improving user experience.
[0040] It should be noted that the data involved in this application (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0041] The data deduplication method of the object storage system disclosed in the present invention will be described in detail below in conjunction with specific embodiments.
[0042] refer to Figure 2 , Figure 2 A flowchart of a method for deduplicating data in an object storage system provided in an embodiment of the present disclosure is provided. The method of this embodiment can be applied in a terminal device or a server. The method for deduplicating data in an object storage system includes:
[0043] S201. Determine a target granularity for deduplication of an object storage system, where the target granularity is an object granularity or a slice granularity, wherein the object granularity takes the entire actual data of the object as a data processing unit, and the slice granularity takes the slice data corresponding to the actual data of the object as a data processing unit.
[0044] In this embodiment, the object storage system may be any system capable of performing object storage, such as an object storage cloud service.
[0045] When it is determined that the object storage system needs to be deduplicated, the deduplication strategy can be determined first. The most important item in the deduplication strategy is the target granularity of deduplication, that is, at what granularity (data processing unit) to filter duplicate data and perform deduplication. The target granularity can be object granularity or slice granularity. The object granularity is to filter and delete the duplicate actual data of the object based on the actual data of the object as a whole. The slice granularity is to divide the actual data of the object into multiple slices of data, filter out duplicate slices of data and delete them. The slice granularity is adopted because in some scenarios, such as updating the actual data of the object, it is possible that only a part of the actual data of the object is updated. The actual data before and after the update of the same object is not duplicate data, but some slice data may be duplicate. Therefore, compared with the object granularity, the slice granularity can improve the deduplication effect and save more storage costs.
[0046] Optionally, when determining the target granularity, the benefit calculation method may be used to determine whether the benefit of using the object granularity is greater or the benefit of using the slice granularity is greater. The details are as follows:
[0047] Determine a first benefit of deduplicating the object storage system at an object granularity and a second benefit of deduplicating the object storage system at a slice granularity; and determine a target granularity of deduplicating the object storage system based on the first benefit and the second benefit.
[0048] In this embodiment, the first benefit and the second benefit can be measured by the total storage cost saved, that is, the first benefit can be the total storage space saved when deduplicating data in the object storage system at the object granularity (including the storage space recovered when deduplicating data, and the new storage space occupied during the processing, such as adding soft links in metadata in subsequent embodiments, resulting in an increase in the storage space occupied by metadata, etc.), and the second benefit can be the total storage space saved when deduplicating data in the object storage system at the slice granularity (including the storage space recovered when deduplicating data, and the new storage space occupied during the processing, such as adding soft links in metadata in subsequent embodiments, resulting in an increase in the storage space occupied by metadata, and the new storage space occupied by non-duplicate slice data rewritten as new slice data in subsequent embodiments. In addition, other factors can be combined with the first benefit and the second benefit, such as the consumption of performance, etc., which are not listed here one by one. Optionally, the first benefit and the second benefit can be determined by sampling and analyzing the object storage system, or can also be determined based on some historical deduplication processes with reference value (such as data of the same or similar users, etc.).
[0049] Further, when determining the target granularity of deduplication of the object storage system according to the first benefit and the second benefit, it may specifically include:
[0050] If the difference between the first benefit and the second benefit is less than a preset threshold (for example, 10%), the target granularity is determined to be the object granularity, that is, when the difference between the first benefit and the second benefit is not large, a larger granularity is adopted to reduce the amount of data, reduce performance consumption, and reduce time consumption; or, if the difference between the first benefit and the second benefit is not less than a preset threshold, the granularity corresponding to the maximum benefit between the first benefit and the second benefit is determined as the target granularity, that is, if the difference between the first benefit and the second benefit is large, a larger benefit granularity is adopted to obtain a greater benefit.
[0051] Of course, if the maximum benefit of the first benefit and the second benefit is less than the preset benefit threshold, that is, the first benefit and the second benefit are not large, the deduplication function of the object storage system can be turned off. In addition, the opening or closing of the deduplication function of the object storage system can also be controlled by the user to meet the user's needs. In addition, when the user turns on the deduplication function of the object storage system, if the maximum benefit of the first benefit and the second benefit is less than the preset benefit threshold, the deduplication function of the object storage system can also be turned off, otherwise the deduplication can continue.
[0052] Optionally, when determining the second benefit of deduplicating the object storage system at a slice granularity, it is taken into account that the splitting method and / or splitting length of the actual data of the object may affect the benefit. Therefore, the second benefit of deduplicating the object storage system at a slice granularity when different splitting methods and / or different splitting lengths are used can be determined respectively. Then, based on the first benefit and the second benefits in multiple different situations, the target granularity of deduplicating the object storage system can be determined. If the target granularity is determined to be the slice granularity, and the second benefit in a certain situation is the largest, the splitting method and / or splitting length corresponding to the largest second benefit can also be determined as the target splitting method and / or target splitting length, so that the actual data of the object can be split based on the target splitting method and / or target splitting length in subsequent embodiments. That is, the target splitting method and / or target splitting length can also be used as a strategy in the deduplication strategy. Optionally, the segmentation method may include but is not limited to content-defined-chunking (CDC) or fixed-length segmentation, etc., wherein the CDC method can ensure that most of the slice data contents of the same object are the same, so as to better save storage costs. For example, the Rabin Fingerprinting algorithm is an algorithm F in the CDC method.
[0053] S202: Screening duplicate data for each data of target granularity in the object storage system.
[0054] In this embodiment, the first check value of each data of the target granularity in the object storage system can be obtained, wherein the first check value can be a SHA1 (Secure Hash Algorithm 1) code, or a CRC (Cyclic Redundancy Check) code. Of course, other types of check codes can also be used. If the data of two target granularities are the same data, the first check codes are the same. Therefore, the first check code can be used to quickly screen the duplicate data of the target granularity in the object storage system.
[0055] Optionally, the first verification value may be calculated by a specific algorithm, or may be pre-calculated and stored in the metadata of the object, and obtained from the metadata, where the metadata may be maintained by a metadata management component of the object storage system; or the first verification value may also be obtained by other means.
[0056] S203: Take one data in any group of duplicate data as the benchmark data, add a soft link in the metadata of other data in any group of duplicate data except the benchmark data to point to the benchmark data, and reclaim the storage space of the other data.
[0057] In this embodiment, after filtering out duplicate data, one of the data in any group of duplicate data is selected as the benchmark data, which can be retained instead of being deleted, while other data that is duplicated with the benchmark data needs to be deleted. The deletion process is to add a soft link pointing to the benchmark data in the metadata of other data that is duplicated with the benchmark data, and reclaim the storage space of other data, that is, the metadata of other data is retained, and only the actual data of other data is deleted. The purpose of the soft link is that the benchmark data can be found through the soft link when the actual data of these other data is subsequently searched based on the metadata.
[0058] The deduplication method of the object storage system of this embodiment determines the target granularity of deduplication of the object storage system, the target granularity is the object granularity or the slice granularity, wherein the object granularity is to use the actual data of the object as a whole as the data processing unit, and the slice granularity is to use the slice data corresponding to the actual data of the object as the data processing unit; duplicate data screening is performed on each data of the target granularity in the object storage system; one data in any group of duplicate data is used as the reference data, and a soft link is added to the metadata of other data in any group of duplicate data except the reference data to point to the reference data, and the storage space of the other data is reclaimed. By determining the target granularity of deduplication data to maximize the benefits of deduplication, and by screening duplicate data, and pointing to the reference data through soft links in the metadata of duplicate data and reclaiming the storage space of duplicate data, the performance overhead of deduplication in the object storage system is reduced, and the stability and security are improved; and there is no need to add complex data management and operation and maintenance operations to customers, thereby improving the user experience.
[0059] Based on any of the above embodiments, in the above embodiments, a first check value of each data of a target granularity in the object storage system is obtained, and data with the same first check value is screened out from the object storage system to be determined as duplicate data, specifically as follows: Figure 3 As shown, it may include:
[0060] S2021, traversing each data of the target granularity in the object storage system, obtaining a first check value of the currently traversed data during the traversal process, and determining whether the first check value of the currently traversed data matches a historical first check value in the first set;
[0061] The first set stores the first historical verification values and the identifiers of the corresponding data obtained for the first time during the traversal process;
[0062] If it is determined that there is a historical first check value in the first set that is the same as the first check value of the currently traversed data, execute S2022; otherwise, execute S2023;
[0063] S2022: If it is determined that there is a historical first check value in the first set that is the same as the first check value of the currently traversed data, determine that the currently traversed data is duplicate data;
[0064] S2023. If it is determined that there is no historical first verification value in the first set that is the same as the first verification value of the currently traversed data, it is determined that the currently traversed data is not duplicate data, and the first verification value of the currently traversed data is used as the historical first verification value, and is associated with the identifier of the currently traversed data and stored in the first set.
[0065] In this embodiment, a first set can be created, and the first set is used to store the first verification values and the identifiers of the corresponding data obtained for the first time during the traversal process, for example, in the form of key-value pairs. For example, in the object granularity, the key-value pair <first verification value, object name> can be used to store it in the first set, and in the slice granularity, the key-value pair <first verification value, slice data name> can be used to store it in the first set.
[0066] Furthermore, in the process of traversing each data of the target granularity in the object storage system, the first check value of the currently traversed data can be used, and the first check value of the currently traversed data can be judged to match with the historical first check value in the first set. If it is determined that there is a historical first check value in the first set that is the same as the first check value of the currently traversed data, it means that the currently traversed data and the data corresponding to the historical first check value are duplicate data; if it is determined that there is no historical first check value in the first set that is the same as the first check value of the currently traversed data, it means that the currently traversed data is not duplicated with any traversed data (but data that is duplicated with the currently traversed data may be found in subsequent traversal processes), and the first check value of the currently traversed data is used as the historical first check value, associated with the identifier of the currently traversed data and stored in the first set as the matching basis for the subsequent traversal process, and then the traversal of the data of the next target granularity is continued.
[0067] Among them, after determining that there is a historical first verification value in the first set that is identical to the first verification value of the currently traversed data, thereby determining that the currently traversed data and the data corresponding to the historical first verification value are duplicate data, the data corresponding to the historical first verification value can be used as benchmark data, and a soft link can be added to the metadata of the currently traversed data to point to the benchmark data, and the storage space of the currently traversed data can be reclaimed, thereby completing the traversal of the currently traversed data and continuing to traverse the data of the next target granularity.
[0068] Based on the above embodiments, there are some differences in the processing process when the target granularity is the object granularity and the slice granularity. The two granularities will be described separately below.
[0069] In one embodiment, the target granularity is the object granularity, and the overall process of deduplication can be as follows: Figure 4 In the above embodiment, traversing each data of the target granularity in the object storage system and obtaining the first check value of the currently traversed data during the traversal process may specifically include:
[0070] Traverse each object in the object storage system, and during the traversal, obtain the second check value of the currently traversed object from the metadata of the currently traversed object, and query the occurrence times of the second check value of the currently traversed object in the second set, where the second set includes each second check value and the corresponding occurrence times among the second check values of all objects in the object storage system;
[0071] If the occurrence times of the second check value of the currently traversed object are greater than once, obtain the actual data of the currently traversed object, and obtain the first check value of the actual data of the currently traversed object.
[0072] In this embodiment, the second check value of each object in the object storage system may be pre-stored in the metadata of each object. The second check value may be a CRC code or an MD5 (Message-Digest Algorithm 5), and of course, it may also be an SHA1 code or other check values. Similarly, if the actual data of two objects is the same, the second check values of the two objects are the same. Based on this, the second check values of all objects can be pre-obtained from the metadata of all objects in the object storage system, and the occurrence times (reference count) of each second check value among the second check values of all objects are determined. A second set is constructed according to each second check value among the second check values of all objects and the corresponding occurrence times. For example, it can be stored in the second set in the form of key-value pairs. The key-value pairs can be <CRC, occurrence times> or <Md5, occurrence times>. Optionally, the specific process of constructing the second set may be to traverse the metadata of all objects in the object storage system. During the traversal, obtain the second check value of the currently traversed object, and query whether there is the same second check value in the second set. If not, write it into the second set in the above key-value pair manner and record the occurrence times as 1; if it exists, increment the occurrence times of the second check value in the second set to update the occurrence times. In this embodiment, the second set can be updated in an atomic manner to ensure the accuracy of the update process.
[0073] On the basis of having obtained the second set, in this embodiment, each object in the object storage system can be traversed. In specific implementation, the object metadata list of the object storage system can be traversed. During the traversal process, the second check value of the current traversed object is obtained from the metadata of the current traversed object, and the number of occurrences of the second check value of the current traversed object is queried from the second set; if the number of occurrences of the second check value of the current traversed object is greater than once, it means that the current traversed object is duplicate data, but at this time it is not clear with whom the current traversed object is duplicated, and whether the data duplicated with the current traversed object has been traversed and processed. Therefore, at this time, the actual data of the current traversed object can be obtained, wherein the actual data of the object can be maintained by the actual data storage component of the object storage system, and the first check value of the actual data of the current traversed object is obtained, and the search process based on the first check value in the above embodiment is performed. In this embodiment, if the number of occurrences of the second check value of the current traversed object is not greater than once, there is no need to obtain the first check value and other subsequent processes, which greatly reduces performance consumption and time consumption.
[0074] Furthermore, after obtaining the first check value of the actual data of the currently traversed object, it can be determined whether the first check value of the currently traversed object matches the historical first check value in the first set; if it is determined that there is a historical first check value in the first set that is identical to the first check value of the currently traversed object, then it is determined that the currently traversed object is duplicate data, and the object corresponding to the historical first check value in the first set that is identical to the first check value of the currently traversed data is used as a reference object, a soft link is added to the metadata of the currently traversed object to point to the reference object, and the storage space of the actual data of the currently traversed object is reclaimed; or, if it is determined that there is no historical first check value in the first set that is identical to the first check value of the currently traversed object, then it is determined that the currently traversed object is not duplicate data, and the first check value of the currently traversed object is used as the historical first check value, and is associated with the identifier of the currently traversed object and stored in the first set.
[0075] In another embodiment, the target granularity is the slice granularity, and the overall process of deduplication can be as follows: Figure 5 In the above embodiment, traversing each data of the target granularity in the object storage system and obtaining the first check value of the currently traversed data during the traversal process may specifically include:
[0076] Traversing each object in the object storage system, acquiring actual data of the currently traversed object during the traversal process, and slicing the actual data of the currently traversed object to obtain slice data;
[0077] Obtain a first check value of the slice data of the currently traversed object.
[0078] In this embodiment, since the second check value in the second set in the above example is obtained based on the actual data of the object as a whole, the slice data in this embodiment is no longer applicable. In this embodiment, each object in the object storage system is traversed. In the specific implementation, the object metadata list of the object storage system can be traversed, and the actual data of the current traversed object is obtained during the traversal process, and the actual data of the current traversed object is segmented using the segmentation method and / or segmentation length determined in the above embodiment to obtain slice data, and then the first check value of the slice data is obtained for the slice data of the current traversed object, where the first check value is the same as in the above embodiment and is not limited here.
[0079] Further, when duplicate data is screened according to the first set, if it is determined that there exists in the first set a historical first check value identical to the first check value of any slice data of the currently traversed object, then it is determined that the slice data corresponding to the historical first check value is duplicated; if it is determined that there does not exist in the first set a historical first check value identical to the first check value of any slice data, then it can be determined that any slice data is not duplicate data, and the any slice data is rewritten into the object storage system as a new slice data, occupying a new storage space, and the first check value of the new slice data is used as the historical first check value, and is associated with the identifier of the new slice data and stored in the first set.
[0080] Furthermore, after the above process, the first check value of each slice data of the currently traversed object has the same historical first check value in the first set (including slice data that is determined not to be duplicate data, which is also the same as the historical first check value of new slice data newly added to the first set), that is, at this time, each slice data of the currently traversed object can find slice data that is duplicated with each other in the first set, and the slice data corresponding to all slice data included in the currently traversed object in the first set (that is, the slice data with the same first check value as each other) are used as benchmark data, and soft links pointing to each benchmark data are added to the metadata of the currently traversed object, and each soft link points to each slice data used as benchmark data, and the storage space occupied by the actual data of the currently traversed object as a whole is reclaimed.
[0081] Furthermore, after the processing of the currently traversed object is completed, the next object is traversed until all objects are processed.
[0082] On the basis of any of the above embodiments, considering that the above object storage system deduplication method can be to dedupe all data in the object storage system, or to dedupe part of the data in the object storage system, in this embodiment, a target range for deduplication in the object storage system can be first determined to dedupe within the target range, that is, the processing processes in the above embodiments are all performed within the target range of the object storage system. Specifically, a target granularity for deduplication within the target range of the object storage system is determined, a first check value of each data of the target granularity within the target range of the object storage system is obtained, and data with the same first check value are screened out from the target range of the object storage system to be determined as duplicate data; one data in any group of duplicate data is used as the reference data, and a soft link is added to the metadata of other data except the reference data in any group of duplicate data to point to the reference data, and the storage space of other data is reclaimed, which will not be described in detail here.
[0083] Optionally, the object storage system may include multiple data buckets, so the target range for deduplication in the object storage system may be one or more target data buckets among all data buckets in the object storage system, wherein the target data bucket may optionally be a data bucket of the same user, or a data bucket selected from all data buckets in the object storage system by other means.
[0084] Assuming that there are multiple target data buckets, whether to use one target data bucket or multiple target data buckets as the target range for deduplication can be determined in the following way.
[0085] Obtain a second check value for each object in a plurality of target data buckets; determine a first duplication ratio of objects in each target data bucket and a second duplication ratio of all objects in the plurality of target data buckets based on the second check value; determine whether a target range for deduplication in the object storage system is one target data bucket or a plurality of target data buckets based on the first duplication ratio and the second duplication ratio.
[0086] In this embodiment, by obtaining the second verification value of all objects in multiple target data buckets (for example, obtained from metadata in the above embodiment), a preliminary screening of duplicate data can be performed, and then the amount of duplicate data in the single target data bucket can be determined based on the second verification value of each object in the single target data bucket, and the first duplication ratio of the objects in the single target data bucket can be determined. Similarly, the amount of duplicate data in the single target data bucket can be determined based on the second verification value of each object in the multiple target data buckets, and the second duplication ratio of the objects in the multiple target data buckets can be determined. Then, based on the first duplication ratio and the second duplication ratio, it can be determined whether the target range has a greater benefit from one target data bucket or multiple target data buckets. When determining the second repetition ratio of all objects in multiple target data buckets, the second repetition ratio of objects in different data bucket combinations in the multiple target data buckets can be determined. For example, assuming that the multiple target data buckets include target data bucket 1, target data bucket 2, and target data bucket 3, the second repetition ratio of all objects in target data bucket 1 and target data bucket 2, the second repetition ratio of all objects in target data bucket 1 and target data bucket 3, the second repetition ratio of all objects in target data bucket 2 and target data bucket 3, and the second repetition ratio of all objects in target data bucket 1, target data bucket 2, and target data bucket 3 can be determined. Accordingly, the target range may be a single target data bucket or any of the above-mentioned data bucket combinations.
[0087] Optionally, when determining a target range for deduplication in the object storage system according to the first duplication ratio and the second duplication ratio, the method may specifically include:
[0088] If the second duplication ratio exceeds the preset ratio, it means that there are more duplicate data in multiple target data buckets, and the target range can be determined to be multiple target data buckets, which has greater benefits; or, if the second duplication ratio does not exceed the preset ratio, and the first duplication ratio corresponding to any target data bucket exceeds the preset ratio, it means that there may not be much duplicate data between multiple target data buckets, and there are more duplicate data in a single target data bucket, and the target range is determined to be any target data bucket. Of course, there may be more than one target data bucket whose corresponding first duplication ratio exceeds the preset ratio. Each target data bucket whose first duplication ratio exceeds the preset ratio can be used as a target range to delete duplicate data in each target range.
[0089] Optionally, the method for deduplicating data in an object storage system of any of the above embodiments can perform deduplication of data offline in the background, without affecting the normal operation of the object storage system, thereby improving the security and stability of the system.
[0090] Corresponding to the method for deduplicating data in the object storage system of the above embodiment, Figure 6A structural block diagram of a data deduplication device for an object storage system provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 6 The deduplication device 600 of the object storage system in this embodiment includes: a policy setting unit 601, a screening unit 602, and a deleting unit 603.
[0091] The policy setting unit 601 is used to determine the target granularity of deduplication of the object storage system, wherein the target granularity is the object granularity or the slice granularity, wherein the object granularity is to use the actual data of the object as a whole as a data processing unit, and the slice granularity is to use the slice data corresponding to the actual data of the object as a data processing unit;
[0092] A screening unit 602 is used to screen duplicate data for each data of a target granularity in the object storage system;
[0093] The deleting unit 603 is used to use one data in any set of duplicate data as the benchmark data, add a soft link in the metadata of other data in any set of duplicate data except the benchmark data to point to the benchmark data, and reclaim the storage space of the other data.
[0094] In one or more embodiments of the present disclosure, when determining the target granularity of deduplication of the object storage system, the policy setting unit 601 is used to:
[0095] Determining a first benefit of deduplicating data for the object storage system at an object granularity and a second benefit of deduplicating data for the object storage system at a slice granularity;
[0096] A target granularity of deduplication of the object storage system is determined according to the first benefit and the second benefit.
[0097] In one or more embodiments of the present disclosure, when determining the target granularity of deduplication of the object storage system according to the first benefit and the second benefit, the policy setting unit 601 is configured to:
[0098] If the difference between the first benefit and the second benefit is less than a preset threshold, determining the target granularity to be the object granularity; or
[0099] If the difference between the first income and the second income is not less than a preset threshold, a granularity corresponding to a maximum income between the first income and the second income is determined as the target granularity.
[0100] In one or more embodiments of the present disclosure, when the screening unit 602 screens duplicate data of each data of a target granularity in the object storage system, it is configured to:
[0101] A first check value of each data of a target granularity in the object storage system is obtained, and data with the same first check value is screened out from the object storage system to be determined as duplicate data.
[0102] In one or more embodiments of the present disclosure, when the screening unit 602 obtains the first check value of each data of the target granularity in the object storage system and screens out data with the same first check value from the object storage system to determine as duplicate data, it is configured to:
[0103] Traversing each data of a target granularity in the object storage system, obtaining a first check value of the currently traversed data during the traversal process, and determining whether the first check value of the currently traversed data matches a historical first check value in a first set, wherein the first set stores each historical first check value obtained for the first time during the traversal process and an identifier of the corresponding data;
[0104] If it is determined that there is a historical first check value in the first set that is the same as the first check value of the currently traversed data, then the currently traversed data is determined to be duplicate data; or
[0105] If it is determined that there is no historical first verification value in the first set that is the same as the first verification value of the currently traversed data, it is determined that the currently traversed data is not duplicate data, and the first verification value of the currently traversed data is used as the historical first verification value, and is associated with the identifier of the currently traversed data and stored in the first set.
[0106] In one or more embodiments of the present disclosure, when the deleting unit 603 uses one data in any group of duplicate data as the reference data, adds a soft link in the metadata of other data in any group of duplicate data except the reference data to point to the reference data, and reclaims the storage space of the other data, it is used to:
[0107] The historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data is used as the benchmark data, a soft link is added to the metadata of the currently traversed data to point to the benchmark data, and the storage space of the currently traversed data is reclaimed.
[0108] In one or more embodiments of the present disclosure, if the target granularity is an object granularity, the screening unit 602, when traversing each data of the target granularity in the object storage system and obtaining the first check value of the currently traversed data during the traversal process, is used to:
[0109] Traversing each object in the object storage system, during the traversal process, obtaining a second check value of the currently traversed object from metadata of the currently traversed object, and querying the number of occurrences of the second check value of the currently traversed object from a second set, wherein the second set includes each second check value of all objects in the object storage system and the corresponding number of occurrences;
[0110] If the second check value of the currently traversed object appears more than once, the actual data of the currently traversed object is obtained, and the first check value of the actual data of the currently traversed object is obtained.
[0111] In one or more embodiments of the present disclosure, before traversing each object in the object storage system, the screening unit 602 is further configured to:
[0112] Obtaining second check values of all objects from metadata of all objects in the object storage system, and determining the number of occurrences of each second check value among the second check values of all objects;
[0113] The second set is constructed according to each second check value of all objects and the corresponding number of occurrences.
[0114] In one or more embodiments of the present disclosure, if the target granularity is a slice granularity, the screening unit 602, when traversing each data of the target granularity in the object storage system and obtaining the first check value of the currently traversed data during the traversal process, is used to:
[0115] Traversing each object in the object storage system, acquiring actual data of the currently traversed object during the traversal process, and slicing the actual data of the currently traversed object to obtain slice data;
[0116] Obtain a first check value of the slice data of the currently traversed object.
[0117] In one or more embodiments of the present disclosure, if the screening unit 602 determines that there is no historical first check value in the first set that is the same as the first check value of the currently traversed data, it determines that the currently traversed data is not duplicate data, and uses the first check value of the currently traversed data as the historical first check value, and associates it with the identifier of the currently traversed data and stores it in the first set, for:
[0118] If it is determined that there is no historical first check value in the first set that is the same as the first check value of any slice data of the currently traversed object, it is determined that any slice data is not duplicate data, and the any slice data is rewritten into the object storage system as a new slice data, and the first check value of the new slice data is used as the historical first check value, and is associated with the identifier of the new slice data and stored in the first set.
[0119] In one or more embodiments of the present disclosure, the deleting unit 603, when taking the historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data as the reference data, adding a soft link in the metadata of the currently traversed data to point to the reference data, and reclaiming the storage space of the currently traversed data, is used to:
[0120] The slice data corresponding to the first set of all the slice data included in the currently traversed object are used as reference data, soft links pointing to the respective reference data are added in the metadata of the currently traversed object, and the storage space occupied by the actual data of the currently traversed object is reclaimed.
[0121] In one or more embodiments of the present disclosure, when the screening unit 602 divides the actual data of the currently traversed object to obtain slice data, it is used to:
[0122] Split the actual data of each object based on the content in an indefinite length segmentation method; or
[0123] The actual data of each object is segmented using a fixed-length segmentation method.
[0124] In one or more embodiments of the present disclosure, when determining the second benefit of deduplication of the object storage system at the slice granularity, the policy setting unit 601 is used to:
[0125] Determine respectively a second benefit of performing data deduplication of the object storage system at a slice granularity when different slicing methods and / or different slicing lengths are adopted;
[0126] Accordingly, after determining the target granularity of deduplication of the object storage system, the policy setting unit 601 is further configured to:
[0127] If the target granularity is determined to be the slice granularity, the segmentation method and / or segmentation length corresponding to the maximum second benefit is determined as the target segmentation method and / or target segmentation length.
[0128] In one or more embodiments of the present disclosure, before determining the target granularity of deduplication of the object storage system, the policy setting unit 601 is further configured to:
[0129] Determine a target range for deduplication in the object storage system to perform deduplication within the target range;
[0130] Accordingly, when determining the target granularity of deduplication of the object storage system, the policy setting unit 601 is used to:
[0131] A target granularity for deduplication within the target range is determined.
[0132] In one or more embodiments of the present disclosure, when determining a target range for deduplication in the object storage system, the policy setting unit 601 is configured to:
[0133] Obtaining a second check value for each object in a plurality of target data buckets in the object storage system;
[0134] Determine a first repetition ratio of objects in each target data bucket and a second repetition ratio of objects in multiple target data buckets according to the second check value;
[0135] A target range for deduplication in the object storage system is determined according to the first duplication ratio and the second duplication ratio.
[0136] In one or more embodiments of the present disclosure, when determining a target range for deduplication in the object storage system according to the first duplication ratio and the second duplication ratio, the policy setting unit 601 is configured to:
[0137] If the second repetition ratio exceeds a preset ratio, determining the target range to be the plurality of target data buckets; or
[0138] If the second repetition ratio does not exceed the preset ratio, and the first repetition ratio corresponding to any target data bucket exceeds the preset ratio, the target range is determined to be any target data bucket.
[0139] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment, and maximizes the benefits of deduplication by determining the target granularity of duplicate data to be deleted, and reduces the performance overhead of deduplication within the object storage system and improves stability and security by screening duplicate data and pointing to baseline data through soft links in the metadata of the duplicate data. The implementation principle and technical effects are similar and will not be repeated in this embodiment.
[0140] refer to Figure 7, which shows a schematic diagram of the structure of an electronic device 700 suitable for implementing the embodiment of the present disclosure, and the electronic device 700 may be a terminal device or a server. The terminal device may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0141] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 to a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0142] Typically, the following devices may be connected to the I / O interface 705: input devices 706 such as a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 such as a magnetic tape, a hard disk, etc.; and communication devices 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0143] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 709, or installed from a storage device 708, or installed from a ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0144] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0145] The computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
[0146] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0147] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0149] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of a unit does not limit the unit itself in some cases. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses".
[0150] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0151] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0152] In a first aspect, according to one or more embodiments of the present disclosure, a method for deduplicating data in an object storage system is provided, comprising:
[0153] Determine a target granularity for deduplication of the object storage system, where the target granularity is an object granularity or a slice granularity, wherein the object granularity takes the entire actual data of the object as a data processing unit, and the slice granularity takes the slice data corresponding to the actual data of the object as a data processing unit;
[0154] Performing duplicate data screening on each data of target granularity in the object storage system;
[0155] One data in any set of duplicate data is used as the reference data, and soft links are added to metadata of other data in any set of duplicate data except the reference data to point to the reference data, and storage space of the other data is reclaimed.
[0156] According to one or more embodiments of the present disclosure, determining a target granularity of deduplication of the object storage system includes:
[0157] Determining a first benefit of deduplicating data for the object storage system at an object granularity and a second benefit of deduplicating data for the object storage system at a slice granularity;
[0158] A target granularity of deduplication of the object storage system is determined according to the first benefit and the second benefit.
[0159] According to one or more embodiments of the present disclosure, determining a target granularity of deduplication of the object storage system according to the first benefit and the second benefit includes:
[0160] If the difference between the first benefit and the second benefit is less than a preset threshold, determining the target granularity to be the object granularity; or
[0161] If the difference between the first income and the second income is not less than a preset threshold, a granularity corresponding to a maximum income between the first income and the second income is determined as the target granularity.
[0162] According to one or more embodiments of the present disclosure, the step of screening duplicate data for each data of a target granularity in the object storage system includes:
[0163] A first check value of each data of a target granularity in the object storage system is obtained, and data with the same first check value is screened out from the object storage system to be determined as duplicate data.
[0164] According to one or more embodiments of the present disclosure, obtaining a first check value of each data of a target granularity in the object storage system, and screening out data with the same first check value from the object storage system to determine as duplicate data, includes:
[0165] Traversing each data of a target granularity in the object storage system, obtaining a first check value of the currently traversed data during the traversal process, and determining whether the first check value of the currently traversed data matches a historical first check value in a first set, wherein the first set stores each historical first check value obtained for the first time during the traversal process and an identifier of the corresponding data;
[0166] If it is determined that there is a historical first check value in the first set that is the same as the first check value of the currently traversed data, then the currently traversed data is determined to be duplicate data; or
[0167] If it is determined that there is no historical first verification value in the first set that is the same as the first verification value of the currently traversed data, it is determined that the currently traversed data is not duplicate data, and the first verification value of the currently traversed data is used as the historical first verification value, and is associated with the identifier of the currently traversed data and stored in the first set.
[0168] According to one or more embodiments of the present disclosure, taking one data in any group of duplicate data as the reference data, adding a soft link in the metadata of other data in any group of duplicate data except the reference data to point to the reference data, and reclaiming the storage space of the other data, includes:
[0169] The historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data is used as the benchmark data, a soft link is added to the metadata of the currently traversed data to point to the benchmark data, and the storage space of the currently traversed data is reclaimed.
[0170] According to one or more embodiments of the present disclosure, if the target granularity is an object granularity, traversing each data of the target granularity in the object storage system and obtaining a first check value of the currently traversed data during the traversal process includes:
[0171] Traversing each object in the object storage system, obtaining a second check value of the currently traversed object from metadata of the currently traversed object during the traversal process, and querying the number of occurrences of the second check value of the currently traversed object from a second set, wherein the second set includes each second check value of all objects in the object storage system and the corresponding number of occurrences;
[0172] If the second check value of the currently traversed object appears more than once, the actual data of the currently traversed object is obtained, and the first check value of the actual data of the currently traversed object is obtained.
[0173] According to one or more embodiments of the present disclosure, before traversing each object in the object storage system, the method further includes:
[0174] Obtaining second check values of all objects from metadata of all objects in the object storage system, and determining the number of occurrences of each second check value among the second check values of all objects;
[0175] The second set is constructed according to each second check value of all objects and the corresponding number of occurrences.
[0176] According to one or more embodiments of the present disclosure, if the target granularity is a slice granularity, traversing each data of the target granularity in the object storage system and obtaining a first check value of the currently traversed data during the traversal process includes:
[0177] Traversing each object in the object storage system, acquiring actual data of the currently traversed object during the traversal process, and slicing the actual data of the currently traversed object to obtain slice data;
[0178] Obtain a first check value of the slice data of the currently traversed object.
[0179] According to one or more embodiments of the present disclosure, if it is determined that there is no historical first check value in the first set that is the same as the first check value of the currently traversed data, determining that the currently traversed data is not duplicate data, and using the first check value of the currently traversed data as the historical first check value, and storing it in the first set in association with the identifier of the currently traversed data, includes:
[0180] If it is determined that there is no historical first check value in the first set that is the same as the first check value of any slice data of the currently traversed object, it is determined that any slice data is not duplicate data, and the any slice data is rewritten into the object storage system as a new slice data, and the first check value of the new slice data is used as the historical first check value, and is associated with the identifier of the new slice data and stored in the first set.
[0181] According to one or more embodiments of the present disclosure, the step of using the historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data as the reference data, adding a soft link in the metadata of the currently traversed data to point to the reference data, and reclaiming the storage space of the currently traversed data includes:
[0182] The slice data corresponding to the first set of all the slice data included in the currently traversed object are used as reference data, soft links pointing to the respective reference data are added in the metadata of the currently traversed object, and the storage space occupied by the actual data of the currently traversed object is reclaimed.
[0183] According to one or more embodiments of the present disclosure, the actual data of the currently traversed object is segmented to obtain slice data, including:
[0184] Split the actual data of each object based on the content in an indefinite length segmentation method; or
[0185] The actual data of each object is segmented using a fixed-length segmentation method.
[0186] According to one or more embodiments of the present disclosure, determining a second benefit of performing data deduplication of an object storage system at a slice granularity includes:
[0187] Determine respectively a second benefit of performing data deduplication of the object storage system at a slice granularity when different slicing methods and / or different slicing lengths are adopted;
[0188] Correspondingly, after determining the target granularity of deduplication of the object storage system, the method further includes:
[0189] If the target granularity is determined to be the slice granularity, the segmentation method and / or segmentation length corresponding to the maximum second benefit is determined as the target segmentation method and / or target segmentation length.
[0190] According to one or more embodiments of the present disclosure, before determining the target granularity of deduplication of the object storage system, the method further includes:
[0191] Determine a target range for deduplication in the object storage system to perform deduplication within the target range;
[0192] Accordingly, determining the target granularity of deduplication of the object storage system includes:
[0193] A target granularity for deduplication within the target range is determined.
[0194] According to one or more embodiments of the present disclosure, determining a target range for deduplication in the object storage system includes:
[0195] Obtaining a second check value for each object in a plurality of target data buckets in the object storage system;
[0196] Determine a first repetition ratio of objects in each target data bucket and a second repetition ratio of objects in multiple target data buckets according to the second check value;
[0197] A target range for deduplication in the object storage system is determined according to the first duplication ratio and the second duplication ratio.
[0198] According to one or more embodiments of the present disclosure, determining a target range for deduplication in the object storage system according to the first duplication ratio and the second duplication ratio includes:
[0199] If the second repetition ratio exceeds a preset ratio, determining the target range to be the plurality of target data buckets; or
[0200] If the second repetition ratio does not exceed the preset ratio, and the first repetition ratio corresponding to any target data bucket exceeds the preset ratio, the target range is determined to be any target data bucket.
[0201] In a second aspect, according to one or more embodiments of the present disclosure, a data deduplication device for an object storage system is provided, comprising:
[0202] A policy setting unit, used to determine a target granularity of deduplication of the object storage system, wherein the target granularity is an object granularity or a slice granularity, wherein the object granularity is to use the actual data of the object as a whole as a data processing unit, and the slice granularity is to use the slice data corresponding to the actual data of the object as a data processing unit;
[0203] A screening unit, used for screening duplicate data of each data of target granularity in the object storage system;
[0204] The deleting unit is used to use one data in any set of duplicate data as the benchmark data, add a soft link in the metadata of other data in any set of duplicate data except the benchmark data to point to the benchmark data, and reclaim the storage space of the other data.
[0205] According to one or more embodiments of the present disclosure, when determining the target granularity of deduplication of the object storage system, the policy setting unit is used to:
[0206] Determining a first benefit of deduplicating data for the object storage system at an object granularity and a second benefit of deduplicating data for the object storage system at a slice granularity;
[0207] A target granularity of deduplication of the object storage system is determined according to the first benefit and the second benefit.
[0208] According to one or more embodiments of the present disclosure, when determining the target granularity of deduplication of the object storage system according to the first benefit and the second benefit, the policy setting unit is configured to:
[0209] If the difference between the first benefit and the second benefit is less than a preset threshold, determining the target granularity to be the object granularity; or
[0210] If the difference between the first income and the second income is not less than a preset threshold, a granularity corresponding to a maximum income between the first income and the second income is determined as the target granularity.
[0211] According to one or more embodiments of the present disclosure, when the screening unit screens duplicate data of each data of a target granularity in the object storage system, it is configured to:
[0212] A first check value of each data of a target granularity in the object storage system is obtained, and data with the same first check value is screened out from the object storage system to be determined as duplicate data.
[0213] According to one or more embodiments of the present disclosure, when the screening unit obtains the first check value of each data of the target granularity in the object storage system and screens out data with the same first check value from the object storage system to determine as duplicate data, it is used to:
[0214] Traversing each data of a target granularity in the object storage system, obtaining a first check value of the currently traversed data during the traversal process, and determining whether the first check value of the currently traversed data matches a historical first check value in a first set, wherein the first set stores each historical first check value obtained for the first time during the traversal process and an identifier of the corresponding data;
[0215] If it is determined that there is a historical first check value in the first set that is the same as the first check value of the currently traversed data, then the currently traversed data is determined to be duplicate data; or
[0216] If it is determined that there is no historical first verification value in the first set that is the same as the first verification value of the currently traversed data, it is determined that the currently traversed data is not duplicate data, and the first verification value of the currently traversed data is used as the historical first verification value, and is associated with the identifier of the currently traversed data and stored in the first set.
[0217] According to one or more embodiments of the present disclosure, when the deletion unit uses one data in any group of duplicate data as the reference data, adds a soft link in the metadata of other data in the any group of duplicate data except the reference data to point to the reference data, and reclaims the storage space of the other data, it is used to:
[0218] The historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data is used as the benchmark data, a soft link is added to the metadata of the currently traversed data to point to the benchmark data, and the storage space of the currently traversed data is reclaimed.
[0219] According to one or more embodiments of the present disclosure, if the target granularity is an object granularity, the screening unit, when traversing each data of the target granularity in the object storage system and obtaining a first check value of the currently traversed data during the traversal process, is used to:
[0220] Traversing each object in the object storage system, obtaining a second check value of the currently traversed object from metadata of the currently traversed object during the traversal process, and querying the number of occurrences of the second check value of the currently traversed object from a second set, wherein the second set includes each second check value of all objects in the object storage system and the corresponding number of occurrences;
[0221] If the second check value of the currently traversed object appears more than once, the actual data of the currently traversed object is obtained, and the first check value of the actual data of the currently traversed object is obtained.
[0222] According to one or more embodiments of the present disclosure, before traversing each object in the object storage system, the screening unit is further configured to:
[0223] Obtaining second check values of all objects from metadata of all objects in the object storage system, and determining the number of occurrences of each second check value among the second check values of all objects;
[0224] The second set is constructed according to each second check value of all objects and the corresponding number of occurrences.
[0225] According to one or more embodiments of the present disclosure, if the target granularity is a slice granularity, the screening unit, when traversing each data of the target granularity in the object storage system and obtaining a first check value of the currently traversed data during the traversal process, is used to:
[0226] Traversing each object in the object storage system, acquiring actual data of the currently traversed object during the traversal process, and slicing the actual data of the currently traversed object to obtain slice data;
[0227] Obtain a first check value of the slice data of the currently traversed object.
[0228] According to one or more embodiments of the present disclosure, if the screening unit determines that there is no historical first check value in the first set that is the same as the first check value of the currently traversed data, it determines that the currently traversed data is not duplicate data, and uses the first check value of the currently traversed data as the historical first check value, and associates it with the identifier of the currently traversed data and stores it in the first set, for:
[0229] If it is determined that there is no historical first check value in the first set that is the same as the first check value of any slice data of the currently traversed object, it is determined that any slice data is not duplicate data, and the any slice data is rewritten into the object storage system as a new slice data, and the first check value of the new slice data is used as the historical first check value, and is associated with the identifier of the new slice data and stored in the first set.
[0230] According to one or more embodiments of the present disclosure, when the deletion unit uses the historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data as the reference data, adds a soft link in the metadata of the currently traversed data to point to the reference data, and reclaims the storage space of the currently traversed data, it is used to:
[0231] The slice data corresponding to the first set of all the slice data included in the currently traversed object are used as reference data, soft links pointing to the respective reference data are added in the metadata of the currently traversed object, and the storage space occupied by the actual data of the currently traversed object is reclaimed.
[0232] According to one or more embodiments of the present disclosure, when the screening unit divides the actual data of the currently traversed object to obtain slice data, it is used to:
[0233] Split the actual data of each object based on the content in an indefinite length segmentation method; or
[0234] The actual data of each object is segmented using a fixed-length segmentation method.
[0235] According to one or more embodiments of the present disclosure, when determining the second benefit of deduplication of the object storage system at a slice granularity, the policy setting unit is used to:
[0236] Determine respectively a second benefit of performing data deduplication of the object storage system at a slice granularity when different slicing methods and / or different slicing lengths are adopted;
[0237] Correspondingly, after determining the target granularity of deduplication of the object storage system, the policy setting unit is further configured to:
[0238] If the target granularity is determined to be the slice granularity, the segmentation method and / or segmentation length corresponding to the maximum second benefit is determined as the target segmentation method and / or target segmentation length.
[0239] According to one or more embodiments of the present disclosure, before determining the target granularity of deduplication of the object storage system, the policy setting unit is further configured to:
[0240] Determine a target range for deduplication in the object storage system to perform deduplication within the target range;
[0241] Accordingly, when determining the target granularity of deduplication of the object storage system, the policy setting unit is used to:
[0242] A target granularity for deduplication within the target range is determined.
[0243] According to one or more embodiments of the present disclosure, when determining a target range for deduplication in the object storage system, the policy setting unit is configured to:
[0244] Obtaining a second check value for each object in a plurality of target data buckets in the object storage system;
[0245] Determine a first repetition ratio of objects in each target data bucket and a second repetition ratio of objects in multiple target data buckets according to the second check value;
[0246] A target range for deduplication in the object storage system is determined according to the first duplication ratio and the second duplication ratio.
[0247] According to one or more embodiments of the present disclosure, when the policy setting unit determines a target range for deduplication in the object storage system according to the first duplication ratio and the second duplication ratio, it is configured to:
[0248] If the second repetition ratio exceeds a preset ratio, determining the target range to be the plurality of target data buckets; or
[0249] If the second repetition ratio does not exceed the preset ratio, and the first repetition ratio corresponding to any target data bucket exceeds the preset ratio, the target range is determined to be any target data bucket.
[0250] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;
[0251] The memory stores computer-executable instructions;
[0252] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method for deduplication of the object storage system as described in the first aspect and various possible designs of the first aspect.
[0253] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer execution instructions. When a processor executes the computer execution instructions, the method for deduplication of the object storage system as described in the first aspect and various possible designs of the first aspect is implemented.
[0254] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising computer execution instructions. When a processor executes the computer execution instructions, the method for deduplication of the object storage system as described in the first aspect and various possible designs of the first aspect is implemented.
[0255] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0256] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0257] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A method for deduplicating data in an object storage system, characterized in that: include: Determine the target granularity of deduplication of the object storage system according to the execution benefits of the object granularity and the slice granularity, wherein the target granularity is the object granularity or the slice granularity, wherein the object granularity takes the actual data of the object as a whole as a data processing unit, and the slice granularity takes the slice data corresponding to the actual data of the object as a data processing unit; the execution benefits at least include storage cost benefits; Performing duplicate data screening on each data of target granularity in the object storage system; One data in any set of duplicate data is used as the reference data, and soft links are added to the metadata of other data in any set of duplicate data except the reference data to point to the reference data, and storage space of the other data is reclaimed.
2. The method according to claim 1, characterized in that Determining the target granularity of deduplication of the object storage system includes: Determining a first storage cost benefit of performing deduplication of the object storage system at an object granularity and a second storage cost benefit of performing deduplication of the object storage system at a slice granularity; A target granularity of data deduplication of the object storage system is determined according to the first storage cost benefit and the second storage cost benefit.
3. The method according to claim 2, characterized in that The step of determining a target granularity of deduplication of the object storage system according to the first storage cost benefit and the second storage cost benefit includes: If the difference between the first storage cost benefit and the second storage cost benefit is less than a preset threshold, determining the target granularity to be the object granularity with the highest performance benefit; or If the difference between the first storage cost benefit and the second storage cost benefit is not less than a preset threshold, the granularity corresponding to the maximum benefit between the first storage cost benefit and the second storage cost benefit is determined as the target granularity.
4. The method according to claim 1, characterized in that: The performing duplicate data screening on each data of target granularity in the object storage system includes: A first check value of each data of a target granularity in the object storage system is obtained, and data with the same first check value is screened out from the object storage system to be determined as duplicate data.
5. The method according to claim 4, characterized in that The obtaining of a first check value of each data of a target granularity in the object storage system, and screening out data having the same first check value from the object storage system to determine as duplicate data, includes: Traversing each data of a target granularity in the object storage system, obtaining a first check value of the currently traversed data during the traversal process, and determining whether the first check value of the currently traversed data matches a historical first check value in a first set, wherein the first set stores each historical first check value obtained for the first time during the traversal process and an identifier of the corresponding data; If it is determined that there is a historical first check value in the first set that is the same as the first check value of the currently traversed data, then the currently traversed data is determined to be duplicate data; or If it is determined that there is no historical first verification value in the first set that is the same as the first verification value of the currently traversed data, it is determined that the currently traversed data is not duplicate data, and the first verification value of the currently traversed data is used as the historical first verification value, and is associated with the identifier of the currently traversed data and stored in the first set.
6. The method according to claim 5, characterized in that The method of taking one data in any group of duplicate data as the reference data, adding a soft link in the metadata of other data in any group of duplicate data except the reference data to point to the reference data, and reclaiming the storage space of the other data includes: The historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data is used as the benchmark data, a soft link is added to the metadata of the currently traversed data to point to the benchmark data, and the storage space of the currently traversed data is reclaimed.
7. The method according to claim 6, characterized in that If the target granularity is an object granularity, traversing each data of the target granularity in the object storage system and obtaining a first check value of the currently traversed data during the traversal process includes: Traversing each object in the object storage system, obtaining a second check value of the currently traversed object from metadata of the currently traversed object during the traversal process, and querying the number of occurrences of the second check value of the currently traversed object from a second set, wherein the second set includes each second check value of all objects in the object storage system and the corresponding number of occurrences; If the second check value of the currently traversed object appears more than once, the actual data of the currently traversed object is obtained, and the first check value of the actual data of the currently traversed object is obtained.
8. The method according to claim 6, characterized in that If the target granularity is a slice granularity, traversing each data of the target granularity in the object storage system and obtaining a first check value of the currently traversed data during the traversal process includes: Traversing each object in the object storage system, acquiring actual data of the currently traversed object during the traversal process, and slicing the actual data of the currently traversed object to obtain slice data; Obtain a first check value of the slice data of the currently traversed object.
9. The method according to claim 8, characterized in that If it is determined that there is no historical first check value in the first set that is the same as the first check value of the currently traversed data, determining that the currently traversed data is not duplicate data, and taking the first check value of the currently traversed data as the historical first check value, and storing it in the first set in association with the identifier of the currently traversed data, includes: If it is determined that there is no historical first check value in the first set that is the same as the first check value of any slice data of the currently traversed object, it is determined that any slice data is not duplicate data, and the any slice data is rewritten into the object storage system as a new slice data, and the first check value of the new slice data is used as the historical first check value, and is associated with the identifier of the new slice data and stored in the first set.
10. The method according to claim 9, characterized in that The method of using the historical first check value corresponding data in the first set that is the same as the first check value of the currently traversed data as the reference data, adding a soft link in the metadata of the currently traversed data to point to the reference data, and reclaiming the storage space of the currently traversed data includes: The slice data corresponding to the first set of all the slice data included in the currently traversed object are used as reference data, soft links pointing to the respective reference data are added in the metadata of the currently traversed object, and the storage space occupied by the actual data of the currently traversed object is reclaimed.
11. The method according to claim 8, characterized in that The actual data of the currently traversed object is sliced to obtain slice data, including: Split the actual data of each object based on the content in an indefinite length segmentation method; or The actual data of each object is segmented using a fixed-length segmentation method.
12. The method according to claim 2, characterized in that: Identify secondary benefits of deduplication in object storage systems at the slice granularity, including: Determine respectively a second benefit of performing data deduplication of the object storage system at a slice granularity when different slicing methods and / or different slicing lengths are adopted; Correspondingly, after determining the target granularity of deduplication of the object storage system, the method further includes: If the target granularity is determined to be the slice granularity, the segmentation method and / or segmentation length corresponding to the maximum second benefit is determined as the target segmentation method and / or target segmentation length.
13. The method according to any one of claims 1 to 12, characterized in that: Before determining the target granularity of deduplication of the object storage system, the method further includes: Determine a target range for deduplication in the object storage system to perform deduplication within the target range; Accordingly, determining the target granularity of deduplication of the object storage system includes: A target granularity for deduplication within the target range is determined.
14. A data deduplication device for an object storage system, characterized in that: include: A policy setting unit is used to determine a target granularity of deduplication of the object storage system according to the execution benefits of the object granularity and the slice granularity, wherein the target granularity is the object granularity or the slice granularity, wherein the object granularity takes the actual data of the object as a whole as a data processing unit, and the slice granularity takes the slice data corresponding to the actual data of the object as a data processing unit; the execution benefits at least include storage cost benefits; A screening unit, used for screening duplicate data of each data of target granularity in the object storage system; The deleting unit is used to use one data in any set of duplicate data as the reference data, add a soft link in the metadata of other data in any set of duplicate data except the reference data to point to the reference data, and reclaim the storage space of the other data.
15. An electronic device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the method according to any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 13 is implemented.
17. A computer program product, characterized in that The method comprises computer-executable instructions, and when a processor executes the computer-executable instructions, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Method and device for optimizing big data of memory system
CN105511812A
Method, device and system for deduplication of repeated data of cloud storage system
CN106020722A