Processing method, evaluation method, device, electronic device and storage medium
By performing fine-grained slicing and counter vector map mapping of super-large data objects, the problem of long response time in cloud storage platforms is solved, efficient data deduplication operations are achieved, and the performance of the storage platform is improved.
Patent Information
- Application Number
- CN202110270071.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-12
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-03-12
AI Technical Summary
In the process of unique verification of super-large data objects in the cloud storage platform, there are problems such as long operation time and frequent database queries, resulting in too long system overhead and response time.
The data object is fine-grained, the identity identification code of each data shard is calculated, and hash mapped to the counter vector diagram. The duplicate data shard is determined by mapping vectors for deduplication operations, reducing the response time to the storage platform.
Through the mapping of fine-grained slicing and counter vector graph, the response time of the cloud storage platform is reduced, the access frequency of databases is reduced, and the performance of the storage platform is improved.
Smart Images

Figure CN115079930B_ABST
Abstract
Description
Background Art
[0002] With the development of technologies such as cloud computing, 5G, and AI, and the increasing storage demand of users for services such as high-definition audio and video, ultra-large data objects in units of GB or TB have gradually become the mainstream data objects. Whether it is a mobile device or a PC, the limited storage space cannot meet the growing storage demand. Therefore, storing ultra-large data objects in the cloud platform has become an effective solution.
[0003] In related technologies, in order to save the storage space of the cloud platform, before uploading a data object, the cloud platform will verify the uniqueness of the data object. Once a duplicate data object is matched, only the link pointing to the data object is retained, and redundant storage is no longer performed. For example, taking the MD5 value of the ultra-large data object (Message-Digest Algorithm 5, which is a widely used cryptographic hash function that can generate a 128-bit (16-byte) hash value for ensuring the integrity and consistency of information transmission) as the "fingerprint" information, and verifying the uniqueness of the data object through database query. This solution has the following defects:
[0004] On the one hand, since the operation time of MD5 is positively correlated with the size of the data object, when the data object is large, applications sensitive to response time can hardly tolerate the corresponding operation time. On the other hand, frequent database query operations will also restrict the performance of the cloud storage platform.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a method for deduplicating data objects, an evaluation method, a device, an electronic device, and a computer-readable storage medium, which can at least to some extent improve the system overhead caused by the methods in related technologies and the problem of too long data request time.
[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a method for deduplicating data objects is provided, including: performing fine-grained segmentation on a specified data object to obtain a plurality of data slices; calculating an identity identification code for each of the data slices; hashing and mapping the identity identification code to a counter vector diagram to obtain a mapping vector; determining duplicate data slices among the plurality of data slices based on the mapping vector, and performing a deduplication operation based on the duplicate data slices.
[0009] In one embodiment, the fine-grained segmentation of the specified data object to obtain multiple data shards includes: determining the seek time and write speed of the first storage medium for storing the data object; determining the segmentation granularity based on the seek time and the write speed; and performing fine-grained segmentation on the specified data object based on the segmentation granularity to obtain multiple data shards.
[0010] In one embodiment, the calculation of the identity identification code for each data shard includes: performing message digest calculation on each data shard based on multiple threads to obtain the corresponding identity identification code.
[0011] In one embodiment, the hashing and mapping of the identity identification code to the counter vector graph to obtain the mapping vector includes: mapping the identity identification code to corresponding multiple counting positions in the counter vectorizer based on multiple hash functions to obtain the mapping vector.
[0012] In one embodiment, the mapping of the identity identification code to corresponding multiple counting positions in the counter vectorizer based on multiple hash functions to obtain the mapping vector includes: performing hash calculation on the identity identification code based on multiple hash functions to obtain corresponding multiple hash values, and respectively mapping them to the corresponding counting positions in the counter vector graph to obtain the mapping vector; when detecting that at least one of the corresponding multiple hash values is 0, determining the corresponding identity identification code as a unique identity; when detecting that all of the corresponding multiple hash values are greater than 0, performing cumulative counting at the corresponding counting positions in the counter vector graph to determine the corresponding data shard as a duplicate data shard.
[0013] In one embodiment, determining the duplicate data shards among the multiple data shards based on the mapping vector and performing a deduplication operation based on the duplicate data shards includes: discarding the duplicate data shards when detecting the existence of duplicate data shards among the multiple data shards based on the mapping vector.
[0014] In one embodiment, it further includes: in response to a deletion instruction for the data shard, determining the corresponding counting position of the data shard to be deleted in the counter vector graph; and performing a decrement operation on the corresponding counting position to delete the data shard to be deleted.
[0015] In one embodiment, before the fine-grained segmentation of the specified data object to obtain multiple data shards, it further includes: presetting the target misjudgment rate of the deduplication process; and configuring the number of data shards, the number of counting positions in the counter vector graph, and the number of hash functions based on the target misjudgment rate.
[0016] In one embodiment, before hashing and mapping the identity identification code to a counter vector graph to obtain a mapping vector, it further includes: constructing the counter vector graph in a second storage medium based on a long integer vector structure.
[0017] According to another aspect of the present disclosure, there is provided a method for evaluating data object deduplication processing, including: performing fine-grained segmentation on a specified data object to obtain a plurality of data slices; calculating an identity identification code for each of the data slices; hashing and mapping the identity identification code to a counter vector graph to obtain a mapping vector; determining duplicate data slices among the plurality of data slices based on the mapping vector, and performing a deduplication operation based on the duplicate data slices; obtaining an evaluation index for the deduplication operation based on the number of the data slices and the number of counting positions in the counter vector graph.
[0018] In one embodiment, the hashing and mapping the identity identification code to a counter vector graph to obtain a mapping vector includes: mapping the identity identification code to corresponding multiple counting positions of the counter vectorizer based on a plurality of hash functions to obtain the mapping vector; the obtaining the evaluation index for the deduplication operation based on the number of the data slices and the number of counting positions in the counter vector graph includes: obtaining a misjudgment rate for the deduplication operation based on the number of the data slices, the number of counting positions in the counter vector graph, and the number of hash functions, and using the misjudgment rate as the evaluation index.
[0019] In one embodiment, the obtaining the misjudgment rate for the deduplication operation based on the number of the data slices, the number of counting positions in the counter vector graph, and the number of hash functions includes: determining a first probability that a counting position does not perform cumulative counting based on the number of the counting positions; determining a second probability that none of the plurality of hash functions performs cumulative counting on a counting position based on the first probability and the number of hash functions; determining a third probability that cumulative counting is performed after inserting the data slice into the counter vector graph based on the second probability and the number of the data slices; determining the misjudgment rate based on the third probability and the number of the counting positions.
[0020] According to still another aspect of the present disclosure, there is provided a data object deduplication processing apparatus, including: a segmentation module for performing fine-grained segmentation on a specified data object to obtain a plurality of data slices; a calculation module for calculating an identity identification code for each of the data slices; a mapping module for hashing and mapping the identity identification code to a counter vector graph to obtain a mapping vector; a deduplication module for determining duplicate data slices among the plurality of data slices based on the mapping vector, and performing a deduplication operation based on the duplicate data slices.
[0021] According to another aspect of the present disclosure, there is provided an evaluation device for data object deduplication processing, including:
[0022] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the data object deduplication processing method and / or the evaluation method for data object deduplication processing of any one of the above by executing the executable instructions.
[0023] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data object deduplication processing method and / or the evaluation method for data object deduplication processing of any one of the above.
[0024] The data object deduplication processing solution provided by the embodiments of the present disclosure performs fine-grained segmentation on a large data object to obtain multiple data shards, calculates the identity identification code of each data shard, and hashes and maps the identity identification code to a counter vector graph to complete the verification of the uniqueness of the data shard. This verification method can reduce the response time of the cloud storage platform when storing a specified data object.
[0025] Further, after obtaining the identity identification code, based on the hash mapping operation, the identity identification code is hashed and mapped into a preset counter vector graph, and deduplication operation is performed based on the mapping result to achieve fine-grained deduplication, which can reduce the access frequency to the database in the storage platform, thereby being beneficial to improving the performance of the storage platform.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0028] Figure 1 A schematic diagram showing the structure of a data object deduplication processing system in an embodiment of the present disclosure;
[0029] Figure 2 A flowchart showing a data object deduplication processing method in an embodiment of the present disclosure;
[0030] Figure 3 A flowchart showing another data object deduplication processing method in an embodiment of the present disclosure;
[0031] Figure 4 Shows a framework schematic diagram of a data object deduplication processing solution in an embodiment of the present disclosure;
[0032] Figure 5 Shows a flowchart of yet another data object deduplication processing method in an embodiment of the present disclosure;
[0033] Figure 6 Shows a flowchart of an evaluation method for data object deduplication processing in an embodiment of the present disclosure;
[0034] Figure 7 Shows a flowchart of another evaluation method for data object deduplication processing in an embodiment of the present disclosure;
[0035] Figure 8 Shows a schematic diagram of a data object deduplication processing device in an embodiment of the present disclosure;
[0036] Figure 9 Shows a schematic diagram of an evaluation device for data object deduplication processing in an embodiment of the present disclosure;
[0037] Figure 10 Shows a schematic diagram of an electronic device in an embodiment of the present disclosure. Detailed implementation manners
[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0039] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0040] The solution provided by this application performs fine-grained segmentation on a large data object to obtain multiple data shards, calculates the identity identification code of each data shard, and hashes and maps the identity identification code to a counter vector diagram to complete the verification of the uniqueness of the data shards. This verification method can reduce the response time of the cloud storage platform when storing a specified data object.
[0041] For ease of understanding, several terms related to the present application are first explained below.
[0042] Hash function: Hash, generally translated as hash or hashing, or transliterated as hash, is a function that transforms an input of any length (also called pre-image) into an output of a fixed length through a hashing algorithm. This output is the hash value. This transformation is a compression mapping, that is, the space of hash values is usually much smaller than the space of the input. Different inputs may hash to the same output, so it is impossible to determine the unique input value from the hash value. Simply put, it is a function that compresses a message of any length into a message digest of a certain fixed length.
[0043] Coarse-grained: It represents the category level, that is, only the type of object is considered, and a specific instance of the object is not considered. For example, in user management, creation and deletion are treated equally for all users, without distinguishing the specific object instances of the operations.
[0044] Fine-grained: It represents the instance level, that is, the instance of the specific object needs to be considered. Of course, the fine-grained level is considered after considering the object category at the coarse-grained level. For example, in contract management, listing and deletion need to distinguish whether the contract instance was created by the current user.
[0045] The average seek time refers to the average time required for the magnetic head to move from the start to the track where the data is located after the disk receives a system instruction. It is the time required for the computer to issue an addressing command until the corresponding target data is found, and the unit is milliseconds (ms).
[0046] The solution provided in the embodiments of the present application involves technologies such as face recognition and machine learning, and is specifically described through the following embodiments.
[0047] Figure 1 The structural schematic diagram of a data object deduplication processing system in an embodiment of the present disclosure is shown, including a plurality of terminals 120 and a server cluster 140.
[0048] The terminal 120 can be a mobile terminal such as a mobile phone, a game console, a tablet computer, an e - book reader, smart glasses, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart home device, an AR (Augmented Reality) device, a VR (Virtual Reality) device, etc. Or, the terminal 120 can also be a personal computer (PC), such as a laptop computer and a desktop computer, etc.
[0049] Among them, an application program for providing data object deduplication processing can be installed in the terminal 120.
[0050] The terminal 120 is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0051] The server cluster 140 is a single server, or consists of several servers, or is a virtualization platform, or is a cloud computing service center. The server cluster 140 is used to provide background services for the data object deduplication processing application program. Optionally, the server cluster 140 undertakes the main computing work and the terminal 120 undertakes the secondary computing work; or, the server cluster 140 undertakes the secondary computing work and the terminal 120 undertakes the main computing work; or, the terminal 120 and the server cluster 140 adopt a distributed computing architecture for collaborative computing.
[0052] In some alternative embodiments, the server cluster 140 is used to store data object deduplication processing models, etc.
[0053] Optionally, the clients of the application programs installed in different terminals 120 are the same, or the clients of the application programs installed on two terminals 120 are clients of the same type of application program on different control system platforms. Based on the differences in the terminal platforms, the specific forms of the clients of the application program can also be different. For example, the client of the application program can be a mobile phone client, a PC client, or a World Wide Web (Web) client, etc.
[0054] Those skilled in the art can know that the number of the above - mentioned terminals 120 can be more or less. For example, there can be only one of the above - mentioned terminals, or dozens or hundreds of the above - mentioned terminals, or even more. The embodiments of the present application do not limit the number and device types of the terminals.
[0055] Optionally, the system can also include a management device ( Figure 1(not shown), the management device is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0056] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or proprietary data communication technologies can also be used to replace or supplement the above data communication technologies.
[0057] Next, each step in the data object deduplication processing method in the present exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.
[0058] Figure 2 The flowchart of a data object deduplication processing method in an embodiment of the present disclosure is shown. The method provided by the embodiment of the present disclosure can be executed by any electronic device with computing and processing capabilities, such as Figure 1 the terminal 120 and / or the server cluster 140 in. In the following illustrative examples, the terminal 120 is used as the execution subject for illustrative purposes.
[0059] As Figure 2 shown, the terminal 120 executes the data object deduplication processing method, including the following steps:
[0060] Step S202, perform fine-grained segmentation on the specified data object to obtain a plurality of data shards.
[0061] Among them, the specified data object can be understood as an ultra-large data object, that is, a data object with a size larger than a preset data threshold. The data threshold can be set according to the processing speed of the electronic device. For example, the data threshold can be 1GB (gigabyte) or 2GB, etc.
[0062] In addition, fine-grained segmentation, which is a segmentation method relative to coarse-grained segmentation, can be understood as a byte-level data segmentation method. By performing fine-grained segmentation on the ultra-large specified data object, multiple smaller data shards can be obtained. For example, a 2.4GB specified data object is segmented into multiple 64MB data shards.
[0063] Step S204, calculate the identity identification code of each data shard.
[0064] Specifically, a specific implementation manner of step S204 for calculating the identity identification code of each data shard includes:
[0065] Perform message digest calculation on each data shard based on multiple threads to obtain the corresponding identity identification code.
[0066] Among them, by using the distributed multi-thread technology, the identity identification codes of the data shards are calculated in parallel to reduce the time overhead generated by the operation of the identity identification codes.
[0067] In addition, the identity identification code can be the result obtained by performing message digest calculation on the data shard using MD5. Specifically, the data shard is passed through an irreversible string transformation algorithm to generate a unique MD5 information digest as the identity identification code.
[0068] For example, the time required to calculate the MD5 value of a 2.4GB specified data object is 19s. After the 2.4GB specified data object is segmented into 39 64MB data shards, the MD5 values of each data shard are calculated in parallel based on multiple threads, and the operation time is approximately 19s / 39 = 0.49s. Compared with 19s, the time overhead of the MD5 value operation is greatly reduced.
[0069] Step S206, hash-map the identity identification code to the counter vector diagram to obtain a mapping vector.
[0070] Among them, the counter vector diagram is a very long integer data vector structure. Different positions in the vector structure are used to represent the storage position of the data shard and to count the data shard.
[0071] Through hash mapping, a definite corresponding relationship is established between the storage locations of data shards and the identification codes. The storage locations of data shards correspond to different counting positions in the counter vector diagram. Through the corresponding relationship, the identification codes of data shards are hash-mapped to determine the mapping positions. Based on the mapping positions, the counters at the corresponding positions in the counter vector diagram are counted, and the corresponding mapping vectors are obtained to complete the uniqueness verification of data shards based on the mapping vectors.
[0072] Step S208: Determine the duplicate data shards among multiple data shards based on the mapping vectors, and perform deduplication operations based on the duplicate data shards.
[0073] Among them, if there are multiple counts at the counting positions of the mapping vectors, it indicates that there are duplicate data shards. Performing deduplication operations on the duplicate data shards can reduce the disk I / O frequency caused by data queries.
[0074] Specifically, for duplicate data shards, only the links pointing to their storage locations are retained, and redundant storage is no longer performed, so as to greatly save storage space.
[0075] In this embodiment, by finely dividing a large data object, multiple data shards are obtained, and the identification codes of each data shard are calculated, and the identification codes are hash-mapped onto the counter vector diagram to complete the verification of the uniqueness of the executed data shards. This verification method can reduce the response time of the cloud storage platform when storing a specified data object.
[0076] Furthermore, after obtaining the identification codes, based on the hash mapping operation, the identification codes are hash-mapped into a preset counter vector diagram, and deduplication operations are performed based on the mapping results to achieve fine-grained deduplication, which can reduce the access frequency to the database in the storage platform, thereby being beneficial to improving the performance of the storage platform.
[0077] As Figure 3 shown, in one embodiment, step S202: A specific implementation manner of finely dividing a specified data object to obtain multiple data shards includes:
[0078] Step S302: Determine the seek time and write speed of the first storage medium for storing the data object.
[0079] Among them, the first storage medium can specifically be the storage medium used by the cloud storage platform, specifically such as SATA hard disks, SAS hard disks, solid state disks, and mechanical disks. Taking a mechanical disk as an example, the average seek time of mechanical disk data is about 10 ms, and the data write speed of the mechanical disk interface is about 50 MB / s.
[0080] Step S304: Determine the segmentation granularity based on the seek time and the writing speed.
[0081] The optimal data transfer time corresponding to the seek time is 10 ms / 0.01 = 1 s, and the corresponding optimal data block size is (50 MB / s) / (1 s) = 50 MB. Therefore, the specified data object is segmented at a granularity of 2 6 = 64 MB per slice.
[0082] Step S306: Perform fine-grained segmentation on the specified data object based on the segmentation granularity to obtain multiple data shards.
[0083] In this embodiment, since the performance of the first storage medium determines its working efficiency, in order to ensure that the first storage medium is in the best operating state, the optimal data transfer time corresponding to the seek time of the first storage medium is determined. Based on the optimal data transfer time and the data writing speed of the first storage medium, the size of the corresponding optimal segmentation granularity can be determined. By segmenting the data object based on the optimal segmentation granularity, the reliability of the segmentation operation and the access efficiency of the first storage medium can be ensured.
[0084] In one embodiment, the identity identification code is hashed and mapped to a counter vector diagram, and the obtained mapping vector includes: mapping the identity identification code to a corresponding plurality of counting positions in the counter vector through a plurality of hash functions to obtain a mapping vector.
[0085] Specifically, the elements in the counter vector diagram are determined by the hash function. Taking the identity identification code as the independent variable, the value calculated through the hash function is the storage address of the corresponding data block. The storage address corresponds to the counting position in the counter vector diagram, and the mapping is realized based on this corresponding relationship to obtain a mapping vector.
[0086] Among them, the construction methods of the hash function include but are not limited to the following types:
[0087] (1) Division-remainder method: Divide the key x by M (usually taking the length of the hash table), and take the remainder as the hash address. The corresponding hash function is: h(x) = x mod M.
[0088] (2) Multiply-remainder and truncation method: First, multiply the key key by a constant A (0 < A < 1), and extract the fractional part of the product. Then, multiply this value by the integer n and truncate the result downward as the hash address. The hash function is: hash(key) = _LOW(n × (A × key % 1)). Among them, "A × key % 1" represents taking the fractional part of A × key, that is, A × key % 1 = A × key - _LOW(A × key), and _LOW(X) represents truncating X downward.
[0089] (3) Mid - square method: Since integer division usually runs slower than multiplication, consciously avoiding the use of the division - remainder operation can improve the running time of the hashing algorithm. The specific implementation of the mid - square method is as follows: First, find the square value of the key code to expand the difference between similar numbers, and then take several middle digits (usually take binary bits) according to the table length as the hash function value. Because several middle digits of a product are related to each digit of the multiplier, the resulting hash addresses are relatively uniform.
[0090] (4) Digital analysis method: Set n d - digit numbers, and each digit may have r different symbols. The frequencies of these r different symbols appearing in each digit are not necessarily the same. They may be evenly distributed in some digits, with an equal probability of each symbol appearing; in some digits, they are not evenly distributed, and only certain symbols often appear. According to the size of the hash table, several digits with uniform distribution of various symbols can be selected as the hash address.
[0091] (5) Radix conversion method: Regard the key code value as a number in another number system and then convert it back to the original number system, and then select several digits as the hash address.
[0092] (6) Folding method: Sometimes the key code contains a large number of digits, and it is too complex to calculate using the mid - square method. Then the key code can be divided into several parts with the same number of digits (the number of digits of the last part can be different), and then the sum of these parts (discarding the carry) is taken as the hash address. This method is called the folding method.
[0093] (7) ELFhash string hashing function: The ELFhash function is used in the "Executable and Linking Format" (ELF) in UNIX System V Release 4. The ELF file format is used to store executable files and object files. The ELFhash function is a hash of strings. It is effective for both long and short strings. Each character in the string has the same effect. It cleverly calculates the ASCII encoding value of the characters, and the ELFhash function can evenly distribute the strings in the hash table.
[0094] In this embodiment, by using multiple hash functions to perform hash mapping on the same data block, the probability of misjudging duplicate data shards due to collisions between different hash functions can be reduced, which is beneficial to improving the accuracy of the deduplication process.
[0095] In addition, there is no coupling relationship between hash functions, which is convenient for parallel implementation by the hardware system. The deduplication operation depends on the "trace" of the data shard, rather than the data shard itself stored, and is suitable for application scenarios with strict confidentiality requirements.
[0096] In one embodiment, mapping the identity identification code to the corresponding multiple counting positions of the counter vectorizer based on multiple hash functions to obtain a mapping vector includes: performing a hash calculation on the identity identification code based on the multiple hash functions to obtain the corresponding multiple hash values, and respectively mapping them to the corresponding counting positions of the counter vector graph to obtain a mapping vector; when detecting that at least one of the corresponding multiple hash values is 0, determining that the corresponding identity identification code is a unique identity; when detecting that all of the corresponding multiple hash values are greater than 0, performing an accumulative count on the corresponding counting positions of the counter vector graph to determine that the corresponding data shard is a duplicate data shard.
[0097] Among them, the mapping process of the identity identification code of the data shard can be understood as the query process of the "fingerprint" of the data shard, and the query process can specifically be understood as the calculation process of the hash value of the identity identification code. If the calculated hash value is 0, it indicates that the data shard has a unique fingerprint. If the calculated hash value is 1, it indicates that the data object stores duplicate data shards.
[0098] As Figure 4 shown, assume that each data shard uses 3 hash functions for hash mapping. After the specified data object 402 is finely sliced, n data shards 404 are obtained, namely data shard 1, data shard 2, data shard 3, and data shard n, etc. Calculate them respectively to obtain the corresponding F-M5 values 406. The identity identification code of data shard 1 is F-MD51, the identity identification code of data shard 2 is F-MD52, the identity identification code of data shard 3 is F-MD53, and the identity identification code of data shard n is F-MD5n.
[0099] The counter vector graph 408 includes m counting positions, namely Count1, Count2, Count3, Count4, Count5, Count6, Counti, and Countm, etc.
[0100] As Figure 4 shown, taking the identity identification code F-MD51 of the data shard as an example, after 3 hash operations, it is finally mapped to the Count1, Count4, and Counti counting positions of the counter vector graph. When performing the "fingerprint" query operation of the data shard, if one Count value is equal to 0, the queried object definitely does not exist; if all three Count values are greater than 0, the queried object can be regarded as existing.
[0101] Specifically, when using a single hash function for hash mapping, since there is a possibility of obtaining the same identity identification code from two different data shards, different identity identification codes are actually mapped to the same storage location. This phenomenon is called a collision. To reduce the probability of collisions, multiple different hash functions are used to implement hash mapping.
[0102] Then, multiple hash function mappings are performed on the counter vector diagram in the memory. Finally, high-performance fine-grained deduplication of ultra-large data objects in the cloud storage platform is achieved.
[0103] Hash mapping refers to using a hash function to map the identity identification code to different positions on the counter vector diagram.
[0104] In one embodiment, determining duplicate data shards among multiple data shards based on the mapping vector and performing a deduplication operation based on the duplicate data shards includes:
[0105] When duplicate data shards are detected among multiple data shards based on the mapping vector, the duplicate data shards are discarded.
[0106] In this embodiment, when it is determined based on the mapping vector that duplicate data shards are generated, the duplicate data shards are discarded to prevent them from being uploaded to the first storage medium in the background again, ensuring the uniqueness of the stored data shards. This method can only retain the link pointing to the storage location without redundant storage, thus greatly saving storage space.
[0107] In one embodiment, it further includes: in response to a deletion instruction for a data shard, determining the corresponding counting position of the data shard to be deleted in the counter vector diagram; performing a decrement operation on the corresponding counting position to delete the data shard to be deleted.
[0108] Specifically, by performing a decrement operation on the corresponding counting position in the counter vector diagram, the purpose is to delete the stored data. When a data deletion operation needs to be performed, by performing a decrement operation on the corresponding position of the counter vector diagram, the unique stored data shard is deleted to release the storage space.
[0109] In this embodiment, by setting the counter vector diagram, when duplicate data shards need to be deleted, a decrement operation needs to be performed on the corresponding counting position in the counter vector diagram, without any change to the data structure space of the counter vector diagram itself, which has good robustness.
[0110] In addition, when a data shard expires or is actively deleted by the user, only a decrement operation needs to be performed on the corresponding counter in the counter vector diagram. Another advantage of the counter in the counter vector diagram is that it supports the deletion operation of data shards in the cloud storage platform to further release precious storage space.
[0111] As Figure 5 shown, in one embodiment, before step S202 of performing fine-grained segmentation on a specified data object to obtain multiple data shards, it further includes:
[0112] Step S502, preset the target misjudgment rate for duplicate removal processing.
[0113] Step S504, configure the number of data shards, the number of counting positions in the counter vector diagram, and the number of hash functions based on the target misjudgment rate.
[0114] Specifically, the calculation formula for the misjudgment rate p is expressed as:
[0115]
[0116] where m is the number of counting positions in the vector diagram, k is the number of hash functions, and n is the number of data shards.
[0117] In this embodiment, by presetting the target misjudgment rate, based on the known misjudgment rate, the number of data shards, the number of counting positions in the counter vector diagram, and the number of hash functions are reasonably set to ensure that the duplicate removal processing has high reliability.
[0118] In one embodiment, before hashing and mapping the identity identification code to the counter vector diagram to obtain a mapping vector, it further includes: constructing a counter vector diagram in the second storage medium based on the long integer vector structure.
[0119] In this embodiment, the second storage medium can be understood as a storage medium with better performance. For example, the second storage medium can be the memory of a cloud storage platform to complete the "fingerprint" query work based on the memory. By setting a long integer counter vector diagram based on the memory, multiple hash function mappings are performed on the vector diagram to achieve fast duplicate removal of duplicate data shards.
[0120] As Figure 6 shown, the terminal 120 executes an evaluation method for data object duplicate removal processing, including the following steps:
[0121] Step S602, perform fine-grained segmentation on a specified data object to obtain multiple data shards.
[0122] Step S604, calculate the identity identification code of each data shard.
[0123] Step S606, hash and map the identity identification code to the counter vector diagram to obtain a mapping vector.
[0124] Step S608, determine the duplicate data shards among the multiple data shards based on the mapping vector, and perform duplicate removal operations based on the duplicate data shards.
[0125] Among them, for the specific descriptions of steps S602 to S608, reference can be made to steps S202 to S208.
[0126] Step S610: Obtain an evaluation index for the deduplication operation based on the number of data shards and the number of counting positions in the counter vector diagram.
[0127] In this embodiment, by pre-determining the number of data shards and the counter vector diagram structure, fine-grained segmentation of the specified data object is performed based on the number of data shards to obtain multiple data shards, and hash mapping of the data shards is completed based on the counter vector diagram, thereby completing the fine-grained deduplication of the data object. An evaluation index for the deduplication operation is obtained based on the number of data shards and the number of counting positions in the counter vector diagram, realizing the evaluation of the feasibility and practicality of the deduplication process of this solution.
[0128] In one embodiment, hashing the identity identification code to the counter vector diagram to obtain a mapping vector includes: mapping the identity identification code to a corresponding plurality of counting positions in the counter vector through a plurality of hash functions to obtain a mapping vector; obtaining an evaluation index for the deduplication operation based on the number of data shards and the number of counting positions in the counter vector diagram includes: obtaining the false positive rate of the deduplication operation based on the number of data shards, the number of counting positions in the counter vector diagram, and the number of hash functions, and using the false positive rate as the evaluation index.
[0129] In one embodiment, obtaining the false positive rate of the deduplication operation based on the number of data shards, the number of counting positions in the counter vector diagram, and the number of hash functions includes:
[0130] Step S702: Determine a first probability that a counting position has not performed cumulative counting based on the number of counting positions.
[0131] Step S704: Determine a second probability that none of the multiple hash functions have performed cumulative counting on a counting position based on the first probability and the number of hash functions.
[0132] Step S706: Determine a third probability that cumulative counting is performed after inserting a data shard into the counter vector diagram based on the second probability and the number of data shards.
[0133] Step S708: Determine the false positive rate based on the third probability and the number of counting positions.
[0134] Assume that the total number of data shards is n, the number of counting positions in the vector diagram is m, the number of hash functions is k, and the false positive rate is p. Then for a specific counter, when a specific data shard is inserted by a hash function, the probability of not performing an increment operation on it is The probability that none of the k hash functions increments it is For the insertion of n data shards, the probability that none of the hash functions increments it is Therefore, the probability that the counter in the counter vector diagram increments this data shard is According to the analysis in the above part, misjudgment occurs only when the values of all k counters for querying the data shard are greater than 0, and the corresponding misjudgment rate p is:
[0135]
[0136] According to L'Hopital's rule There is:
[0137]
[0138] Since the number of counters m in the counter vector diagram is very large, that is, there is According to the above formula, formula (1) can be equivalently expressed as:
[0139]
[0140] Analyzing formula (3), it can be found that when the total number of data shards n decreases, or the number of counting positions m in the vector diagram increases, the misjudgment rate will decrease, which conforms to our intuitive feeling. Further, we regard the misjudgment rate p as a function f(k) of the number of hash functions k, and let That is, there is:
[0141] f(k) = (1 - b -k ) k (4)
[0142] Taking the logarithm of both sides of formula (4) and differentiating simultaneously, we get:
[0143]
[0144] Let f′(k) = 0 in formula (5), then there is:
[0145] (1 - b -k )·ln(1 - b -k ) + k·b -k ·lnb = 0 (6)
[0146] Solving formula (6) gives 1 - b -k = b -k That is,
[0147] Therefore, we can obtain the following equivalent relationship regarding the total number of data shards n, the number of counting positions m in the vector diagram, and the number of hash functions k:
[0148]
[0149] Substituting the above results into formula (3), we can obtain:
[0150]
[0151] To more intuitively explain the inventive method, taking the misjudgment rate p as 1%, it can be calculated according to formula (8) that That is, the number of counting positions required to store one data shard is 10. Since the counter is of integer data type, the memory space required to store one data shard is 40B; the number of hash functions required Calculated with a data volume of 100 million, the required memory space is only 3.7GB, which is completely acceptable for the cloud storage platform, proving the practicability and feasibility of this solution.
[0152] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0153] Those skilled in the art to which the present invention pertains can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "system".
[0154] Next, with reference to Figure 8 to describe the data object deduplication processing apparatus 800 according to this embodiment of the present invention. Figure 8 The data object deduplication processing apparatus 800 shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0155] The data object deduplication processing apparatus 800 is presented in the form of a hardware module. The components of the data object deduplication processing apparatus 800 may include but are not limited to: a splitting module 802 for performing fine-grained splitting on a specified data object to obtain a plurality of data shards; a calculating module 804 for calculating the identity identification code of each of the data shards; a mapping module 806 for hashing and mapping the identity identification code to a counter vector diagram to obtain a mapping vector; and a deduplication module 808 for determining duplicate data shards among the plurality of data shards based on the mapping vector and performing a deduplication operation based on the duplicate data shards.
[0156] The following will refer to Figure 9 to describe the data object deduplication processing device 900 according to this embodiment of the present invention. Figure 9 The data object deduplication processing device 900 shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0157] The data object deduplication processing device 900 is embodied in the form of a hardware module. The components of the data object deduplication processing device 900 may include but are not limited to: a segmentation module 902 for performing fine-grained segmentation on a specified data object to obtain a plurality of data shards; a calculation module 904 for calculating the identity identification code of each of the data shards; a mapping module 906 for hashing and mapping the identity identification code to a counter vector diagram to obtain a mapping vector; a deduplication module 908 for determining duplicate data shards among the plurality of data shards based on the mapping vector and performing a deduplication operation based on the duplicate data shards; and an evaluation module 910 for obtaining an evaluation index of the deduplication operation based on the number of the data shards and the number of counting positions in the counter vector diagram.
[0158] The following will refer to Figure 10 to describe the electronic device 1000 according to this embodiment of the present invention. Figure 10 The shown electronic device 1000 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0159] As Figure 10 shown, the electronic device 1000 is embodied in the form of a general-purpose computing device. The components of the electronic device 1000 may include but are not limited to: at least one of the above-mentioned processing units 1010, at least one of the above-mentioned storage units 1020, and a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010).
[0160] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit 1010 can execute steps S202 to S208 as Figure 2 shown, as well as other steps defined in the data object deduplication processing method of the present disclosure.
[0161] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202, and may further include a read-only storage unit (ROM) 10203.
[0162] The storage unit 1020 may also include a program / utility 10204 having a set (at least one) of program modules 10205. Such program modules 10205 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0163] The bus 1030 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0164] The electronic device 1000 may also communicate with one or more external devices 1060 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device, and / or may communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 1050. Moreover, the electronic device 1000 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1050. As shown in the figure, the network adapter 1050 communicates with other modules of the electronic device 1000 through the bus 1030. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0165] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0166] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above methods of this specification is stored. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0167] The program product for implementing the above method according to an embodiment of the present invention can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0168] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0169] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0170] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0171] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0172] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0173] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present disclosure.
[0174] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A method for deduplicating data objects, characterized in that, Including: Performing fine-grained segmentation on a specified data object to obtain multiple data shards; Calculating an identity identification code for each of the data shards; Hash-mapping the identity identification code to a counter vector diagram to obtain a mapping vector, including: performing hash calculation on the identity identification code based on multiple hash functions to obtain corresponding multiple hash values, and respectively mapping them to corresponding counting positions in the counter vector diagram to obtain the mapping vector; Determining duplicate data shards among the multiple data shards based on the mapping vector, and performing a deduplication operation based on the duplicate data shards. Among them, when detecting that at least one of the corresponding multiple hash values is 0, determining the corresponding identity identification code as a unique identity; when detecting that all of the corresponding multiple hash values are greater than 0, performing an accumulative count at the corresponding counting position in the counter vector diagram to determine the corresponding data shard as the duplicate data shard.
2. The data object deduplication method according to claim 1, wherein The performing fine-grained segmentation on a specified data object to obtain multiple data shards includes: Determining the seek time and write speed of a first storage medium for storing the data object; Determining a segmentation granularity based on the seek time and the write speed; Performing fine-grained segmentation on the specified data object based on the segmentation granularity to obtain multiple data shards.
3. The data object deduplication method according to claim 1, characterized in that, The calculating an identity identification code for each of the data shards includes: Performing message digest calculation on each of the data shards based on multiple threads to obtain the corresponding identity identification code.
4. The data object deduplication method according to claim 1, wherein The determining duplicate data shards among the multiple data shards based on the mapping vector, and performing a deduplication operation based on the duplicate data shards includes: When detecting duplicate data shards among the multiple data shards based on the mapping vector, discarding the duplicate data shards.
5. The data object deduplication method according to claim 1, wherein Also including: In response to a deletion instruction for the data shard, determining a corresponding counting position of the data shard to be deleted in the counter vector diagram; Performing a decrement operation on the corresponding counting position to delete the data shard to be deleted.
6. The data object deduplication method according to claim 1, characterized in that Before performing fine-grained segmentation on a specified data object to obtain multiple data shards, it further includes: Presetting a target misjudgment rate for the deduplication process; Configuring the number of data shards, the number of counting positions in the counter vector diagram, and the number of hash functions based on the target misjudgment rate.
7. The data object deduplication processing method according to any one of claims 1 to 6, characterized in that, Before hash-mapping the identity identification code to a counter vector diagram to obtain a mapping vector, it further includes: Constructing the counter vector diagram in a second storage medium based on a long integer vector structure.
8. An evaluation method for data object deduplication processing, characterized in that Including: Performing fine-grained segmentation on a specified data object to obtain multiple data shards; Calculating an identity identification code for each of the data shards; Hash-mapping the identity identification code to a counter vector diagram to obtain a mapping vector, including: performing hash calculation on the identity identification code based on a hash function to obtain corresponding multiple hash values, and respectively mapping them to corresponding counting positions in the counter vector diagram to obtain the mapping vector; Determine duplicate data shards among the multiple data shards based on the mapping vector, and perform a deduplication operation based on the duplicate data shards. Among them, when at least one of the corresponding multiple hash values is detected to be 0, determine the corresponding identity code as the unique identity; when all of the corresponding multiple hash values are detected to be greater than 0, perform an accumulation count at the corresponding count position in the counter vector diagram to determine the corresponding data shard as the duplicate data shard; Obtain an evaluation index of the deduplication operation based on the number of data shards and the number of count positions in the counter vector diagram.
9. The evaluation method for duplicate data object removal according to claim 8, characterized in that, The step of hashing and mapping the identity code to the counter vector diagram to obtain a mapping vector includes: Map the identity code to the corresponding multiple count positions of the counter vector through multiple hash functions to obtain the mapping vector; The step of obtaining an evaluation index of the deduplication operation based on the number of data shards and the number of count positions in the counter vector diagram includes: Obtain the misjudgment rate of the deduplication operation based on the number of data shards, the number of count positions in the counter vector diagram, and the number of hash functions, and use the misjudgment rate as the evaluation index.
10. The evaluation method for data object duplicate removal processing according to claim 9, characterized in that, The step of obtaining the misjudgment rate of the deduplication operation based on the number of data shards, the number of count positions in the counter vector diagram, and the number of hash functions includes: Determine a first probability that no accumulation count is performed at one of the count positions based on the number of count positions; Determine a second probability that none of the multiple hash functions perform an accumulation count at one of the count positions based on the first probability and the number of hash functions; Determine a third probability that an accumulation count is performed after inserting the data shard into the counter vector diagram based on the second probability and the number of data shards; Determine the misjudgment rate based on the third probability and the number of count positions.
11. An evaluation device for deduplication processing of data objects, characterized in that, It includes: A slicing module for performing fine-grained slicing on a specified data object to obtain multiple data shards; A calculation module for calculating the identity code of each data shard; A mapping module for hashing and mapping the identity code to the counter vector diagram to obtain a mapping vector, including: performing a hash calculation on the identity code through multiple hash functions to obtain corresponding multiple hash values, and respectively mapping them to the corresponding count positions of the counter vector diagram to obtain the mapping vector; A deduplication module for determining duplicate data shards among the multiple data shards based on the mapping vector, and performing a deduplication operation based on the duplicate data shards. Among them, when at least one of the corresponding multiple hash values is detected to be 0, determine the corresponding identity code as the unique identity; when all of the corresponding multiple hash values are detected to be greater than 0, perform an accumulation count at the corresponding count position in the counter vector diagram to determine the corresponding data shard as the duplicate data shard.
12. An evaluation device for deduplication processing of data objects, characterized in that, It includes: A slicing module for performing fine-grained slicing on a specified data object to obtain multiple data shards; A calculation module for calculating the identity code of each data shard; A mapping module, configured to hash-map the identity identification code to a counter vector diagram to obtain a mapping vector, including: performing hash calculation on the identity identification code based on a hash function to obtain a corresponding plurality of hash values, and respectively mapping them to corresponding counting positions of the counter vector diagram to obtain the mapping vector; A deduplication module, configured to determine duplicate data slices among the plurality of data slices based on the mapping vector, and perform a deduplication operation based on the duplicate data slices. Wherein, when it is detected that at least one of the corresponding plurality of hash values is 0, it is determined that the corresponding identity identification code is a unique identity; when it is detected that all of the corresponding plurality of hash values are greater than 0, cumulative counting is performed at the corresponding counting position of the counter vector diagram to determine that the corresponding data slice is the duplicate data slice; An evaluation module, configured to obtain an evaluation index of the deduplication operation based on the number of data slices and the number of counting positions in the counter vector diagram.
13. An electronic device, characterized in that, Comprising: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the data object deduplication processing method according to any one of claims 1 to 7 and / or the evaluation method of the data object deduplication processing according to any one of claims 8 to 10 by executing the executable instructions.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data object deduplication processing method according to any one of claims 1 to 7 and / or the evaluation method of the data object deduplication processing according to any one of claims 8 to 10.
Citation Information
Patent Citations
Data deduplication storage method and device based on sliding window and storage medium
CN109582640A
Data processing method and data processing device
US20140172795A1