Repeated data detection method and device, storage medium and computer program product

The preset bit array and hash table are constructed through multiple preset hash functions, and the duplicate data blocks are detected layer by layer, solving the problem of inefficient detection of duplicate data in the existing technology, and achieving efficient and accurate duplicate data recognition.

CN120296008AActive Publication Date: 2025-07-11INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510771673.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-11
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In the prior art, duplicate data detection methods are inefficient in large-scale data and have excessive memory usage, so they cannot efficiently identify and remove duplicate data.

Method used

A preset bit array is constructed using multiple preset hash functions, initially filtering out possible duplicate data blocks, and finely detecting it through preset hash tables, and repeating data detection is carried out layer by layer.

Benefits of technology

It improves the efficiency of repeated data detection, reduces system resource consumption, and ensures the accuracy and flexibility of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296008A_ABST
    Figure CN120296008A_ABST
Patent Text Reader

Abstract

The invention discloses a duplicated data detection method and device, a storage medium and a computer program product, and relates to the technical field of data processing, a preset bit array is constructed through a plurality of preset functions and stored data blocks to serve as the basis of data duplicated detection of to-be-detected data blocks, and the data duplicated detection efficiency is improved. The to-be-detected data blocks which may be repeated are preliminarily screened out through the preset bit array, and then the to-be-detected data blocks which may be repeated are further subjected to refined detection through the preset hash table, so that layered detection of data is realized, and compared with hash table look-up traversal comparison of all the to-be-detected data blocks, the detection efficiency is effectively improved, and the detection time is shortened. And meanwhile, the detection accuracy is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a method, device, storage medium, and computer program product for detecting duplicate data. Background Art

[0002] Removing duplicate data is one of the common steps in the data processing process. Duplicate data may cause problems such as model training deviation and resource waste.

[0003] Generally, to detect duplicate data, a hash table needs to be established for an existing database to record the hash values of each piece of data in it. When new data arrives, calculate the hash value of the new data and traverse the hash table to check for the same hash value. If the same hash value is found, it means the new data is duplicate with the existing data in the database.

[0004] However, since the maintenance and traversal of the hash table both consume a large amount of system resources, this method of detecting duplicate data often has problems such as low efficiency and excessive memory occupation when dealing with large-scale data. Summary of the Invention

[0005] This application provides a method, device, storage medium, and computer program product for detecting duplicate data to at least solve the problems of low efficiency and excessive memory occupation in detecting duplicate data in related technologies.

[0006] This application provides a method for detecting duplicate data, including: Performing hash calculations on the hash values of at least one stored data block respectively based on multiple preset hash functions to determine the addresses of multiple first target element positions corresponding to the at least one stored data block; Setting the multiple first target element positions in a preset bit array to non-empty values; Performing hash calculations on the hash value of a data block to be detected respectively based on the multiple preset hash functions to determine the addresses of multiple second target element positions corresponding to the data block to be detected; If there are no empty values among the multiple second target element positions in the preset bit array, determine whether the data block to be detected is a duplicate data block according to a preset hash table.

[0007] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above methods for detecting duplicate data when executing the computer program.

[0008] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above methods for detecting duplicate data are implemented.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any one of the above-mentioned duplicate data detection methods when executed by a processor.

[0010] The present application constructs a preset bit array through multiple preset functions and the stored data blocks as the basis for detecting data duplicates of the data blocks to be detected. The data blocks to be detected that may be duplicates are initially screened out through the preset bit array, and then the data blocks to be detected that may be duplicates are further refined and detected through a preset hash table, thereby realizing hierarchical detection of data. Compared with the hash table traversal comparison of all data blocks to be detected, the detection efficiency is effectively improved, and the accuracy of detection is ensured at the same time. Description of the Drawings

[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a flowchart of a duplicate data detection method provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a duplicate data detection device provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0014] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0015] To enable those skilled in the art of the present technology to better understand the solution of this application, the following further elaborates on this application in conjunction with the accompanying drawings and specific embodiments.

[0016] An embodiment of this application provides a duplicate data detection method. In combination with the execution process of the duplicate data detection method, the method will be described in detail.

[0017] Figure 1 It is a flowchart of a duplicate data detection method provided by an embodiment of this application. This method can be applied to any one of multiple computing nodes, where each node can be any electronic device with data processing capabilities, such as a server, a smart phone, a personal digital assistant, a tablet computer, a desktop computer, a laptop computer, an all-in-one computer, etc. It can be understood that the duplicate data detection method provided by the embodiments of this application can also be applied in other scenarios.

[0018] The following Figure 1 introduces the shown duplicate data detection method, and the specific steps included in this method are as follows: S101: Based on multiple preset hash functions, perform hash calculations on the hash values of at least one stored data block respectively to determine the addresses of multiple first target element bits corresponding to at least one stored data block.

[0019] The multiple preset hash functions are a set of mutually independent hash functions selected for constructing a preset bit array, where there is a low correlation between each preset hash function, so as to effectively disperse the spatial distribution of the calculation results of the multiple preset hash functions.

[0020] Optionally, the number k of preset hash functions is determined by the length w of the preset bit array and the number v of expected inserted elements:

[0021] Among them, the number of expected inserted elements refers to the sum of the number of stored data blocks and the number of non-repeating data blocks to be detected.

[0022] The stored data blocks are the data blocks added to the storage space at historical moments. When the electronic device obtains the stored data blocks, hash calculations are performed on them to obtain the hash values of the stored data blocks. That is, each stored data block has its own corresponding hash value.

[0023] Optionally, the hash function used to calculate the hash value of the stored data block can be a hash function different from the multiple preset hash functions.

[0024] The first target element position refers to the mapping position of the calculation results obtained by performing hash calculations on the hash values of at least one stored data block by multiple preset hash functions in a preset bit array. According to the address of the first target element position, the corresponding first target element position can be determined in the preset bit array.

[0025] Specifically, for each stored data block, based on multiple preset hash functions, hash calculations are performed on the hash value of the stored data block to obtain the addresses of the first target element positions corresponding to each stored data block respectively. Among them, the number of addresses of the first target element positions corresponding to each stored data block is the same as the number of multiple preset hash functions, that is, for a single stored data block, the address of the first target element position corresponds one-to-one with the preset hash function.

[0026] Furthermore, the set of the first target element positions corresponding to each stored data block respectively is the address of the multiple first target element positions corresponding to at least one stored data block.

[0027] S102. Set multiple first target element positions in the preset bit array to non-empty values.

[0028] The preset bit array includes multiple element positions. In the initial state, the element value in each element position is an empty value.

[0029] Optionally, the initial capacity of the preset bit array is set to 1.2 times the estimated data volume to avoid frequent expansion; the false positive rate of the preset bit array is adjusted according to the memory resources and the requirements of the judgment accuracy. The false positive rate of the preset bit array refers to the probability of determining a non-repetitive data block as a possibly repetitive data block.

[0030] Based on the addresses of the multiple first target element positions, determine the corresponding multiple first target element positions in the preset bit array, and set the multiple first target element positions to non-empty values to obtain the preset bit array corresponding to the current at least one stored data block.

[0031] Specifically, successively modify the element value of the corresponding first target element position according to the address of the first target element position corresponding to each stored data block respectively. When a certain first target element position has already been set to a non-empty value, there is no need to modify it again.

[0032] Optionally, the preset bit array is a binary array with an initial state of 0. For example, if the addresses of the multiple first target element positions corresponding to a certain stored data block are h1, h2, and h3, then set the element values at the three positions h1, h2, and h3 in the preset bit array to 1.

[0033] Among them, the preset bit array is a Bloom filter.

[0034] S103. Calculate the hash values of the hash values of the data block to be detected respectively based on multiple preset hash functions, and determine the addresses of multiple second target element bits corresponding to the data block to be detected.

[0035] The data block to be detected refers to a new data unit that is newly input and needs to be detected for data duplication.

[0036] When the data block to be detected is obtained, the electronic device calculates the hash value of the data block to be detected by using the same hash function as that for calculating the hash value of the stored data block.

[0037] After obtaining the hash value corresponding to the data block to be detected, perform hash calculation on the hash value corresponding to the data block to be detected based on the above multiple preset hash functions to obtain the addresses of multiple second target element bits. Among them, the addresses of the second target element bits correspond one by one to the preset hash functions.

[0038] The multiple preset functions adopted in this step are the same as the multiple preset functions used in the process of constructing the preset bit array according to the stored data blocks.

[0039] S104. If there are no null values in multiple second target element bits in the preset bit array, determine whether the data block to be detected is a duplicate data block according to the preset hash table.

[0040] According to the addresses of multiple second target element bits, determine multiple second target element bits in the preset bit array, and detect whether the element values in each second target element bit are null values.

[0041] If the element values in each second target element bit are not null values, it means that there may be multiple first target element bits corresponding to the stored data block that are the same as multiple target element bits corresponding to the data block to be detected. Since the multiple preset hash functions used to determine the first target element bit or the second element bit are the same, the hash value of the data block to be detected may be the same as the hash value of a certain stored data block, and the data block to be detected may be duplicated with the stored data block. At this time, it is necessary to further determine whether the data block to be detected is a duplicate data block in combination with the preset hash table.

[0042] If there are null values in multiple second target element bits, determine that the data block to be detected is not a duplicate data block.

[0043] Specifically, if there is at least one second target element bit whose element value is not a null value, it means that multiple second target element bits corresponding to the data block to be detected are not exactly the same as multiple first target element bits corresponding to any stored data block. Correspondingly, the hash value of the data block to be detected is not the same as the hash value of any stored data block, and it can be determined that the data block to be detected is not duplicated with any stored data block.

[0044] When there are null values among multiple second target element bits, or after determining that the data block to be detected is not a duplicate data block according to a preset hash table, update the preset bit array according to the multiple second target element bits; update the preset hash table according to the data block to be detected and the hash value of the data block to be detected. The preset hash table records the stored data blocks and the hash values corresponding to each stored data block respectively.

[0045] Specifically, set the element values of the multiple second target element bits to non-null values; write the data block to be detected and the hash value of the data block to be detected into the preset hash table.

[0046] In the embodiments of the present application, a preset bit array is constructed with multiple preset functions and the stored data blocks as the basis for detecting data duplication of the data blocks to be detected. The data blocks to be detected that may be duplicates are initially screened out through the preset bit array, and then the data blocks to be detected that may be duplicates are further refined and detected through the preset hash table, so as to achieve hierarchical detection of data. Compared with the hash table traversal comparison of all data blocks to be detected, the detection efficiency is effectively improved, and the accuracy of detection is ensured at the same time.

[0047] In some embodiments, before determining the addresses of the multiple second target element bits corresponding to the data block to be detected by performing hash calculation on the hash value of the data block to be detected based on multiple preset hash functions, the method further includes: dividing the data to be detected into blocks according to the type of the data to be detected to obtain multiple data blocks to be detected; calculating the hash value of the data block to be detected.

[0048] Dividing the data to be detected into blocks is a process of splitting the data to be detected into discrete processing units. For different data types, different data block splitting strategies are selected respectively.

[0049] Among them, dividing the data to be detected into blocks according to the type of the data to be detected to obtain multiple data blocks to be detected includes: if the data to be detected is unstructured data, move a first sliding window in the data to be detected with a first preset step length, and calculate the data feature value in the first sliding window in real time; when the data feature value meets the preset condition, determine the block splitting position according to the position of the first sliding window to obtain at least one block splitting position; according to the block splitting position, split the data to be detected into multiple data blocks to be detected.

[0050] When the data to be detected is unstructured data, such as data of web page content, etc., trigger the sliding window block splitting mechanism.

[0051] Taking the data to be detected as the processing object, using the first sliding window, sliding the window according to the preset step length, and calculating the data feature value in each window, such as the hash value of the data in the first sliding window.

[0052] When the data characteristic value meets the preset condition, the block position is determined according to the position of the current first sliding window.

[0053] For example, the starting position of the current first sliding window is used as the block position; or, the ending position of the current first sliding window is used as the block position; or, according to the ending position of the current first sliding window and the step size of window sliding, the ending position of the first sliding window before this window sliding operation is determined, so as to determine the block position.

[0054] Furthermore, the data to be detected before the block position is divided into a data block to be detected.

[0055] Optionally, if the difference between the data characteristic value corresponding to the current first sliding window and the data characteristic value in the previous window is greater than the preset threshold, it means that the content of the data in the current first sliding window changes greatly compared with the data in the previous window, and it is determined as the semantic boundary, then the starting position of the current first sliding window is used as the block position.

[0056] Correspondingly, calculate the hash value of the data block to be detected, including: determining that the data hash value corresponding to the first sliding window at the initial position in the data block to be detected is the hash value of the data block to be detected.

[0057] Specifically, for the string s[0..m-1] of the first sliding window at the initial position, calculate the hash value:

[0058] Where d is the base of the hash function; q is the modulus of the hash function; s[i] represents the character encoding in the string, usually the ASCII code.

[0059] Or, the hash calculation can also be performed on the data block to be detected segmented based on the first sliding window to obtain the hash value of the data block to be detected.

[0060] Or, if the data to be detected is structured data, the data to be detected is divided into multiple data blocks to be detected according to the preset number of bytes, and the number of bytes of each data block to be detected is the preset number of bytes.

[0061] When the data to be detected is structured data, such as a log file, etc., the data to be detected is divided into multiple data blocks to be detected with equal number of bytes according to the preset number of bytes.

[0062] The hash calculation is performed on each of the multiple data blocks to be detected obtained by segmentation to obtain the hash value corresponding to each data block to be detected.

[0063] In this embodiment, a data chunking strategy is adaptively determined according to the data type of the data to be detected. For structured data, fixed - byte - count chunking is used, and for unstructured data, dynamic chunking is used, which conforms to the data characteristics of the corresponding data types and improves the flexibility and accuracy of data chunking.

[0064] In some embodiments, determining whether a data block to be detected is a duplicate data block according to a preset hash table includes: searching for a target data block with the same hash value as the data block to be detected in the preset hash table; comparing the data block to be detected and the target data block to obtain a comparison result; and determining whether the data block to be detected is a duplicate data block according to the comparison result.

[0065] If there are no null values in multiple second target element bits in the preset bit array, it is necessary to perform further refined comparison on the data block to be detected according to the stored data blocks and their hash values recorded in the preset hash table.

[0066] First, search in the preset hash table according to the hash value of the data block to be detected to determine the target data block in the stored data blocks that has the same hash value as the data block to be detected, and then perform refined comparison between the data block to be detected and the target data block. The target data block is the data block that may be duplicate with the data block to be detected.

[0067] Among them, comparing the data block to be detected and the target data block to obtain a comparison result includes: moving a second sliding window in the data block to be detected and the target data block, and calculating the feature value to be detected and the target feature value corresponding to each window movement operation in real - time. The feature value to be detected is the data feature value within the second sliding window of the data block to be detected, and the target feature value is the data feature value within the second sliding window of the target data block; comparing the feature value to be detected and the target feature value corresponding to the same window movement operation to obtain a comparison result.

[0068] The second sliding window has a preset window size and a preset moving step. Based on the second sliding window, the data block to be detected and the target data block are processed in parallel, and the feature value to be detected of the data in the second sliding window of the data block to be detected and the target feature value of the data in the second sliding window of the target data block corresponding to each window sliding operation are calculated.

[0069] If the feature value to be detected and the target feature value corresponding to at least one window movement operation are different, it is determined that the data block to be detected and the target data block are not duplicates; otherwise, it is determined that the data block to be detected and the target data block are duplicates.

[0070] Compare the feature value to be detected corresponding to each window movement operation with the target feature value. When the feature value to be detected is the same as the target feature value, it means that the data in the second sliding window of the current data block to be detected is the same as the data in the second sliding window of the target data block. Continue with the next window movement operation, and compare the feature value to be detected and the target feature value corresponding to the next window movement operation. If, after processing the data block to be detected and the target data block in parallel based on the second sliding window, the feature value to be detected corresponding to each window sliding operation is the same as the target feature value, it means that the data in the second sliding window of each data block to be detected is the same as the data in the second sliding window of the target data block, that is, the data block to be detected and the target data block are duplicates.

[0071] When the feature value to be detected is different from the target feature value, it means that the data in the second sliding window of the current data block to be detected is different from the data in the second sliding window of the target data block, that is, the data block to be detected and the target data block are not duplicates.

[0072] When there is a difference between the feature value to be detected and the target feature value corresponding to one window movement operation, it can be determined that the data block to be detected and the target data block are not duplicates. At this time, stop the window movement operation, and update the preset bit array and the preset hash table according to the data block to be detected.

[0073] In the embodiments of the present disclosure, the detection data block and the target data block are processed in parallel through the second sliding window, and it is determined whether the data block to be detected and the target data block are duplicates according to the comparison result of the feature value to be detected and the target feature value. Once there is a difference between the feature value to be detected and the target feature value corresponding to one window movement operation, it can be determined that the data block to be detected and the target data block are not duplicates, and there is no need to perform a full-byte traversal comparison on the data block to be detected and the target data block, reducing the data calculation amount and further improving the efficiency of duplicate data detection.

[0074] In some embodiments, when moving the second sliding window in the data block to be detected and the target data block, and calculating the feature value to be detected and the target feature value corresponding to each window movement operation in real time, it includes: moving the second sliding window in the preset data block, and determining the removed characters and the added characters in the second sliding window during a single movement. The preset data block is the data block to be detected or the target data block; according to the hash contribution of the removed characters and the added characters, update the first window hash value to obtain the second window hash value. The first window hash value is the hash value of the data in the second sliding window before a single movement operation, and the second window hash value is the hash value of the data in the second sliding window after a single movement operation.

[0075] The removed characters refer to the characters that disappear from the second sliding window due to the movement operation of the second sliding window, and the added characters refer to the characters that newly enter the second sliding window due to the movement operation of the second sliding window.

[0076] The first window hash value is the data feature value in the second sliding window before the current window movement operation. The second window hash value is the data feature value in the current second sliding window. The second window hash value is obtained by updating the first window hash value, rather than directly calculated by hashing the data in the current second sliding window.

[0077] Updating the first window hash value includes subtracting the hash contribution of the removed character from the first window hash value and adding the hash contribution of the newly added character.

[0078] Specifically, multiplying the base value by the removed character to obtain the hash contribution of the removed character. The base value is calculated based on the radix, modulus, and window size of the second sliding window of the hash function; determining the difference between the first window hash value and the hash contribution of the removed character, and calculating the product of the difference and the radix; taking the modulus of the sum of the product and the newly added character based on the modulus to obtain the second window hash value.

[0079] Before the window movement operation of the second sliding window, perform pre-calculation based on the radix, modulus, and window size of the second sliding window of the hash function to obtain the base value. The specific calculation process is as follows:

[0080] Where h is the base value, d is the radix of the hash function, n is the window size of the second sliding window, and q is the modulus of the hash function.

[0081] The calculation process of updating the first window hash value to obtain the second window hash value is expressed as:

[0082] is the first window hash value, is the second window hash value, s[x] is the character encoding of the removed character, and s[x + n] is the character encoding of the newly added character. represents the hash contribution of the removed character in the first window hash value.

[0083] Calculate the difference between the first window hash value and the hash contribution of the removed character, further calculate the product of the difference and the radix d for exponent alignment; further take the modulus of the sum of the product and the newly added character based on the modulus to add the hash contribution of the newly added character to obtain the second window hash value.

[0084] Optionally, the radix of the hash function needs to cover the range of character encodings. For example, for ASCII character encoding, the radix is set to 256; for Unicode character encoding, the radix is set to 65536. The modulus of the hash function should be set to a large prime number to reduce hash collisions; the window size is adjusted according to the data type. For example, for text data, it is set to 64 - 256 characters; for binary data, it is set to 4KB - 1MB.

[0085] In the embodiment of the present application, the second window hash value is obtained by updating the first window hash value, avoiding recalculating the hash value of the entire window during the second sliding window process, thereby effectively reducing the time complexity and improving the comparison efficiency between the data block to be detected and the target data block.

[0086] In some embodiments, among multiple nodes, there are a master node and multiple local nodes. The master node is communicatively connected to each local node through a network or other means. Each node respectively maintains its own preset bit array and preset hash table, and the master node maintains a global bit array and a global hash table. The global bit array and the global hash table are constructed based on the set of stored data blocks of multiple local nodes. The specific construction method is the same as that of the preset bit array and the preset hash table, which will not be elaborated here.

[0087] The duplicate data detection method further includes: detecting the non - duplicate data blocks to be detected in each local node based on the global bit array and the global hash table to determine whether the data blocks to be detected are duplicate data blocks.

[0088] Specifically, the non - duplicate data blocks to be detected in each local node are the data blocks that have been detected based on the duplicate data detection process in the above - mentioned embodiment and are determined not to be duplicate with the stored data blocks of that local node. For such data blocks, it is determined whether there are null values in multiple second target element bits in the global bit array.

[0089] If there are no null values in multiple second target element bits in the global bit array, then it is determined whether the data block to be detected is a duplicate data block according to the global hash table. Alternatively, if there are null values in multiple second target element bits in the global bit array, it is determined that the data block to be detected is not a duplicate data block, stored in the distributed file system or database, and the global bit array and the global hash table are updated for subsequent use.

[0090] Optionally, the master node is further configured to fragment the data to be detected, and transmit the fragmented data to each local node through a network in the form of a message queue or a distributed file system for repeated detection, so as to reduce the processing burden of a single node and improve the overall processing capacity of multiple nodes. For example, fragmentation can be performed according to the range of text hash values; or, for web page data, fragmentation can be performed according to the Uniform Resource Locator (URL) of the web page.

[0091] Optionally, for each node, the multi-threading technology can be further used to accelerate the data processing capacity, and the data block to be detected is allocated to multiple duplicate data detection threads for processing. Among them, the task parallelism between multiple threads can be set to 2-3 times the number of central processing unit cores on this node, so as to make full use of the cluster resources.

[0092] In the embodiments of the present disclosure, the preset bit array and the global bit array are used to perform collaborative processing on the module to be detected, and the non-duplicate data blocks of the local node are re-checked in the global bit array and the global hash table, further improving the accuracy of duplicate data detection; at the same time, multi-threaded parallel processing is used on each node to accelerate the duplicate data detection, further improving the efficiency of duplicate data detection.

[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0094] Figure 2 This is a schematic structural diagram of the duplicate data detection device provided by the embodiments of the present application. The duplicate data detection device provided by the embodiments of the present application can execute the processing flow provided by the embodiments of the duplicate data detection method. As Figure 2 shown, the duplicate data detection device 20 includes a first calculation module 21, a setting module 22, a second calculation module 23, and a determination module 24; the first calculation module 21 is configured to perform hash calculation on the hash values of at least one stored data block respectively based on a plurality of preset hash functions, and determine the addresses of a plurality of first target element bits corresponding to at least one stored data block; the setting module 22 is configured to set the plurality of first target element positions in the preset bit array to non-null values; the second calculation module 23 is configured to perform hash calculation on the hash value of the data block to be detected respectively based on a plurality of preset hash functions, and determine the addresses of a plurality of second target element bits corresponding to the data block to be detected; the determination module 24 is configured to, if there are no null values in the plurality of second target element bits in the preset bit array, determine whether the data block to be detected is a duplicate data block according to the preset hash table.

[0095] Optionally, the duplicate data detection device 20 further includes: a chunking module, configured to chunk the data to be detected according to the type of the data to be detected, obtaining a plurality of data chunks to be detected; and a third calculation module, configured to calculate the hash value of the data chunk to be detected.

[0096] Optionally, if the data to be detected is unstructured data, the chunking module is configured to move a first sliding window in the data to be detected with a first preset step length, and calculate the data feature value in the first sliding window in real time; when the data feature value meets a preset condition, determine the chunking position according to the position of the first sliding window, obtaining at least one chunking position; and divide the data to be detected into a plurality of data chunks to be detected according to the chunking position.

[0097] Optionally, the third calculation module is further configured to determine that the data hash value corresponding to the initial position of the first sliding window in the data chunk to be detected is the hash value of the data chunk to be detected.

[0098] Optionally, if the data to be detected is structured data, the chunking module is configured to chunk the data to be detected according to a preset number of bytes, obtaining a plurality of data chunks to be detected, and the number of bytes of each data chunk to be detected is the preset number of bytes.

[0099] Optionally, the duplicate data detection device 20 further includes an update module, configured to determine that the data chunk to be detected is not a duplicate data chunk if there are null values in a plurality of second target element bits; update a preset bit array according to the plurality of second target element bits; and update a preset hash table according to the data chunk to be detected and the hash value of the data chunk to be detected.

[0100] Optionally, the determination module 24 includes a search unit, a comparison unit, and a determination unit; the search unit is configured to search for a target data chunk with the same hash value as the data chunk to be detected in a preset hash table; the comparison unit is configured to compare the data chunk to be detected and the target data chunk, obtaining a comparison result; and the determination unit is configured to determine whether the data chunk to be detected is a duplicate data chunk according to the comparison result.

[0101] Optionally, the comparison unit is configured to move a second sliding window in the data chunk to be detected and the target data chunk, and calculate the data feature value to be detected and the target feature value corresponding to each window movement operation in real time, where the data feature value to be detected is the data feature value in the second sliding window in the data chunk to be detected, and the target feature value is the data feature value in the second sliding window in the target data chunk; and compare the data feature value to be detected and the target feature value corresponding to the same window movement operation, obtaining a comparison result.

[0102] Optionally, the comparison unit is further configured to move a second sliding window in a preset data block, determine the removed characters and added characters in the second sliding window during a single movement, where the preset data block is a data block to be detected or a target data block; update the first window hash value according to the hash contribution of the removed characters and the added characters to obtain a second window hash value, where the first window hash value is the hash value of the data in the second sliding window before a single movement operation, and the second window hash value is the hash value of the data in the second sliding window after a single movement operation.

[0103] Optionally, the comparison unit is further configured to multiply the base value by the removed characters to obtain the hash contribution of the removed characters, where the base value is calculated based on the radix, modulus, and window size of the second sliding window; determine the difference between the first window hash value and the hash contribution of the removed characters, and calculate the product of the difference and the radix; perform modulo operation on the sum of the product and the added characters based on the modulus to obtain the second window hash value.

[0104] Optionally, the determination unit is further configured to determine that the data block to be detected and the target data block are not duplicate if the eigenvalue to be detected corresponding to at least one window movement operation is different from the target eigenvalue; otherwise, determine that the data block to be detected and the target data block are duplicate.

[0105] Optionally, the update module is further configured to update the preset bit array according to multiple second target element bits if the data block to be detected is not a duplicate data block; update the preset hash table according to the data block to be detected and the hash value of the data block to be detected.

[0106] Figure 2 The duplicate data detection device in the illustrated embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0107] Figure 3 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device provided by the embodiment of the present application can execute the processing flow provided by the duplicate data detection method embodiment, as Figure 3 shown, the electronic device 30 includes: a memory 31, a processor 32, a computer program, and a communication interface 33; wherein, the computer program is stored in the memory 31 and is configured to be executed by the processor 32 to perform the duplicate data detection method as described above.

[0108] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above duplicate data detection method embodiments when running.

[0109] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0110] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the duplicate data detection method.

[0111] Another embodiment of the present application also provides a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the duplicate data detection method.

[0112] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0113] The above has introduced in detail a duplicate data detection method, device, storage medium, and computer program product provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for detecting duplicate data, characterized in that, Including: Based on multiple preset hash functions, respectively perform hash calculations on the hash values of at least one stored data block to determine the addresses of multiple first target element bits corresponding to the at least one stored data block; Set the multiple first target elements in the preset bit array to non-null values; Based on the multiple preset hash functions, respectively perform hash calculations on the hash value of the data block to be detected to determine the addresses of multiple second target element bits corresponding to the data block to be detected; If there are no null values among the multiple second target element bits in the preset bit array, determine whether the data block to be detected is a duplicate data block according to a preset hash table.

2. The duplicate data detection method according to claim 1, wherein Before the step of based on the multiple preset hash functions, respectively perform hash calculations on the hash value of the data block to be detected to determine the addresses of multiple second target element bits corresponding to the data block to be detected, the method further includes: Chunk the data to be detected according to the type of the data to be detected to obtain multiple data blocks to be detected; Calculate the hash value of the data block to be detected.

3. The duplicate data detection method according to claim 2, wherein The step of chunk the data to be detected according to the type of the data to be detected to obtain multiple data blocks to be detected includes: If the data to be detected is unstructured data, move a first sliding window in the data to be detected with a first preset step length, and calculate the data feature value in the first sliding window in real time; When the data feature value meets a preset condition, determine the chunking positions according to the positions of the first sliding window to obtain at least one chunking position; According to the chunking positions, split the data to be detected into multiple data blocks to be detected.

4. The method according to claim 3, wherein The step of calculate the hash value of the data block to be detected includes: Determine that the data hash value corresponding to the initial position of the first sliding window in the data block to be detected is the hash value of the data block to be detected.

5. The deduplication detection method according to claim 2, wherein The step of chunk the data to be detected according to the type of the data to be detected to obtain multiple data blocks to be detected includes: If the data to be detected is structured data, chunk the data to be detected according to a preset number of bytes to obtain multiple data blocks to be detected, and the number of bytes of each data block to be detected is the preset number of bytes.

6. The deduplication method according to claim 1, wherein After the step of based on the multiple preset hash functions, respectively perform hash calculations on the hash value of the data block to be detected to determine the addresses of multiple second target element bits corresponding to the data block to be detected, the method further includes: If there are null values among the multiple second target element bits, determine that the data block to be detected is not a duplicate data block; Update the preset bit array according to the multiple second target element bits; Update the preset hash table according to the data block to be detected and the hash value of the data block to be detected.

7. The duplicate data detection method according to claim 1, characterized in that, The step of determine whether the data block to be detected is a duplicate data block according to a preset hash table includes: Search for a target data block with the same hash value as the data block to be detected in the preset hash table; Compare the data block to be detected and the target data block to obtain a comparison result; According to the comparison result, determine whether the data block to be detected is a duplicate data block.

8. The duplicate data detection method according to claim 7, wherein Performing comparison on the data block to be detected and the target data block to obtain a comparison result, including: Moving a second sliding window in the data block to be detected and the target data block, and calculating in real time the to-be-detected feature value and the target feature value corresponding to each window movement operation, where the to-be-detected feature value is the data feature value within the second sliding window in the data block to be detected, and the target feature value is the data feature value within the second sliding window in the target data block; Comparing the to-be-detected feature value and the target feature value corresponding to the same window movement operation to obtain the comparison result.

9. The duplicate data detection method according to claim 8, wherein The moving a second sliding window in the data block to be detected and the target data block and calculating in real time the to-be-detected feature value and the target feature value corresponding to each window movement operation includes: Moving a second sliding window in a preset data block, and determining the removed characters and the added characters in the second sliding window during a single movement, where the preset data block is the data block to be detected or the target data block; Updating a first window hash value according to the hash contribution of the removed characters and the added characters to obtain a second window hash value, where the first window hash value is the hash value of the data within the second sliding window before a single movement operation, and the second window hash value is the hash value of the data within the second sliding window after the single movement operation.

10. The duplicate data detection method according to claim 9, characterized in that, The updating a first window hash value according to the hash contribution of the removed characters and the added characters to obtain a second window hash value includes: Multiplying a base value by the removed characters to obtain the hash contribution of the removed characters, where the base value is calculated based on the radix of the hash function, the modulus, and the window size of the second sliding window; Determining the difference between the first window hash value and the hash contribution of the removed characters, and calculating the product of the difference and the radix; Taking the modulus of the sum of the product and the added characters based on the modulus to obtain the second window hash value.

11. The deduplication method according to claim 8, wherein The determining whether the data block to be detected is a duplicate data block according to the comparison result includes: If the to-be-detected feature value and the target feature value corresponding to at least one window movement operation are different, determining that the data block to be detected and the target data block are not duplicates; Otherwise, determining that the data block to be detected and the target data block are duplicates.

12. The duplicate data detection method according to claim 1, wherein The method further includes: If the data block to be detected is not a duplicate data block, updating the preset bit array according to the multiple second target element bits; Updating the preset hash table according to the data block to be detected and the hash value of the data block to be detected.

13. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the duplicate data detection method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the duplicate data detection method according to any one of claims 1 to 12 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the duplicate data detection method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • A fuzzy matching-supporting cloud storage data dereplication method

    CN105868305A

  • Data deduplication processing method, device and equipment and storage medium

    CN112559452A

  • Cloud storage deduplication method and device based on sliding window block optimization algorithm

    CN114185850A

  • Document detection processing method and device, storage medium and electronic equipment

    CN114444464A

  • Fixed chunk size deduplication with variable-size chunking

    US20210073178A1

Cited By

  • Data detection method and device, equipment and storage medium

    CN121210449A