Data duplicate checking method and device, equipment, storage medium and program product
Through the bucket storage method of high-bit data and low-bit data, the problem of high memory continuity requirements of Bloom filters is solved, and efficient data plagiarism checking is achieved on non-continuous memory addresses, improving the universality and efficiency of data plagiarism checking.
Patent Information
- Application Number
- CN202510472430.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-05
AI Technical Summary
In big data scenarios, the existing Bloom filter has high requirements for device memory continuity, resulting in the memory being unable to meet the data reliance requirements, and the existing methods are inefficient under large data volumes.
The bucket storage method of high-bit data and low-bit data is adopted to calculate the high-bit data and obtain the target bucket through a hash algorithm, and the preset element value of the low-bit data is judged in the bucket, reducing the continuity requirement of the memory space.
It realizes data latch checking on non-continuous memory addresses, improves the universality and efficiency of data latch checking, and is suitable for more devices.
Smart Images

Figure CN120429288A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to a method, apparatus, device, storage medium and program product for data duplication checking. Background Art
[0002] In big data scenarios, due to the large volume of data and the low density of effective data, the equipment needs to clean and filter massive amounts of data during the data processing process.
[0003] To ensure processing efficiency, the device's local memory and central processing unit (CPU) resources can be used to clean data. The device can deploy a local Bloom filter to determine whether the received data to be processed is duplicate data. However, in scenarios with large data volumes, Bloom filters require a very long bit array and require the device's memory addresses to be continuous, which the device's memory cannot meet. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, device, storage medium and program product for data duplication checking, which can reduce the continuity requirements for the memory space of the device and improve the universality of data duplication checking.
[0005] In a first aspect, an embodiment of the present application provides a method for checking for duplicate data, comprising:
[0006] Receive data to be processed;
[0007] Calculating high-order data and low-order data corresponding to the data to be processed;
[0008] Based on the correspondence between the high-order data and the bucket position, a target bucket corresponding to the high-order data is obtained, where the target bucket is obtained from the target bucket position corresponding to the high-order data;
[0009] Determine whether the target bucket stores a preset element value corresponding to the low-order data;
[0010] In a case where the preset element value is stored in the target bucket, it is determined that the data to be processed is duplicate data.
[0011] In a possible implementation, determining whether the target bucket stores a preset element value corresponding to the low-order data includes:
[0012] In a case where the target bucket is a first type bucket, the low-order data is used as the preset element value, and the low-order data is searched for in segments in the target bucket;
[0013] When the target bucket is a second type bucket, performing remainder processing on a preset factor using the low-order data to obtain a target index number;
[0014] According to the preset correspondence between the index number and the index position, searching for the target index position corresponding to the target index number;
[0015] Determine whether the value at the target index position is the preset element value.
[0016] In a possible implementation, the data corresponding to the preset element value is non-fixed data; after determining whether the preset element value corresponding to the low-order data is stored in the target bucket, the method further includes:
[0017] If the preset element value is not stored in the target bucket and the target bucket is the first type of bucket, the low-order data is written into the target bucket in the order of the numerical value of the low-order data and the numerical value of the element value in the target bucket;
[0018] If the preset element value is not stored in the target bucket and the target bucket is the second-type bucket, the value of the target index position is set to the preset element value.
[0019] In a possible implementation, when the target bucket is a first-type bucket, searching for the low-order data in the target bucket in segments includes:
[0020] When the target bucket is the first type bucket, determining whether the number of elements stored in the first type bucket is less than a quantity threshold;
[0021] When the number of elements stored in the first-type bucket is less than the number threshold, the low-order data is searched for in segments in the target bucket.
[0022] In a possible implementation, the method further includes:
[0023] When the number of elements stored in the first type bucket is equal to the number threshold, construct a second type bucket containing a preset number of elements;
[0024] According to the order of the elements in the first type of bucket, using each element to perform a modulo process on the preset factor to obtain an index number corresponding to each element;
[0025] According to the preset correspondence between the index number and the index position, the value of the index position corresponding to each element is set to the preset element value;
[0026] Performing remainder processing on the preset factor using the low-order data to obtain the target index number;
[0027] Searching for the target index position according to a preset correspondence between the index number and the index position;
[0028] Determine whether the value at the target index position is a preset element value.
[0029] In a possible implementation, the method further includes:
[0030] When the data corresponding to the element value stored in each bucket is not the specified data and the target bucket is not obtained based on the corresponding relationship, a first type bucket is created;
[0031] The low-order data is written into the first-type bucket.
[0032] In a second aspect, an embodiment of the present application provides a device for checking data duplication, comprising:
[0033] A receiving module, used for receiving data to be processed;
[0034] A calculation module, used for calculating the high-order data and the low-order data corresponding to the data to be processed;
[0035] a judgment module, configured to, upon obtaining a target bucket corresponding to the high-order data based on a correspondence between the high-order data and the bucket position, determine whether the target bucket stores a preset element value corresponding to the low-order data, wherein the target bucket is obtained from the target bucket position corresponding to the high-order data;
[0036] The judging module is further configured to judge whether the target bucket stores a preset element value corresponding to the low-order data;
[0037] A search module is used to determine that the data to be processed is duplicate data when the preset element value is stored in the target bucket.
[0038] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions;
[0039] When the processor executes the computer program instructions, the method for checking data duplication as described in the first aspect is implemented.
[0040] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method for checking for duplicate data as in the first aspect is implemented.
[0041] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the method for data duplication checking as in the first aspect.
[0042] A method, apparatus, device, storage medium and program product for checking for duplicate data in an embodiment of the present application calculates the high-order data and low-order data corresponding to the data to be processed after receiving the data to be processed. The local storage has a correspondence between the bucket position and the high-order data. Based on the correspondence, the target bucket position corresponding to the data to be processed can be found, and then the target bucket can be obtained from the target bucket position. The device can check for duplicate data in the target bucket. In this way, by splitting the array storing the preset element values according to the bucket position, each bucket can correspond to less data to be processed. Therefore, the data duplication checking process can be completed using discontinuous memory addresses in the device, which reduces the use threshold of the device and improves the universality of data duplication checking. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 is an exemplary schematic diagram of the working principle of a Bloom filter in the prior art;
[0045] Figure 2 This is a flow chart of a method for checking for duplicate data provided by an embodiment of the present application;
[0046] Figure 3 This is an exemplary schematic diagram of a data duplication checking method provided in an embodiment of the present application;
[0047] Figure 4 This is a flowchart of a bucket construction method provided in an embodiment of the present application.
[0048] Figure 5 This is an exemplary schematic diagram of a bucket update method provided in an embodiment of the present application;
[0049] Figure 6 This is an exemplary schematic diagram of another data duplication checking method provided in an embodiment of the present application;
[0050] Figure 7 This is an exemplary schematic diagram of another bucket update method provided in an embodiment of the present application;
[0051] Figure 8This is an exemplary schematic diagram of another method for checking data duplication provided in an embodiment of the present application;
[0052] Figure 9 This is a structural diagram of a data duplication checking device provided in an embodiment of the present application;
[0053] Figure 10 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0055] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0056] Currently, when checking for duplicate data in big data, the following three solutions can be used:
[0057] The first solution: After receiving data, the electronic device generates a key value (KEY) based on the data content and stores the KEY and corresponding time information in a database or middleware such as Redis. After receiving the data to be processed, the electronic device calculates the corresponding KEY according to the same rules and then searches the middleware for the same KEY. If the same KEY exists, it indicates that the data to be processed is a duplicate.
[0058] However, in big data scenarios, since the KEY of each original data needs to be saved, the local memory of the electronic device may not be able to support the storage of the full amount of data. Therefore, the electronic device can call the disk to read and write data. As the number of keys written increases, the number of disk read and write times increases, thereby increasing network overhead. In addition, due to the bottleneck of the disk's read and write performance, the efficiency of data duplication checking is low.
[0059] The second solution: The electronic device generates a KEY based on the data content of the received data, and writes the KEY into a Set type variable in the local memory. Subsequently, after the electronic device receives the data to be processed, it generates a KEY corresponding to the data to be processed according to the same rules, and then searches the Set type variable for the same KEY. If so, it indicates that the data to be processed is duplicate data.
[0060] With the second solution, since Set type variables are stored in the local memory of the electronic device, when the number of KEYs is large, a large amount of memory resources of the electronic device are occupied, and some electronic devices may have insufficient memory resources.
[0061] The third solution: Electronic devices use Bloom filters to check for data duplication. First, a fixed-length bit array is generated and each bit in the array is set to 0. Multiple hash values are then calculated for the received data using multiple hash algorithms. For each hash value, the length of the bit array is modulo the hash value to obtain the array index corresponding to that hash value. The bit corresponding to each array index is then set to 1.
[0062] After receiving the data to be processed, the same hash algorithm and modulo method are used to calculate multiple target indexes corresponding to the data to be processed. The value of the bit corresponding to each target index is determined to be 1. If the value of the bit corresponding to each target index is 1, the data to be processed is determined to be duplicate data.
[0063] Specifically, such as Figure 1 As shown in the figure, bitArray is a bit array with a length of m. Each element in the array occupies one bit, and the value of each element can be 0 or 1. When initializing the bit array, the value of each element is set to 0.
[0064] When an electronic device receives data to be processed, namely "Zhang San", it first performs hash calculations on the data using multiple hash algorithms to obtain multiple hash values corresponding to the data to be processed. In an embodiment of the present application, the electronic device uses three hash algorithms to calculate three hash values corresponding to the data to be processed. For each hash value, the remainder corresponding to each hash value is calculated by calculating the remainder m-1 using the hash value. The remainders obtained are 2, 6, and 11, respectively.
[0065] Using the remainder obtained from the calculation above as an index, search the bit array for bits numbered 2, 6, and 11. If all three bits are 1, it means that the data to be processed, "Zhang San," may already exist. If at least one of the three bits is 0, it means that the data to be processed, "Zhang San," definitely does not exist.
[0066] When one of the three bit values is 0, it means that the data to be processed "Zhang San" does not exist, and the values of the three bits are all set to 1.
[0067] It should be noted that the Bloom filter has a certain misjudgment rate, which can be calculated using the following formula:
[0068] p=pow(1-exp(-k / (m / n)),k);
[0069] Among them, p represents the false positive rate, k is the number of hash functions, m is the length of the bit array, and n is the number of different elements inserted.
[0070] Taking the bit array length of 100 million and inserting 50 million elements as an example, the false positive rate of the Bloom filter is shown in Table 1:
[0071] Table 1
[0072] k (number of hashes) m (bit array length) n (number of inserted elements) p(false positive rate) 1 100000000 50000000 39.35% 2 100000000 50000000 39.96% 3 100000000 50000000 46.89% 4 100000000 50000000 55.90% 5 100000000 50000000 65.16% 6 100000000 50000000 73.61% 7 100000000 50000000 80.68% 8 100000000 50000000 86.25% 9 100000000 50000000 90.43% 10 100000000 50000000 93.46%
[0073] Taking the bit array length as 100 million and inserting 10 million elements as an example, the false positive rate of the Bloom filter is shown in Table 2:
[0074] Table 2
[0075]
[0076]
[0077] Taking the bit array length of 100 million and inserting 5 million elements as an example, the false positive rate of the Bloom filter is shown in Table 3:
[0078] Table 3
[0079] k (number of hashes) m (bit array length) n (number of inserted elements) p(false positive rate) 1 100000000 5000000 4.8771% 2 100000000 5000000 0.9056% 3 100000000 5000000 0.2703% 4 100000000 5000000 0.1080% 5 100000000 5000000 0.0530% 6 100000000 5000000 0.0303% 7 100000000 5000000 0.0196% 8 100000000 5000000 0.0140% 9 100000000 5000000 0.0108% 10 100000000 5000000 0.0089% 11 100000000 5000000 0.0078% 12 100000000 5000000 0.0071% 13 100000000 5000000 0.0068% 14 100000000 5000000 0.0067% 15 100000000 5000000 0.0068% 16 100000000 5000000 0.0071% 17 100000000 5000000 0.0076% 18 100000000 5000000 0.0083% 19 100000000 5000000 0.0092% 20 100000000 5000000 0.0104%
[0080] As shown in Tables 1, 2, and 3, when the bit array length is fixed, the Bloom filter's false positive rate increases as more data elements are written—that is, as more 1s are present in the bit array. This means that while maintaining a low false positive rate for Bloom filtering, the length of the bit array increases, increasing memory consumption. Furthermore, the bit array occupies consecutive memory addresses, placing high demands on the device's memory continuity.
[0081] In order to solve the problems of the prior art, the embodiments of the present application provide a method, apparatus, device, storage medium and program product for checking for duplicate data. The following first describes the method for checking for duplicate data provided by the embodiments of the present application.
[0082] like Figure 2 As shown, the method is applied to an electronic device, and the method includes:
[0083] S201: Receive data to be processed.
[0084] S202: Calculate the high-order data and low-order data corresponding to the data to be processed.
[0085] The high-order data and the low-order data are in one-to-one correspondence, and each data to be processed may correspond to a combination of multiple groups of high-order data and low-order data.
[0086] Regarding the above S202, calculating the high-order data and low-order data corresponding to the data to be processed, the high-order data and low-order data corresponding to each data to be processed can be calculated in the following manner, specifically:
[0087] Step 1: Use the preset hash algorithm to calculate the hash value corresponding to the data to be processed.
[0088] There are multiple preset hash algorithms, and the electronic device uses different preset hash algorithms to calculate multiple hash values for the same data to be processed.
[0089] Step 2: For each hash value, convert the hash value according to a preset base number to obtain the target data corresponding to each hash value.
[0090] The preset base numbers are pre-set based on actual business needs.
[0091] Step 3: For each target data, use the preset segmentation rule to segment the target data to obtain the high-order data and low-order data corresponding to each hash value.
[0092] In one example, using a preset hash algorithm, the electronic device calculates a hash value of 774889, and the preset base number is hexadecimal. The calculated hash value is converted to hexadecimal, resulting in a hexadecimal number of 000BD2E9. The electronic device then uses a preset segmentation rule to obtain the high-order data as 000B and the low-order data as D2E9.
[0093] In actual implementation, different preset hash algorithms and different segmentation rules can be determined according to the amount of data to be processed, and no specific restrictions are made here.
[0094] S203: Based on the correspondence between the high-order data and the bucket positions, obtain the target bucket corresponding to the high-order data.
[0095] The target bucket is obtained from the target bucket location corresponding to the high-order data. The bucket location is the storage address of each bucket in local memory. The bucket can be an array, and each element in the array is used to store data.
[0096] Specifically, after splitting to obtain the high-order data, the electronic device can search for the storage location of the target bucket corresponding to the high-order data based on the correspondence between the high-order data and the bucket position, and then obtain the target bucket at the storage location of the target bucket.
[0097] S204: Determine whether the target bucket stores a preset element value corresponding to the low-order data.
[0098] It's understandable that for the data to be written, the element values at the corresponding position of the same data in different types of buckets are different. Continuing with the above example, if the bucket is a shortArray array and the low-order data is D2E9, when this low-order data is written to the shortArray array, the preset element value at the corresponding position of this low-order data is D2E9. If the bucket is a bitArray array and the low-order data is D2E9, when this low-order data is written to the bitArray array, the preset element value corresponding to this low-order data is 1.
[0099] In the case where there are multiple preset hash algorithms, the electronic device can obtain multiple target buckets and then determine whether each target bucket stores a preset element value corresponding to the data to be processed.
[0100] S205: When the preset element value is stored in the target bucket, determine that the data to be processed is duplicate data.
[0101] Using the above method, after receiving the data to be processed, the high-order data and low-order data corresponding to the data to be processed are calculated. Among them, the local storage has a correspondence between the bucket position and the high-order data. Based on the correspondence, the target bucket position corresponding to the data to be processed can be found, and then the target bucket is obtained from the target bucket position. The device can check for duplicate data in the target bucket. In this way, by splitting the array storing the preset element values according to the bucket position, each bucket can correspond to less data to be processed, so each bucket can occupy less memory, that is, each bucket occupies fewer memory addresses. Therefore, the data duplication check process can be completed using discontinuous memory addresses in the device, which lowers the use threshold of the device, is applicable to more devices, and improves the universality of data duplication check.
[0102] Regarding the above S204, determining whether the target bucket stores the preset element value corresponding to the low-order data, due to different types of target buckets, there are two specific cases:
[0103] Case 1: When the target bucket is the first type of bucket, the low-order data is used as the preset element value, and the low-order data is searched for in segments in the target bucket.
[0104] The first type of bucket is used to store low-order data of the written data. The first type of bucket can be a shortArray array.
[0105] Case 2: If the target bucket is the second type of bucket, use the low-order data to perform a modulo operation on the preset factor to obtain the target index number. Based on the preset correspondence between index numbers and index positions, find the target index position corresponding to the target index number. Determine whether the value at the target index position is the preset element value.
[0106] Among them, the second type of bucket can be a bitArray array. For the data to be stored, the electronic device calculates the index number corresponding to the data to be stored, searches for the corresponding bit position in the bitArray array according to the index number, and sets the element of the found bit position to the preset element value, thereby completing the writing of the data to be stored.
[0107] Using the method provided in the embodiment of the present application, when the target bucket is a first type of bucket, it means that the target bucket is a shortArray array, that is, the specific numerical value of the data is stored in the target bucket. Therefore, the low-order bits of the data to be processed can be used as preset element values, and the target bucket is searched for the same value as the low-order data of the data to be processed. If it exists, it means that the data to be processed is duplicate data. When the target bucket is a second type of bucket, it means that the target bucket is a bitArray array. Therefore, the electronic device calculates the index number corresponding to the data to be processed, and thus searches for the element value of the corresponding position in the bitArray array to determine whether the data to be processed is stored in the target bucket. In this way, two different types of buckets are used to store data. When the number of elements is small, directly storing specific numerical values will not introduce additional space overhead. When the number of elements is large, using the bitArray array can effectively compress the storage requirements.
[0108] It should be noted that the data stored in the above-mentioned buckets can be divided into fixed data and non-fixed data. Non-fixed data means that the electronic device can update the data in the bucket when it determines that the data to be processed is non-duplicate data. Fixed data means that the data in the bucket is in read-only mode. The electronic device only uses the bucket to determine whether the data to be processed is duplicate data and cannot update the data in the bucket. Among them, if the data corresponding to the preset element value in the bucket is non-fixed data, the electronic device can update the data in the bucket. Specifically, the method for the electronic device to update the data is as follows:
[0109] Case 1: When the preset element value is not stored in the target bucket and the target bucket is a first type bucket, the low-order data is written into the target bucket in the order of the numerical value of the low-order data and the numerical value of the element value in the target bucket.
[0110] Case 2: When the preset element value is not stored in the target bucket and the target bucket is a second type bucket, the value of the target index position is set to the preset element value.
[0111] Using the method provided in the embodiment of the present application, when the target bucket is a first type of bucket, it indicates that the target bucket is a shortArray array, which stores specific numerical values. Therefore, the electronic device can write the numerical value of the low-order data into the target bucket in the order of the numerical size of the low-order data and the numerical size of the element value in the target bucket. After the low-order data is written in the order of the numerical size of the low-order data and the numerical size of the element value in the target bucket, it is convenient for the electronic device to search for the subsequently received data to be processed, thereby improving the search efficiency. In the case where the preset element value is not stored in the target bucket, and the target bucket is a second type of bucket, it indicates that the target bucket is a bitArray array, which stores 0 or 1 identifiers. Therefore, the electronic device needs to set the bit value of the value to the preset element value after calculating the target index position corresponding to the low-order data, so as to facilitate the electronic device to search for the subsequently received data to be processed.
[0112] The following combination Figure 3 Introducing a data duplication checking method provided by the embodiment of the present application, such as Figure 3 As shown:
[0113] The electronic device receives the data to be processed "Zhang San", calculates the hash value of "Zhang San" as 774889, converts it to hexadecimal 000BD2E9, and then splits it according to the preset segmentation rules to obtain high-order data 000B and low-order data D2E9, and then obtains the target bucket according to the correspondence between the high-order data and the bucket position.
[0114] Among them, the target bucket corresponding to 000B is a bitArray array, and the length of the bitArray array is m. The electronic device uses the low-order data D2E9 to calculate the remainder of m-1 to obtain the target index number corresponding to the low-order data. The target index number is 237. According to the index number corresponding to each bit position in the target bucket, the bit position corresponding to the target index number is searched. The element value in the bit position is 1, indicating that the data to be processed is duplicate data.
[0115] It should be noted that there are two types of buckets: bitArray and shortArray. When the target bucket is a shortArray, due to the small number of storage elements, the shortArray is used to store the data. The electronic device determines whether the low-order data already exists in the shortArray. If not, the low-order data is inserted into the shortArray. The specific method of inserting the low-order data is described in the above embodiment and will not be repeated here.
[0116] When the target bucket found is a bitArray array, a Bloom filter is used to determine whether the data to be processed is duplicate data.
[0117] It should be noted that when the target bucket is a first-type bucket, it means that the target bucket is a shortArray array. This shortArray array is used to store specific values and is suitable for scenarios with small amounts of data. As data is continuously written, the amount of data stored in the shortArray array increases. To reduce memory consumption, the shortArray array can be converted to a bitArray array, and the received data can be stored in the bitArray array. Based on this, for the above situation 1, when the target bucket is a first-type bucket, the low-order data is used as the preset element value, and the low-order data is searched segmented in the target bucket, which can be implemented as follows:
[0118] When the target bucket is a first type bucket, it is determined whether the number of elements stored in the first type bucket is less than a quantity threshold.
[0119] When the number of elements stored in the first type bucket is less than the quantity threshold, the low-order data is searched for in segments in the target bucket.
[0120] In this way, if the number of elements stored in the first-type bucket is less than the threshold, it indicates that the current data stored in memory is small. Therefore, the first-type bucket can be used to store specific data values. Furthermore, because the first-type bucket is sorted by the value of each element, the segmented search method can improve data query efficiency.
[0121] Accordingly, if Figure 4 As shown in FIG, when the number of elements stored in the first type bucket is equal to the quantity threshold, the method for the electronic device to construct the second type bucket is as follows:
[0122] S401: When the number of elements stored in the first type bucket is equal to a quantity threshold, construct a second type bucket containing a preset number of elements.
[0123] Among them, the quantity threshold is pre-set based on experience. When the number of elements stored in the first type of bucket is equal to the quantity threshold, it means that there is a lot of data stored in the current memory. In order to compress the storage space, a second type of bucket based on the Bloom filter is constructed.
[0124] S402 : According to the order of the elements in the first type bucket, a modulo process is performed on a preset factor using each element to obtain an index number corresponding to each element.
[0125] S403 : According to the preset correspondence between the index number and the index position, the value at the index position corresponding to each element is set to a preset element value.
[0126] It can be understood that each bit in the second type of bucket corresponds to an index number. After calculating the index number corresponding to the element in the first type of bucket, the value of the index position corresponding to the index number can be directly set to the preset element value. By setting the value of the corresponding bit to 1, it indicates that the electronic device has stored the original data corresponding to the bit.
[0127] S404: Perform remainder processing on the preset factor using the low-order data to obtain a target index number.
[0128] S405: Search for the target index position according to the preset correspondence between the index number and the index position.
[0129] S406: Determine whether the value at the target index position is a preset element value.
[0130] Using the method provided in the embodiment of the present application, when the number of elements stored in the first type of bucket is equal to the quantity threshold, it means that there is a lot of data written in the first type of bucket. Since new data may continue to be written subsequently, a second type of bucket is constructed, and the elements in the first type of bucket are filled into the corresponding positions in the second type of bucket. In this way, when a large amount of data is written, the specific values of the low-order data are avoided from being stored, and the corresponding data writing is represented by the preset element value in the second type of bucket, thereby compressing the data storage space and reducing memory consumption.
[0131] In some embodiments of the present application, after the electronic device obtains the high-order data corresponding to the data to be processed and determines the target bucket position, the corresponding target bucket may not yet be stored in the target bucket position, that is, the electronic device has not yet written the data at the current moment. Therefore, the data corresponding to the element value stored in each bucket is non-fixed data, and based on the corresponding relationship, in order to obtain the target bucket, a first type of bucket is created, and the low-order data is written into the first type of bucket. In this way, when the electronic device determines that the data in each bucket is non-fixed data, the electronic device can update the data in the bucket. Therefore, the electronic device writes the above-mentioned low-order data into the first type of bucket to realize dynamic writing of the data to be processed.
[0132] It should be noted that, when the data stored in the bucket is non-fixed data, the electronic device can execute a preset asynchronous timer task for the bucket stored in the local memory of the electronic device, generate a new bucket, each bucket is empty, and update the bucket according to a preset period. Figure 5 As shown:
[0133] The electronic device executes an asynchronous timer task at a preset period, initially generating a completely empty bucket and replacing the currently active bucket in memory with the newly generated bucket. After receiving the data to be processed, the electronic device uses the replaced bucket for logical comparison to check for duplicate data.
[0134] The length and storage location of the newly generated bucket are the same as those of the currently used bucket.
[0135] In the case where the data stored in each bucket is non-fixed data, the electronic device can update the bucket based on the received data to be processed. Based on this, the following describes the data processing flow of non-fixed data in combination with 6, such as Figure 6 As shown, the method includes:
[0136] S601: Receive data to be processed.
[0137] S602: Check whether the format is passed.
[0138] If it passes, execute S603; if it fails, end the process.
[0139] The electronic device first filters data with abnormal formats using preset filtering rules.
[0140] S603: Check data for duplicates.
[0141] The method for performing data duplication checking is described in the above embodiments and will not be repeated here.
[0142] S604: Determine whether it is possible to include.
[0143] If yes, the process ends; if no, execute S605.
[0144] Among them, after the above-mentioned electronic device performs data duplication checking, if the target bucket is found to be the first type of bucket, the electronic device determines that the memory includes the data to be processed after finding the same element value as the low-order data of the data to be processed. If the target bucket is found to be the second type of bucket, the electronic device calculates multiple low-order data of the data to be processed, calculates the index number corresponding to each low-order data respectively, and searches for the value of the corresponding bit according to each index number. If the value of the bit corresponding to each index number is a preset element value, it is determined that the memory of the electronic device may include the data to be processed.
[0145] S605: Perform business logic processing.
[0146] Among them, if the electronic device determines that the memory does not include the data to be processed, the data to be processed is processed. For example, the business logic can be data storage, and the business logic can be set according to actual business needs. In the embodiment of this application, there is no specific limitation on the business logic.
[0147] By adopting the method provided in the embodiment of the present application, after receiving the data to be processed, the data to be processed is first verified according to the format, and data with abnormal formats is filtered out, which can improve the data processing efficiency. Then, the data to be processed is checked for duplicates, and after determining that the data to be processed is not stored in the memory, the data to be processed is processed for business logic processing. In this way, when the data stored in the bucket is non-fixed data, the electronic device can dynamically interpolate and write to the bucket. Moreover, in the process of real-time big data processing, due to the large amount of data and the low density of effective data, most of the data is filtered out in the data duplication check link, avoiding business logic processing of invalid data, thereby saving computing resources.
[0148] like Figure 7 As shown, when the data stored in the bucket is fixed data, the electronic device executes an asynchronous timing task according to a preset period, generates a new bucket corresponding to each bucket position according to each bucket position and each bucket size, and then obtains the latest fixed data, writes the latest fixed data into the new bucket, and then replaces the bucket in use with the new bucket.
[0149] It should be noted that the replaced bucket is in read-only mode, and the electronic device cannot modify the data stored in the replaced bucket.
[0150] When the data stored in the bucket is fixed data, the electronic device can perform a logical comparison based on the received data to be processed to check for duplicates, and use different business logic to process the data that is the same as the fixed data. Based on this, the following introduces the data processing flow of fixed data in combination with 8, such as Figure 8 As shown, the method includes:
[0151] S801: Receive data to be processed.
[0152] S802: Check whether the format is passed.
[0153] If it passes, execute S803; if it fails, end the process.
[0154] S803: Check data for duplicates.
[0155] Specifically, the method of checking for data duplication is described in the above embodiments and will not be repeated here.
[0156] S804: Determine whether it is possible to include.
[0157] If yes, execute S805; if no, execute S807.
[0158] S805: Query original data.
[0159] The original data is pre-set fixed data.
[0160] After using the above-mentioned data duplication checking method to determine that the memory may include data to be processed, the electronic device needs to accurately determine whether the data to be processed is fixed data. Therefore, in order to further improve the accuracy of duplication checking, the original data is used to determine whether the data to be processed is fixed data.
[0161] In one example, the fixed data can be a mobile phone number for number portability or a mobile phone number on a text message blacklist. If the data to be processed is a mobile phone number, the fixed data is stored in a bucket. The electronic device first uses the aforementioned data duplication checking method to determine whether the fixed data may contain the mobile phone number. If the electronic device determines that the fixed data may contain the mobile phone number, it compares the original fixed data with the mobile phone number to further determine whether the original data contains the mobile phone number.
[0162] S806: Determine whether it is included.
[0163] If yes, execute S808; if no, execute S807.
[0164] S807. Process using business logic 1.
[0165] S808: Use business logic 2 for processing.
[0166] Using the method provided in the embodiment of the present application, for the data to be processed that is not duplicated with the original data, the business logic 1 is directly used to process the data to be processed. For the electronic device, the Bloom filter is used to determine that the data to be processed may be duplicate data, and the original data is further obtained. The original data is used to determine whether the data to be processed is duplicate data, thereby improving the accuracy of the judgment result. In addition, depending on whether the data stored in the bucket is fixed data and whether a certain amount of false positives are received, different processing processes are adopted for different scenarios, which can be applicable to different application scenarios.
[0167] Based on the same concept, the present application provides a device for checking data duplication, such as Figure 9 As shown, the device includes:
[0168] Receiving module 901, used to receive data to be processed;
[0169] A calculation module 902 is used to calculate the high-order data and the low-order data corresponding to the data to be processed;
[0170] A judgment module 903 is configured to obtain a target bucket corresponding to the high-order data based on a correspondence between the high-order data and the bucket position, where the target bucket is obtained from the target bucket position corresponding to the high-order data;
[0171] The judging module 903 is further configured to judge whether the target bucket stores a preset element value corresponding to the low-order data;
[0172] The search module 904 is configured to determine that the data to be processed is duplicate data when the preset element value is stored in the target bucket.
[0173] In a possible implementation, the determination module 903 is specifically configured to:
[0174] In a case where the target bucket is a first type bucket, the low-order data is used as the preset element value, and the low-order data is searched for in segments in the target bucket;
[0175] When the target bucket is a second type bucket, performing remainder processing on a preset factor using the low-order data to obtain a target index number;
[0176] According to the preset correspondence between the index number and the index position, searching for the target index position corresponding to the target index number;
[0177] Determine whether the value at the target index position is the preset element value.
[0178] In a possible implementation, the data corresponding to the preset element value is non-fixed data; the device further includes:
[0179] a writing module, configured to, if the preset element value is not stored in the target bucket and the target bucket is the first type of bucket, write the low-order data into the target bucket in the order of the numerical value of the low-order data and the numerical value of the element value in the target bucket;
[0180] The writing module is further configured to set the value of the target index position to the preset element value when the preset element value is not stored in the target bucket and the target bucket is the second type bucket.
[0181] In a possible implementation, the determination module 903 is specifically configured to:
[0182] When the target bucket is the first type bucket, determining whether the number of elements stored in the first type bucket is less than a quantity threshold;
[0183] When the number of elements stored in the first-type bucket is less than the number threshold, the low-order data is searched for in segments in the target bucket.
[0184] In a possible implementation, the device further includes:
[0185] A construction module, configured to construct a second type of bucket containing a preset number of elements when the number of elements stored in the first type of bucket is equal to the number threshold;
[0186] a remainder calculation module, configured to perform remainder calculation on the preset factor using each element in the order of the elements in the first type of buckets to obtain an index number corresponding to each element;
[0187] The writing module is further configured to set the value of the index position corresponding to each element to the preset element value according to a preset correspondence between the index number and the index position;
[0188] The remainder calculation module is further configured to perform remainder processing on the preset factor using the low-order data to obtain the target index number;
[0189] The search module 904 is further configured to search for the target index position according to a preset correspondence between the index number and the index position;
[0190] The judging module 903 is further configured to judge whether the value at the target index position is a preset element value.
[0191] In a possible implementation, the method further includes:
[0192] A creation module, configured to create a first type of bucket when the data corresponding to the element value stored in each bucket is not the specified data and the target bucket is not obtained based on the corresponding relationship;
[0193] The writing module is further configured to write the low-order data into the first-type bucket.
[0194] It should be noted that the device for checking data duplication is a device corresponding to the above-mentioned method for checking data duplication. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of the device and can achieve the same technical effects.
[0195] Figure 10 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0196] The electronic device may include a processor 1001 and a memory 1002 storing computer program instructions.
[0197] Specifically, the processor 1001 may include a central processing unit (CPU) or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0198] The memory 1002 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 1002 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 1002 may include removable or non-removable (or fixed) media. Where appropriate, the memory 1002 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 1002 is a non-volatile solid-state memory.
[0199] In certain embodiments, the memory 1002 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0200] The processor 1001 implements any one of the data storage methods in the above embodiments by reading and executing computer program instructions stored in the memory 1002 .
[0201] In one example, the electronic device may further include a communication interface 1003 and a bus 1004. Figure 10 As shown, the processor 1001, the memory 1002, and the communication interface 1003 are connected via a bus 1004 and communicate with each other.
[0202] The communication interface 1003 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0203] The bus 1004 includes hardware, software, or both that couples components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Super Transmission (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 1004 may include one or more buses. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.
[0204] In addition, in conjunction with the data duplication checking method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the data duplication checking methods in the above embodiments is implemented.
[0205] An embodiment of the present application also provides a computer program product, including a computer program, which, when processed and executed, implements any one of the data duplication checking methods in the above embodiments.
[0206] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0207] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (erasable read-only memory, EROM), floppy disks, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical discs, hard disks, optical fiber media, radio frequency (Radio Frequency, RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0208] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0209] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0210] The above is only a specific implementation method of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited to this. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application.
Claims
1. A method for checking data duplication, characterized in that: include: Receive data to be processed; Calculating high-order data and low-order data corresponding to the data to be processed; Based on the correspondence between the high-order data and the bucket position, a target bucket corresponding to the high-order data is obtained, where the target bucket is obtained from the target bucket position corresponding to the high-order data; Determine whether the target bucket stores a preset element value corresponding to the low-order data; In a case where the preset element value is stored in the target bucket, it is determined that the data to be processed is duplicate data.
2. The method according to claim 1, characterized in that The determining whether the target bucket stores a preset element value corresponding to the low-order data includes: In a case where the target bucket is a first type bucket, the low-order data is used as the preset element value, and the low-order data is searched for in segments in the target bucket; When the target bucket is a second type bucket, performing remainder processing on a preset factor using the low-order data to obtain a target index number; According to the preset correspondence between the index number and the index position, searching for the target index position corresponding to the target index number; Determine whether the value at the target index position is the preset element value.
3. The method according to claim 2, characterized in that The data corresponding to the preset element value is non-fixed data; after determining whether the preset element value corresponding to the low-order data is stored in the target bucket, the method further includes: If the preset element value is not stored in the target bucket and the target bucket is the first type of bucket, the low-order data is written into the target bucket in the order of the numerical value of the low-order data and the numerical value of the element value in the target bucket; If the preset element value is not stored in the target bucket and the target bucket is the second-type bucket, the value of the target index position is set to the preset element value.
4. The method according to claim 2, characterized in that When the target bucket is a first type bucket, searching for the low-order data in the target bucket in segments includes: When the target bucket is the first type bucket, determining whether the number of elements stored in the first type bucket is less than a quantity threshold; When the number of elements stored in the first-type bucket is less than the number threshold, the low-order data is searched for in segments in the target bucket.
5. The method according to claim 4, characterized in that Also includes: When the number of elements stored in the first type bucket is equal to the number threshold, construct a second type bucket containing a preset number of elements; According to the order of the elements in the first type of bucket, using each element to perform a modulo process on the preset factor to obtain an index number corresponding to each element; According to the preset correspondence between the index number and the index position, the value of the index position corresponding to each element is set to the preset element value; Performing remainder processing on the preset factor using the low-order data to obtain the target index number; Searching for the target index position according to a preset correspondence between the index number and the index position; Determine whether the value at the target index position is a preset element value.
6. The method according to claim 1, wherein Also includes: When the data corresponding to the element value stored in each bucket is not the specified data and the target bucket is not obtained based on the corresponding relationship, a first type bucket is created; The low-order data is written into the first-type bucket.
7. A device for checking data duplication, characterized in that: include: A receiving module, used for receiving data to be processed; A calculation module, used for calculating the high-order data and the low-order data corresponding to the data to be processed; A judgment module is used to obtain a target bucket corresponding to the high-order data based on the correspondence between the high-order data and the bucket position, wherein the target bucket is obtained from the target bucket position corresponding to the high-order data; The judging module is further configured to judge whether the target bucket stores a preset element value corresponding to the low-order data; A search module is used to determine that the data to be processed is duplicate data when the preset element value is stored in the target bucket.
8. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the method for checking data duplication as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method for checking for duplicate data according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the data duplication checking method as described in any one of claims 1 to 6.