Data processing method and device, electronic equipment, storage medium and program product

By using the combination method of Bloom filter and database deduplication table in a distributed environment, the problem that messages may be delivered multiple times is solved, and idempotence and efficient query of data processing are achieved.

CN120067085APending Publication Date: 2025-05-30KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510142423.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In a distributed environment, performing the same operation multiple times may lead to messages being delivered multiple times, resulting in unwarranted consumption of system resources, and the prior art is difficult to effectively ensure message idempotence.

Method used

By receiving the target service data sent by the service service at the preset deduplication verification service cluster node, determining its deduplication verification mark, and querying it using the target Bloom filter and the preset database deduplication table to determine whether the data exists. If it does not exist, it will be processed and stored in the database and the Bloom filter.

Benefits of technology

It effectively guarantees the idempotence of business data processing, avoids repeated data processing, reduces database query pressure, and improves query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067085A_ABST
    Figure CN120067085A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, electronic equipment, a storage medium and a program product, and the method comprises the steps: determining a deduplication verification identifier of target business data when the target business data sent by any business service is received; on the basis of the deduplication verification identifier of the target business data, whether the target business data exists or not is determined by querying the target Bloom filter and the preset database deduplication table; and under the condition that the target business data does not exist, after the target business processing is performed on the target business data, storing the target business data into the preset database deduplication table and writing the target business data into the target Bloom filter, so that the idempotence of business data processing is ensured, and the deduplication efficiency is improved under the condition that the Bloom filter queries the deduplication table. The frequency of querying the deduplication table can be greatly reduced, the query pressure of the database can be effectively reduced, whether the data exists or not can be quickly judged, and the query efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to data processing technology, and in particular to a data processing method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] Data idempotence means that in a distributed environment, the results of executing the same operation multiple times are consistent. In other words, no matter how many times a user initiates the same request, the final state of the system is the same, and there will be no different results due to multiple operations.

[0003] Nowadays, data idempotency issues are often encountered in software systems. For example, many MQ (Message Queue) systems (such as RocketMQ and Kafka) use an "at least once delivery" mechanism to ensure that messages are not lost. However, this also leads to new problems - messages may be delivered multiple times. In this case, if the idempotence of messages cannot be guaranteed, messages will be processed repeatedly, resulting in unnecessary consumption of system resources. Summary of the invention

[0004] To solve the technical problems in the related art, the embodiments of the present disclosure provide a data processing method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] According to a first aspect of an embodiment of the present disclosure, a data processing method is provided, the method being applied to a preset deduplication verification service cluster node, the method comprising:

[0006] Receive target business data sent by any business service, and determine a deduplication verification identifier of the target business data;

[0007] Based on the deduplication verification flag of the target business data, determining whether the target business data exists in the preset database deduplication table by querying a target Bloom filter and a preset database deduplication table;

[0008] In the case where it is determined that the target business data does not exist in the preset database deduplication table, performing target business processing on the target business data;

[0009] The target business data that has completed the target business processing is stored in the preset database deduplication table and written into the target Bloom filter.

[0010] As an optional embodiment, the determining whether the target business data exists in the preset database deduplication table by querying a target Bloom filter and a preset database deduplication table based on the deduplication verification flag of the target business data includes:

[0011] Determine whether the target Bloom filter is normal;

[0012] If the target Bloom filter is normal, query the target Bloom filter through the deduplication verification identifier;

[0013] If it is determined that the deduplication verification identifier does not exist by querying the target Bloom filter, determine that the target service data does not exist in the preset database deduplication table;

[0014] If it is determined that the deduplication verification identifier exists by querying the target Bloom filter, query the preset database deduplication table a second time to determine whether the target service data exists in the preset database deduplication table;

[0015] If it is determined that the target service data exists by querying the preset database deduplication table, determine that the target service data exists in the preset database deduplication table.

[0016] As an alternative embodiment, determining whether the target service data exists in the preset database deduplication table by querying the target Bloom filter and the preset database deduplication table based on the deduplication verification identifier of the target service data includes:

[0017] In the case where the target Bloom filter is abnormal, query the preset database deduplication table to determine whether the target service data exists in the preset database deduplication table.

[0018] As an alternative embodiment, receiving the target service data sent by any business service and determining the deduplication verification identifier of the target service data includes:

[0019] Receive the target service data sent by any business service and determine whether the target service data includes a preset identifier field;

[0020] If the target service data includes a preset identifier field, determine the preset identifier field of the target service data as the deduplication verification identifier of the target service data;

[0021] If the target service data does not have a preset identifier field, convert the target service data into a preset data type;

[0022] Perform a hash calculation on the target service data of the preset data type to obtain the hash value corresponding to the target service data;

[0023] Determine the hash value corresponding to the target service data as the deduplication verification identifier of the target service data.

[0024] As an alternative embodiment, the method further includes:

[0025] By determining the false positive rate and the length of the bit array of the Bloom filter, the target Bloom filter is obtained.

[0026] As an alternative embodiment, the obtaining the target Bloom filter by determining the false positive rate and the length of the bit array of the Bloom filter includes:

[0027] Receiving the value of the false positive rate of the Bloom filter and the maximum value range of the amount of service data of the service input by the user;

[0028] According to the value of the false positive rate and the maximum value range of the amount of service data of the service, determining the length of the bit array of the Bloom filter according to a preset calculation formula;

[0029] Based on the value of the false positive rate and the calculated length of the bit array of the Bloom filter, obtaining the target Bloom filter.

[0030] As an alternative embodiment, the obtaining the target Bloom filter by determining the false positive rate and the length of the bit array of the Bloom filter includes:

[0031] Receiving the component information of the Bloom filter, the value of the false positive rate, and the maximum number of service data of the service input by the user;

[0032] Determining the Bloom filter components according to the component information of the Bloom filter;

[0033] Inputting the value of the false positive rate and the maximum number of service data of the service into the Bloom filter components to obtain the length of the bit array and the number of hash functions of the Bloom filter;

[0034] Obtaining the target Bloom filter according to the value of the false positive rate, the length of the bit array of the Bloom filter, and the number of hash functions.

[0035] As an alternative embodiment, the method further includes:

[0036] Determining whether the service data of any service meets the preset data conditions;

[0037] When it is determined that the preset data conditions are met, partitioning the service data of the service according to a preset range to obtain a plurality of data ranges;

[0038] Creating a corresponding target Bloom filter for each data range, so that when the target service data of this range is received and it is confirmed that the target service data of this range does not exist, writing the target service data into the target Bloom filter.

[0039] As an alternative embodiment, the method further includes:

[0040] Determine whether the target Bloom filter is abnormal;

[0041] In the case where it is determined that the target Bloom filter is abnormal, add an abnormality identification flag at a preset position of the deduplication table where the target service data in the preset database deduplication table is located;

[0042] In the mode of querying whether the target service data exists through the preset database deduplication table, traverse the data rows with the abnormality identification flag in the preset database deduplication table;

[0043] Write the target service data corresponding to the abnormality identification flag into the target Bloom filter.

[0044] As an alternative embodiment, the method further includes:

[0045] In the case where it is determined that the target Bloom filter is abnormal, send the information of the abnormality of the target Bloom filter to the configuration center, so that the configuration center changes the preset configuration item identifier to notify other preset deduplication verification service cluster nodes of the abnormal working state of the target Bloom filter, and the preset configuration item identifier is used to mark the working state of the target Bloom filter.

[0046] According to the second aspect of the embodiments of the present disclosure, a data processing device is provided. The device is configured in a preset deduplication verification service cluster node, and the device includes:

[0047] A verification identifier determination module, configured to receive target service data sent by any business service and determine the deduplication verification identifier of the target service data;

[0048] A query module, configured to determine whether the target service data exists in the preset database deduplication table by querying a target Bloom filter and a preset database deduplication table based on the deduplication verification identifier of the target service data;

[0049] A service processing module, configured to perform target service processing on the target service data in the case where it is determined that the target service data does not exist in the preset database deduplication table;

[0050] A data writing module, configured to store the target service data that has completed the target service processing into the preset database deduplication table and write it into the target Bloom filter.

[0051] According to the third aspect of the embodiments of the present disclosure, an electronic device is provided, including:

[0052] A memory, configured to store a computer program product;

[0053] A processor for executing a computer program product stored in the memory, and when the computer program product is executed, implementing the method according to any one of the above embodiments.

[0054] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, implementing the method according to any one of the above embodiments.

[0055] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including computer program instructions, and when the computer program instructions are executed by a processor, implementing the method according to any one of the above embodiments.

[0056] Through the technical solution of the embodiments of the present disclosure, when receiving target service data sent by a service, the de-duplication check identifier of the target service data can be used to query the target Bloom filter and the preset database de-duplication table to determine whether the target service data exists in the preset database de-duplication table. If the target service data does not exist, after performing target service processing on the target service data, the target service data is stored in the preset database de-duplication table and written into the target Bloom filter. In this way, before data writing or storage, the method of querying by using a Bloom filter and / or the method of querying the de-duplication table through the preset database de-duplication table is used to query whether the data already exists, effectively ensuring the idempotency of service data processing and avoiding the problem of data duplication. In addition, in the case of querying through a Bloom filter, since the Bloom filter can determine whether the target service data exists only through its bit array without having to query a large amount of data in the de-duplication table, the frequency of querying the de-duplication table can be greatly reduced, the query pressure on the database can be effectively reduced, and it can quickly determine whether the data already exists, improving the query efficiency.

[0057] The technical solution of the present disclosure will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.

[0059] Referring to the drawings, the present disclosure can be more clearly understood from the following detailed description, wherein:

[0060] Figure 1 FIG. is a schematic diagram of the system architecture to which the data processing method according to an embodiment of the present disclosure is applied.

[0061] Figure 2 FIG. is a schematic diagram of the initial state of the Bloom filter in the data processing method according to an embodiment of the present disclosure.

[0062] Figure 3 Schematic diagram of the state of adding data to the Bloom filter in the data processing method provided by an embodiment of the present disclosure.

[0063] Figure 4 Schematic diagram of the state of querying data from the Bloom filter in the data processing method provided by an embodiment of the present disclosure.

[0064] Figure 5 Flowchart of the data processing method provided by an embodiment of the present disclosure.

[0065] Figure 6 Flowchart of the data processing method provided by another embodiment of the present disclosure.

[0066] Figure 7 Flowchart of the data processing method provided by still another embodiment of the present disclosure.

[0067] Figure 8 Block diagram of the structure of the data processing device provided by an embodiment of the present disclosure.

[0068] Figure 9 Block diagram of the structure of a data processing device provided by another embodiment of the present disclosure.

[0069] Figure 10 Block diagram of the structure of the electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0070] In the related art, common methods for ensuring the idempotency of data processing include the following:

[0071] 1. Database unique constraint

[0072] In the application of this method, database operations may become a bottleneck. Especially in high-concurrency scenarios, frequent insert operations may lead to performance degradation. In addition, the unique constraint of the database may cause deadlocks or other concurrency problems.

[0073] 2. Deduplication table

[0074] When processing a request, first check whether the data ID already exists in the deduplication table. If it exists, it means that the data has been processed and can be skipped; if it does not exist, the data can be processed and the data ID can be inserted into the deduplication table. The disadvantage is that it will cause a large pressure on the database in high-concurrency scenarios.

[0075] In summary, the methods commonly used in the related art to ensure the idempotency of data processing are difficult to fully ensure the idempotency of data processing. In addition, in the process of ensuring the idempotency of data processing, a large data processing pressure will be imposed on the database, thereby affecting the data processing efficiency.

[0076] To solve some or all of the technical problems in the related art, embodiments of the present disclosure propose a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. The technical solutions of the embodiments of the present disclosure will be introduced in detail below with reference to the accompanying drawings.

[0077] Figure 1 Schematic diagram of the system architecture to which a data processing method provided by an embodiment of the present disclosure is applied. As Figure 1 shown, the system architecture of the embodiments of the present disclosure may include multiple business services (such as three business services A, B, and C in the figure), a deduplication verification service cluster (hereinafter referred to as a preset deduplication verification service cluster; which may include multiple nodes), a configuration center (such as the Zookeeper cluster in the figure), a Bloom filter cluster (such as the Redis cluster in the figure; hereinafter referred to as the target Bloom filter), and a database (such as the MySQL master database and the MySQL slave database in the figure).

[0078] Among them, a business service may be, for example, a software system capable of performing data processing and generating business data, and it may send a deduplication verification request to the deduplication verification service cluster (for example, when business data is generated and needs to be deduplicated and verified).

[0079] The deduplication verification service cluster can provide an interface externally to implement the core verification logic. Multiple nodes are used to ensure the high availability of the system and can be horizontally scaled. After any one of the nodes receives the deduplication verification request, it performs deduplication verification processing using the data processing method of the embodiments of the present disclosure. When it is determined that the business data does not exist, it responds to the deduplication verification request, writes the business data into the Bloom filter in the Bloom filter cluster, and stores it in the database. On the other hand, if any one of the nodes in the re-verification service cluster determines that the Bloom filter cluster is abnormal, it can report the information of the abnormal Bloom filter cluster to the configuration center, so that the configuration center modifies the relevant configuration items. Other nodes in the re-verification service cluster obtain the information of the abnormal Bloom filter cluster by reading the information of the configuration items, and thus switch the deduplication verification mode (for example, switch from the Bloom filter deduplication verification mode to the database deduplication verification mode).

[0080] The Redis cluster is used to store the Bloom filter or the Bloom filter cluster. For example, the RedisBloom plugin provided by versions of Redis after 4.0 can be used to implement the Bloom filter. Specifically, common Bloom filter implementation solutions include, but are not limited to, the RedisBloom plugin provided by versions of Redis after 4.0, the BloomFilter class in Guava, etc. Among them, RedisBloom is a module of Redis that provides implementations of various data structures such as Bloom filters, counters, and Top-K. By storing these data structures in Redis, it provides efficient data processing and query functions for application programs or software systems. It can utilize the high-performance characteristics of Redis to provide faster query speeds and better concurrency processing capabilities, and is particularly suitable for scenarios that require efficiently determining whether an element exists in a set, such as software systems in a distributed environment, etc.; while the BloomFilter in Guava is a component in an open-source library that uses a series of hash functions and bit arrays to determine whether an element is in a set, and can efficiently determine the existence of an element, but there is a certain false recognition rate. It is more suitable for scenarios that require efficient queries but can tolerate a certain false recognition rate, such as determining hot data in a cache, etc. To adapt to the problem of data idempotency in a distributed environment, the RedisBloom in this embodiment of the present disclosure is taken as the preferred embodiment.

[0081] The configuration center or the ZooKeeper cluster can create persistent nodes to store the global deduplication verification mode flag (database query or Bloom filter query). Each service node in the deduplication verification service cluster listens for data changes in the configuration item of this ZooKeeper cluster node. When the data of this node configuration item changes, it can notify each service node in the deduplication verification service cluster in real time to update the local deduplication verification mode field to determine whether to use a database query or a Bloom filter query.

[0082] The database, taking the MySQL database as an example, is mainly used to store the deduplication table to store the business data of the business service.

[0083] To enable those skilled in the art to accurately understand the technical solutions of the embodiments of the present disclosure, the Bloom filter will be introduced in detail below with reference to the accompanying drawings.

[0084] Figure 2 It is a schematic diagram of the initial state of the Bloom filter in the data processing method provided by an embodiment of the present disclosure. The Bloom filter is composed of a binary vector or bitmap of a fixed size and a series of mapping functions (hash functions or Hash functions), as Figure 2As shown, assume the initial state of a Bloom filter with m = 7. In the initial state, for the bit array of length 7, all its bits are set to 0.

[0085] Figure 3 This is a schematic diagram of the state of adding data to the Bloom filter in the data processing method provided by an embodiment of the present disclosure. Still taking Figure 2 the Bloom filter with the median array length of 7 as an example, as Figure 3 shown, when storing two data x and y in the Bloom filter, for example, K mutually independent Hash functions are used for calculation, and then the K positions obtained by the Hash mapping of the two data are all set to 1. Figure 3 Taking K = 3 as an example in , for data x, assuming the hash values (or Hash values) calculated by the 3 Hash functions are 1, 4, and 7 respectively, then the corresponding positions of the bit array with bit position numbers 1, 4, and 7 are set to 1. Similarly, for data y, assuming the hash values (or Hash values) calculated by the 3 Hash functions are 2, 4, and 6 respectively, then the corresponding positions of the bit array with bit position numbers 2, 4, and 6 are set to 1. Among them, the position with bit position number 4 coincides with that of x, so only one 1 needs to be retained.

[0086] Correspondingly, when querying or detecting whether a certain data exists, these K Hash functions are still used to calculate K positions. If all these K positions are 1, it indicates that the data exists; otherwise, it does not exist. Figure 4 This is a schematic diagram of the state of querying data in the Bloom filter in the data processing method provided by an embodiment of the present disclosure. As Figure 4 shown, still taking Figure 3 the Bloom filter with the array length of 7 and storing data in as an example, query whether data x and w exist. For data x, assuming K = 3, then the hash values (or Hash values) calculated by the 3 Hash functions are 1, 4, and 7 respectively. Query Figure 4 the Bloom filter in . The bits corresponding to the positions of the bit array with bit position numbers 1, 4, and 7 are 1, so x may exist (because the position with bit position number 4 of y's bit array is also 1, so it is not necessarily x). Similarly, for data w, assuming K = 3, then the hash values (or Hash values) calculated by the 3 Hash functions are 1, 5, and 7 respectively. Among the bits corresponding to the positions of the bit array with bit position numbers 1, 5, and 7, only the bits corresponding to the positions with bit position numbers 1 and 7 of the bit array are 1, and the bit corresponding to the position with bit position number 5 of the bit array is 0, so they are not all 1, then it is determined that w does not exist.

[0087] Combined withFigure 1 The system structure and Figure 4 In the Bloom filter, compared with the data deduplication verification in the related art, the Bloom filter only requires a bit array and several hash functions (or hash functions). Compared with traditional data structures such as databases, its space occupancy is much smaller. Moreover, the query operation of the Bloom filter only needs to query the bit array, and the calculation of the hash function is also very fast. While database queries usually require retrieving a large amount of data. Therefore, especially when dealing with a large amount of data, the efficiency of the Bloom filter is significantly higher, so the query speed is very fast. In summary, the embodiments of the present disclosure propose a data processing method and device. By using the Bloom filter to query the deduplication table, the frequency of querying the deduplication table can be greatly reduced, the query pressure on the database can be effectively reduced, and the query efficiency of the deduplication table can be improved.

[0088] Figure 5 The flowchart of the data processing method provided by an embodiment of the present disclosure. As Figure 5 shown, the data processing method provided by the embodiments of the present disclosure can be applied to any node of a preset deduplication verification service cluster, and specifically includes the following steps:

[0089] Step 501, receive the target service data sent by any business service, and determine the deduplication verification identifier of the target service data.

[0090] The business service can refer to Figure 1 the relevant descriptions in it. It can generate a large amount of business data. To ensure the idempotency of business data processing, the business service sends the business data it generates to any node of the preset deduplication verification service cluster. In the embodiments of the present disclosure, the business data to be subjected to deduplication verification received by the preset deduplication verification service cluster is referred to as the target service data.

[0091] As an optional embodiment of the present disclosure, according to the data characteristics of the target service data, some identifying data in the target service data can be used as the deduplication verification identifier. Specifically, when receiving the target service data sent by any business service, it is determined whether the target service data includes a preset identification field. If the target service data includes a preset identification field, then the preset identification field of the target service data is determined as the deduplication verification identifier of the target service data. Exemplarily, if the target service data includes fields with uniqueness such as order numbers, user IDs, database primary keys, etc., any field with uniqueness can be selected as the deduplication verification identifier.

[0092] As another alternative embodiment of the present disclosure, if the target business data does not have a unique field such as an order number, a user ID, a database primary key, etc., the target business data can be processed, such as hash calculation, etc., and the finally processed value is used as the deduplication verification identifier. Specifically, when receiving the target business data sent by any business service, it is determined whether the target business data includes a preset identifier field. If the target business data does not have a preset identifier field, the target business data is converted into a preset data type, and then the target business data of the preset data type is subjected to hash calculation to obtain the hash value corresponding to the target business data, and the hash value corresponding to the target business data is determined as the deduplication verification identifier of the target business data. Exemplarily, the entire data content of the target business data can be converted into a JSON (JavaScript Object Notation) string, and then the JSON string of the target business data is subjected to hash calculation to obtain the MD5 (Message-Digest Algorithm 5) value as the deduplication verification identifier.

[0093] Step 502, based on the deduplication verification identifier of the target business data, determine whether the target business data exists in the preset database deduplication table by querying the target Bloom filter and the preset database deduplication table.

[0094] According to the embodiments of the present disclosure, the target Bloom filter can be determined first. Specifically, the false positive rate and the length of the bit array of the Bloom filter can be determined to obtain the target Bloom filter.

[0095] In some embodiments, the value of the false positive rate of the Bloom filter input by the user and the maximum value range of the number of business data of the business service can be received. According to the value of the false positive rate and the maximum value range of the number of business data of the business service, the length of the bit array of the Bloom filter is determined according to a preset calculation formula, and then based on the value of the false positive rate and the calculated length of the bit array of the Bloom filter, the target Bloom filter is obtained. Exemplarily, in the embodiments of the present disclosure, the false positive rate can be between 1% and 10%. Assume that the false positive rate P fp = 5% in the embodiments of the present disclosure. From the perspective of the actual application of the business service, the value range of the deduplication verification identifier (such as the user ID) is determined according to the user level. For example, the user ID range is 1 to 3,000,000, that is, the maximum value range of the number of business data. Then, the expected number n of user IDs placed in the target Bloom filter is 3,000,000; finally, the length m of the bit array of the Bloom filter is calculated according to the following formula to be 18,705,673, and the occupied memory space is 18,705,673 bit / (8 * 1024 * 1024) ≈ 2.23 MB. It can be seen that the occupied memory space is very small.

[0096] Exemplarily, the calculation formula for the Bloom filter bit array is, for example:

[0097]

[0098] where m is the Bloom filter bit array, n is the maximum value range of the number of service data, and P fp is the misjudgment rate value.

[0099] In some other embodiments, it is also possible to receive the component information of the Bloom filter, the value of the misjudgment rate, and the maximum number value of the service data of the service from the user input. Determine the Bloom filter components according to the component information of the Bloom filter, input the value of the misjudgment rate and the maximum number value of the service data of the service into the Bloom filter components to obtain the length of the bit array and the number of hash functions of the Bloom filter, and obtain the target Bloom filter according to the value of the misjudgment rate, the length of the bit array of the Bloom filter, and the number of hash functions. In this embodiment, the user can select the component information of the Bloom filter through the input device to determine the open-source Bloom filter component to be selected or used, and input the misjudgment rate, the maximum data volume value of the service data, etc. into the selected open-source Bloom filter component, then the information such as the length of the bit array and the number of hash functions of the Bloom filter can be obtained, so as to determine the target Bloom filter.

[0100] Those skilled in the art can understand that the lower the misjudgment rate, the longer the required bit array and the larger the storage space occupied. At the same time, if the misjudgment rate is too high, it will lead to a large number of false positives (judging that it exists but actually does not), and the value of using the Bloom filter will be lost. Therefore, in the embodiments of the present disclosure, the misjudgment rate is determined according to the actual application of the service. For example, when the misjudgment rate takes P fp = 5%, it can not only meet the actual application requirements of the service, but also ensure that the length of the required bit array is within a suitable range, does not occupy a large storage space, and the misjudgment rate is also within the suitable range of the Bloom filter, avoiding too many false positives, ensuring the deduplication verification efficiency of the Bloom filter while also ensuring the accuracy of its deduplication verification.

[0101] After determining the target Bloom filter, it is possible to combine the target Bloom filter and the preset database deduplication table to determine whether the target service data exists, that is, perform deduplication verification on the target service data, so as to ensure the idempotency of service data processing. According to the embodiments of the present disclosure, the target Bloom filter is used as the pre-mode of deduplication verification, that is, in the case where the target Bloom filter is unavailable, the mode of querying the preset database deduplication table is switched for deduplication verification.

[0102] Specifically, as an alternative embodiment, it is possible to first determine whether the target Bloom filter is normal. If the target Bloom filter is determined to be normal, the target Bloom filter is queried through the deduplication verification identifier to determine whether the target business data exists in the preset database deduplication table. Exemplarily, to implement querying the target Bloom filter to determine whether the target business data exists can be achieved using a Lua script (which is a lightweight and compact scripting language, written in standard C language and open-sourced in source code form, and is designed to be embedded in an application to provide flexible extension and customization functions for the application), ensuring the atomicity of the operation.

[0103] Preferably, it is possible to first determine whether the queried key (i.e., the target business data) exists in the target Bloom filter. If it does not exist, return 0, and store the key (i.e., the target business data) in the target Bloom filter. If it exists, then return 1.

[0104] After verifying that the target Bloom filter exists (returns 1), it is also necessary to query the preset database deduplication table for confirmation. If the target business data does not exist in the deduplication table, the target business data needs to be written into the deduplication table. Specifically, if it is determined that the target business data exists by querying the target Bloom filter, the preset database deduplication table is queried a second time to determine whether the target business data exists in the preset database deduplication table. Further, if it is determined that the target business data exists by querying the preset database deduplication table, it is determined that the target business data exists in the preset database deduplication table. For this embodiment, first determine whether the target business data exists through the target Bloom filter (specifically, the query process can be as shown in Figure 4 and will not be elaborated here). If it determines that the target business data exists, then make a second determination by querying the preset database deduplication table. This can reduce the number of times of querying the preset database deduplication table, reduce the query pressure on the preset database deduplication table, and compared with traversing a large amount of business data in the preset database deduplication table to determine whether the target business data exists, the Bloom filter only needs to query its bit array, and the query efficiency is higher; and in the case where the Bloom filter determines that the business data exists, since the Bloom filter may produce false positives (judging as existing but actually not existing), the database query method is used for secondary confirmation to ensure the accuracy of deduplication verification, thereby ensuring the idempotency of data processing.

[0105] As another alternative embodiment, in the case where the target Bloom filter is abnormal, it is determined whether the target service data exists by querying the preset database de-duplication table. Due to the complexity of the distributed system, special situations such as Redis cluster downtime or network anomalies may occur, resulting in the unavailability of the target Bloom filter. To ensure the availability and result accuracy of the de-duplication verification service, the embodiments of the present disclosure can be switched to the preset database de-duplication table mode, that is, directly query the preset database de-duplication table to confirm whether the target service data exists in the preset database de-duplication table.

[0106] Step 503, in the case where it is determined that the target service data does not exist in the preset database de-duplication table, perform target service processing on the target service data.

[0107] In the embodiments of the present disclosure, according to the specific service request corresponding to the target service data, the target service data can be subjected to target service processing through, for example, the server or processor corresponding to the service request to respond to the service request. For example, the target service data can be used to create a form for user order information. Correspondingly, in response to the target service data, that is, perform form creation processing on the user order information (target service data); for another example, the target service data is user order payment, then in response to the target service data, that is, complete the payment operation of the user order.

[0108] Step 504, store the target service data that has completed the target service processing into the preset database de-duplication table and write it into the target Bloom filter.

[0109] In the embodiments of the present disclosure, if it is determined that the target service data does not exist, it is necessary to store the target service data into the preset database de-duplication table and write it into the target Bloom filter to provide an accurate de-duplication verification data basis and ensure the idempotency of data processing. If the target Bloom filter determines that the target service data does not exist, first write it into the preset database de-duplication table, and then write it into the target Bloom filter. In the case of service anomalies (such as downtime), only the preset database de-duplication table will be written and not the target Bloom filter. As a preferred solution of the embodiments of the present disclosure, as long as it is determined that the target service data does not exist in the target Bloom filter, first write it into the target Bloom filter, and then store it in the preset database de-duplication table. In this way, the target Bloom filter may have misjudgments, but there will be no data errors, ensuring that the data is correctly stored, thereby ensuring the accuracy of the data and ensuring the idempotency of data processing.

[0110] In Figure 5 Based on the method embodiments shown, Figure 6The flowchart of the data processing method provided by an embodiment of the present disclosure. It can be understood that the business data of business services may increase gradually, or the value range of business data may be relatively large. In these cases, the length of the bit array of the target Bloom filter calculated will be too long, which may cause the Big Key problem in the Redis cluster (that is, the value of a certain key is too large, resulting in occupying a large amount of memory space and affecting the system performance). In view of this, the embodiment of the present disclosure provides another embodiment of a data processing method as shown in Figure 6 to achieve the segmentation of the Bloom filter and create target Bloom filters for multiple data ranges.

[0111] As shown in Figure 6 , the data processing method provided by the embodiment of the present disclosure may further include the following steps 601 to 603. Among them, steps 601 to 603 may be before the steps 501 to 503 as shown in Figure 5 , that is, first determine the target Bloom filter, and on the basis of the determined target Bloom filter, continue to implement the steps 501 to 503 as shown in Figure 5 . Specifically, steps 601 to 603 are implemented as follows:

[0112] Step 601, determine whether the business data of any business service meets the preset data condition.

[0113] In the embodiment of the present disclosure, the preset data condition is, for example, that the business data of any business service has an increasing trend, or for another example, there is an obvious discriminative field in the business data of any business service. Specifically, for example, the business data of any business service has an increasing trend, or a certain field (such as user ID) in the business data has an increasing trend, or there is an obvious discriminative field such as provinces, cities or large administrative regions in the business data.

[0114] Step 602, when it is determined that the preset data condition is met, divide the business data of the business service according to a preset range to obtain multiple data ranges.

[0115] As an embodiment of the present disclosure, if the business data of any business service has an increasing trend or a certain field in the business data has an increasing trend, it can be segmented according to the data range, for example, segmented according to the user ID range or according to the creation time range, and divided into multiple data ranges.

[0116] As another embodiment of the present disclosure, if there is an obvious discriminative field in the business data, such as provinces, cities or large administrative regions, the range can be divided according to each province, city or large administrative region to obtain multiple data ranges.

[0117] Step 603: Create a corresponding target Bloom filter for each data range.

[0118] Exemplarily, if the data is segmented by time range, such as the business data of a certain business service in one year, it can be divided into four data ranges corresponding to spring, summer, autumn, and winter according to seasons, and four target Bloom filters corresponding to the four seasons are created respectively. Or, if it is divided by obvious differentiating fields such as provinces and cities, assuming there is business data in three provinces A, B, and C, three target Bloom filters corresponding to these three provinces are created respectively.

[0119] As an optional embodiment of the present disclosure, in order to ensure the persistent storage of the target Bloom filter, the expiration time setting of the target Bloom filter is cancelled. For example, no expiration time is set for the created Redis Bloom filter.

[0120] After the target Bloom filters are created for each data range, the following can be executed Figure 5 In the illustrated embodiment, that is, when the target business data of this data range is received, duplicate elimination verification is preferentially performed through the target Bloom filter corresponding to the data range, that is, it is confirmed whether the target business data exists, and when it is confirmed that the target business data in this range does not exist, the target business data is written into the target Bloom filter.

[0121] In Figure 5 and Figure 6 Based on the method embodiments shown, Figure 7 This is a flowchart of a data processing method provided by another embodiment of the present disclosure. Redis cluster anomalies may cause the target business data to be unable to be inserted into the target Bloom filter. After the subsequent Redis cluster is restored, due to the target Bloom filter missing the new business data during the Redis cluster anomaly period, there will inevitably be a situation where business data is misjudged as non-existent. To solve this technical problem, the embodiments of the present disclosure identify the business data written into the preset database duplicate elimination table during the Redis cluster anomaly period, and after the Redis cluster is restored, these business data are compensated and inserted into the target Bloom filter.

[0122] As Figure 7 shown, the data processing method provided by the embodiments of the present disclosure may further include the following steps:

[0123] Step 701: Determine whether the target Bloom filter is abnormal.

[0124] In the Redis cluster, special situations such as downtime or network anomalies may occur, resulting in the unavailability of the target Bloom filter. At this time, when any node in the preset duplicate-checking service cluster queries the target Bloom filter or writes the target service data to the target Bloom filter, failures or errors will occur. At this time, the node of the preset duplicate-checking service cluster can determine that the target Bloom filter is in an abnormal state.

[0125] Step 702, in the case of determining that the target Bloom filter is abnormal, add an anomaly identification flag at the preset position of the duplicate-checking table where the target service data in the preset database duplicate-checking table is located.

[0126] As described in the foregoing embodiments, in the case of an abnormal target Bloom filter, switch to the method of querying the preset database duplicate-checking table to determine whether the target service data exists. If it does not exist, the preset database duplicate-checking table will store the target service data in its duplicate-checking table. In the embodiments of the present disclosure, an identification field can be added to the duplicate-checking table. Each time a target service data is stored, an anomaly identification flag is added to the identification field of the corresponding data row to be used to identify whether the target service data is written to the target Bloom filter. Exemplarily, a new field bfFlag can be added to the preset database duplicate-checking table, where bfFlag = false is the anomaly identification flag, used to indicate that the current data row has not been written to the target Bloom filter, and bfFlag = true is the non-anomaly identification flag, used to indicate that the current data row has been written to the target Bloom filter.

[0127] Step 703, in the mode of querying whether the target service data exists through the preset database duplicate-checking table, traverse the data rows in the preset database duplicate-checking table that have the anomaly identification flag.

[0128] In the embodiments of the present disclosure, a timing task can be set. When the timing task runs, first determine whether the current duplicate-checking mode is the query mode of the preset database duplicate-checking table. If so, traverse or scan the data rows at the preset position in the preset database duplicate-checking table or the data rows with bfFlag = false in the field to determine which target service data has not been written to the target Bloom filter.

[0129] Step 704, write the target service data of the data row corresponding to the anomaly identification flag to the target Bloom filter.

[0130] After confirming the data rows at the preset position in the preset database duplicate-checking table or the data rows with bfFlag = false in the field, write the target service data corresponding to these data rows to the target Bloom filter. After writing to the target Bloom filter, bfFlag = true can be modified until there are no data rows with bfFlag = false.

[0131] It can be understood that the data with bfFlag = false in the preset database deduplication table are those that were not added to the target Bloom filter during the abnormal period. Therefore, the embodiments of the present disclosure provide a technical solution for data compensation of the target Bloom filter. By identifying bfFlag = false, the target service data corresponding to the data row is compensated and written into the target Bloom filter. If the write is successful, it is determined that the target Bloom filter resumes normal operation. In subsequent deduplication verification processing, the query mode of the target Bloom filter is still preferentially used, so that the data in the target Bloom filter maintains a certain data accuracy, providing guarantee for the query accuracy rate of the target Bloom filter, reducing the false positive rate and error rate of the Bloom filter, and providing data guarantee for the idempotency of data processing.

[0132] When the preset deduplication verification service cluster node communicates with the target Bloom filter, if the target Bloom filter is abnormal, the preset deduplication verification service cluster node may still be writing the target service data or querying the target service data, etc., which will cause a large number of data processing failures. In view of this, the embodiments of the present disclosure provide a technical solution for a global notification mechanism for the deduplication verification mode, so that when communication between a certain node of the preset deduplication verification service cluster and the Redis cluster (or the target Bloom filter) is abnormal, the data of this node cannot be written into the target Bloom filter. At this time, the verification result of the target Bloom filter is inaccurate. Therefore, while switching this node to the mode of querying data from the preset database deduplication table, it is also necessary to notify other nodes of the preset deduplication verification service cluster to switch to the database mode. Specifically, in the case of determining that the target Bloom filter is abnormal, the information about the abnormality of the target Bloom filter is sent to the configuration center, so that the configuration center changes the preset configuration item identifier to notify other nodes of the preset deduplication verification service cluster of the abnormal working state of the target Bloom filter, where the preset configuration item identifier is used to mark the working state of the target Bloom filter. Exemplarily, a configuration item identifier for data deduplication verification mode (i.e., the preset configuration item identifier) can be added to the configuration center (such as a ZooKeeper cluster). All nodes of the preset deduplication verification service cluster listen for changes to this configuration item to determine whether to query the target Bloom filter for deduplication verification. If an abnormality occurs when a certain node of the preset deduplication verification service cluster writes to the target Bloom filter, this configuration item in the configuration center is modified, and the configuration center will notify all nodes in the preset deduplication verification service cluster in real time to switch to the preset database deduplication table for data deduplication verification.

[0133] It can be understood that after the target Bloom filter resumes normal working state, for example, when executing Figure 6If step 704 shown above is successful, it is determined that the target Bloom filter resumes normal operation, and the nodes of the corresponding preset duplicate-checking service cluster also send the information that the target Bloom filter resumes normal operation to the configuration center, so that the configuration center changes its preset configuration item identifier and notifies other nodes of the preset duplicate-checking service cluster, causing other nodes to also switch to the target Bloom filter for data duplicate-checking.

[0134] Through the above embodiments, the abnormality of the target Bloom filter is avoided, which affects the accuracy of the duplicate-checking result. After the target Bloom filter is normal, even if the data duplicate-checking mode is switched to the target Bloom filter, the efficiency of duplicate-checking is guaranteed, and the query pressure on the duplicate table in the preset database is reduced.

[0135] In summary, through the technical solution of the embodiments of the present disclosure, when receiving the target service data sent by the service, the target Bloom filter and the preset database duplicate table can be queried through the duplicate-checking identifier of the target service data to determine whether the target service data exists in the preset database duplicate table. If the target service data does not exist, after performing the target service processing on the target service data, the target service data is stored in the preset database duplicate table and written into the target Bloom filter. In this way, before data is written or stored, the method of querying through the Bloom filter and / or the method of querying the duplicate table through the preset database duplicate table are used to query whether the data already exists, effectively ensuring the idempotency of business data processing and avoiding the problem of data duplication. In addition, in the case of querying through the Bloom filter, since the Bloom filter can determine whether the target service data exists only through its bit array without having to query a large amount of data in the duplicate table, the frequency of querying the duplicate table can be greatly reduced, the query pressure on the database can be effectively reduced, and it can quickly determine whether the data already exists, improving the query efficiency.

[0136] Correspondingly, the embodiments of the present disclosure also provide a corresponding apparatus embodiment to the foregoing method embodiment. Figure 8 It is a structural block diagram of a data processing apparatus provided by an embodiment of the present disclosure. As Figure 8 shown, the data processing apparatus can be configured in the nodes of the preset duplicate-checking service cluster. Specifically, the apparatus can include a check identifier determination module 801, a query module 802, a service processing module 803, and a data writing module 804.

[0137] The check identifier determination module 801 can be used to receive the target service data sent by any service and determine the duplicate-checking identifier of the target service data.

[0138] The query module 802 can be used to determine whether the target service data exists in the preset database de-duplication table by querying the target Bloom filter and the preset database de-duplication table based on the de-duplication verification identifier of the target service data.

[0139] The service processing module 803 can be used to perform target service processing on the target service data when it is determined that the target service data does not exist in the preset database de-duplication table.

[0140] The data writing module 804 can be used to store the target service data that has completed the target service processing into the preset database de-duplication table and write it into the target Bloom filter.

[0141] Through the technical solution of the embodiments of the present disclosure, when receiving the target service data sent by the service, the de-duplication verification identifier of the target service data can be used to query the target Bloom filter and the preset database de-duplication table to determine whether the target service data exists in the preset database de-duplication table. If the target service data does not exist, after performing target service processing on the target service data, the target service data is stored in the preset database de-duplication table and written into the target Bloom filter. In this way, before data writing or storage, the method of querying using a Bloom filter and / or querying the de-duplication table through the preset database de-duplication table is adopted to query whether the data already exists, effectively ensuring the idempotency of service data processing and avoiding the problem of data duplication. In addition, in the case of querying through the Bloom filter, since the Bloom filter can determine whether the target service data exists only through its bit array without having to query a large amount of data in the de-duplication table, the frequency of querying the de-duplication table can be greatly reduced, the query pressure on the database can be effectively reduced, and it can quickly determine whether the data already exists, improving the query efficiency.

[0142] Furthermore, Figure 9 is a structural block diagram of a data processing device provided by another embodiment of the present disclosure. As Figure 9 shown, in the data processing device provided by the embodiments of the present disclosure, it may further include:

[0143] The verification identifier determination module 801 may include an identifier field determination unit 8011, a verification identifier first determination unit 8012, a data type conversion unit 8013, a hash calculation unit 8014, and a verification identifier second determination unit 8015, where:

[0144] The identifier field determination unit 8011 is configured to receive the target service data sent by any service and determine whether the target service data includes a preset identifier field;

[0145] The verification identifier first determination unit 8012 is configured to, if the target service data includes a preset identifier field, determine the preset identifier field of the target service data as the deduplication verification identifier of the target service data;

[0146] The data type conversion unit 8013 is configured to, if the target service data does not have a preset identifier field, convert the target service data into a preset data type;

[0147] The hash calculation unit 8014 is configured to perform hash calculation on the target service data of the preset data type to obtain a hash value corresponding to the target service data;

[0148] The verification identifier second determination unit 8015 is configured to determine the hash value corresponding to the target service data as the deduplication verification identifier of the target service data.

[0149] The query module 802 may include a first determination unit 8021, a first query unit 8022, a second query unit 8023, and a second determination unit 8024, where

[0150] The first determination unit 8021 is configured to determine whether the target Bloom filter is normal.

[0151] The first query unit 8022 is configured to, if the target Bloom filter is normal, query the target Bloom filter through the deduplication verification identifier.

[0152] The second determination unit 8023 is configured to, if it is determined through querying the target Bloom filter that the target service data does not exist in the preset database deduplication table, query the preset database deduplication table again to determine whether the target service data exists in the preset database deduplication table.

[0153] The second query unit 8024 is configured to, if it is determined through querying the preset database deduplication table that the target service data exists, determine that the target service data exists in the preset database deduplication table.

[0154] In some embodiments, the second query unit 8023 may also be configured to, when the target Bloom filter is abnormal, determine whether the target service data exists in the preset database deduplication table by querying the preset database deduplication table.

[0155] In other embodiments, a data processing device according to an embodiment of the present disclosure may further include a Bloom filter determination module 804, which is configured to obtain the target Bloom filter by determining the false positive rate and the bit array length of the Bloom filter.

[0156] As an alternative embodiment, the Bloom filter determination module 804 may include a first receiving unit 8041, a bit array length determination unit 8042, and a Bloom filter determination unit 8043, where:

[0157] The first receiving unit 8041 is configured to receive the value of the false positive rate of the Bloom filter and the maximum value range of the number of service data of the service from user input;

[0158] The bit array length determination unit 8042 is configured to determine the bit array length of the Bloom filter according to the value of the false positive rate and the maximum value range of the number of service data of the service according to a preset calculation formula;

[0159] The first Bloom filter determination unit 8043 is configured to obtain the target Bloom filter based on the value of the false positive rate and the calculated bit array length of the Bloom filter.

[0160] As another alternative embodiment, the Bloom filter determination module 804 may further include

[0161] The second receiving unit 8044 is configured to receive the component information of the Bloom filter, the value of the false positive rate, and the maximum number value of the service data of the service from user input;

[0162] The component determination unit 8045 is configured to determine the Bloom filter components according to the component information of the Bloom filter;

[0163] The Bloom parameter determination unit 8046 is configured to input the value of the false positive rate and the maximum number value of the service data of the service into the Bloom filter components to obtain the bit array length and the number of hash functions of the Bloom filter;

[0164] The second Bloom filter determination unit 8047 is configured to obtain the target Bloom filter according to the value of the false positive rate, the bit array length of the Bloom filter, and the number of hash functions.

[0165] As another alternative embodiment, a data processing device according to an embodiment of the present disclosure may further include:

[0166] The data condition judgment module 805 is configured to determine whether the service data of any service satisfies a preset data condition;

[0167] The data range division module 806 is configured to, when it is determined that the preset data condition is satisfied, divide the service data of the service into data ranges according to a preset range to obtain a plurality of data ranges;

[0168] A creation module 807 is configured to create a corresponding target Bloom filter for each data range, so as to write the target service data into the target Bloom filter when the target service data in this range is received and it is confirmed that the target service data in this range does not exist.

[0169] As another alternative embodiment, a data processing device according to an embodiment of the present disclosure may further include:

[0170] An exception determination module 808 is configured to determine whether the target Bloom filter is abnormal;

[0171] An exception identification module 809 is configured to, when it is determined that the target Bloom filter is abnormal, add an exception identification flag at a preset position in the deduplication table where the target service data in the preset database deduplication table is located;

[0172] A deduplication table traversal module 810 is configured to traverse the data rows with the exception identification flag in the preset database deduplication table in a mode of querying whether the target service data exists through the preset database deduplication table;

[0173] Based on the foregoing modules, a data writing module 803 may further be configured to write the target service data of the data row corresponding to the exception identification flag into the target Bloom filter.

[0174] As another alternative embodiment, a data processing device according to an embodiment of the present disclosure may further include:

[0175] An exception information sending module 811 is configured to, when it is determined that the target Bloom filter is abnormal, send the information that the target Bloom filter is abnormal to a configuration center, so that the configuration center changes a preset configuration item flag to notify other preset deduplication verification service cluster nodes of the working state that the target Bloom filter is abnormal, and the preset configuration item flag is used to mark the working state of the target Bloom filter.

[0176] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0177] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.

[0178] Next, reference Figure 10 is made to describe an electronic device according to an embodiment of the present disclosure. The electronic device can be any one or both of the first device and the second device, or a stand-alone device independent of them. The stand-alone device can communicate with the first device and the second device to receive the input signals collected from them.

[0179] Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated. As Figure 10 shown, the electronic device includes one or more processors and a memory.

[0180] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.

[0181] The memory can store one or more computer program products. The memory can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products can be stored on the computer-readable storage media, and the processor can run the computer program products to implement the data processing methods of the various embodiments of the present disclosure described above and / or other desired functions.

[0182] In one example, the electronic device can further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0183] In addition, the input device can further include, for example, a keyboard, a mouse, and so on.

[0184] The output device can output various information to the outside, including the determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.

[0185] Of course, for simplicity, Figure 10 only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device can further include any other appropriate components.

[0186] In addition to the above methods and devices, the embodiments of the present disclosure can also be a computer program product, which includes computer program instructions. When the computer program instructions are run by the processor, the processor is caused to execute the steps in the data processing methods according to the various embodiments of the present disclosure described in the above part of this specification.

[0187] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0188] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the data processing method according to various embodiments of the present disclosure described in the foregoing part of this specification.

[0189] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0190] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-described specific details are only for illustrative and facilitating understanding purposes and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0191] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with each other.

[0192] The methods and apparatuses of the present disclosure may be implemented in many ways. For example, the methods and apparatuses of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustration only, and the steps of the methods of the present disclosure are not limited to the order specifically described above unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.

[0193] It should also be noted that in the apparatuses, devices, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0194] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0195] The above description has been presented for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A data processing method, characterized in that: The method is applied to a preset deduplication verification service cluster node, and the method includes: Receive target business data sent by any business service, and determine a deduplication verification identifier of the target business data; Based on the deduplication verification flag of the target business data, determining whether the target business data exists in the preset database deduplication table by querying a target Bloom filter and a preset database deduplication table; In the case where it is determined that the target business data does not exist in the preset database deduplication table, performing target business processing on the target business data; The target business data that has completed the target business processing is stored in the preset database deduplication table and written into the target Bloom filter.

2. The method according to claim 1, characterized in that The deduplication verification flag based on the target business data, determining whether the target business data exists in the preset database deduplication table by querying a target Bloom filter and a preset database deduplication table, includes: Determining whether the target Bloom filter is normal; If the target Bloom filter is normal, querying the target Bloom filter through the deduplication verification identifier; If it is determined by querying the target Bloom filter that the target business data does not exist in the preset database deduplication table, querying the preset database deduplication table for a second time to determine whether the target business data exists in the preset database deduplication table; If it is determined that the target business data exists by querying the preset database deduplication table, it is determined that the target business data exists in the preset database deduplication table.

3. The method according to claim 1, characterized in that The deduplication verification flag based on the target business data, determining whether the target business data exists in the preset database deduplication table by querying a target Bloom filter and a preset database deduplication table, includes: In the case where the target Bloom filter is abnormal, whether the target business data exists in the preset database deduplication table is determined by querying the preset database deduplication table.

4. The method according to claim 1, characterized in that: The receiving target business data sent by any business service and determining a deduplication verification identifier of the target business data includes: Receive target business data sent by any business service, and determine whether the target business data includes a preset identification field; If the target service data includes a preset identification field, determining the preset identification field of the target service data as a deduplication verification identification of the target service data; If the target service data does not have a preset identification field, converting the target service data into a preset data type; Performing hash calculation on the target business data of the preset data type to obtain a hash value corresponding to the target business data; The hash value corresponding to the target business data is determined as a deduplication verification identifier of the target business data.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: The target Bloom filter is obtained by determining the false positive rate and the bit array length of the Bloom filter.

6. The method according to claim 5, characterized in that The step of obtaining the target Bloom filter by determining the false positive rate and the bit array length of the Bloom filter includes: Receiving a value of the false positive rate of the Bloom filter and a maximum value range of the amount of business data of the business service inputted by a user; Determine the bit array length of the Bloom filter according to a preset calculation formula based on the value of the false positive rate and the maximum value range of the amount of business data of the business service; The target Bloom filter is obtained based on the value of the false positive rate and the calculated bit array length of the Bloom filter.

7. The method according to claim 5, characterized in that The step of obtaining the target Bloom filter by determining the false positive rate and the bit array length of the Bloom filter includes: Receiving component information of the Bloom filter, a value of the false positive rate, and a maximum number of business data of the business service inputted by a user; Determine a Bloom filter component according to component information of the Bloom filter; Inputting the value of the false positive rate and the maximum number of business data of the business service into the Bloom filter component to obtain the bit array length and the number of hash functions of the Bloom filter; The target Bloom filter is obtained according to the value of the false positive rate, the bit array length of the Bloom filter and the number of hash functions.

8. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Determine whether the business data of any business service meets the preset data conditions; When it is determined that the preset data condition is met, the business data of the business service is divided according to the preset range to obtain multiple data ranges; A corresponding target Bloom filter is created for each data range, so that when target service data in the range is received and it is confirmed that the target service data in the range does not exist, the target service data is written into the target Bloom filter.

9. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Determining whether the target Bloom filter is abnormal; When it is determined that the target Bloom filter is abnormal, an abnormal identification mark is added at a preset position of the deduplication table where the target business data of the preset database deduplication table is located; In a mode of querying whether the target business data exists through the preset database deduplication table, traversing the data rows having the abnormal identification identifier in the preset database deduplication table; The target business data of the data row corresponding to the abnormal identification identifier is written into the target Bloom filter.

10. The method according to claim 9, characterized in that The method further comprises: When it is determined that the target Bloom filter is abnormal, information about the target Bloom filter abnormality is sent to the configuration center, so that the configuration center changes the preset configuration item identifier to notify other preset deduplication verification service cluster nodes of the abnormal working status of the target Bloom filter, and the preset configuration item identifier is used to mark the working status of the target Bloom filter.

11. A data processing device, characterized in that: The device is configured in a preset deduplication verification service cluster node, and the device includes: A verification mark determination module is used to receive target business data sent by any business service and determine a deduplication verification mark of the target business data; A query module, configured to determine whether the target business data exists in the preset database deduplication table by querying a target Bloom filter and a preset database deduplication table based on the deduplication verification identifier of the target business data; A business processing module, configured to perform target business processing on the target business data when it is determined that the target business data does not exist in the preset database deduplication table; A data writing module is used to store the target business data that has completed the target business processing into the preset database deduplication table and write it into the target Bloom filter.

12. An electronic device, characterized in that: include: a memory for storing a computer program product; A processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, it implements the method described in any one of claims 1 to 10.

13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 10 is implemented.

14. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 10 is implemented.