Data storage and real-time duplicate removal method and device, equipment and storage medium
By using the Bitmap data structure and multi-threaded concurrency technology in Redis, the problems of low list deduplication efficiency and repeated data storage in existing technologies are solved, and efficient and conflict-free list deduplication is achieved, which is suitable for high-concurrency scenarios and new business expansion.
Patent Information
- Application Number
- CN202510747167.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-05
AI Technical Summary
The existing list deduplication solutions based on Spring Batch and database queries have low timeliness, concurrency control issues, and insufficient scalability, resulting in duplicate data entry and performance bottlenecks in high-concurrency scenarios.
The Redis key-value storage database is used, and the Bitmap data structure and the "cut head and tail" strategy are used to build the index. Multi-threaded concurrency and Pipeline technology are combined for deduplication, and atomic instructions are used to achieve zero-conflict deduplication and avoid duplicate writes.
It achieves efficient list deduplication, supports large-scale list storage, is suitable for high-concurrency scenarios, and has flexible expansion capabilities to adapt to new business needs.
Smart Images

Figure CN120596482A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data storage technology, and in particular to methods, devices, equipment, and storage media for data storage and real-time deduplication. Background Art
[0002] Currently, most list deduplication solutions are based on Spring Batch (a batch processing framework) and database query solutions. The core modules typically include a list acquisition and verification module, a database deduplication query module, and a database storage module. The list acquisition and verification module consumes Kafka streaming data or reads batch list files to obtain the customer list to be stored, and performs basic field verification on the list (such as format and integrity). The database deduplication query module queries the database for the existence of the customer list item by item and inserts the non-duplicate lists into the database in batches. The database storage module uses a relational database to store the list data, relying on the database's unique index or transaction mechanism to ensure data consistency.
[0003] However, this approach has certain limitations. First, it suffers from low timeliness. Single database queries result in high latency. Each list entry requires a separate database query. The cumulative network latency and database response time result in a long time-consuming process for entering large lists. Second, there are concurrency control issues. When querying the existence of a list, if duplicate lists exist in the same batch, they may be incorrectly detected as "not existing" and written repeatedly because they have not yet been inserted into the database during the writer phase, creating the risk of duplicate entries within the same batch. If lists from different batches overlap in processing time (for example, batch B begins querying while batch A is not yet inserted), this can lead to duplicate entries across batches, creating the risk of cross-batch duplication. In high-concurrency scenarios, when multiple threads / processes attempt to insert the same list simultaneously, the lack of atomicity between SELECT and INSERT operations can lead to "dirty writes," resulting in duplicate entries. Third, scalability is insufficient. Due to its reliance on relational database query and write operations, the database becomes a performance bottleneck in high-concurrency, high-data-volume scenarios, preventing horizontal scalability to meet growing business demands.
[0004] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a data storage and real-time deduplication method, device, equipment and storage medium, aiming to solve the technical problems in the prior art that traditional deduplication solutions have low efficiency and data is easily duplicated.
[0006] To achieve the above objectives, the present application provides a data storage and real-time deduplication method, the method comprising:
[0007] Obtain an existing valid list in a target database, and determine index key data and an offset of the index key data based on a customer number in the existing valid list;
[0008] Based on the index key data and the offset of the index key data, storing the existing valid list in a key-value storage database;
[0009] Obtain a list of customers to be stored in real time, and determine the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list;
[0010] Deduplication of the list to be stored is performed based on the offset of the key data to be stored and the offset of the index key data to determine a non-duplicate list;
[0011] The non-repeated list is stored in the target database, and the key-value storage database is updated.
[0012] In one embodiment, the step of obtaining an existing valid list in the target database and determining index key data and an offset of the index key data based on a customer number in the existing valid list includes:
[0013] Determine the truncated length and the trimmed length based on the customer number of the existing valid list, and determine the middle offset based on the truncated length and the trimmed length;
[0014] Extracting a first index key from the customer numbers in the existing valid list based on the truncated length, and extracting a second index key corresponding to the first index key from the customer numbers in the existing valid list based on the truncated length;
[0015] Concatenate the first index key with the second index key corresponding to the first index key to obtain the index key data;
[0016] Based on the middle offset bit, the offset of the index key data is extracted from the customer number of the existing valid list, and the offset of the index key data is updated to the first value.
[0017] In one embodiment, the step of obtaining a list of customers to be stored in real time and determining the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list of customers to be stored includes:
[0018] Extracting a first waiting-for-warehouse key from the customer number on the waiting-for-warehouse list based on the truncated length, and extracting a second waiting-for-warehouse key corresponding to the first waiting-for-warehouse key from the customer number on the waiting-for-warehouse list based on the truncated length;
[0019] Concatenate the first to-be-warehouse key and the second to-be-warehouse key corresponding to the first to-be-warehouse key to obtain the to-be-warehouse key data;
[0020] Based on the middle offset bit, the offset of the key data to be stored is extracted from the customer number of the list to be stored, and the offset of the key data to be stored is updated to the first value.
[0021] In one embodiment, the step of deduplicating the list to be stored based on the offset of the key data to be stored and the offset of the index key data to determine a non-duplicate list includes:
[0022] When matching key data of the key data to be stored exists in the index key data, obtaining an offset of the matching key data;
[0023] When the offset of the matching key data is different from the offset of the to-be-stored key data, determining that the to-be-stored list corresponding to the to-be-stored key data is a non-duplicate list;
[0024] When the offset of the matching key data is the same as the offset of the to-be-stored key data, it is determined that the to-be-stored list corresponding to the to-be-stored key data is a duplicate list.
[0025] In one embodiment, after the step of storing the non-duplicate list in the target database and updating the key-value storage database, the step further includes:
[0026] Obtain the invalid list of the target database on that day, and update the offset of the index key data corresponding to the invalid list of the day in the key-value storage database to a second value.
[0027] In one embodiment, the method further comprises:
[0028] Calculating a cyclic redundancy check value of the key data to be stored;
[0029] Determine a target value based on a cyclic redundancy check value of the key data to be stored and the number of hash slots;
[0030] Determining a target hash slot based on the target value;
[0031] The storage node where the target hash slot is located is used as the target node of the list to be stored;
[0032] Based on the target node of the to-be-entered list, a corresponding execution thread is allocated for deduplication of the to-be-entered list, wherein the to-be-entered list of the same target node uses the same execution thread for deduplication.
[0033] In one embodiment, after the step of storing the non-duplicate list in the target database and updating the key-value storage database, the step further includes:
[0034] Record the number of lists to be stored, the number of non-duplicate lists, and performance indicators, wherein the performance indicators include at least processing time and memory usage;
[0035] Determine the number of duplicate lists to be removed based on the number of the lists to be stored and the number of the non-duplicate lists;
[0036] Determine a deduplication rate based on the deduplication quantity and the quantity of the list to be stored;
[0037] When the deduplication rate is greater than a preset deduplication threshold or the performance indicator meets the performance optimization condition, a warning message is output.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a data storage and real-time deduplication device, which includes:
[0039] A data deduplication module is used to obtain an existing valid list in the target database, and determine index key data and an offset of the index key data based on the customer number of the existing valid list;
[0040] The data deduplication module is further configured to store the existing valid list into a key-value storage database based on the index key data and the offset of the index key data;
[0041] The data deduplication module is further used to obtain a list of customers to be stored in real time, and determine the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list of customers to be stored;
[0042] The data deduplication module is further configured to deduplicate the list to be stored based on the offset of the key data to be stored and the offset of the index key data, and determine a non-duplicate list;
[0043] The data writing module is used to store the non-repetitive list into the target database and update the key-value storage database.
[0044] In addition, to achieve the above-mentioned purpose, the present application also proposes a data storage and real-time deduplication device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the data storage and real-time deduplication method as described above.
[0045] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the data storage and real-time deduplication method as described above are implemented.
[0046] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the data storage and real-time deduplication method as described above.
[0047] The present application provides a data storage and real-time deduplication method, which obtains an existing valid list in a target database, determines index key data and an offset of the index key data based on the customer number of the existing valid list; stores the existing valid list in a key-value storage database based on the index key data and the offset of the index key data; obtains a list to be stored in real time, determines key data to be stored and an offset corresponding to the key data to be stored based on the customer number of the list to be stored; deduplicates the list to be stored based on the offset of the key data to be stored and the offset of the index key data, and determines a non-duplicate list; stores the non-duplicate list in the target database, and updates the key-value storage database. This application uses the Bitmap data structure for storage in Redis (key-value storage database), and uses the "cut-off" strategy to construct Redis key names to reduce storage space. After the index is established, atomic instructions can be used to achieve zero-conflict deduplication and avoid repeated writing. At the same time, multi-threaded concurrency and Pipeline technology are used to improve deduplication efficiency, which can support the deduplication operation of large-scale lists in the warehousing and can be applied to high-concurrency scenarios. In addition, the template mode is used to encapsulate the deduplication logic, which can be flexibly expanded and can be used out of the box for new business scenarios, solving the technical problems of low deduplication efficiency of traditional solutions and easy repeated warehousing of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0050] Figure 1 This is a flow chart of the first embodiment of the data storage and real-time deduplication method of the present application;
[0051] Figure 2This is a schematic diagram of the multi-threaded Pipeline query optimization of the data storage and real-time deduplication method provided in Example 1 of the present application;
[0052] Figure 3 This is a flow chart of the second embodiment of the data storage and real-time deduplication method of this application;
[0053] Figure 4 A schematic diagram of a simplified flow chart of the data storage and real-time deduplication method provided in Example 2 of the present application;
[0054] Figure 5 This is a schematic diagram of the module structure of the data storage and real-time deduplication device according to an embodiment of the present application;
[0055] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the data storage and real-time deduplication method in the embodiment of the present application.
[0056] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0057] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0058] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0059] The main solution of the embodiment of the present application is: obtain the existing valid list in the target database, determine the index key data and the offset of the index key data based on the customer number of the existing valid list; store the existing valid list in the key-value storage database based on the index key data and the offset of the index key data; obtain the list to be stored in real time, determine the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list to be stored; deduplicate the list to be stored based on the offset of the key data to be stored and the offset of the index key data, and determine the non-duplicate list; store the non-duplicate list in the target database, and update the key-value storage database.
[0060] Currently, each list in the stored list file needs to query the database, which makes the list storage process time-consuming and unable to meet the business's timeliness requirements for the list; concurrency control problems may occur in the same batch, different batches, and high-concurrency scenarios; when new scenarios are added, the amount of data continues to increase, and scalability is insufficient.
[0061] This application provides a solution that uses the Bitmap data structure for storage in Redis, uses the "cut head and tail" strategy to construct Redis key names, reduces storage space, and can use atomic instructions to achieve zero-conflict deduplication after indexing to avoid repeated writing. At the same time, it improves deduplication efficiency through multi-threaded concurrency and Pipeline technology, can support the deduplication operation of large-scale lists in the warehousing, and can be applied to high-concurrency scenarios. In addition, the template mode is used to encapsulate the deduplication logic, which can be flexibly expanded and can be used out of the box for new business scenarios, solving the technical problems of low deduplication efficiency of traditional solutions and easy repeated warehousing of data.
[0062] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of performing the aforementioned functions, data storage and real-time deduplication equipment, etc., and this embodiment does not specifically limit this. The following uses data storage and real-time deduplication equipment as an example to illustrate this embodiment and the following embodiments.
[0063] The present invention provides a method for data storage and real-time deduplication. Figure 1 , Figure 1 This is a flow chart of the first embodiment of the data storage and real-time deduplication method of the present application.
[0064] In this embodiment, the data storage and real-time deduplication method includes steps S10 to S50:
[0065] Step S10, obtaining an existing valid list in the target database, and determining index key data and an offset of the index key data based on the customer number in the existing valid list;
[0066] It should be noted that the target database, i.e., the database originally used to store the list, is typically a relational database, and this is not specifically limited. The existing valid list is the currently valid list. The number of existing valid lists is determined based on actual circumstances, and this embodiment does not specifically limit this. Generally speaking, each existing valid list contains some customer information (fields), such as customer number, customer type, etc., and this embodiment does not specifically limit this. The existing valid list in the target database can be obtained through Gemini Quick Transfer, and this embodiment does not specifically limit this.
[0067] In addition, it should be noted that in order to avoid duplicate data entry, when new lists are entered into the database, it is usually necessary to establish corresponding indexes in order to deduplicate these new lists. This embodiment requires writing the existing valid lists into Redis (key-value storage database) to establish a Bitmap index, and using Redis's Bitmap data structure to implement atomic deduplication query and storage. Redis uses key-value pairs to store data. The index key data, that is, the key (Key) used when the existing valid list is stored in Redis, can be used for indexing. The index value data, that is, the value (Value) used when the existing valid list is stored in Redis, is usually the customer information contained in each existing valid list.
[0068] It is understandable that the storage of the existing valid list in Redis adopts a "cut-off" strategy, extracting the bits in front of and behind the customer number according to a certain length as the Redis key, that is, the index key data, and the remaining middle bits as the (bit) offset, thereby reducing storage space and avoiding single keys being too long, distributing the keys to different Redis nodes, and improving read and write concurrency capabilities.
[0069] Furthermore, after the step of obtaining the existing valid list in the target database, the step also includes: performing an integrity check on the existing valid list, and after the existing valid list passes the integrity check, executing the step of determining the index key data and the offset of the index key data based on the customer number of the existing valid list.
[0070] It is understandable that after obtaining the existing valid list, the integrity of the fields needs to be verified. Different fields can be verified in different ways, for example: using regular expressions to verify the format of the customer number, checking whether the business type is within the legal range, etc. Each field in the list is verified one by one. If all fields pass the verification, the next step of data storage is carried out. If the verification fails, the business person who entered the list needs to be notified for processing.
[0071] Step S20: storing the existing valid list into a key-value storage database based on the index key data and the offset of the index key data;
[0072] It should be noted that according to the set index key data and the offset of the index key data, the existing valid list is written into Redis to create a Bitmap index.
[0073] Step S30: obtaining a list of customers to be stored in real time, and determining the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list;
[0074] It should be noted that the pending list is a list of new items that need to be added to the warehouse, which is obtained in real time. The number of pending lists is also determined based on actual conditions and is not specifically limited in this embodiment. Accordingly, each pending list also contains some customer information, such as customer number and customer type. The pending list can be obtained through Kafka (a distributed stream processing platform), which is not specifically limited in this embodiment.
[0075] It's understood that the pending key data, or the key used when storing the pending list in Redis, can be used for deduplication. The pending value data, or the value used when storing the pending list in Redis, is typically the customer information included in each pending list. The pending list also uses a "cut-off-the-head-and-tail" strategy to determine the pending value data and its offset, allowing for data deduplication using the index key data and its offset in Redis.
[0076] Furthermore, after the step of obtaining the list of items to be stored in real time, the step also includes: performing an integrity check on the list of items to be stored, and after the list of items to be stored passes the integrity check, executing a step of determining the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list of items to be stored.
[0077] It is understandable that after obtaining the list to be stored, it is necessary to verify the integrity of the fields. If all fields pass the verification, the next step of deduplication will be carried out. If the verification fails, it is necessary to notify the business personnel who entered the list for processing.
[0078] Step S40, based on the offset of the to-be-stored key data and the offset of the index key data, deduplication is performed on the to-be-stored list to determine a non-duplicate list;
[0079] It should be noted that a non-duplicate list is a list that does not duplicate the existing valid list stored in Redis, and a duplicate list is a list that duplicates the existing valid list stored in Redis.
[0080] It can be understood that when the list of items to be stored is obtained, the Setbit instruction of Redis is called to set the offset of the key data to be stored to 1, and the offset of the index key data in Redis is used to determine whether each list of items to be stored is repeated. If it is repeated, it is a repeated list, and if it is not repeated, it is a non-duplicate list.
[0081] It should be understood that this embodiment retains the original duplicate detection interface of the target database, so that it can be used as an alternative solution when Redis is unavailable.
[0082] In a feasible implementation, step S40 may include: when matching key data of the key data to be stored exists in the index key data, obtaining the offset of the matching key data; when the offset of the matching key data is different from the offset of the key data to be stored, determining that the to-be-stored list corresponding to the key data to be stored is a non-duplicate list; when the offset of the matching key data is the same as the offset of the key data to be stored, determining that the to-be-stored list corresponding to the key data to be stored is a duplicate list.
[0083] It is understood that the matching key data is the same index key data in Redis as the to-be-entered key data. If the offset of the matching key data and the offset of the to-be-entered key data are both 1, then the to-be-entered list corresponding to the to-be-entered key data is considered to be a duplicate list in Redis, and the to-be-entered list corresponding to the to-be-entered key data is a duplicate list. If the offset of the to-be-entered key data is 1 and the offset of the matching key data is 0, then the to-be-entered list corresponding to the to-be-entered key data is considered to be a non-duplicate list.
[0084] In another feasible implementation, step S40 may include: when there is no matching key data of the to-be-stored key data in the index key data, determining that the to-be-stored list corresponding to the to-be-stored key data is a non-duplicate list.
[0085] It should be understood that if there is no index key data in Redis that is the same as the key data to be stored, it can be directly determined that the list to be stored corresponding to the key data to be stored is not repeated with the list in Redis. At this time, the list to be stored corresponding to the key data to be stored is a non-duplicate list.
[0086] Furthermore, this embodiment uses multithreading combined with Pipeline to optimize the deduplication process.
[0087] In a feasible implementation, the data storage and real-time deduplication method may further include steps A11 to A15:
[0088] Step A11, calculating the cyclic redundancy check value of the key data to be stored;
[0089] It should be noted that Redis Cluster must first be adapted, usually by extending JedisCluster (Redis client library) to support Pipeline.
[0090] It is understandable that the cyclic redundancy check value (CRC16 value) of the key data to be stored is calculated using CRC16 (Cyclic Redundancy Check 16-bit, 16-bit cyclic redundancy check algorithm). In this case, the cyclic redundancy check value is a 16-bit check value.
[0091] Step A12: determining a target value based on the cyclic redundancy check value of the key data to be stored and the number of hash slots;
[0092] It's understood that the number of hash slots refers to the number of hash slots in Redis. This embodiment uses the default number of 16384. The number of hash slots is moduloed based on the cyclic redundancy check value of the key data to be stored. Modulo essentially calculates the remainder. This modulo mapping allows the cyclic redundancy check value of the key data to be stored to be within the range of 0-16383, with each value in the range corresponding to a specific hash slot. The target value is the value obtained after the modulo calculation.
[0093] Step A13: determining a target hash slot based on the target value;
[0094] It should be noted that the target hash slot is the hash slot corresponding to the target value.
[0095] Step A14: Use the storage node where the target hash slot is located as the target node of the list to be stored;
[0096] It should be noted that Redis Cluster pre-assigns 16,384 hash slots to each Redis node, with different nodes responsible for handling different ranges of hash slots. Once the hash slot to which the key data to be stored belongs is determined, the corresponding Redis node, i.e., the target node, can be found based on a predefined mapping relationship. For example, if the CRC16 value of a key is 5000 after modulo, and it is known that node A is responsible for hash slots 3000-6000, then the target node for this key is node A.
[0097] Step A15: Based on the target node of the to-be-stored list, a corresponding execution thread is allocated for deduplication of the to-be-stored list, wherein the to-be-stored list of the same target node uses the same execution thread for deduplication.
[0098] It is understandable that the execution thread is the execution pipeline. Multiple deduplication commands for the same target node are sent through a single pipeline, reducing network overhead such as TCP (Transmission Control Protocol) handshakes. Figure 2, assuming that the target nodes corresponding to the waiting list 1 (customer 1) and the waiting list 2 (customer 2) are both node A, and the target nodes corresponding to the waiting list 3 (customer 3) and the waiting list 4 (customer 4) are both node B, then the deduplication of the waiting list 1 and the waiting list 2 uses thread 1 and is sent using one Pipeline, and the deduplication of the waiting list 3 and the waiting list 4 uses thread 2 and is sent using another Pipeline, thereby achieving batch deduplication and obtaining the final result, that is, whether the waiting lists 1-4 are non-duplicate lists.
[0099] It should be understood that this embodiment improves throughput through multi-threaded parallel processing and utilizes pipeline technology to achieve batch deduplication and improve deduplication efficiency.
[0100] Step S50: storing the non-duplicate list in the target database and updating the key-value storage database.
[0101] It should be noted that the obtained non-duplicate list is stored in the target database to achieve the storage of deduplicated data.
[0102] It is understandable that to ensure the accuracy of the next deduplication run, the current non-duplicate list also needs to be written to Redis, and the Bitmap index in Redis updated. This embodiment utilizes storage optimization, concurrent queries, and batch operations to achieve millisecond-level deduplication response, second-level processing of millions of lists, and zero data conflicts in concurrent scenarios.
[0103] Furthermore, in a feasible implementation manner, the invalid list of the day of the target database is obtained, and the offset of the index key data corresponding to the invalid list of the day in the key-value storage database is updated to a second value.
[0104] It should be noted that a daily expiration list is a list that has expired that day and must be removed from the existing valid list. Generally speaking, a list that has expired and is re-entered into the database is not considered a duplicate list. The daily expiration list is usually obtained once a day, and the time of retrieval can be set according to actual needs.
[0105] It is understandable that the second value is usually set to 0. Considering the life cycle of the list, in order to ensure the data consistency between the target database and Redis, the bit position of the offset corresponding to the invalid list of the day is set to 0 on a daily basis.
[0106] Furthermore, after step S50, it also includes: recording the number of the lists to be stored, the number of non-duplicate lists and performance indicators, the performance indicators at least including processing time and memory usage; determining the number of duplicates to be removed based on the number of the lists to be stored and the number of non-duplicate lists; determining the deduplication rate based on the deduplication number and the number of the lists to be stored; when the deduplication rate is greater than the preset deduplication threshold or the performance indicator meets the performance optimization conditions, outputting a warning message.
[0107] It should be noted that performance metrics include at least processing time and memory usage. The number of lists to be added, the number of unique lists, and performance metrics can be stored in a dedicated monitoring database for subsequent analysis and processing. Information such as the number of lists to be added, the number of unique lists, and performance metrics will also be promptly communicated to relevant business personnel and developers, allowing them to be informed of the status of list additions.
[0108] It can be understood that the number of duplicates removed can be calculated according to the number of lists to be stored and the number of non-duplicate lists, that is, the number of duplicates removed = the number of lists to be stored - the number of non-duplicate lists. The deduplication rate can be calculated according to the number of lists to be stored and the number of duplicates removed, that is, the deduplication rate = the number of duplicates removed / the number of lists to be stored. The preset deduplication threshold is the threshold of the deduplication rate, for example: 30%, which can be flexibly adjusted according to actual conditions, and there is no specific limitation on this. The performance optimization condition is the condition that is met when the performance needs to be optimized, for example: the memory usage rate is greater than the memory usage threshold, and there is no specific limitation on this. Among them, the memory usage threshold can be flexibly adjusted according to actual conditions, for example: 80%.
[0109] It should be understood that if the deduplication rate is greater than the preset deduplication threshold, an early warning message will be output to the corresponding business personnel / developers to analyze and process the current situation. If the performance indicators meet the performance optimization conditions, an early warning message will be output to the developers to optimize the deduplication process. For example, the non-duplicate lists are batched and stored in batches. This embodiment does not specifically limit this. In the specific implementation, a designated App (application) can be used to notify business personnel / developers.
[0110] This embodiment provides a data storage and real-time deduplication method, which obtains an existing valid list in a target database, determines index key data and an offset of the index key data based on the customer number of the existing valid list; stores the existing valid list in a key-value storage database based on the index key data and the offset of the index key data; obtains a list of items to be stored in real time, determines key data to be stored and an offset corresponding to the key data to be stored based on the customer number of the list to be stored; deduplicates the list to be stored based on the offset of the key data to be stored and the offset of the index key data, and determines a non-duplicate list; stores the non-duplicate list in the target database, and updates the key-value storage database. This embodiment adopts Bitmap data structure for storage in Redis (key-value storage database), and uses the "cut-off head and tail" strategy to construct Redis key names to reduce storage space. After the index is established, atomic instructions can be used to achieve zero-conflict deduplication to avoid repeated writing. At the same time, multi-threaded concurrency and Pipeline technology are used to improve deduplication efficiency, which can support the deduplication operation of large-scale lists in the warehousing and can be applied to high-concurrency scenarios. In addition, the template mode is used to encapsulate the deduplication logic, which can be flexibly expanded and can be used out of the box for new business scenarios.
[0111] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 , step S10 may include steps S101 to S104:
[0112] Step S101, determining the truncated length and the trimmed length based on the customer number of the existing valid list, and determining the middle offset based on the truncated length and the trimmed length;
[0113] It should be noted that the truncated length is the bit length extracted from the front part of the customer number, the truncated length is the bit length extracted from the back part of the customer number, and the middle offset is the bit used as the offset, that is, the middle bits remaining after removing the bits corresponding to the truncated length and the bits corresponding to the truncated length.
[0114] It can be understood that the length of the prefix and the length of the tail can be flexibly adjusted according to the length of the customer number. For example, if the customer number is 10 digits, the length of the prefix is set to 2 digits and the length of the tail is set to 1 digit. At this time, the middle offset bit is the middle 7 bits; if the customer number is 16 digits, the length of the prefix is set to 3 digits and the length of the tail is set to 2 digits. At this time, the middle offset bit is the middle 11 bits.
[0115] Step S102: extracting a first index key from the customer numbers in the existing valid list based on the truncated length, and extracting a second index key corresponding to the first index key from the customer numbers in the existing valid list based on the truncated length;
[0116] It should be noted that the first part of the index key data, namely the first index key, is extracted according to the truncated length, and the second part of the index key data, namely the second index key, is extracted according to the truncated length.
[0117] It is understood that when extracting the first index key, it is usually extracted from the customer number in order from the front to the back. For example, assuming the customer number is 0101010670 and the length of the first digit is 2, the first index key extracted is the first two digits, i.e., 01. When extracting the second index key, it is usually extracted from the customer number in order from the back to the front. For example, assuming the customer number is 0101010670 and the length of the last digit is 1, the second index key extracted is the last digit, i.e., 0.
[0118] Step S103: concatenate the first index key with the second index key corresponding to the first index key to obtain the index key data;
[0119] It is understandable that the corresponding index key data can be obtained by concatenating the first index key and the second index key. For example, assuming the customer number is 0101010670, which has a total of 10 digits, and the leading digit length is set to 2 digits and the trailing digit length is set to 1 digit, then the first two digits (01) are extracted as the first index key, and the last digit (0) is extracted as the second index key. In this case, the index key data is 010.
[0120] Step S104: extracting the offset of the index key data from the customer number in the existing valid list based on the middle offset bit, and updating the offset of the index key data to a first value.
[0121] It can be understood that according to the middle offset bit, the corresponding value is extracted from the customer number in the existing valid list as the offset of the index key data. For example, assuming the customer number is 0101010670, the leading length is set to 2 bits and the trailing length is set to 1 bit, then the middle offset bit is the middle 7 bits. At this time, the index key data is 010, and the offset corresponding to the index key data 010 is 0101067, which can reduce the storage space by 33%. The first value is usually set to 1. After determining the index key data and the offset of the index key data, the offset of the index key data is set to 1.
[0122] Further, in a feasible implementation, step S30 may include: extracting the first to-be-entered key from the customer number in the to-be-entered list based on the truncated length, and extracting the second to-be-entered key corresponding to the first to-be-entered key from the customer number in the to-be-entered list based on the truncated length; splicing the first to-be-entered key and the second to-be-entered key corresponding to the first to-be-entered key to obtain the to-be-entered key data; extracting the offset of the to-be-entered key data from the customer number in the to-be-entered list based on the middle offset bit, and updating the offset of the to-be-entered key data to the first value.
[0123] It should be noted that the first part of the key data to be stored is extracted according to the length of the truncated head, that is, the first key to be stored, and the second part of the key data to be stored is extracted according to the length of the truncated tail, that is, the second key to be stored.
[0124] It is understood that, similar to the first index key, the first pending storage key is typically retrieved sequentially from the customer number in order from front to back. Similar to the second index key, the second pending storage key is typically retrieved sequentially from the customer number in order from back to front. The corresponding pending storage key data can be obtained by concatenating the first pending storage key and the second pending storage key.
[0125] It should be understood that, according to the middle offset, the corresponding value is extracted from the customer number in the waiting list as the offset of the waiting key data. After the waiting key data and the offset of the waiting key data are determined, the offset of the waiting key data is set to 1.
[0126] Furthermore, when there is a large amount of customer number data of different business types, different prefix values can be designed to classify the customer numbers according to the business type. This can effectively avoid key conflicts and improve Redis's read and write performance. For example, if the prefix value corresponding to the business type is 2, and the key data to be stored according to the "cut off the head and tail" strategy is 010, then the prefix value 2 is added to this, that is, 2010.
[0127] This embodiment provides a data storage and real-time deduplication method, which determines the truncated length and the trimmed length based on the customer number of the existing valid list, and determines the middle offset based on the truncated length and the trimmed length; extracts the first index key from the customer number of the existing valid list based on the truncated length, and extracts the second index key corresponding to the first index key from the customer number of the existing valid list based on the trimmed length; concatenates the first index key and the second index key corresponding to the first index key to obtain index key data; extracts the offset of the index key data from the customer number of the existing valid list based on the middle offset, and updates the offset of the index key data to the first value. This embodiment uses a Bitmap data structure for storage in Redis (key-value storage database), uses a "truncated length" strategy to construct Redis key names, reduces storage space, and can use atomic instructions to achieve zero-conflict deduplication after indexing to avoid duplicate writes.
[0128] For example, in order to help understand the implementation process of the data storage and real-time deduplication method obtained by combining this embodiment with the above-mentioned embodiment 2, please refer to Figure 4 , Figure 4 A brief flowchart of a data storage and real-time deduplication method is provided. Specifically:
[0129] a) First, write the existing valid list in the database to Redis to create a Bitmap index. Bitmap storage optimization uses a "cut the head and tail" strategy to reduce storage space and prevent single keys from being too long. Keys are distributed across different Redis nodes, improving read and write concurrency.
[0130] b) When the list is stored, the Redis setbit command is called to set the bit corresponding to the customer number offset to 1 and return the original value of the bit. The original value is used to determine if the list is a duplicate. If not, the list is stored. The original database's duplicate detection interface is also retained as a fallback if Redis is unavailable. Multithreading and pipeline technologies are also used to optimize deduplication.
[0131] c) Taking into account the life cycle of the list (a list that has expired and is re-entered into the database is not considered a duplicate list), to ensure data consistency, the bit position corresponding to the offset of the list that has expired on that day is set to 0 on a daily basis.
[0132] d) When a list file is put into the database, the total number of lists, the number of duplicates removed, and other information in the list file will be counted, and the business and development personnel will be notified in a timely manner so that they can be aware of the list storage status in a timely manner.
[0133] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the data storage and real-time deduplication method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0134] This application also provides a data storage and real-time deduplication device, please refer to Figure 5 , the data storage and real-time deduplication device includes:
[0135] The data deduplication module 10 is used to obtain an existing valid list in the target database, and determine index key data and an offset of the index key data based on the customer number of the existing valid list;
[0136] The data deduplication module 10 is further configured to store the existing valid list into a key-value storage database based on the index key data and the offset of the index key data;
[0137] The data deduplication module 10 is further configured to obtain a list of customers to be stored in real time, and determine the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list of customers to be stored;
[0138] The data deduplication module 10 is further configured to deduplicate the list to be stored based on the offset of the key data to be stored and the offset of the index key data, and determine a non-duplicate list;
[0139] The data writing module 20 is used to store the non-repetitive list into the target database and update the key-value storage database.
[0140] In a feasible implementation manner, the data deduplication module 10 is further configured to determine the header length and the tail length based on the customer number of the existing valid list, and determine the middle offset bit based on the header length and the tail length.
[0141] Extracting a first index key from the customer numbers in the existing valid list based on the truncated length, and extracting a second index key corresponding to the first index key from the customer numbers in the existing valid list based on the truncated length;
[0142] Concatenate the first index key with the second index key corresponding to the first index key to obtain the index key data;
[0143] Based on the middle offset bit, the offset of the index key data is extracted from the customer number of the existing valid list, and the offset of the index key data is updated to the first value.
[0144] In a feasible embodiment, the data deduplication module 10 is further configured to extract a first to-be-stored key from the customer number on the to-be-stored list based on the length of the first truncated number, and to extract a second to-be-stored key corresponding to the first to-be-stored key from the customer number on the to-be-stored list based on the length of the last truncated number;
[0145] Concatenate the first to-be-warehouse key and the second to-be-warehouse key corresponding to the first to-be-warehouse key to obtain the to-be-warehouse key data;
[0146] Based on the middle offset bit, the offset of the key data to be stored is extracted from the customer number of the list to be stored, and the offset of the key data to be stored is updated to the first value.
[0147] In a feasible implementation manner, the data deduplication module 10 is further configured to obtain an offset of the matching key data when matching key data of the key data to be stored exists in the index key data;
[0148] When the offset of the matching key data is different from the offset of the to-be-stored key data, determining that the to-be-stored list corresponding to the to-be-stored key data is a non-duplicate list;
[0149] When the offset of the matching key data is the same as the offset of the to-be-stored key data, it is determined that the to-be-stored list corresponding to the to-be-stored key data is a duplicate list.
[0150] In a feasible implementation, the data deduplication module 10 is further configured to obtain a daily invalidation list of the target database, and update the offset of the index key data corresponding to the daily invalidation list in the key-value storage database to a second value.
[0151] In a feasible implementation manner, the data deduplication module 10 is further configured to calculate a cyclic redundancy check value of the key data to be stored;
[0152] Determine a target value based on a cyclic redundancy check value of the key data to be stored and the number of hash slots;
[0153] Determining a target hash slot based on the target value;
[0154] The storage node where the target hash slot is located is used as the target node of the list to be stored;
[0155] Based on the target node of the to-be-entered list, a corresponding execution thread is allocated for deduplication of the to-be-entered list, wherein the to-be-entered list of the same target node uses the same execution thread for deduplication.
[0156] In a feasible embodiment, the monitoring and warning module 30 is used to record the number of lists to be stored, the number of non-duplicate lists, and performance indicators, wherein the performance indicators at least include processing time and memory usage;
[0157] Determine the number of duplicate lists to be removed based on the number of the lists to be stored and the number of the non-duplicate lists;
[0158] Determine a deduplication rate based on the deduplication quantity and the quantity of the list to be stored;
[0159] When the deduplication rate is greater than a preset deduplication threshold or the performance indicator meets the performance optimization condition, a warning message is output.
[0160] In a feasible implementation, the list verification module 40 is used to perform integrity verification on the existing valid list. After the existing valid list passes the integrity verification, the steps of determining the index key data and the offset of the index key data based on the customer number of the existing valid list are executed.
[0161] In a feasible implementation, the list verification module 40 is also used to perform integrity verification on the list to be stored. After the list to be stored passes the integrity verification, a step is executed to determine the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list to be stored.
[0162] The data storage and real-time deduplication device provided by this application adopts the data storage and real-time deduplication method of the above-mentioned embodiment, which can solve the technical problems of low deduplication efficiency and easy duplication of data in traditional solutions. Compared with the existing technology, the beneficial effects of the data storage and real-time deduplication device provided by this application are the same as the beneficial effects of the data storage and real-time deduplication method provided by the above-mentioned embodiment, and the other technical features of the data storage and real-time deduplication device are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.
[0163] The present application provides a data storage and real-time deduplication device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data storage and real-time deduplication method in the above-mentioned embodiment one.
[0164] Reference below Figure 6, which shows a schematic diagram of the structure of a data storage and real-time deduplication device suitable for implementing the embodiments of the present application. The data storage and real-time deduplication device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The data storage and real-time deduplication device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0165] like Figure 6 As shown, the data storage and real-time deduplication device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 to RAM (Random Access Memory) 1004. Various programs and data required for the operation of the data storage and real-time deduplication device are also stored in RAM 1004. The processing device 1001, ROM 1002 and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the data storage and real-time deduplication device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a data storage and real-time deduplication device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.
[0166] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0167] The data storage and real-time deduplication device provided by this application, which utilizes the data storage and real-time deduplication method of the above-mentioned embodiment, can solve the technical problems of low deduplication efficiency and easy duplication of data in traditional solutions. Compared with the prior art, the beneficial effects of the data storage and real-time deduplication device provided by this application are the same as the beneficial effects of the data storage and real-time deduplication method provided by the above-mentioned embodiment, and the other technical features of the data storage and real-time deduplication device are the same as those disclosed in the method of the previous embodiment, which will not be repeated here.
[0168] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0169] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0170] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the data storage and real-time deduplication method in the above-mentioned embodiment.
[0171] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0172] The computer-readable storage medium may be included in the data storage and real-time deduplication device; or it may exist independently without being assembled into the data storage and real-time deduplication device.
[0173] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the data storage and real-time deduplication device, the data storage and real-time deduplication device: obtains the existing valid list in the target database, and determines the index key data and the offset of the index key data based on the customer number of the existing valid list; stores the existing valid list in the key-value storage database based on the index key data and the offset of the index key data; obtains the list to be stored in real time, and determines the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list to be stored; deduplicates the list to be stored based on the offset of the key data to be stored and the offset of the index key data, and determines the non-duplicate list; stores the non-duplicate list in the target database, and updates the key-value storage database.
[0174] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0175] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0176] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0177] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned data storage and real-time deduplication method. This computer-readable storage medium can solve the technical problems of low deduplication efficiency and easy duplication of data in traditional solutions. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the data storage and real-time deduplication method provided in the above-mentioned embodiment, and will not be elaborated here.
[0178] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned data storage and real-time deduplication method when executed by a processor.
[0179] The computer program product provided in this application can address the technical issues of low deduplication efficiency and the tendency for data to be repeatedly stored in the database due to conventional solutions. Compared to the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the data storage and real-time deduplication methods provided in the aforementioned embodiments, and are not further elaborated here.
[0180] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A data storage and real-time deduplication method, characterized in that: The method comprises: Obtain an existing valid list in a target database, and determine index key data and an offset of the index key data based on a customer number in the existing valid list; Based on the index key data and the offset of the index key data, storing the existing valid list in a key-value storage database; Obtain a list of customers to be stored in real time, and determine the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list; Deduplication of the list to be stored is performed based on the offset of the key data to be stored and the offset of the index key data to determine a non-duplicate list; The non-repeated list is stored in the target database, and the key-value storage database is updated.
2. The method according to claim 1, wherein The step of obtaining an existing valid list in the target database and determining index key data and an offset of the index key data based on the customer number of the existing valid list includes: Determine the truncated length and the trimmed length based on the customer number of the existing valid list, and determine the middle offset based on the truncated length and the trimmed length; Extracting a first index key from the customer numbers in the existing valid list based on the truncated length, and extracting a second index key corresponding to the first index key from the customer numbers in the existing valid list based on the truncated length; Concatenate the first index key with the second index key corresponding to the first index key to obtain the index key data; Based on the middle offset bit, the offset of the index key data is extracted from the customer number of the existing valid list, and the offset of the index key data is updated to the first value.
3. The method according to claim 1, wherein The step of obtaining a list of customers to be stored in real time and determining key data to be stored and an offset corresponding to the key data to be stored based on the customer number of the list of customers to be stored comprises: Extracting a first waiting-for-warehouse key from the customer number on the waiting-for-warehouse list based on the truncated length, and extracting a second waiting-for-warehouse key corresponding to the first waiting-for-warehouse key from the customer number on the waiting-for-warehouse list based on the truncated length; Concatenate the first to-be-warehouse key and the second to-be-warehouse key corresponding to the first to-be-warehouse key to obtain the to-be-warehouse key data; Based on the middle offset bit, the offset of the key data to be stored is extracted from the customer number of the list to be stored, and the offset of the key data to be stored is updated to the first value.
4. The method according to claim 1, wherein The step of removing duplicates from the list to be stored based on the offset of the key data to be stored and the offset of the index key data to determine a non-duplicate list includes: When matching key data of the key data to be stored exists in the index key data, obtaining an offset of the matching key data; When the offset of the matching key data is different from the offset of the to-be-stored key data, determining that the to-be-stored list corresponding to the to-be-stored key data is a non-duplicate list; When the offset of the matching key data is the same as the offset of the to-be-stored key data, it is determined that the to-be-stored list corresponding to the to-be-stored key data is a duplicate list.
5. The method according to claim 1, wherein After the step of storing the non-duplicate list in the target database and updating the key-value storage database, the method further includes: Obtain the invalid list of the target database on that day, and update the offset of the index key data corresponding to the invalid list of the day in the key-value storage database to a second value.
6. The method according to claim 1, wherein The method further comprises: Calculating a cyclic redundancy check value of the key data to be stored; Determine a target value based on a cyclic redundancy check value of the key data to be stored and the number of hash slots; Determining a target hash slot based on the target value; The storage node where the target hash slot is located is used as the target node of the list to be stored; Based on the target node of the to-be-entered list, a corresponding execution thread is allocated for deduplication of the to-be-entered list, wherein the to-be-entered list of the same target node uses the same execution thread for deduplication.
7. The method according to any one of claims 1 to 6, characterized in that After the step of storing the non-duplicate list in the target database and updating the key-value storage database, the method further includes: Record the number of lists to be stored, the number of non-duplicate lists, and performance indicators, wherein the performance indicators include at least processing time and memory usage; Determine the number of duplicate lists to be removed based on the number of the lists to be stored and the number of the non-duplicate lists; Determine a deduplication rate based on the deduplication quantity and the quantity of the list to be stored; When the deduplication rate is greater than a preset deduplication threshold or the performance indicator meets the performance optimization condition, a warning message is output.
8. A data storage and real-time deduplication device, characterized in that: The device comprises: A data deduplication module is used to obtain an existing valid list in the target database, and determine index key data and an offset of the index key data based on the customer number of the existing valid list; The data deduplication module is further configured to store the existing valid list into a key-value storage database based on the index key data and the offset of the index key data; The data deduplication module is further used to obtain a list of customers to be stored in real time, and determine the key data to be stored and the offset corresponding to the key data to be stored based on the customer number of the list of customers to be stored; The data deduplication module is further configured to deduplicate the list to be stored based on the offset of the key data to be stored and the offset of the index key data, and determine a non-duplicate list; The data writing module is used to store the non-repetitive list into the target database and update the key-value storage database.
9. A data storage and real-time deduplication device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data storage and real-time deduplication method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data storage and real-time deduplication method according to any one of claims 1 to 7 are implemented.