Data processing method and device, computer equipment and storage medium
By identifying and managing the number and times of repeated data blocks in the database, the problem of low reliability of junk data recognition is solved, and efficient utilization of data storage space and balanced data reliability is achieved.
Patent Information
- Application Number
- CN202510421521.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
The reliability of the garbage data identification method in the prior art is low, resulting in garbage data recovery errors and affecting the healthy development of the data storage system.
By identifying the number of duplicate data blocks in the database, generating the number of repetitions, and fine-grained management based on the preset number of times when deletion is required, ensuring the accurate deletion of data blocks and the reasonable retention of the number of copies.
Improve data storage space utilization, prevent data loss, improve deduplication efficiency, and balance deduplication effect and data reliability.
Smart Images

Figure CN120276677A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and data processing technologies, and particularly to a data processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of the medical and financial industries, servers have accumulated a vast amount of data. However, with the continuous growth of the data volume, the problem of garbage data has become increasingly prominent, seriously affecting the healthy development of the industry. The data recovery work is extremely urgent. The purpose of garbage data recovery is to recycle some unnecessary data blocks. The main method is to specify the deletion of the corresponding later version numbers by the user, read the corresponding metadata information, find the unnecessary data blocks in the data blocks, and recycle them at an appropriate time. The main process includes reading metadata, finding and recycling data blocks, and recycling data.
[0003] In related technologies, garbage data identification generally adopts the reference counting method, which uses reference count values to replace duplicate data and store it in the database, and judges whether the data blocks in the database are garbage data blocks by whether the reference count value is 0. The main disadvantage of this method is low reliability. Any repeated update or delayed update of the reference count value will cause the value to be incorrect, resulting in inconsistency between the stored / referenced data blocks and the count value in the system, and causing garbage data recovery errors. Summary of the Invention
[0004] In view of this, this application provides a data processing method, apparatus, computer device, and storage medium to achieve a more refined deletion operation for duplicate data blocks.
[0005] In a first aspect, a data processing method is provided, including:
[0006] In response to receiving data to be stored, storing at least one data block containing the data to be stored in a database;
[0007] Determining duplicate data blocks in the database;
[0008] Based on the number of mutually duplicate data blocks, marking the mutually duplicate data blocks to generate the repetition times of the duplicate data blocks;
[0009] In response to a deduplication instruction of the database, obtaining the repetition times of the current duplicate data blocks;
[0010] If the repetition times of the current duplicate data blocks are greater than a preset number of times, deleting the current duplicate data blocks, and reducing the repetition times of other duplicate data blocks that are mutually duplicate with the current duplicate data blocks by 1 until the repetition times are equal to the preset number of times.
[0011] In a second aspect, a data processing apparatus is provided, including:
[0012] A storage module, configured to store at least one data block including the data to be stored in a database in response to receiving the data to be stored;
[0013] A duplicate detection module, configured to determine duplicate data blocks in the database; and,
[0014] Based on the number of the mutually duplicate data blocks, mark the mutually duplicate data blocks to generate the number of repetitions of the duplicate data blocks;
[0015] An information acquisition module, configured to acquire the number of repetitions of the current duplicate data blocks in response to a deduplication instruction of the database;
[0016] A data cleaning module, configured to, if the number of repetitions of the current duplicate data blocks is greater than a preset number, delete the current duplicate data blocks, and decrease the number of repetitions of other duplicate data blocks that are mutually duplicate with the current duplicate data blocks by 1 until the number of repetitions is equal to the preset number.
[0017] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above data processing method are implemented.
[0018] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above data processing method are implemented.
[0019] In the solution implemented by the above data processing method, apparatus, computer device, and storage medium, after new data to be stored is deposited into the database, the data blocks in the database are checked for duplicates to mark the duplicate data blocks in the database. The number of duplicate data blocks that duplicate each other is counted, and this number is used as the number of repetitions of the duplicate data blocks that duplicate each other. When duplicate data needs to be deleted, the number of repetitions of the duplicate data blocks is obtained by reading the metadata of the data blocks that have been marked as duplicates. If the number of repetitions is greater than the preset number, it means that there are still data blocks in the database that duplicate the duplicate data block, so the duplicate data block can be deleted. At the same time, the number of repetitions of the data blocks that duplicate the duplicate data block is reduced by 1 to ensure that the number of repetitions can be updated in real time with the deletion operation, ensuring that the preset minimum number of copies is always maintained. On the contrary, if the number of repetitions is equal to the preset number, it means that the remaining number of data blocks that duplicate each other in the database has met the required number and there is no need to delete, so the duplicate data block can be skipped. On the one hand, by identifying and deleting duplicate data blocks through the number of repetitions, necessary copies can be accurately retained while deleting data, preventing data loss caused by accidental deletion, balancing the deduplication effect and data reliability, effectively reducing the space occupied by data storage, and improving the utilization rate of storage resources. On the other hand, when the deduplication instruction is triggered, the data blocks marked as duplicate data blocks are processed, avoiding the overhead of full-scale scanning and improving the deduplication efficiency.
[0020] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically described below. Brief Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic diagram of an application environment of the data processing method in an embodiment of the present application;
[0023] Figure 2 It is a schematic flowchart of the data processing method in an embodiment of the present application;
[0024] Figure 3 is Figure 2 A schematic flowchart of a specific implementation manner of step S20 in
[0025] Figure 4It is a schematic structural diagram of a data processing device in an embodiment of the present application;
[0026] Figure 5 It is a schematic structural diagram of a computer device in an embodiment of the present application;
[0027] Figure 6 It is another schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0029] The data processing method provided by the embodiments of the present application can be applied in an application environment such as Figure 1 where the client communicates with the server through a network. The server can receive the data to be stored through the client and store at least one data block containing the data to be stored in the database; determine the duplicate data blocks in the database; based on the number of mutually duplicate data blocks, mark the mutually duplicate data blocks to generate the number of repetitions of the duplicate data blocks; when the client issues a deduplication instruction or the server automatically triggers a deduplication instruction, the server responds to the deduplication instruction of the database and obtains the number of repetitions of the current duplicate data blocks; if the number of repetitions of the current duplicate data blocks is greater than the preset number, delete the current duplicate data blocks and reduce the number of repetitions of other duplicate data blocks that are mutually duplicate with the current duplicate data blocks by 1 until the number of repetitions is equal to the preset number. In the present application, on the one hand, by identifying and deleting duplicate data blocks through the number of repetitions, necessary copies can be accurately retained while deleting data, preventing data loss caused by accidental deletion, balancing the deduplication effect and data reliability, effectively reducing the space occupied by data storage, and improving the utilization rate of storage resources. On the other hand, when the deduplication instruction is triggered, the data blocks marked as duplicate data blocks are processed, avoiding the overhead of full-scale scanning and improving the deduplication efficiency. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments below.
[0030] Please refer to Figure 2 as shown in Figure 2 which is a flowchart of a data processing method provided by an embodiment of the present application, including the following steps:
[0031] S10: In response to receiving data to be stored, store at least one data block containing the data to be stored in a database;
[0032] Among them, data is the smallest information unit, which can be binary 0 / 1, characters, numbers, etc. A data block (Block / Chunk) is composed of continuous data and is the smallest logical unit managed by the storage system (such as a 4KB block).
[0033] For step S10, the data to be stored can be written into at least one data block according to a preset rule and then stored in the database in the form of data blocks. Thereby, the data to be stored is reasonably allocated to different data blocks, making the organization of the data to be stored more orderly, avoiding data storage fragmentation, and thus improving the space utilization rate of the storage device and the efficiency of data management.
[0034] The data processing method provided in this embodiment can be widely applied to backup systems (such as ZFS), cloud storage systems, database storage optimization systems, etc. The above systems can be implemented through a server, which can receive and store the data to be stored provided by the client in real time. For example, in the field of insurance applications, the electronic contracts signed by users often need to be backed up in a database. However, due to the different salespersons involved each time, the electronic contracts may be frequently modified or archived. By identifying and deleting duplicate data blocks, the space utilization rate of the storage device and the efficiency of data management can be improved.
[0035] In a possible implementation manner, in the medical field, the data to be stored can be medical data, such as Electronic Healthcare Record, electronic personal health records, including a series of electronic records with the value of being preserved for future reference, such as medical records, electrocardiograms, medical images, etc. Medical images refer to internal tissues obtained in a non-invasive manner for medical or medical research purposes. For example, images of the stomach, abdomen, heart, knee, and brain, such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), US (ultrasonic), X-ray images, electroencephalograms, and optical photographs generated by medical instruments. In the financial field, the data to be stored can be transaction data, payment data, business data, purchase data, etc.
[0036] It can be understood that the preset rules can be divided according to data sources, data volumes, data formats, acquisition times, etc., and the embodiments of the present application do not make specific limitations. Taking the preset rule of dividing according to data sources as an example, the data of user A is written into a data block A, and the data of user B is written into a data block B. Taking the preset rule of dividing according to data volume as an example, in this update, 10KB of data from client A and 3KB of data from client B are obtained, and the upper limit of the data volume of the data block is 50KB, then all the data from client A and client B can be stored in the same data block.
[0037] S20: Determine duplicate data blocks in the database;
[0038] In some embodiments of the present application, as Figure 3 shown, a specific entity alignment scheme is provided. In step S20, that is, to determine duplicate data blocks in the database, it specifically includes the following steps:
[0039] S21: Based on fingerprint information, determine the duplication ratio of duplicate data in any two data blocks in the database;
[0040] Among them, rolling hash (such as Rabin Fingerprint) or fixed block (such as CDC) can be used to calculate block fingerprints.
[0041] In a specific entity alignment scheme, step S21, that is, based on fingerprint information, determine the duplication ratio of duplicate data in any two data blocks in the database, specifically includes the following steps S211 - S214:
[0042] S211: Obtain the first fingerprint information of the data in the first data block and the second fingerprint information of the data in the second data block;
[0043] Among them, the first data block and the second data block are any two data blocks in the database. Any two data blocks can be data blocks newly generated based on the data to be stored, or data blocks that have been previously stored in the database.
[0044] S212: Compare the first fingerprint information with the second fingerprint information to determine the similarity between the first fingerprint information and the second fingerprint information;
[0045] S213: If the similarity is greater than the preset similarity, then use the data corresponding to the first fingerprint information or the second fingerprint information as duplicate data;
[0046] Among them, the preset similarity can be reasonably set according to the duplicate checking accuracy. For example, the preset similarity of 100% means that the two data are exactly the same, and the preset similarity of 80% means that the two data are partially repeated.
[0047] S214: Calculate the quotient of the data volume of the duplicate data and the data volume of the first data block or the second data block respectively to obtain the duplication ratio.
[0048] In this embodiment, considering that data with the same content essence usually has differences in aspects such as format and storage method, fingerprints with high similarity but different are generated. By comparing the fingerprint information instead of directly comparing the original data, potential duplicate relationships between data can be found more accurately, avoiding missing duplicate data due to surface differences in the data, and when dealing with large-scale data, comparing fingerprint information can significantly reduce the consumption of computing resources and improve the speed of data processing. After determining possible duplicate data pairs, further calculate the duplication ratio of the duplicate data volume to the total data block volume, avoiding misjudgment that may occur based on a single similarity index and improving the accuracy of duplicate data detection.
[0049] S22: If the duplication ratio of any two data blocks is greater than or equal to the preset ratio, then regard any two data blocks as duplicate data blocks that are mutually duplicate;
[0050] Among them, the preset ratio can be reasonably set according to the duplicate checking accuracy, and the embodiments of the present application do not make specific limitations.
[0051] For step S22, when the duplication ratio of the duplicate data in two data blocks in the database reaches or exceeds the preset ratio, these two data blocks are determined as duplicate data blocks, and the two data blocks are mutually duplicate. Thus, misjudgment due to a small amount of duplicate data is avoided, ensuring the accuracy of the deduplication operation, effectively reducing redundant data, and releasing a large amount of storage space.
[0052] S23: If the duplication ratio of one of any two data blocks is greater than or equal to the preset ratio, and the duplication ratio of the other of any two data blocks is less than the preset ratio, then based on the duplicate data, divide the data block with a duplication ratio less than the preset ratio among any two data blocks into multiple sub-data blocks;
[0053] S24: Regard the data block with a duplication ratio greater than or equal to the preset ratio among any two data blocks and the sub-data blocks containing duplicate data as duplicate data blocks that are mutually duplicate.
[0054] For steps S23 - S24, when the repetition situations of two data blocks are not completely symmetric, by splitting the data block with a lower repetition ratio, at least two sub - data blocks are formed. Let part of the at least two sub - data blocks contain repeated data, while the other part does not contain repeated data. In this way, the repeated part can be more accurately identified and processed. The data source with a higher repetition ratio and the sub - data blocks containing repeated data after splitting are regarded as mutually repeated data blocks. Thus, it is possible to avoid ignoring local repeated data due to the overall repetition ratio not meeting the standard, improve the fineness of data deduplication, and further effectively reduce redundant data storage, which helps to optimize the use of storage space and reduce storage costs.
[0055] Exemplarily, taking a medical case file as an example, data block A records the relevant data of file A, including a complete template + case description (a total of 1000 words, with the template accounting for 950 words). Data block B records the relevant data of file B, including a complete template + case description (a total of 1200 words, with the template accounting for 950 words). The repetition ratio of data block A is 95%, and the repetition ratio of data block B is 79%. If the preset ratio is set to 80%, the system considers that the two do not meet the repetition condition. At this time, data block B is split into sub - blocks B1 and B2 according to the content type. Among them, sub - data block B1 is used to record the template content (950 words) that is repeated with data block A. Sub - data block B2 is used to record the unique content of data block B (350 words). Compared with data block A, the repetition ratio of sub - data block B1 is 100%, and the repetition ratio of data block A is still 95%. The two meet the repetition condition. Even if the overall repetition rate is insufficient, high - repetition sub - blocks (such as template content) can still be extracted for deduplication.
[0056] Furthermore, in some embodiments of the present application, after step S23, the data processing method further includes: associating the data blocks with a repetition ratio greater than or equal to the preset ratio in any two data blocks with the sub - data blocks that do not contain repeated data; in response to a read instruction for the data blocks with a repetition ratio less than the preset ratio in any two data blocks, splicing the data blocks with a repetition ratio greater than or equal to the preset ratio and the sub - data blocks that do not contain repeated data in any two data blocks, or splicing multiple sub - data blocks belonging to the same data block, to generate a target data block; outputting the target data block.
[0057] In this embodiment, after splitting the data block with a lower repetition ratio, the data block with a higher repetition ratio is associated with the sub - data blocks that do not contain repeated data. Even after data deduplication, when the data block with a lower repetition ratio or the sub - data blocks containing repeated data are deleted, the system can read and integrate data from multiple scattered locations through the association relationship. Only the differences are retained in the database, but users can always obtain complete data. While avoiding data loss or errors and ensuring the quality and usability of the data, the space required for data storage is greatly reduced, and the utilization rate of storage resources is improved.
[0058] S30: Based on the number of mutually repeated duplicate data blocks, mark the mutually repeated duplicate data blocks to generate the repetition times of the duplicate data blocks.
[0059] For step S30, count the number of mutually repeated duplicate data blocks and use it as the repetition times of the mutually repeated duplicate data blocks, so as to accurately identify the duplicate data blocks containing the same data during data deduplication and improve the accuracy of data deletion.
[0060] S40: In response to the deduplication instruction of the database, obtain the repetition times of the current duplicate data blocks.
[0061] Specifically, the deduplication instruction of the database can be issued by the user through the client, so that the timing and scope of deduplication can be flexibly selected according to specific situations. It can also be triggered by an automatic deduplication program based on specific conditions or time intervals, so as to reduce manual intervention and improve the efficiency and timeliness of data processing.
[0062] S50: If the repetition times of the current duplicate data blocks are greater than the preset times, delete the current duplicate data blocks and reduce the repetition times of other duplicate data blocks that are mutually repeated with the current duplicate data blocks by 1 until the repetition times are equal to the preset times.
[0063] Among them, the preset times can be reasonably set according to the number of copies to be retained. For example, if the preset times are set to 1, only one data block is retained among the mutually repeated data blocks. If the preset times are set to 2, two data blocks are retained among the mutually repeated data blocks.
[0064] In this embodiment, after new data to be stored is deposited into the database, duplicate checking is performed on the data blocks in the database to mark the duplicate data blocks existing in the database. The number of duplicate data blocks that are mutually duplicate is counted, and this number is used as the number of repetitions of the mutually duplicate data blocks. When duplicate data needs to be deleted, the number of repetitions of the duplicate data blocks is obtained by reading the metadata of the data blocks that have been marked as duplicates. If the number of repetitions is greater than the preset number, it indicates that there are still data blocks in the database that are mutually duplicate with this duplicate data block, then this duplicate data block can be deleted, and at the same time, the number of repetitions of the data blocks that are mutually duplicate with this duplicate data block is decreased by 1 to ensure that the number of repetitions can be updated in real time with the deletion operation and ensure that the preset minimum number of copies is always maintained. On the contrary, if the number of repetitions is equal to the preset number, it indicates that the remaining number of mutually duplicate data blocks in the database has met the required number and there is no need to delete, then this duplicate data block can be skipped. On the one hand, by identifying and deleting duplicate data blocks through the number of repetitions, necessary copies can be accurately retained while deleting data, preventing data loss caused by accidental deletion, balancing the deduplication effect and data reliability, effectively reducing the space occupied by data storage, and improving the utilization rate of storage resources. On the other hand, when the deduplication instruction is triggered, the data blocks marked as duplicate data blocks are processed, avoiding the overhead of full-scale scanning and improving the deduplication efficiency.
[0065] Exemplarily, in the database, the documents uploaded by the user are divided into 3 data blocks: block A1, block A2, and block A3, and are stored in the database. The database already stores block B1, block B2, block B3, block B4, block C1, block C2, and block C3. Through duplicate detection, it is determined that block A1 and block B2 are mutually duplicate data blocks, and block A3, block B3, and block C3 are mutually duplicate data blocks. The system records block A1, block B2, block A3, block B3, and block C3 as duplicate data blocks, assigns the number of repetitions of block A1 and block B2 as 2, assigns the number of repetitions of block A3, block B3, and block C3 as 3, and the number of repetitions of the non-duplicate blocks A2, block B1, block B4, and block C1, block C2 is the initial value 1. When executing the data deduplication instruction, the number of repetitions of block A1 is greater than the preset value 1, block A1 is deleted, and the number of repetitions of block B2 that is mutually duplicate with block A1 is decreased by 1. At this time, the number of repetitions of block B2 is updated to 1. When the metadata of block B2 is read, since the number of repetitions is updated to 1 and no longer meets the condition of being greater than the preset value 1, block B2 will not be deleted.
[0066] It can be understood that before performing the operation of deleting duplicate data blocks, the metadata of this database can be fed back to the corresponding client for confirmation, and the deletion operation is performed after the client user replies. So as to more accurately release more server space.
[0067] In some embodiments of the present application, a specific entity alignment scheme is provided, and the data processing method further includes the following steps:
[0068] S50: Periodically obtain the most recent reference time point of the data block in the database;
[0069] S60: If the reference time point is before the preset time point, delete the data block or transfer the data block to a backup database.
[0070] Among them, the reference time point is the time when the data block is accessed, which is used to reflect the usage frequency or the latest degree of the data.
[0071] In this embodiment, by deleting the data blocks whose reference time points are before the preset time point, the storage space occupied by these data blocks that are no longer used or rarely used can be released. Thus, the storage structure of the main database is optimized, and the waste of storage space is avoided.
[0072] It can be understood that the data blocks that meet the condition that the reference time point is before the preset time point can be duplicate data blocks or data blocks without duplicates.
[0073] In some embodiments of the present application, step S50, that is, obtaining the most recent reference time point of the data block in the database, specifically includes the following steps:
[0074] S51: If the data block is a duplicate data block, compare the reference time points of the duplicate data blocks that are mutually duplicate;
[0075] S52: If the reference time points of the first duplicate data blocks are all after the reference time points of the second duplicate data blocks, update the reference time point of the second duplicate data block to the reference time point of the first duplicate data block.
[0076] Among them, the first duplicate data block is any one of the duplicate data blocks that are mutually duplicate, and the second duplicate data block is the other data blocks among the duplicate data blocks that are mutually duplicate except the first duplicate data block.
[0077] In this embodiment, when there are multiple mutually duplicate data blocks, the reference time of the most recently accessed duplicate data block is used as the reference time of the other data blocks, so as to update the earlier reference time point to the later reference time point, realize the time synchronization of duplicate content, facilitate more reasonable operations such as data deletion or transfer to a backup database, reduce the possibility of misdeleting important data, and make the database more accurate and reliable in data management.
[0078] In some embodiments of the present application, after step S20, the data processing method further includes: if the data sources of the data in the mutually duplicate data blocks are different, establish an index relationship between the data source and the mutually duplicate data blocks.
[0079] In this embodiment, for duplicate data blocks with different sources but substantially the same data content, an index relationship can be established between the data sources and the mutually duplicate data blocks. In this way, if the data in a certain data source changes or needs to be corrected, even if the duplicate data blocks are deleted, all related duplicate data blocks can be quickly found through the index relationship, ensuring consistent updates of the data throughout the system, thereby improving the data query speed and reducing the query response time.
[0080] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0081] In one embodiment, a data processing device is provided, which corresponds one-to-one to the data processing method in the above embodiment. As Figure 4 shown, the data processing device includes a storage module 401, a duplicate detection module 402, an information acquisition module 403, and a data cleaning module 404. The detailed description of each functional module is as follows:
[0082] The storage module 401 is configured to store at least one data block including the data to be stored in the database in response to receiving the data to be stored;
[0083] The duplicate detection module 402 is configured to determine duplicate data blocks in the database; and mark the mutually duplicate data blocks based on the number of mutually duplicate data blocks to generate the duplicate times of the duplicate data blocks;
[0084] The information acquisition module 403 is configured to acquire the duplicate times of the current duplicate data blocks in response to the deduplication instruction of the database;
[0085] The data cleaning module 404 is configured to, if the duplicate times of the current duplicate data blocks are greater than the preset times, delete the current duplicate data blocks and reduce the duplicate times of other duplicate data blocks that are mutually duplicate with the current duplicate data blocks by 1 until the duplicate times are equal to the preset times.
[0086] In one embodiment, the duplicate detection module 402 is specifically configured to: determine the duplicate ratio of duplicate data in any two data blocks in the database based on the fingerprint information; if the duplicate ratio of any two data blocks is greater than or equal to a preset ratio, then regard any two data blocks as duplicate data blocks that are duplicates of each other; if the duplicate ratio of one of any two data blocks is greater than or equal to the preset ratio, and the duplicate ratio of the other of any two data blocks is less than the preset ratio, then divide the data block with a duplicate ratio less than the preset ratio among any two data blocks into multiple sub-data blocks based on the duplicate data; and regard the data block with a duplicate ratio greater than or equal to the preset ratio among any two data blocks and the sub-data blocks containing the duplicate data as duplicate data blocks that are duplicates of each other.
[0087] In one embodiment, the duplicate detection module 402 is specifically configured to obtain the first fingerprint information of the data in the first data block and the second fingerprint information of the data in the second data block, where the first data block and the second data block are any two data blocks in the database; compare the first fingerprint information with the second fingerprint information to determine the similarity between the first fingerprint information and the second fingerprint information; if the similarity is greater than a preset similarity, then regard the data corresponding to the first fingerprint information or the second fingerprint information as duplicate data; and calculate the quotient of the data volume of the duplicate data and the data volume of the first data block or the second data block respectively to obtain the duplicate ratio.
[0088] In one embodiment, the data processing device further includes:
[0089] An association module (not shown in the figure), configured to associate the data blocks with a duplicate ratio greater than or equal to a preset ratio among any two data blocks with the sub-data blocks that do not contain duplicate data;
[0090] A feedback module (not shown in the figure), configured to, in response to a read instruction for a data block with a duplicate ratio less than a preset ratio among any two data blocks, splice the data blocks with a duplicate ratio greater than or equal to the preset ratio among any two data blocks and the sub-data blocks that do not contain duplicate data based on the association relationship, or splice multiple sub-data blocks belonging to the same data block to generate a target data block; and output the target data block.
[0091] In one embodiment, the information acquisition module 403 is further configured to periodically acquire the most recent reference time point of the data blocks in the database;
[0092] The data cleaning module 404 is further configured to, if the reference time point is before a preset time point, delete the data block or transfer the data block to a standby database.
[0093] In one embodiment, the information acquisition module 403 is specifically configured to compare the reference time points of duplicate data blocks that are duplicates of each other if the data block is a duplicate data block; if the reference time points of the first duplicate data blocks are all after the reference time point of the second duplicate data block, update the reference time point of the second duplicate data block to the reference time point of the first duplicate data block, where the first duplicate data block is any one of the duplicate data blocks that are duplicates of each other, and the second duplicate data block is any other data block among the duplicate data blocks that are duplicates of each other except the first duplicate data block.
[0094] In one embodiment, the data processing device further includes:
[0095] An indexing module (not shown in the figure), configured to establish an index relationship between the data source and the duplicate data blocks that are duplicates of each other if the data sources of the data in the duplicate data blocks that are duplicates of each other are different.
[0096] This application provides a data processing device. After new data to be stored is deposited into the database, the data blocks in the database are checked for duplicates to mark the duplicate data blocks in the database. The number of duplicate data blocks that are duplicates of each other is counted, and this number is used as the number of repetitions of the duplicate data blocks that are duplicates of each other. When duplicate data needs to be deleted, the number of repetitions of the duplicate data blocks is obtained by reading the metadata of the data blocks that have been marked as duplicates. If the number of repetitions is greater than the preset number, it means that there are still data blocks in the database that are duplicates of this duplicate data block, then this duplicate data block can be deleted, and at the same time, the number of repetitions of the data blocks that are duplicates of this duplicate data block is reduced by 1, so as to ensure that the number of repetitions can be updated in real time with the deletion operation, and ensure that the preset minimum number of copies is always maintained. On the contrary, if the number of repetitions is equal to the preset number, it means that the remaining number of mutually duplicate data blocks in the database has met the required number and there is no need to delete, then this duplicate data block can be skipped. On the one hand, by identifying and deleting duplicate data blocks through the number of repetitions, necessary copies can be accurately retained while deleting data, preventing data loss caused by accidental deletion, balancing the deduplication effect and data reliability, effectively reducing the space occupied by data storage, and improving the utilization rate of storage resources. On the other hand, when the deduplication instruction is triggered, the data blocks marked as duplicate data blocks are processed, avoiding the overhead of full-scale scanning and improving the deduplication efficiency.
[0097] For the specific limitations of the data processing device, reference can be made to the limitations on the data processing method in the above text, which will not be elaborated here. Each module in the above data processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0098] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as shown in Figure 5 . The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a data processing method.
[0099] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as shown in Figure 6 . The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a data processing method.
[0100] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0101] In response to receiving data to be stored, store at least one data block containing the data to be stored in the database;
[0102] Determine duplicate data blocks in the database;
[0103] Based on the number of mutually duplicate data blocks, mark the mutually duplicate data blocks to generate the number of repetitions of the duplicate data blocks;
[0104] In response to a deduplication instruction of the database, obtain the number of repetitions of the current duplicate data blocks;
[0105] If the number of repetitions of the current duplicate data blocks is greater than a preset number, delete the current duplicate data blocks, and reduce the number of repetitions of other duplicate data blocks that are mutually duplicate with the current duplicate data blocks by 1 until the number of repetitions is equal to the preset number.
[0106] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0107] In response to receiving data to be stored, store at least one data block including the data to be stored in a database;
[0108] Determine duplicate data blocks in the database;
[0109] Based on the number of mutually duplicate data blocks, mark the mutually duplicate data blocks to generate the number of repetitions of the duplicate data blocks;
[0110] In response to a deduplication instruction of the database, obtain the number of repetitions of the current duplicate data blocks;
[0111] If the number of repetitions of the current duplicate data blocks is greater than a preset number, delete the current duplicate data blocks, and reduce the number of repetitions of other duplicate data blocks that are mutually duplicate with the current duplicate data blocks by 1 until the number of repetitions is equal to the preset number.
[0112] It should be noted that for the functions or steps that can be implemented by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions of the data processing method in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0113] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0114] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0115] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A data processing method, characterized in that, Including: In response to receiving data to be stored, storing at least one data block containing the data to be stored in a database; Determining duplicate data blocks in the database; Based on the number of mutually duplicate data blocks, marking the mutually duplicate data blocks to generate the number of repetitions of the duplicate data blocks; In response to a deduplication instruction of the database, obtaining the number of repetitions of the current duplicate data blocks; If the number of repetitions of the current duplicate data blocks is greater than a preset number, deleting the current duplicate data blocks and reducing the number of repetitions of other duplicate data blocks that are mutually duplicate with the current duplicate data blocks by 1 until the number of repetitions is equal to the preset number.
2. The data processing method according to claim 1, wherein The determining of duplicate data blocks in the database includes: Based on fingerprint information, determining the repetition ratio of duplicate data in any two data blocks in the database; If the repetition ratios of any two data blocks are both greater than or equal to a preset ratio, regarding the any two data blocks as the mutually duplicate data blocks; If the repetition ratio of one of any two data blocks is greater than or equal to the preset ratio and the repetition ratio of the other of the any two data blocks is less than the preset ratio, dividing the data block with a repetition ratio less than the preset ratio among the any two data blocks into multiple sub-data blocks based on the duplicate data; Regarding the data block with a repetition ratio greater than or equal to the preset ratio among the any two data blocks and the sub-data blocks containing the duplicate data as the mutually duplicate data blocks.
3. The data processing method according to claim 2, characterized in that The determining of the repetition ratio of duplicate data in any two data blocks in the database based on fingerprint information includes: Obtaining first fingerprint information of data in a first data block and second fingerprint information of data in a second data block, where the first data block and the second data block are any two data blocks in the database; Comparing the first fingerprint information with the second fingerprint information to determine the similarity between the first fingerprint information and the second fingerprint information; If the similarity is greater than a preset similarity, regarding the data corresponding to the first fingerprint information or the second fingerprint information as the duplicate data; Calculating the quotient of the data volume of the duplicate data and the data volume of the first data block or the second data block respectively to obtain the repetition ratio.
4. The data processing method according to claim 2, wherein The method further includes: Associating the data block with a repetition ratio greater than or equal to the preset ratio among the any two data blocks with the sub-data blocks not containing the duplicate data; In response to a read instruction for a data block with a repetition ratio less than the preset ratio among the any two data blocks, splicing the data block with a repetition ratio greater than or equal to the preset ratio among the any two data blocks and the sub-data blocks not containing the duplicate data based on the association relationship, or splicing multiple sub-data blocks belonging to the same data block to generate a target data block; Outputting the target data block.
5. The data processing method according to any one of claims 1 to 4, characterized in that The method further includes: Periodically obtaining the most recent reference time point of a data block in the database; If the reference time point is before a preset time point, deleting the data block or transferring the data block to a standby database.
6. The data processing method according to any one of claims 1 to 4, wherein the obtaining of the most recent reference time point of the data block in the database comprises: If the data block is the duplicate data block, comparing the reference time points of the duplicate data blocks that are duplicates of each other; If the reference time points of the first duplicate data blocks are all after the reference time point of the second duplicate data block, updating the reference time point of the second duplicate data block to the reference time point of the first duplicate data block, where the first duplicate data block is any one of the duplicate data blocks that are duplicates of each other, and the second duplicate data block is the other data blocks among the duplicate data blocks that are duplicates of each other except the first duplicate data block.
7. The data processing method according to any one of claims 1 to 4, characterized in that The method further comprises: If the data sources of the data in the duplicate data blocks that are duplicates of each other are different, establishing an index relationship between the data sources and the duplicate data blocks that are duplicates of each other.
8. A data processing device, characterized in that, Comprising: A storage module, configured to store at least one data block including the data to be stored in a database in response to receiving the data to be stored; A duplicate detection module, configured to determine duplicate data blocks in the database; And, Based on the number of the duplicate data blocks that are duplicates of each other, marking the duplicate data blocks that are duplicates of each other to generate the number of repetitions of the duplicate data blocks; An information acquisition module, configured to acquire the number of repetitions of the current duplicate data blocks in response to a deduplication instruction of the database; A data cleaning module, configured to, if the number of repetitions of the current duplicate data blocks is greater than a preset number, delete the current duplicate data blocks, and reduce the number of repetitions of the other duplicate data blocks that are duplicates of the current duplicate data blocks by 1 until the number of repetitions is equal to the preset number.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the data processing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.