Data processing method, computing device and computer readable storage medium
By backing up file metadata in the index table, the self-description capability of cloud disk data index structure is realized, the problem of insufficient performance and security is solved, and the data consistency verification speed and user experience are improved.
Patent Information
- Application Number
- CN202410160497.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-04
- Publication Date
- 2025-08-05
AI Technical Summary
The data index structure of existing cloud disk files has insufficient performance and security, and cannot meet customers' access performance and data security requirements.
By backing up the file metadata that references the data fingerprint in the index table, the index table has the ability to self-describe, and can quickly locate the data's belongings and realize consistency verification between the data fingerprint and the actual existing file.
It improves the speed of data consistency verification, improves index performance and ensures data security, reduces system resource consumption, and improves user experience.
Smart Images

Figure CN120429487A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to data processing methods, computing devices, and computer-readable storage media. Background Art
[0002] Cloud disk service is a service that provides storage resources in an on-demand and easily expandable manner via the Internet. Based on cloud disks, storage services such as personal cloud disks, enterprise cloud disks, and cloud photo albums can be realized. Cloud disks provide storage resources for enterprises or individuals, enabling customers to enjoy fast, efficient, and mass-storage-supported services.
[0003] The system for managing files of hundreds of millions of users of cloud disks has the core requirement of storing and accessing a large amount of data safely and efficiently. However, there are some deficiencies in the data index structure of current cloud disk files in terms of performance and data security, which affect the user experience. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a computing device, a computer-readable storage medium, and a computer program product, so as to solve the technical deficiencies in the performance and security of the data index structure in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a data processing method is provided, including: determining a data fingerprint to be verified, where the data fingerprint is an identifier for distinguishing different data contents; obtaining metadata corresponding to the data fingerprint from an index table of the data fingerprint, where the index table is used to store index entries of the data fingerprint and metadata corresponding to the data fingerprint, and the metadata is metadata of a file that references the data fingerprint; using the obtained metadata to locate an actually existing file, and obtaining a verification result of the consistency between the data fingerprint and the actually existing file.
[0006] According to the second aspect of the embodiments of this specification, a computing device is provided, including: a memory and a processor; the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the data processing method described in any embodiment of this specification are implemented.
[0007] According to the third aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the data processing method described in any embodiment of this specification are implemented.
[0008] According to a fourth aspect of the embodiments of the present specification, there is provided a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the data processing method described in any embodiment of the present specification.
[0009] An embodiment of the present specification implements a data processing method, which determines a data fingerprint to be verified, obtains metadata corresponding to the data fingerprint from an index table of the data fingerprint, where the index table is used to store index entries of the data fingerprint and metadata corresponding to the data fingerprint, and the metadata is metadata of a file that references the data fingerprint, and locates an actually existing file by using the obtained metadata, so as to obtain a verification result of the consistency between the data fingerprint and the actually existing file. It can be seen that this method enables the index table to have the self-description ability of the file by backing up the metadata of the file that references the data fingerprint in the index table, that is, the attribution of the data can be located through the index table itself. When the actually existing file is located, the data fingerprint is consistent with the actually existing file, otherwise it is inconsistent. Therefore, the verification result of the consistency between the data fingerprint and the actually existing file can be obtained quickly, realizing faster data consistency verification, effectively improving the index performance and ensuring data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic diagram of a data processing method provided by an embodiment of the present specification in a cloud storage application scenario;
[0011] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present specification;
[0012] Figure 3a is a schematic diagram of an index provided by an embodiment of the present specification;
[0013] Figure 3b is a schematic diagram of the relationship between a file metadata table and an index table provided by an embodiment of the present specification;
[0014] Figure 4 is a flowchart of the processing process of file creation of a data processing method provided by an embodiment of the present specification;
[0015] Figure 5 is a flowchart of the processing process of file deletion of a data processing method provided by an embodiment of the present specification;
[0016] Figure 6 is a flowchart of the processing process of garbage collection of a data processing method provided by an embodiment of the present specification;
[0017] Figure 7It is a flowchart of the processing procedure based on the table content change notification channel of a data processing method provided by an embodiment of this specification;
[0018] Figure 8 It is a schematic structural diagram of a data processing device provided by an embodiment of this specification;
[0019] Figure 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. Specific embodiments
[0020] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.
[0021] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0022] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first can also be referred to as the second, and similarly, the second can also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0023] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0024] First, the noun terms involved in one or more embodiments of this specification are explained.
[0025] A cloud disk provides services such as information storage, reading, and downloading for users via the Internet, and has the characteristic of massive storage.
[0026] Network disks and cloud photo albums provide services such as file and photo storage, reading, and downloading for users via the Internet, and can also implement functions such as automatic upload, automatic synchronization, and easy sharing, with the characteristic of massive storage.
[0027] The metadata of a file contains the mapping relationship between the logical address of the file in the logical space and the actual physical storage location.
[0028] Garbage Collection (GC) is a mechanism used to recycle the space occupied by objects with no references at irregular intervals during idle time, which can release the space occupied by garbage data.
[0029] The file metadata table is a database table used to store the metadata of files. The metadata of a file can include, for example, the file unique identifier, creation time, modification time, creator, file size, file type, etc., which are the basic information of the file.
[0030] A data fingerprint is an identifier calculated based on the data content of a file, used to distinguish different data contents and uniquely identify the corresponding data content. In other words, any piece of data content in a file has a unique data fingerprint, which is different from the data fingerprints of other data contents.
[0031] The index table of data fingerprints is used to store the primary key, which is used to point to the data fingerprint.
[0032] An index entry is a record in the index table that stores the primary key, also known as a sentinel fingerprint index entry, and is used for querying data fingerprints.
[0033] A reference entry is a record in the index table that stores the reference relationship between a file and a data fingerprint. The existence of a reference relationship between a file and a data fingerprint indicates that the file references the file data pointed to by the data fingerprint.
[0034] The reference count is used to indicate how many files reference a data fingerprint.
[0035] The table content change notification channel is a notification mechanism in which any modification operation within a table row will be notified in the order of changes.
[0036] The scale of the file data managed by the cloud disk has reached the level of hundreds of billions. Therefore, efficient access, inspection, and verification of data are very important. Currently, the data index structure of cloud disk files has poor capabilities in terms of verification and cannot meet the requirements of customer access performance and data security. Therefore, there is an urgent need for a data processing method that can efficiently complete the verification of data indexing.
[0037] In view of this, in this specification, a data processing method is provided. By backing up the metadata of the file that references the data fingerprint in the index table, the index table is enabled to have the self-description ability of the file, that is, the attribution of the data can be located through the index table itself, so that the data verification can be quickly implemented, effectively ensuring the customer access performance and data security.
[0038] Specifically, in this specification, a data processing method is provided. This specification also relates to a data processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0039] See Figure 1 , Figure 1 shows a schematic diagram of a data processing method provided according to an embodiment of this specification in a cloud storage application scenario. As Figure 1 shown, the cloud includes a computing cluster and a storage cluster. The computing cluster includes one or more cloud servers. The storage cluster includes one or more storage nodes for supporting the storage of the cloud disks of the cloud servers. The cloud disks of the cloud servers can be used to implement the storage of, for example, network disks and cloud albums. Based on the cloud storage service provided by the cloud, the user side can send a request to access data to the cloud. Any one or more cloud servers in the computing cluster of the cloud access the cloud disk according to the request to obtain the corresponding file data, and return the file data to the user side. A cloud disk is a virtual storage device for file-level data provided by the cloud. The storage of the file-level data provided by the cloud disk can present a storage experience similar to that of a local file system for external access. The underlying layer of the cloud disk can be a physical block storage device such as a disk. Just like a disk, the user can perform operations such as partitioning, formatting, and creating a file system on the cloud disk mounted on the cloud server, and perform persistent storage of data. Different from a disk, the user can achieve persistent storage of data without additional operations.
[0040] The cloud server manages files on the cloud disk based on metadata, and the data content of the file is uniquely identified by a data fingerprint, and the data fingerprint is quickly queried through an index table. The index table is used to save the index items of the data fingerprint. The index item saves the primary key. The primary key is used to point to the data fingerprint. In the method provided in the embodiment of this specification, the index table is not only used to save the index items of the data fingerprint, but also to save the metadata corresponding to the data fingerprint, and the metadata is the metadata of the file that references the data fingerprint. Therefore, the processing process of the cloud server to verify the data may include: determining the data fingerprint to be verified, obtaining the metadata corresponding to the data fingerprint from the index table of the data fingerprint, and the metadata is the metadata of the file that references the data fingerprint, and using the obtained metadata to locate the actual existing file to obtain a verification result of the consistency between the data fingerprint and the actual existing file.
[0041] In the above application scenario, according to the data processing method provided in the embodiment of this specification, by backing up the metadata of the file that references the data fingerprint in the index table, the index table has the self-describing ability of the file, that is, the ownership of the data can be located through the index table itself, so that the data consistency verification can be quickly realized.
[0042] It should be noted that Figure 1 The application scenarios shown are only used to schematically illustrate the methods provided in the embodiments of this specification and do not constitute a limitation on the methods provided in the embodiments of this specification. For example, according to the methods provided in the embodiments of this specification, the cloud disk can be a cloud disk of a cloud server that provides any cloud computing capabilities. The cloud server can be a distributed server cluster including multiple servers or a single server. The services provided by the cloud server may include: cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.
[0043] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0044] Step 202: Determine the data fingerprint to be verified, where the data fingerprint is an identifier used to distinguish different data contents.
[0045] The data fingerprint is a unique identifier calculated based on the data content of a file and is used to uniquely identify the corresponding data content. For example, the data fingerprint can be calculated by algorithms including but not limited to SHA1 (Secure Hash Algorithm 1) and MD5 (Message-Digest Algorithm).
[0046] The data fingerprint is queried through the index entries in the index table. Therefore, determining the data fingerprint to be verified can also be understood as determining the index entry corresponding to the data fingerprint to be verified.
[0047] Step 204: Obtain the metadata corresponding to the data fingerprint from the index table of the data fingerprint. The index table is used to store the index entries of the data fingerprint and the metadata corresponding to the data fingerprint. The metadata is the metadata of the file that references the data fingerprint.
[0048] The index entry is a record in the index table used to store the primary key of the data fingerprint. The index entry is used for querying the data fingerprint, and the primary key is used to point to the data fingerprint. For example: As Figure 3a shown in the index table, an index entry (which can also be called a guard_fp_entry, a sentinel index entry) is used to store the primary key of a data fingerprint. In addition to the index entry, the index table may also include a reference entry, which is used to store the reference relationship between the file and the data fingerprint. The reference relationship between the file and the data fingerprint can be represented as the correspondence between the unique identifier of the file and the primary key of the data fingerprint. The metadata corresponding to the data fingerprint can be stored in the reference entry corresponding to the primary key.
[0049] Since the index table is not only used to store the index entries of the data fingerprint but also used to store the metadata corresponding to the data fingerprint, the metadata corresponding to the data fingerprint can be obtained from the index table of the data fingerprint. For example: As Figure 3a shown in the index table, assuming that the data fingerprint A to be verified is pointed to by the primary key A1, after determining the primary key A1, the metadata recorded in the reference entry corresponding to the primary key A1 can be obtained from the index table.
[0050] Step 206: Use the obtained metadata to locate the actually existing file and obtain the verification result of the consistency between the data fingerprint and the actually existing file.
[0051] Since the metadata of a file can be used to locate the ownership of the data, the obtained metadata can be used to verify the consistency between the data fingerprint and the actually existing file, and the specific implementation manner of the verification is not limited. For example, the data content can be obtained through the metadata, a data fingerprint can be constructed based on the data content, and the constructed data fingerprint can be compared with the data fingerprint corresponding to the index item to achieve the consistency verification; for another example, the consistency between the metadata recorded in the file metadata table and the obtained metadata can be compared to achieve the consistency verification between the data fingerprint and the actually existing file.
[0052] This method enables the index table to have the self-description ability of the file by backing up the metadata of the file that references the data fingerprint in the index table, that is, the ownership of the data can be located through the index table itself, so that the verification of data consistency can be quickly achieved, effectively ensuring the client access performance and data security.
[0053] In one or more embodiments of this specification, the consistency between the data fingerprint and the actually existing file is verified by inspecting the file metadata table. Specifically, the use of the obtained metadata to locate the actually existing file and obtain the verification result of the consistency between the data fingerprint and the actually existing file includes:
[0054] Obtain the metadata to be verified of the file to be verified from the file metadata table, where the file to be verified is the file that references the data fingerprint, and the file metadata table is used to store the metadata of the actually existing file;
[0055] Compare the obtained metadata with the metadata to be verified, and determine whether they are the same metadata;
[0056] If they are not the same, determine that there is an abnormality in the consistency between the data fingerprint and the actually existing file.
[0057] The file to be verified can be any file in the file metadata table. For example: each file metadata row in the file metadata table can be traversed, and the metadata in each file metadata row can be used as the metadata to be verified of the file to be verified, so as to achieve the verification of the consistency between each data fingerprint and the actually existing file.
[0058] In one or more embodiments of this specification, in combination with the above embodiments, a primary key can be constructed based on the data content of the actually existing file to determine the data fingerprint to be verified. Specifically, the determination of the data fingerprint to be verified includes:
[0059] Obtain the data content of the file to be verified for constructing the data fingerprint;
[0060] Based on the data content, construct the primary key of the data fingerprint to be queried;
[0061] Using the primary key of the data fingerprint to be queried, in the index table of the data fingerprint, query the data fingerprint corresponding to the index entry that matches the primary key of the data fingerprint to be queried as the data fingerprint to be verified. The index entry is used to store the primary key of the corresponding data fingerprint.
[0062] The data content used to construct the data fingerprint refers to the data content uniquely identified by the data fingerprint in the file.
[0063] The primary key is used to point to the data fingerprint and can be understood as the primary key of the database table used to store the data fingerprint. The primary key can be calculated based on the data content. For example: The primary key can be calculated based on the SHA1 value, MD5 value, and data size of the data content. Therefore, in the above embodiments, the primary key of the data fingerprint to be queried can be constructed based on the obtained data content.
[0064] In the above embodiments, the relationship between the file metadata table and the index table is as Figure 3b shown in the schematic diagram of the relationship between the file metadata table file and the index table fp: Any row record in the file metadata table file is used to store the metadata of a file. The reference relationship between the file and the data fingerprint is recorded in the index entry of the index table fp. The reference relationship between the file and the data fingerprint can be represented by the corresponding relationship between the file unique identifier and the primary key of the data fingerprint. Specifically, Figure 3b in, file_x represents the metadata row of a single file stored in the file metadata table file; fp_entry_x, that is, the reference item, represents a reference of the file file_x to the data fingerprint; a data fingerprint can be referenced by multiple files to support file-level data deduplication, and multiple identical files will reference the same data fingerprint; gurad_fp_entry, that is, the index entry, which stores the primary key of the data fingerprint, and the primary key is used for the requirement of querying the data fingerprint table. Thus, from Figure 3b the file metadata table file shown, obtain the metadata to be verified of the file to be verified, and construct the primary key based on the metadata to obtain the data content, and then be able to query from the index table fp as shown in Figure 3b the index entry with a matching primary key. The data fingerprint pointed to by the primary key in the index entry is the data fingerprint to be verified.
[0065] After determining the data fingerprint to be verified, the metadata to be verified obtained from the file metadata table can be compared with the metadata obtained from the index table to determine the verification result of the consistency between the data fingerprint and the actually existing file.
[0066] For example: As Figure 3bThe shown index table compares the metadata obtained from the reference items with the metadata to be verified. If the two are consistent, the verification result is passed, indicating that the data fingerprint is consistent with the actually existing file. Otherwise, it is determined that there is an abnormality in the data consistency.
[0067] Taking the example of traversing each file metadata row in the file metadata table file and using the metadata in each file metadata row as the metadata to be verified for the file to be verified to implement the inspection of the file metadata table file, the processing process of the above embodiment may include: scanning the file table row by row and performing the following operations on each row of the file table:
[0068] For the metadata recorded in the current row, construct the corresponding primary key by using the data content corresponding to the metadata;
[0069] According to the constructed primary key, query the matching records in the index table fp;
[0070] Check whether the metadata in the queried index item is consistent with the metadata recorded in the current row of the table file. If it is consistent, it indicates that the data consistency is normal. Otherwise, it indicates that the data consistency is abnormal.
[0071] It should be noted that the metadata recorded in the index item can be partial key metadata or all metadata of the file. This specification does not limit this, as long as the correctness of the data can be verified. For example, the metadata recorded in the index item may include but is not limited to: file unique identifier, creator, creation time, size, etc.
[0072] In one or more embodiments of this specification, the verification of data consistency is implemented by traversing the index table. Specifically, the data fingerprint corresponding to any index item in the index table is the data fingerprint to be verified; the process of using the obtained metadata to locate the actually existing file and obtaining the verification result of the consistency between the data fingerprint and the actually existing file includes:
[0073] Using the obtained metadata, query the records matching the metadata in the file metadata table, and the file metadata table is used to store the metadata of the actually existing file
[0074] Judge whether the obtained metadata is the same as the metadata stored in the record;
[0075] If they are not the same, it is determined that there is an abnormality in the consistency between the data fingerprint and the actually existing file.
[0076] For example: such as Figure 3bFor the index table shown, traverse the index table line by line, compare the metadata obtained from the traversed reference item with the metadata to be verified. If the two are consistent, the verification result is passed, indicating that the data consistency is normal; otherwise, it is determined that there is an abnormality in the data consistency.
[0077] In addition, when traversing the index table, an index item can also be traversed to verify the primary key in the index item. For example, the interface of the metadata index library can be called to obtain the corresponding metadata, and verify whether the SHA1 value is consistent with the primary key in the index item.
[0078] Taking the example of traversing each line of the index table fp to implement the inspection of fp, the processing process of the above embodiment may include: scanning the entire table fp line by line, and performing the following operations on each line of the table fp:
[0079] If the traversed record is an index item, verify whether the primary key of the index item record is correct;
[0080] If the traversed record is a reference item, construct the primary key of the file metadata table according to the metadata of the reference item record;
[0081] Query the records in the file metadata table according to the constructed primary key of the file metadata table;
[0082] Verify whether the metadata of the file metadata table record is consistent with the metadata of the record in the reference item; if it is consistent, it indicates that the data consistency is normal, otherwise it indicates that the data consistency is abnormal.
[0083] In the above embodiment, since the index table contains the metadata corresponding to the data fingerprint and has the ability of self-description of data, therefore, based on the data inspection of the index table and / or the file metadata table, the actual existing files can be accurately located, the data consistency can be verified, the consumption of system resources can be effectively reduced, the stability of the system can be ensured, and the user experience can be improved.
[0084] In one or more embodiments of this specification, the reference relationship between the file and the data fingerprint is also written into the index table, so that the reference relationship and the data fingerprint are saved in the same table. In this way, the creation process of the data fingerprint index item of the file can be realized by accessing one table, reducing the access to the table, reducing the resource consumption, and improving the performance. Specifically, the index table further includes a reference item, and the reference item is used to save the reference relationship between the file and the data fingerprint;
[0085] The method further includes:
[0086] In response to a file creation request, determine whether there is a data fingerprint corresponding to the file to be created in the index table for the file creation request;
[0087] If it exists, write the reference relationship between the to-be-created file and the data fingerprint into the reference entry of the index table;
[0088] If it does not exist, create an index entry corresponding to the data fingerprint referenced by the to-be-created file in the index table, and write the reference relationship between the to-be-created file and the data fingerprint into the reference entry of the index table.
[0089] For example: The file creation request may carry the unique identifier of the to-be-created file. Based on the unique identifier of the to-be-created file, look up the corresponding reference entry in the index table.
[0090] In the above embodiment, since the reference relationship between the file and the data fingerprint is written into the index table, the access to the table is reduced when creating a file, which can effectively alleviate the performance problems caused by hot file sharing.
[0091] Combined with the above embodiment, after writing a new reference relationship to the index table, it is also possible to check whether there are problems such as concurrent creation and deletion operations by reading the index table to ensure the correctness of the data. Specifically, after writing the reference relationship between the to-be-created file and the data fingerprint into the index table, it further includes:
[0092] Judge whether there is an index entry corresponding to the data fingerprint referenced by the to-be-created file in the index table;
[0093] If it exists, judge whether the data fingerprint referenced by the to-be-created file in the index entry has changed;
[0094] If it has changed, delete the reference entry corresponding to the reference relationship between the to-be-created file and the data fingerprint.
[0095] Correspondingly, if it has not changed, the creation is successful.
[0096] For example, the above judgment and deletion steps can be performed through a conditional update mechanism. The conditional update mechanism means that when the judgment condition is met, the update action is executed, otherwise it is rejected. The judgment and update are atomic operations.
[0097] After writing a new reference relationship to the index table in the above embodiment, it is checked whether there are problems such as concurrent creation and deletion operations by reading whether the newly written reference relationship in the index table has changed. If it has changed, it means that there are concurrent problems. By deleting the incorrect reference entry, it is avoided that there are incorrect reference entries in the index table, so as to solve the problem of incorrect data caused by concurrent operations such as creation and deletion, and ensure the correctness of the data when the creation is successful.
[0098] In addition, to prevent the data fingerprint referenced by the newly created file from being deleted due to the garbage collection mechanism, the garbage collection status can also be processed to ensure the availability of the created data fingerprint. Specifically, after writing the reference relationship between the file to be created and the data fingerprint into the reference entry of the index table, the following steps are further included:
[0099] Determine whether there is an index entry in the index table corresponding to the data fingerprint referenced by the file to be created;
[0100] If it exists, determine whether the data fingerprint referenced by the file to be created is in the garbage collection status;
[0101] If it is in the garbage collection status, cancel the garbage collection status of the data fingerprint referenced by the file to be created.
[0102] It can be understood that during the garbage collection process, the data to be garbage collected will be marked as in the garbage collection status and wait to be deleted before being actually deleted. Therefore, the garbage collection status of the data fingerprint means that the data fingerprint has a garbage collection mark and waits to be deleted.
[0103] In the above embodiment, since the garbage collection status of the newly created file is cancelled, the newly created file can normally reference the data fingerprint, preventing the record corresponding to the data fingerprint from being deleted by the background garbage collection mechanism and ensuring the availability of the data fingerprint.
[0104] It should be noted that the above embodiments of concurrent processing of creation and other operations, as well as the embodiments of garbage collection status processing, can be implemented separately or combined. This specification does not limit this. Next, an exemplary description of the processing process of combined implementation will be given.
[0105] Figure 4 The flowchart of the file creation process of a data processing method provided in an embodiment of this specification is shown, which specifically includes the following steps.
[0106] Step 402: Receive a file creation request and obtain the index table fp.
[0107] Step 404: Determine whether there is a data fingerprint in the index table fp corresponding to the file to be created referenced by the file creation request.
[0108] Step 406: If it exists, write the reference relationship between the file to be created and the data fingerprint into the reference entry of the index table.
[0109] Step 408: Obtain the index table fp again.
[0110] Step 410: Determine whether there is a data fingerprint referenced by the file to be created in the index table fp.
[0111] Step 412: If it exists, determine whether the data fingerprint referenced by the to-be-created file in the index entry has changed.
[0112] If it is determined in Step 410 that it does not exist, then proceed to Step 418 to delete the reference entry corresponding to the reference relationship between the to-be-created file and the data fingerprint.
[0113] Step 414: If it is determined in Step 412 that it has not changed, determine whether the data fingerprint referenced by the to-be-created file is in the garbage collection state.
[0114] If it is determined in Step 414 that it is not in the garbage collection state, then determine that the file creation is successful.
[0115] Step 416: If it is determined in Step 414 that it is in the garbage collection state, cancel the garbage collection state of the data fingerprint referenced by the to-be-created file.
[0116] If the cancellation fails, then determine that the file creation fails.
[0117] If the cancellation is successful, then determine that the file creation is successful.
[0118] If it is determined in Step 412 that it has changed, then proceed to Step 418 to delete the reference entry corresponding to the reference relationship between the to-be-created file and the data fingerprint.
[0119] Step 418: If it is determined in Step 410 that it does not exist, or if it is determined in Step 412 that it has changed, then delete the reference entry corresponding to the reference relationship between the to-be-created file and the data fingerprint.
[0120] Step 420: If it is determined in Step 404 that it does not exist, then create an index entry corresponding to the data fingerprint referenced by the to-be-created file in the index table, and write the reference relationship between the to-be-created file and the data fingerprint into the reference entry of the index table.
[0121] Step 422: Determine whether the writing in Step 420 is successful.
[0122] If the writing is successful, then the creation is successful.
[0123] If the writing fails, then the creation fails.
[0124] In the above embodiments, the concurrency problems of file creation and deletion and other operations are judged and solved, and the concurrency problems of file creation and data fingerprint garbage collection are judged and solved, solving the problems of incorrect data and accidental deletion by garbage collection, effectively ensuring the correctness of the data.
[0125] Next, an exemplary description will be given of the specific implementation manner of deleting a file when the reference relationship between the file and the data fingerprint is written into the index table so that the reference relationship and the data fingerprint are saved in the same table. In one or more embodiments of this specification, the method further includes:
[0126] In response to a file deletion request, in the file metadata table, mark the metadata to be deleted corresponding to the file deletion request as deleted. The file metadata table is used to save the metadata of the file;
[0127] Delete the reference item corresponding to the metadata to be deleted in the index table;
[0128] When it is determined that the deletion of the reference relationship is successful, delete the record corresponding to the metadata to be deleted from the file metadata table according to the deletion mark.
[0129] Through the above embodiments, the index table can be effectively updated when a file is deleted, ensuring the correctness of the data.
[0130] Next, an exemplary description will be given of the specific processing process of deleting a file.
[0131] Figure 5 The flowchart of the processing process of file deletion of a data processing method provided by an embodiment of this specification is shown, which specifically includes the following steps.
[0132] Step 502: In response to a file deletion request, in the file metadata table, mark the metadata to be deleted corresponding to the file deletion request as deleted.
[0133] Step 504: Determine whether the marking is successful.
[0134] Step 506: If the marking is successful, delete the reference item corresponding to the metadata to be deleted in the index table.
[0135] If the marking fails, determine that the deletion process fails.
[0136] Step 508: Determine whether the deletion of the reference item is successful.
[0137] Step 510: If the deletion of the reference item is successful, delete the record corresponding to the metadata to be deleted from the file metadata table according to the deletion mark.
[0138] It can be understood that in the case of a remote procedure call timeout or a failure to update the storage, the deletion of the reference item may fail. If the deletion of the reference item fails, it is determined that the deletion process fails.
[0139] Step 512: Determine whether the deletion of the record corresponding to the metadata to be deleted from the file metadata table is successful.
[0140] If successful, the deletion process is completed; otherwise, the deletion fails.
[0141] Next, an exemplary description will be given of the specific implementation manner of maintaining the reference count of a data fingerprint in the case where the reference relationship between a file and the data fingerprint is written into an index table and the reference relationship and the data fingerprint are saved in the same table. Specifically, in one or more embodiments of this specification, the method further includes:
[0142] In response to a change occurring in a record in the index table, determine whether the change that has occurred is a change in the reference relationship between a file and a data fingerprint;
[0143] If so, in the case where the change that has occurred is an addition of a reference relationship between a file and a data fingerprint, increment the reference count corresponding to the data fingerprint; in the case where the change that has occurred is a reduction of the reference relationship between a file and a data fingerprint, decrement the reference count corresponding to the data fingerprint.
[0144] In the above embodiment, since the index table is used to save the reference relationship between a file and a data fingerprint, thus, when a new reference relationship is added, the index table can be updated in a timely manner, so that the reference count of the data fingerprint can be updated in a timely manner accordingly based on the change recorded in the index table, simplifying the maintenance process of the reference count and improving the performance.
[0145] In addition, in order to enable the index table to be updated in a timely manner during garbage collection of data fingerprints and ensure the correctness of the data, the method further includes:
[0146] In response to a change occurring in a record in the index table, determine whether the change that has occurred is that the data fingerprint has completed garbage collection;
[0147] If so, mark the record corresponding to the data fingerprint as deleted in the index table.
[0148] The record corresponding to the data fingerprint includes an index entry and a reference entry corresponding to the data fingerprint.
[0149] In the above embodiment, in the case where a data fingerprint occurs and completes garbage collection, the corresponding index entry and reference entry in the index table can be deleted in a timely manner, which can ensure the correctness of the records in the index table.
[0150] In addition, in combination with the above embodiment of maintaining the reference count of a data fingerprint, in order to be able to delete the data that is no longer referenced by a file through garbage collection in a timely manner and release storage resources, in one or more embodiments of this specification, the method further includes:
[0151] According to the reference count of the data fingerprint in the index table, determine whether the data fingerprint needs to be set to the garbage collection state;
[0152] If so, set the data fingerprint to the garbage collection state.
[0153] Among them, setting the data fingerprint to the garbage collection state may include: setting the index entry and reference entry corresponding to the data fingerprint to the garbage collection state.
[0154] For example: in the above embodiment, the garbage collection of data fingerprints with a reference count of zero can be performed in two parts: foreground and background. Among them, the foreground can refer to the GC service querying the reference count of the data fingerprint and marking the data fingerprint with a reference count of zero as the garbage collection state; the background can refer to the GC service listening to the table content change notification channel, and when consuming a certain notification indicating that the record corresponding to a certain data fingerprint is marked as the garbage collection state, performing the actual garbage collection action.
[0155] Next, an exemplary description of the specific processing procedure for setting the garbage collection state will be given.
[0156] Figure 6 The flowchart of the processing procedure for garbage collection of a data processing method provided by an embodiment of this specification is shown, which specifically includes the following steps.
[0157] Step 602: Traverse the index table to retrieve the index entry with a reference count of zero for the data fingerprint currently traversed.
[0158] Step 604: Determine whether the traversal is over.
[0159] Step 606: If the traversal is not over, determine whether the data fingerprint with a reference count of zero currently traversed is in the normal reference state, that is, not in the garbage collection state or the garbage collection has been completed.
[0160] Step 608: If so, set the data fingerprint with a reference count of zero currently traversed to the garbage collection state.
[0161] If the traversal is not over, traverse the next record in the index table.
[0162] If the traversal is over, exit the processing procedure.
[0163] In the above embodiment, by saving the reference relationship between the file and the data fingerprint to the index table of the data fingerprint and maintaining the reference count based on the update of the index table, it is possible to scan out the record corresponding to the data fingerprint with a reference count of zero as the candidate record to be garbage collected. Therefore, setting the data fingerprint with a reference count of zero to the garbage collection state, in this way, the garbage collection mechanism will delete the record corresponding to the data fingerprint in the garbage collection state, thereby effectively improving the garbage collection efficiency.
[0164] In the above embodiments, the judgment on whether the records in the index table have changed can be made based on any possible implementation, and this specification does not limit it. For example, in one or more embodiments of this specification, the judgment is made based on the mechanism of the table content change notification channel. Specifically, the method further includes:
[0165] Obtain the notification in the table content change notification channel of the index table;
[0166] Judge whether the records in the index table have changed according to the notification;
[0167] When the processing step marked as deletion is successfully executed, return a consumption completion message to the table content change notification channel;
[0168] When the processing step marked as deletion fails, return a consumption failure message to the table content change notification channel and re-enter the step of obtaining the notification in the table content change notification channel of the index table.
[0169] In the above embodiments, the judgment on whether the records in the index table have changed can be completed based on the consumption logic of the table content change notification channel. Through the sequential characteristics of the consumption logic of the table content change notification channel, the subsequent processing process in the case of record change can be carried out more accurately, timely and efficiently, improving the processing efficiency.
[0170] Next, an exemplary description is given of the subsequent processing process in the case of index table record change based on the mechanism of the table content change notification channel in the above embodiments.
[0171] Figure 7 The flowchart of the processing process based on the table content change notification channel of a data processing method provided by an embodiment of this specification is shown, which specifically includes the following steps.
[0172] Step 702: Obtain the notification in the table content change notification channel by listening to the table content change notification channel.
[0173] Step 704: Judge whether the reference item in the index table has changed according to the notification.
[0174] Step 706: If the reference item has changed, judge whether it is the reference relationship between the newly added file and the data fingerprint.
[0175] Step 708: If it is the reference relationship between the newly added file and the data fingerprint, increment the reference count corresponding to the data fingerprint of the reference item.
[0176] Step 710: If it is the reference relationship between the reduced file and the data fingerprint, decrement the reference count corresponding to the data fingerprint of the reference item.
[0177] Step 712: If it is not a change in the reference item, determine whether the change is that the data fingerprint has completed garbage collection.
[0178] Step 714: If garbage collection is completed, mark the record corresponding to the data fingerprint as deleted.
[0179] Step 716: Determine whether the marking is successful.
[0180] Step 718: If the marking is successful, return a consumption completion message to the table content change notification channel.
[0181] Step 720: If the marking fails, return a consumption failure message to the table content change notification channel, and wait to retry to obtain the notification in the table content change notification channel of the index table.
[0182] In the above embodiments, based on the table content change notification channel, by consuming the notifications in the channel, it is possible to achieve the update of consistent reference counting and improve the recycling efficiency of garbage data.
[0183] Corresponding to the above method embodiments, this specification also provides embodiments of a data processing device, Figure 8 showing a schematic structural diagram of a data processing device provided by an embodiment of this specification. As Figure 8 shown, the device includes:
[0184] A verification determination module 802, configured to determine a data fingerprint to be verified, where the data fingerprint is an identifier for distinguishing different data contents;
[0185] An index data acquisition module 804, configured to obtain metadata corresponding to the data fingerprint from an index table of the data fingerprint, where the index table is used to store index items of the data fingerprint and metadata corresponding to the data fingerprint, and the metadata is metadata of a file referring to the data fingerprint;
[0186] A data verification module 806, configured to use the obtained metadata to locate an actually existing file, and obtain a verification result of the consistency between the data fingerprint and the actually existing file.
[0187] In one or more embodiments of this specification, the data verification module is configured to obtain metadata to be verified of a file to be verified from a file metadata table, where the file to be verified is a file referring to the data fingerprint, and the file metadata table is used to store metadata of actually existing files, compare the obtained metadata with the metadata to be verified, determine whether they are the same metadata, and if not, determine that there is an abnormality in the consistency between the data fingerprint and the actually existing file.
[0188] In one or more embodiments of this specification, the verification determination module includes:
[0189] A content acquisition sub-module configured to acquire the data content of the file to be verified for constructing a data fingerprint;
[0190] A primary key construction sub-module configured to construct a primary key of the data fingerprint to be queried based on the data content;
[0191] A data matching sub-module configured to use the primary key of the data fingerprint to be queried to query, in the index table of the data fingerprint, the data fingerprint corresponding to the index entry that matches the primary key of the data fingerprint to be queried as the data fingerprint to be verified, where the index entry is used to store the primary key of the corresponding data fingerprint.
[0192] In one or more embodiments of this specification, the data verification module is configured to use the acquired metadata to query, in the file metadata table, a record that matches the metadata, where the file metadata table is used to store the metadata of the actually existing files, and determine whether the acquired metadata is the same as the metadata stored in the record. If not, it is determined that there is an abnormality in the consistency between the data fingerprint and the actually existing file.
[0193] In one or more embodiments of this specification, the index table further includes a reference entry, where the reference entry is used to store the reference relationship between the file and the data fingerprint. The apparatus further includes:
[0194] A file creation module configured to, in response to a file creation request, determine whether there is a data fingerprint referenced by the file to be created corresponding to the file creation request in the index table;
[0195] A reference writing module configured to, if the file creation module determines that there is, write the reference relationship between the file to be created and the data fingerprint into the reference entry of the index table. If the file creation module determines that there is no, create an index entry corresponding to the data fingerprint referenced by the file to be created in the index table, and write the reference relationship between the file to be created and the data fingerprint into the reference entry of the index table.
[0196] In one or more embodiments of this specification, the apparatus further includes:
[0197] A fingerprint existence judgment module configured to determine whether there is an index entry corresponding to the data fingerprint referenced by the file to be created in the index table;
[0198] A change judgment module configured to, if the fingerprint existence judgment module determines that there is, determine whether the data fingerprint referenced by the file to be created in the index entry has changed;
[0199] A reference deletion module, configured to delete an entry corresponding to the reference relationship between the to-be-created file and the data fingerprint if the change judgment module determines that a change has occurred.
[0200] In one or more embodiments of the present specification, the apparatus further includes:
[0201] A fingerprint existence judgment module, configured to judge whether there is an index entry corresponding to the data fingerprint referenced by the to-be-created file in the index table;
[0202] A garbage collection judgment module, configured to judge whether the data fingerprint referenced by the to-be-created file is in a garbage collection state if the fingerprint existence judgment module determines that it exists;
[0203] A recovery cancellation module, configured to cancel the garbage collection state of the data fingerprint referenced by the to-be-created file if the garbage collection judgment module determines that it is in a garbage collection state.
[0204] In one or more embodiments of the present specification, the apparatus further includes:
[0205] A deletion marking module, configured to, in response to a file deletion request, mark the metadata to be deleted corresponding to the file deletion request as deleted in the file metadata table, where the file metadata table is used to store file metadata;
[0206] A reference deletion module, configured to delete the entry corresponding to the metadata to be deleted in the index table;
[0207] A marked data deletion module, configured to, when it is determined that the deletion of the entry is successful, delete the record corresponding to the metadata to be deleted from the file metadata table according to the deletion mark.
[0208] In one or more embodiments of the present specification, the apparatus further includes:
[0209] A reference change judgment module, configured to judge whether the change that occurs is a change in the reference relationship between a file and a data fingerprint in response to a change in the record in the index table;
[0210] A count update module, configured to, if the reference change judgment module determines that it is, increment the reference count corresponding to the data fingerprint when the change that occurs is an addition of a reference relationship between a file and a data fingerprint, and decrement the reference count corresponding to the data fingerprint when the change that occurs is a reduction of the reference relationship between a file and a data fingerprint.
[0211] In one or more embodiments of the present specification, the apparatus further includes:
[0212] A recycling change judgment module, configured to determine whether the change that occurs is that the data fingerprint completes garbage collection in response to a change in a record in the index table;
[0213] A deletion marking module, configured to mark the record corresponding to the data fingerprint as deleted in the index table if the recycling change judgment module determines that it is;
[0214] In one or more embodiments of this specification, the device further includes:
[0215] A counting judgment module, configured to determine whether the data fingerprint needs to be set to a garbage collection state according to the reference count of the data fingerprint in the index table;
[0216] A recycling setting module, configured to set the data fingerprint to a garbage collection state if the counting judgment module determines that it is;
[0217] In one or more embodiments of this specification, the device further includes:
[0218] A notification acquisition module, configured to acquire a notification in the table content change notification channel of the index table;
[0219] A record change judgment module, configured to determine whether a record in the index table has changed according to the notification;
[0220] A consumption success notification module, configured to return a consumption completion message to the table content change notification channel when the processing step marked as deleted is successfully executed;
[0221] A consumption failure notification module, configured to return a consumption failure message to the table content change notification channel when the processing step marked as deleted fails, and retry the step of acquiring a notification in the table content change notification channel of the index table.
[0222] The above is a schematic solution of a data processing device in this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.
[0223] Figure 9 FIG. shows a structural block diagram of a computing device 900 provided according to an embodiment of this specification. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.
[0224] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0225] In one embodiment of the present specification, the above components of the computing device 900 and Figure 9 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 9 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0226] The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 may also be a mobile or stationary server.
[0227] The processor 920 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.
[0228] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above data processing method.
[0229] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above data processing method when executed by a processor.
[0230] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above data processing method.
[0231] An embodiment of this specification also provides a computer program, which causes a computer to execute the steps of the above data processing method when executed on the computer.
[0232] An embodiment of this specification also provides a computer program product including computer programs / instructions, which implement the steps of the above data processing method when executed by a processor.
[0233] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above data processing method.
[0234] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0235] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunications signals.
[0236] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0237] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0238] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Determining a data fingerprint to be verified, wherein the data fingerprint is an identifier used to distinguish different data contents; Obtaining metadata corresponding to the data fingerprint from an index table of the data fingerprint, wherein the index table is used to store index items of the data fingerprint and metadata corresponding to the data fingerprint, and the metadata is metadata of a file that references the data fingerprint; The obtained metadata is used to locate the actual existing file, and a verification result of the consistency between the data fingerprint and the actual existing file is obtained.
2. The method according to claim 1, wherein locating an actual file using the obtained metadata and obtaining a consistency verification result between the data fingerprint and the actual file comprises: Obtaining metadata of the file to be verified from a file metadata table, wherein the file to be verified is a file that references the data fingerprint, and the file metadata table is used to store metadata of an actual file; Comparing the obtained metadata with the metadata to be verified to determine whether they are identical metadata; If they are not the same, it is determined that there is an anomaly in the consistency between the data fingerprint and the actual existing file.
3. The method according to claim 2, wherein determining the data fingerprint to be verified comprises: Obtaining data content of the file to be verified for constructing a data fingerprint; Based on the data content, construct a primary key of the data fingerprint to be queried; Using the primary key of the data fingerprint to be queried, in the index table of the data fingerprint, the data fingerprint corresponding to the index item that matches the primary key of the data fingerprint to be queried is queried as the data fingerprint to be verified, and the index item is used to store the primary key of the corresponding data fingerprint.
4. The method according to claim 1, wherein the data fingerprint corresponding to any index item in the index table is the data fingerprint to be verified; The method of locating an actual file by using the obtained metadata and obtaining a consistency verification result between the data fingerprint and the actual file includes: Using the obtained metadata, the records matching the metadata are searched in the file metadata table, which is used to store the metadata of the actual files. Determining whether the obtained metadata is the same as the metadata stored in the record; If they are not the same, it is determined that there is an anomaly in the consistency between the data fingerprint and the actual existing file.
5. The method according to claim 1, wherein the index table further comprises a reference item, wherein the reference item is used to store a reference relationship between the file and the data fingerprint; The method further comprises: In response to a file creation request, determining whether a data fingerprint referenced by a to-be-created file corresponding to the file creation request exists in the index table; If so, writing the reference relationship between the file to be created and the data fingerprint into the reference item of the index table; If it does not exist, an index item corresponding to the data fingerprint referenced by the file to be created is created in the index table, and a reference relationship between the file to be created and the data fingerprint is written into the reference item of the index table.
6. The method according to claim 5, further comprising: after writing the reference relationship between the file to be created and the data fingerprint into the reference item of the index table; Determine whether there is an index entry corresponding to the data fingerprint referenced by the file to be created in the index table; If so, determining whether the data fingerprint referenced by the file to be created in the index item has changed; If a change occurs, the reference item corresponding to the reference relationship between the file to be created and the data fingerprint is deleted.
7. The method according to claim 5 or 6, further comprising: after writing the reference relationship between the file to be created and the data fingerprint into the reference item of the index table; Determine whether there is an index entry corresponding to the data fingerprint referenced by the file to be created in the index table; If so, determining whether the data fingerprint referenced by the file to be created is in a garbage collection state; If it is in the garbage collection state, cancel the garbage collection state of the data fingerprint referenced by the file to be created.
8. The method according to claim 7, further comprising: In response to a file deletion request, marking the to-be-deleted metadata corresponding to the file deletion request in a file metadata table as deleted, the file metadata table being used to store metadata of the file; Deleting the reference item corresponding to the metadata to be deleted in the index table; When it is determined that the reference item is deleted successfully, the record corresponding to the metadata to be deleted is deleted from the file metadata table according to the deletion mark.
9. The method according to claim 5, further comprising: In response to a change in a record in the index table, determining whether the change is a change in a reference relationship between a file and a data fingerprint; If so, when the change is a new reference relationship between a file and a data fingerprint, the reference count corresponding to the data fingerprint is incremented; when the change is a reduction in the reference relationship between a file and a data fingerprint, the reference count corresponding to the data fingerprint is decremented.
10. The method according to claim 5 or 9, further comprising: In response to a change in a record in the index table, determining whether the change is a data fingerprint that completes garbage collection; If yes, the record corresponding to the data fingerprint is marked as deleted in the index table.
11. The method according to claim 9, further comprising: Determining whether to set the data fingerprint to a garbage collection state according to a reference count of the data fingerprint in the index table; If so, the data fingerprint is set to a garbage collection state.
12. The method according to claim 9, further comprising: Obtain notifications from a table content change notification channel of the index table; Determining whether a record in the index table has changed according to the notification; If the processing step marked as deleted is successfully executed, a consumption completion message is returned to the table content change notification channel; In the case where the processing step marked as deleted fails, a consumption failure message is returned to the table content change notification channel, and the step of obtaining the notification from the table content change notification channel of the index table is retried.
13. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the data processing method according to any one of claims 1 to 12 are implemented.
14. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program / instruction, which implements the steps of the data processing method according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Power network data quality check method and system
CN107679146A
Storage system data consistency verification method and system, equipment and medium
CN113608692A
Metadata management method for deleting duplicated data
CN115098481A
Method, system and device for checking directory consistency and storage medium
CN116069728A
File leakage risk detection method, equipment, storage medium and device
CN117034360A