Data processing method, computer device, storage medium and program product
By creating a de-deletion storage device group in the data center and adopting data processing strategies, identifying and deleting duplicate data between multiple storage devices, the problem of duplicate data in the data center is solved, and effective storage space savings and storage efficiency improvements are achieved.
Patent Information
- Application Number
- CN202411529280.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-30
AI Technical Summary
In the data center, duplicate data is not recognized and deleted between multiple storage devices, resulting in the inability to effectively save the total storage space.
By creating a de-deletion storage device group, multiple storage devices are added to the group, and the duplicate data modules between a single storage device and a plurality of storage devices are identified and deleted, respectively, using the first and second data processing strategies.
Deduplication of data between multiple storage devices is realized, effectively reducing the space usage of duplicate data between different storage devices and improving storage efficiency.
Smart Images

Figure CN119045747B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of storage technology, and in particular to a data processing method, a computer device, a storage medium and a program product. Background Art
[0002] Data deduplication (hereinafter referred to as "deduplication") is a very important capability in storage systems. It reduces the amount of data by identifying and deleting duplicate data in stored data, thereby meeting the growing demand for data storage. The current data deduplication function is based on the deduplication of a single storage device. In a data center, there are usually multiple sets of storage devices, and there will be duplicate data between different storage devices, which cannot be identified and deleted to save total storage space.
[0003] Therefore, there is an urgent need to propose a data processing method, computer equipment, storage medium and program product that can effectively save total storage space. Summary of the invention
[0004] Based on this, it is necessary to provide a data processing method, computer equipment, storage medium and program product that can effectively save total storage space in response to the above technical problems.
[0005] In a first aspect, a data processing method is provided, the method comprising:
[0006] Creating a deduplication storage device group, adding a plurality of storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein a storage device includes at least one data module;
[0007] According to the first data processing strategy, determining and deleting a first target data module in the first duplicate data module corresponding to the first target storage device;
[0008] According to the second data processing strategy, determining and deleting a second target data module in the second duplicate data modules corresponding to the plurality of storage devices in the target deduplication storage device group;
[0009] The data read and write request is responded to based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0010] Optionally, creating a deduplication storage device group, and adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group includes:
[0011] Create a deduplication storage device group;
[0012] Adding a second target storage device in the data center to the deduplication storage device group, wherein the physical links between the second target storage device and other storage devices in the deduplication storage device group are interconnected;
[0013] Based on the physical link connectivity relationship between the storage devices, an association relationship is established between the second target storage device and other storage devices in the deduplication storage device group.
[0014] Optionally, after establishing an association relationship between the second target storage device and other storage devices in the deduplication storage device group, the method further includes:
[0015] Testing the link connectivity between any two storage devices in the deduplication storage device group;
[0016] In response to detecting that the link connectivity test passes, adding the next storage device in the data center to the deduplication storage device group, and repeatedly performing the association relationship establishment and link connectivity test operations until the target number of storage devices in the data center are added;
[0017] Define the deduplication storage device group to which the target number of storage devices in the data center have been added as the target deduplication storage device group.
[0018] Optionally, after generating the target deduplication storage device group, the method further includes:
[0019] Based on the preset block size, the space occupied by the data in each storage device is divided into blocks to obtain multiple data modules.
[0020] Optionally, according to the first data processing strategy, determining and deleting the first target data module in the first duplicate data module corresponding to the first target storage device includes:
[0021] Traversing the data modules in the first target storage device, and calculating and determining the data fingerprint of the data module;
[0022] In response to detecting the presence of a first duplicate data fingerprint, defining a data module corresponding to the first duplicate data fingerprint as a first duplicate data module;
[0023] The third target data module selected in the first duplicate data module is retained, and the first target data modules except the third target data module in the first duplicate data module are deleted.
[0024] Optionally, the method further includes:
[0025] Pointing the data access address of the first repeated data module to the third target data module, and obtaining first relevant information of the third target data module to generate a one-to-one mapping relationship, wherein the first relevant information at least includes a data fingerprint, a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0026] and / or, obtaining second relevant information of other data modules and generating a one-to-one mapping relationship, wherein the second relevant information at least includes a data fingerprint, a data access address, a number of repetitions, a number of updates, and a number of reads and writes;
[0027] The multiple mapping relationships are saved in a data fingerprint library of the first target storage device.
[0028] Optionally, the method further includes:
[0029] Record the number of repetitions corresponding to the first repetitive data module;
[0030] Based on the number of repetitions, the count of the number of repetitions corresponding to the third target data module in the data fingerprint library is adjusted.
[0031] Optionally, according to the second data processing strategy, determining and deleting the second target data module in the second duplicate data modules corresponding to the plurality of storage devices in the target deduplication storage device group includes:
[0032] Obtaining data fingerprint libraries of multiple storage devices in the target deduplication storage device group, and comparing data fingerprints in different data fingerprint libraries;
[0033] In response to detecting the presence of the second duplicate data fingerprint, comparing the number of the second duplicate data fingerprints with the preset data module storage amount;
[0034] In response to detecting that the number of the second duplicate data fingerprints is greater than the preset data module storage amount, obtaining a deduplication weight of each data module in the second duplicate data module corresponding to the second duplicate data fingerprint in a corresponding storage device;
[0035] Based on the deduplication weight of each data module in the corresponding storage device, determining and deleting a second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint;
[0036] In response to detecting that the number of the second duplicate data fingerprints is less than or equal to the preset data module storage amount, the second duplicate data module corresponding to the second duplicate data fingerprint is retained.
[0037] Optionally, the method for determining the deduplication weight of each data module in the corresponding storage device includes:
[0038] Acquire third related information of each data module in the second repeated data module, wherein the third related information at least includes a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0039] Based on the third related information, a deduplication weight of each data module in a corresponding storage device is calculated and determined, wherein the calculation method includes:
[0040]
[0041] in, Indicates The deduplication weight of each storage device, , and Both represent weight coefficients, Indicates the number of repetitions. Indicates the number of read and write times. Indicates the number of updates. Indicates the number of storage devices;
[0042] In response to detecting that the deduplication weight has been calculated, the update times and the read / write times of the corresponding data module are cleared.
[0043] Optionally, based on the deduplication weight of each data module in the corresponding storage device, determining and deleting the second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint includes:
[0044] Based on the deduplication weights, sorting the data modules in the second duplicate data module in descending order to obtain a sorting result;
[0045] Based on the preset data module storage amount, the data modules in the sorting result are divided to determine the second target data module, and the deletion operation is performed, and the data access address corresponding to the reserved data module is adjusted.
[0046] Optionally, based on the target deduplication storage device group after the first target data module and the second target data module have been deleted, responding to the data read and write request includes:
[0047] In response to receiving a read request, detecting whether a data module corresponding to the read request in the target deduplication storage device group has been deduplicated;
[0048] In response to detecting that the data module corresponding to the read request in the target deduplication storage device group is not deduplicated, responding to the read request based on data in the local storage device;
[0049] In response to detecting that a data module corresponding to the read request in the target deduplication storage device group is deduplicated, adjusting the number of reads corresponding to the data module, and acquiring a data access address corresponding to the data fingerprint of the data module according to a data fingerprint library;
[0050] In response to detecting that the data access address is a local storage device, responding to the read request based on data in the local storage device;
[0051] In response to detecting that the data access address is not a local storage device, the read request is sent to the corresponding storage device, and data is obtained in response to the read request.
[0052] Optionally, based on the target deduplication storage device group after the first target data module and the second target data module have been deleted, responding to the data read and write request includes:
[0053] In response to receiving a write request, detecting whether a data module corresponding to the write request is pointed to by a data access address in a data fingerprint library;
[0054] In response to detecting that the data module corresponding to the write request is not pointed to by the data access address in the data fingerprint library, directly overwrite and write the data module to respond to the write request;
[0055] In response to detecting that the data module corresponding to the write request is pointed to by the data access address in the data fingerprint library, a new data module is written to respond to the write request, and at the same time, the count of the number of writes and the number of repeated occurrences of the data module is adjusted.
[0056] In a second aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0057] Creating a deduplication storage device group, adding a plurality of storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein a storage device includes at least one data module;
[0058] According to the first data processing strategy, determining and deleting a first target data module in the first duplicate data module corresponding to the first target storage device;
[0059] According to the second data processing strategy, determining and deleting a second target data module in the second duplicate data modules corresponding to the plurality of storage devices in the target deduplication storage device group;
[0060] The data read and write request is responded to based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0061] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0062] Creating a deduplication storage device group, adding a plurality of storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein a storage device includes at least one data module;
[0063] According to the first data processing strategy, determining and deleting a first target data module in the first duplicate data module corresponding to the first target storage device;
[0064] According to the second data processing strategy, determining and deleting a second target data module in the second duplicate data modules corresponding to the plurality of storage devices in the target deduplication storage device group;
[0065] The data read and write request is responded to based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0066] In a fourth aspect, a computer program product is provided, the computer program product comprising a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0067] Creating a deduplication storage device group, adding a plurality of storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein a storage device includes at least one data module;
[0068] According to the first data processing strategy, determining and deleting a first target data module in the first duplicate data module corresponding to the first target storage device;
[0069] According to the second data processing strategy, determining and deleting a second target data module in the second duplicate data modules corresponding to the plurality of storage devices in the target deduplication storage device group;
[0070] The data read and write request is responded to based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0071] The above-mentioned data processing method, computer device, storage medium and program product, the method includes: creating a deduplication storage device group, adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein a storage device includes at least one data module; according to a first data processing strategy, determining and deleting a first target data module in a first duplicate data module corresponding to a first target storage device; according to a second data processing strategy, determining and deleting a second target data module in a second duplicate data module corresponding to multiple storage devices in the target deduplication storage device group; based on the target deduplication storage device group after deleting the first target data module and the second target data module, responding to data read and write requests, the present application realizes deduplication of data between multiple sets of storage devices within the data center, which can reduce the space occupied by duplicate data between different storage devices and effectively improve storage efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 A schematic diagram of the overall structure of two sets of storage devices in a data processing method according to an embodiment respectively calculating respective duplicate data;
[0073] Figure 2 An application environment diagram of a data processing method in an embodiment;
[0074] Figure 3 is a flow chart of a data processing method in one embodiment;
[0075] Figure 4 A schematic diagram of the overall structure of calculating duplicate data between storage devices in a data processing method in one embodiment;
[0076] Figure 5 A schematic diagram of a storage device group creation process for a data processing method in one embodiment;
[0077] Figure 6 A schematic diagram of a process flow of calculating and deleting duplicate data in a single storage device of a data processing method in one embodiment;
[0078] Figure 7 A schematic diagram of information storage in a data fingerprint library of a data processing method in one embodiment;
[0079] Figure 8 A schematic diagram of a process flow of calculating and deleting duplicate data between storage devices in a data processing method in one embodiment;
[0080] Fig. 9 A schematic diagram of a host read IO processing flow of a data processing method in one embodiment;
[0081] Fig.10A schematic diagram of a host write IO processing flow of a data processing method in one embodiment;
[0082] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0083] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0084] It should be understood that in the description of the present application, unless the context clearly requires otherwise, words such as "include", "comprises", and the like throughout the specification should be interpreted as including rather than being exclusive or exhaustive; that is, as including but not limited to.
[0085] It should also be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0086] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing the steps, and do not specifically refer to the order or sequence, nor are they used to limit the present application. They are only for the convenience of describing the method of the present application, and cannot be understood as indicating the order of the steps. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0087] The data deduplication function of the storage system is usually implemented by scanning and calculating each data block or file stored in the current storage system and generating a unique identifier (usually called a fingerprint or hash value). If the two identifiers are the same, it means that the data is duplicated, and the system will delete the duplicate data block or file, leaving only a unique copy. The deduplication function is generally divided into two modes: online deduplication and post-deduplication. When new data is written, online deduplication calculates the unique identifier of the newly written data and compares it with the identifier of the existing data block. If they are the same, it is considered to be duplicate data and is deleted before writing. Post-deduplication is generally calculated and deleted after the data is written to the storage system when the system is relatively idle. According to the background technology, the current data deduplication function is based on the deduplication of a single storage device. In a data center, there are usually multiple sets of storage devices, and there will be duplicate data between different storage devices, which cannot be identified and deleted to save total storage space. Figure 1 As shown, fingerprint F is calculated for the data block of the storage device, and data with the same fingerprint are considered to be the same data. Data with the same fingerprint F exists both within the storage device and between storage devices, but only duplicate data within the storage device is identified, and duplicate data between storage devices is not identified.
[0088] To solve the above technical problems, the present application provides a data processing method, apparatus, device and storage medium. The method identifies and deletes duplicate data according to the dimensions of multiple sets of storage devices in a data center, which can more effectively save the total storage space. Multiple sets of storage devices are added to the same storage device group, and duplicate data is calculated and deleted between storage device groups. According to the number of repetitions and read and write frequencies of data modules in different storage devices, the most suitable storage device is selected to retain the duplicate data. When other storage devices that have deleted duplicate data read data blocks, they read from the storage device that retains the duplicate data, so as to meet the balance requirements of reliability, performance and space utilization of the storage system operation.
[0089] The data processing method provided in this application can be applied to Figure 2 In the application environment shown, the terminal 102 communicates with the data processing platform set on the server 104 through the network, wherein the terminal 102 can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0090] In one embodiment, Figure 3 As shown, a data processing method is provided, which is applied to Figure 2 The terminal in is used as an example to illustrate, including the following steps:
[0091] S1: creating a deduplication storage device group, and adding a plurality of storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein one storage device includes at least one data module.
[0092] It should be noted that deduplication refers to duplicate data deletion, abbreviated as deduplication; a deduplication storage device group includes multiple sets of storage devices, that is, it is responsible for adding multiple sets of storage devices to the same deduplication storage device group, completing the mutual identification and interconnection between different storage devices, so as to identify and delete duplicate data between the storage devices in the storage device group; a data module is also called a data block, which is obtained by dividing the data of each storage device into blocks of fixed size, and a storage device can include multiple data blocks.
[0093] In some specific embodiments, Figure 5 As shown, creating a deduplication storage device group, adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group includes:
[0094] Create a deduplication storage device group, that is, add / create a deduplication storage device group in the storage system;
[0095] Adding a second target storage device in the data center to the deduplication storage device group, wherein the physical links between the second target storage device and other storage devices in the deduplication storage device group are interconnected, that is, adding the storage devices in the data center (the physical links between each storage device need to be interconnected) to the deduplication storage device group;
[0096] Based on the physical link connectivity relationship between storage devices, an association relationship is established between the second target storage device and other storage devices in the deduplication storage device group, that is, each storage device is traversed to establish an association relationship between it and other storage devices in the deduplication storage device group to establish an interconnection relationship between different storages.
[0097] In some specific implementations, after establishing an association relationship between the second target storage device and other storage devices in the deduplication storage device group, the method further includes:
[0098] Testing the link connectivity between any two storage devices in the deduplication storage device group, that is, testing the link connectivity between every two sets of storage devices in the deduplication storage device group;
[0099] In response to detecting that the link connectivity test has passed, adding the next storage device in the data center to the deduplication storage device group, and repeatedly performing the association relationship establishment and link connectivity test operations until the target number of storage devices in the data center is added, wherein the target number generally refers to the number of all storage devices in the data center. When a storage device passes the connectivity test, continue to add the next storage device, and repeat the above steps until all storage devices in the data center are added;
[0100] Define the deduplication storage device group to which the target number of storage devices in the data center have been added as the target deduplication storage device group.
[0101] In some specific implementations, after generating the target deduplication storage device group, the method further includes:
[0102] Based on a preset block size, the space occupied by data in each storage device is divided into blocks to obtain multiple data modules, wherein the preset block size can be set according to actual needs.
[0103] Specifically, the data in each storage device is processed in blocks according to a fixed size (i.e., the preset block size). The block size can be configured by the user. The smaller the block, the higher the repetition rate, but the redirection forwarding when reading data will also increase, which will affect the IO (Input / Output) performance. Therefore, the block size can be adjusted according to the actual situation. If the business expects a smaller space and is not sensitive to performance impact, it can be appropriately adjusted to a smaller block, otherwise it will be adjusted to a larger block. The block size of the data module can be adjusted by dynamic configuration, such as setting the block size to 512KB or 1MB.
[0104] In the above implementation, by creating a storage device group and adding multiple sets of storage devices to the same storage device group, mutual identification and interconnection between different storage devices are completed, so as to realize identification and deletion of duplicate data between multiple storage devices in the storage device group.
[0105] S2: According to the first data processing strategy, determine and delete a first target data module in the first duplicate data module corresponding to the first target storage device.
[0106] It should be noted that the first data processing strategy refers to determining the duplicate data in a single storage device by calculating the fingerprint of the data module in a single storage device, and performing corresponding deletion operations to ensure that the data in a single storage device are all different. It triggers the background task (i.e., triggers the task corresponding to the first data processing strategy) based on a set time interval. The time interval can be set according to actual needs, such as one hour; the first target storage device refers to a single storage device in the deduplication storage device for calculating duplicate data; the first target data module refers to a data module that needs to be deleted from multiple duplicate data modules.
[0107] In some specific implementations, according to the first data processing strategy, determining and deleting the first target data module in the first duplicate data module corresponding to the first target storage device includes:
[0108] Traversing the data modules in the first target storage device, calculating and determining the data fingerprint of the data module, wherein the calculation method of the data fingerprint is a common method, such as a hash function, etc., and the specific calculation process thereof is not repeated here;
[0109] In response to detecting the presence of a first duplicate data fingerprint, defining a data module corresponding to the first duplicate data fingerprint as a first duplicate data module, that is, first comparing fingerprints of data blocks in the storage device, identifying data in data blocks with the same fingerprint as duplicate data, and defining the corresponding data module as the first duplicate data module;
[0110] The third target data module selected in the first duplicate data module is retained, and the first target data module except the third target data module in the first duplicate data module is deleted, wherein the third target data module refers to any data module in the first duplicate data module, and the first target data module refers to other data modules in the first duplicate data module. Generally, only one copy of data is retained in a single storage device, and other duplicate data is deleted.
[0111] Among them, the order of fingerprint calculation, comparison and deletion of the data module can be to calculate, compare and delete the data modules in the storage device one by one. For example, after the fingerprint of the current data module has been calculated and compared, the fingerprint of the next data module is calculated and compared with the fingerprint of the calculated data module. If there is a duplicate, any data module is deleted. The fingerprint of all data modules in the storage device can also be calculated at the same time, and the above steps are used to determine the duplicate fingerprints and delete the corresponding data modules. The specific process will not be repeated here.
[0112] In some embodiments, the method further comprises:
[0113] The data access address of the first duplicate data module is pointed to the third target data module, and the first relevant information of the third target data module is obtained to generate a one-to-one mapping relationship, wherein the first relevant information at least includes a data fingerprint, a number of repetitions, a number of updates, and a number of reads and writes, that is, a corresponding mapping relationship is generated based on the data fingerprint, the number of repetitions, the number of updates, the number of reads and writes, and the data access address, and is saved in a mapping table, wherein the number of repetitions, the number of updates, and the number of reads and writes constitute deduplication weight information of the corresponding data block, and the mapping table is as follows: Figure 7 As shown, the data fingerprint is F1...F n , the corresponding number of repetitions is Ndup1...Ndupn, the number of updates is IOPS_w1...IOPS_wn, the number of reads and writes is IOPS_r1...IOPS_rn, the corresponding data access address is Storage1-LUN1-Addr1...Storagen-LUNn-Addrn, IOPS (Input / Output Per Second) refers to the number of IOs processed per second by the storage device;
[0114] and / or, obtaining second relevant information of other data modules and generating a one-to-one mapping relationship, wherein the second relevant information at least includes a data fingerprint, a data access address, a number of repeated occurrences, a number of updates, and a number of read and write times, wherein the other data modules refer to data modules corresponding to non-repeated fingerprints, and if there are data modules corresponding to non-repeated fingerprints, obtaining corresponding information and generating a corresponding mapping relationship, and saving multiple mapping relationships into a mapping table as in the above steps;
[0115] The multiple mapping relationships are saved in the data fingerprint library of the first target storage device, that is, the mapping table generated by the multiple mapping relationships is saved in the data fingerprint library corresponding to the first target storage device.
[0116] In some embodiments, the method further comprises:
[0117] Record the number of repetitions corresponding to the first repeated data module, that is, the number of times the data module is repeated, and the number of repetitions is the corresponding number;
[0118] Based on the number of repetitions, the count of the number of repetitions corresponding to the third target data module in the data fingerprint library is adjusted, that is, if the fingerprint calculation and comparison are performed on the data blocks one by one, the repetition count of the currently processed data block is continuously updated to the fingerprint library; if the fingerprint calculation and comparison are performed on the data blocks at the same time, then in this processing stage, the count of the number of repetitions corresponding to the third target data module is adjusted once according to the statistical results of the number of repetitions.
[0119] Specifically, Figure 6 As shown, the specific steps are as follows: the background task is triggered periodically, all data modules of a single storage device are traversed, and the fingerprint of each data module is calculated; whether it is a repeated fingerprint is determined; if it is a repeated fingerprint, the repeated count of the data module is updated to the fingerprint library, the data block is deleted, and its access address is pointed to the data block of the retained fingerprint; if it is a non-duplicate fingerprint, the fingerprint and data access address are saved in the data fingerprint library; and other relevant information of the data block is saved at the same time.
[0120] In the above implementation, duplicate data in a single storage device is deleted to ensure that data in one storage device are all different, which can effectively improve the utilization rate of the storage device and enhance the processing efficiency of duplicate data between multiple devices.
[0121] S3: According to the second data processing strategy, determine and delete a second target data module in the second duplicate data modules corresponding to the multiple storage devices in the target deduplication storage device group.
[0122] It should be noted that the second data processing strategy refers to determining the duplicate data between multiple storage devices by calculating the fingerprints of the data modules in multiple storage devices, and determining the data modules that need to be deleted based on the weight information of the data blocks corresponding to the duplicate data in their storage devices, and performing the corresponding deletion operations. It triggers the background task (i.e., triggers the task corresponding to the second data processing strategy) based on a set time interval. The time interval can be set according to actual needs, such as one hour, etc.; the second target data module refers to the data module that needs to be deleted among the multiple duplicate data modules.
[0123] In some specific embodiments, Figure 4 As shown, according to the second data processing strategy, determining and deleting the second target data module in the second duplicate data modules corresponding to the multiple storage devices in the target deduplication storage device group includes:
[0124] Obtaining data fingerprint libraries of multiple storage devices in the target deduplication storage device group, and comparing data fingerprints in different data fingerprint libraries;
[0125] In response to detecting the existence of the second duplicate data fingerprint, comparing the number of the second duplicate data fingerprints and the preset data module storage amount, wherein the preset data module storage amount can be set according to actual needs, and the second duplicate data fingerprint refers to the duplicate fingerprints corresponding to the multiple storage devices;
[0126] In response to detecting that the number of the second duplicate data fingerprints is greater than the preset data module storage amount, obtaining a deduplication weight of each data module in the second duplicate data module corresponding to the second duplicate data fingerprint in the corresponding storage device, wherein the deduplication weight is used to describe weight information of retaining the data module corresponding to the duplicate fingerprint in the corresponding storage device;
[0127] Based on the deduplication weight of each data module in the corresponding storage device, determine and delete the second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint, that is, delete the data modules other than the preset data module storage amount;
[0128] In response to detecting that the number of the second duplicate data fingerprints is less than or equal to the preset data module storage amount, the second duplicate data module corresponding to the second duplicate data fingerprint is retained.
[0129] In some specific implementations, a method for determining a deduplication weight of each data module in a corresponding storage device includes:
[0130] Acquire third related information of each data module in the second repeated data module, wherein the third related information at least includes a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0131] Based on the third related information, a deduplication weight of each data module in a corresponding storage device is calculated and determined, wherein the calculation method includes:
[0132]
[0133] in, Indicates The deduplication weight of each storage device, , and Both represent weight coefficients, Indicates the number of repetitions. Indicates the number of read and write times. Indicates the number of updates. Indicates the number of storage devices;
[0134] In response to detecting that the deduplication weight has been calculated, the update times and the read / write times of the corresponding data module are cleared.
[0135] In some specific implementations, based on the deduplication weight of each data module in the corresponding storage device, determining and deleting the second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint includes:
[0136] Based on the deduplication weights, sorting the data modules in the second duplicate data module in descending order to obtain a sorting result;
[0137] Based on the preset data module storage capacity, the data modules in the sorting result are divided to determine the second target data module, and the deletion operation is performed, and the data access address corresponding to the retained data module is adjusted. That is, generally only the number of duplicate data corresponding to the preset data module storage capacity needs to be retained. Assuming that the preset data module storage capacity is M, only the data modules corresponding to the first M copies of the data in the sorting result need to be retained, and the other data modules (i.e., the second target data modules) are deleted in the corresponding storage device, and the data access address stored in the data fingerprint library corresponding to each storage device is updated.
[0138] Specifically, Figure 8 As shown, the fingerprints of data modules between storage devices are compared, and data blocks with the same fingerprints are identified as duplicate data. The corresponding deduplication and retention policies include:
[0139] Only M copies of data are retained, and the data blocks corresponding to the duplicate data are deleted. M can be configured by the user. The higher the M value, the higher the data reliability, but the duplicate data will occupy more space. In order to improve access performance, the number of duplicate occurrences and the frequency of historical read and write data are used to jointly decide which storage devices to select to retain data. The overall selection principles are as follows:
[0140] If the number of duplicate data Ndup on the storage device is greater, the storage device is preferentially selected to retain the duplicate data;
[0141] If the number of IOPS_r of reading the data block within a period of time on the storage device is greater, the storage device is preferentially selected to retain duplicate data, wherein the period of time may be the cycle of each calculation and deletion of duplicate data;
[0142] If the number of IOPS_w times the data block is updated on the storage device within a period of time is greater, it means that the data block changes frequently, and the storage device should not be selected to retain duplicate data;
[0143] According to the above principles, the number of repetitions of the data block of each storage device and the proportion of IO read and write times in all devices are determined, and the weight Weight of each storage device to retain the data block is calculated. Among them, when the deduplication weight calculation is completed, the IOPS_r and IOPS_w of the corresponding data block are cleared, and the influence weights of factors such as Ndup, IOPS_r, IOPS_w are adjusted by K1, K2 and K3. K1, K2 and K3 can be configured according to the requirements of data reading efficiency and space utilization. Finally, each data block uses the M storage devices with the highest weight as the storage devices for retaining data according to the set M value. At the same time, the data access address stored in the data fingerprint library corresponding to each storage device is updated, and the data blocks in the storage devices that do not need to retain data are deleted.
[0144] In the above implementation, based on the duplicate data identification and deletion results within a single storage device, duplicate data between multiple storage devices is identified and deleted. When identifying and deleting duplicate data, the deletion and retention policies are adjusted with priority based on factors such as the number of duplicate occurrences and the read and write access frequency, thereby improving the efficiency of accessing duplicate data, effectively reducing the space occupied by duplicate data between storage devices, and improving storage efficiency. When retaining duplicate data, copies can be retained on multiple devices to achieve redundant protection, thereby meeting different requirements for performance in reading duplicate data, storage space utilization, and reliability, and improving the reliability of storage system operation.
[0145] S4: Responding to the data read and write request based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0146] It should be noted that data reading and writing refers to a request to read data from or write data to a storage device.
[0147] In some specific implementations, based on the target deduplication storage device group after the first target data module and the second target data module have been deleted, responding to the data read and write request includes:
[0148] In response to receiving a read request, detecting whether a data module corresponding to the read request in the target deduplication storage device group has been deduplicated;
[0149] In response to detecting that the data module corresponding to the read request in the target deduplication storage device group is not deduplicated, responding to the read request based on data in the local storage device;
[0150] In response to detecting that a data module corresponding to the read request in the target deduplication storage device group is deduplicated, adjusting the number of reads corresponding to the data module, and acquiring a data access address corresponding to the data fingerprint of the data module according to a data fingerprint library;
[0151] In response to detecting that the data access address is a local storage device, responding to the read request based on data in the local storage device;
[0152] In response to detecting that the data access address is not a local storage device, the read request is sent to the corresponding storage device, and data is obtained in response to the read request.
[0153] Specifically, Fig. 9 As shown, the host sends a read IO request; determines whether the data corresponding to the request has been deduplicated; if not, reads the data directly from the local storage device and returns the host IO request; if it has been deduplicated, updates the read IOPS_r++ (i.e., the number of reads) of this data block, and obtains the access address corresponding to the fingerprint of the requested data block according to the data fingerprint library; if the access address is local, reads the data directly from the local storage device and returns the host IO request; if the access address is on other storage devices, sends a request to the corresponding storage device to obtain data, and then returns the host IO request.
[0154] In some specific implementations, based on the target deduplication storage device group after the first target data module and the second target data module have been deleted, responding to the data read and write request includes:
[0155] In response to receiving a write request, detecting whether a data module corresponding to the write request is pointed to by a data access address in a data fingerprint library;
[0156] In response to detecting that the data module corresponding to the write request is not pointed to by the data access address in the data fingerprint library, directly overwrite and write the data module to respond to the write request;
[0157] In response to detecting that the data module corresponding to the write request is pointed to by the data access address in the data fingerprint library, a new data module is written to respond to the write request, and at the same time, the count of the number of writes and the number of repeated occurrences of the data module is adjusted.
[0158] Specifically, Fig.10 As shown, the host sends a write IO request; it determines whether the data block corresponding to the address space of the IO write request is still referenced (that is, it is pointed to by the data access address in the data fingerprint library); if it is not referenced, it directly overwrites the original data block and returns the IO write completion to the host; if it is referenced, it writes a new data block, updates the write IOPS_w++ of the original data block, updates the duplicate count Ndup-- of the original data block, and returns the IO write completion to the host.
[0159] Among them, in the IO write processing flow, if the original duplicate data block changes, only the duplicate count Ndup of the original data block is updated, and the background calculation and deduplication data module running above automatically deletes the original data block that has no valid reference according to the calculated deduplication weight.
[0160] In the above implementation, for the read processing of duplicate data blocks, the data space in the corresponding storage device is accessed according to the data access address stored in the data fingerprint library. For the write processing of duplicate data blocks, if the access address of the original duplicate data has no reference, the original data space is overwritten. If the access address of the original duplicate data still has a reference, a new data space is applied for to save the written data, and the original duplicate data is processed to facilitate the calculation of the deduplication weight, thereby further improving the efficiency of duplicate data processing.
[0161] In some embodiments, the method further comprises:
[0162] Based on a preset time period, obtaining the number of deleted data modules of the target storage device in the target deduplication storage device group and the corresponding data module identifier, and defining the number as a first target number, wherein the preset time period can be set according to actual needs, such as one day, etc.;
[0163] Based on the number of occurrences of the target data module identifier, determine the number of deletions of the data module corresponding to the target data module identifier, and define it as a second target number;
[0164] According to the first target number and the second target number, weight information corresponding to the data module corresponding to the target data module identifier is determined, and the calculation method includes:
[0165]
[0166] in, Represents weight information, represents the first target number, Indicates the second target number, , All represent weight coefficients;
[0167] In response to detecting that the value of the weight information is greater than a preset threshold, when it is detected that the storage devices to which the storage modules corresponding to the duplicate data between multiple storage devices belong include the target storage device, and the number of duplicate storage modules is greater than the preset data module storage capacity, the storage modules in the target storage device are directly included in the deduplication plan without the need to recalculate and determine the deduplication weight of the data module in the corresponding storage device, wherein the preset threshold can be set according to actual needs.
[0168] In the above implementation, based on two dimensions, namely, the first target number of deleted data modules of the target storage device and the second target number of corresponding data modules of the target data module identifier, the data modules that can be directly deleted when duplicate data are detected between multiple storage devices are determined, thereby further improving the duplicate data processing efficiency.
[0169] In the above-mentioned data processing method, the method includes: creating a deduplication storage device group, adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein one storage device includes at least one data module; according to a first data processing strategy, determining and deleting a first target data module in a first duplicate data module corresponding to a first target storage device; according to a second data processing strategy, determining and deleting a second target data module in a second duplicate data module corresponding to multiple storage devices in the target deduplication storage device group; based on the target deduplication storage device group after deleting the first target data module and the second target data module, responding to a data read and write request, the present application calculates duplicate data between storage devices based on the dimensions between multiple sets of storage devices in a data center, realizes duplicate data deletion between multiple sets of storage, effectively reduces the space occupied by duplicate data between storage devices, and improves storage efficiency; according to index factors such as the number of repetitions and read and write frequencies of data blocks in different storage devices, calculates multiple optimal storage devices for retaining duplicate data, and adjusts the weights of different factors to meet different requirements for reading duplicate data performance, storage space utilization, and reliability.
[0170] It should be understood that although Figure 3-Figure 10 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 3-Figure 10 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0171] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.11As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a data processing method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0172] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0173] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0174] S1: creating a deduplication storage device group, adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein one storage device includes at least one data module;
[0175] S2: According to the first data processing strategy, determine and delete the first target data module in the first duplicate data module corresponding to the first target storage device;
[0176] S3: according to the second data processing strategy, determine and delete the second target data module in the second duplicate data modules corresponding to the multiple storage devices in the target deduplication storage device group;
[0177] S4: Responding to the data read and write request based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0178] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0179] Create a deduplication storage device group;
[0180] Adding a second target storage device in the data center to the deduplication storage device group, wherein the physical links between the second target storage device and other storage devices in the deduplication storage device group are interconnected;
[0181] Based on the physical link connectivity relationship between the storage devices, an association relationship is established between the second target storage device and other storage devices in the deduplication storage device group.
[0182] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0183] Testing the link connectivity between any two storage devices in the deduplication storage device group;
[0184] In response to detecting that the link connectivity test passes, adding the next storage device in the data center to the deduplication storage device group, and repeatedly performing the association relationship establishment and link connectivity test operations until the target number of storage devices in the data center are added;
[0185] Define the deduplication storage device group to which the target number of storage devices in the data center have been added as the target deduplication storage device group.
[0186] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0187] Based on the preset block size, the space occupied by the data in each storage device is divided into blocks to obtain multiple data modules.
[0188] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0189] Traversing the data modules in the first target storage device, and calculating and determining the data fingerprint of the data module;
[0190] In response to detecting the presence of a first duplicate data fingerprint, defining a data module corresponding to the first duplicate data fingerprint as a first duplicate data module;
[0191] The third target data module selected in the first duplicate data module is retained, and the first target data modules except the third target data module in the first duplicate data module are deleted.
[0192] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0193] Pointing the data access address of the first repeated data module to the third target data module, and obtaining first relevant information of the third target data module to generate a one-to-one mapping relationship, wherein the first relevant information at least includes a data fingerprint, a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0194] and / or, obtaining second relevant information of other data modules and generating a one-to-one mapping relationship, wherein the second relevant information at least includes a data fingerprint, a data access address, a number of repetitions, a number of updates, and a number of reads and writes;
[0195] The multiple mapping relationships are saved in a data fingerprint library of the first target storage device.
[0196] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0197] Record the number of repetitions corresponding to the first repetitive data module;
[0198] Based on the number of repetitions, the count of the number of repetitions corresponding to the third target data module in the data fingerprint library is adjusted.
[0199] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0200] Obtaining data fingerprint libraries of multiple storage devices in the target deduplication storage device group, and comparing data fingerprints in different data fingerprint libraries;
[0201] In response to detecting the presence of the second duplicate data fingerprint, comparing the number of the second duplicate data fingerprints with the preset data module storage amount;
[0202] In response to detecting that the number of the second duplicate data fingerprints is greater than the preset data module storage amount, obtaining a deduplication weight of each data module in the second duplicate data module corresponding to the second duplicate data fingerprint in a corresponding storage device;
[0203] Based on the deduplication weight of each data module in the corresponding storage device, determine and delete a second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint;
[0204] In response to detecting that the number of the second duplicate data fingerprints is less than or equal to the preset data module storage amount, the second duplicate data module corresponding to the second duplicate data fingerprint is retained.
[0205] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0206] Acquire third related information of each data module in the second repeated data module, wherein the third related information at least includes a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0207] Based on the third related information, a deduplication weight of each data module in a corresponding storage device is calculated and determined, wherein the calculation method includes:
[0208]
[0209] in, Indicates The deduplication weight of each storage device, , and Both represent weight coefficients, Indicates the number of repetitions. Indicates the number of read and write times. Indicates the number of updates. Indicates the number of storage devices;
[0210] In response to detecting that the deduplication weight has been calculated, the update times and the read / write times of the corresponding data module are cleared.
[0211] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0212] Based on the deduplication weights, sorting the data modules in the second duplicate data module in descending order to obtain a sorting result;
[0213] Based on the preset data module storage amount, the data modules in the sorting result are divided to determine the second target data module, and the deletion operation is performed, and the data access address corresponding to the reserved data module is adjusted.
[0214] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0215] In response to receiving a read request, detecting whether a data module corresponding to the read request in the target deduplication storage device group has been deduplicated;
[0216] In response to detecting that the data module corresponding to the read request in the target deduplication storage device group is not deduplicated, responding to the read request based on data in the local storage device;
[0217] In response to detecting that a data module corresponding to the read request in the target deduplication storage device group is deduplicated, adjusting the number of reads corresponding to the data module, and acquiring a data access address corresponding to the data fingerprint of the data module according to a data fingerprint library;
[0218] In response to detecting that the data access address is a local storage device, responding to the read request based on data in the local storage device;
[0219] In response to detecting that the data access address is not a local storage device, the read request is sent to the corresponding storage device, and data is obtained in response to the read request.
[0220] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0221] In response to receiving a write request, detecting whether a data module corresponding to the write request is pointed to by a data access address in a data fingerprint library;
[0222] In response to detecting that the data module corresponding to the write request is not pointed to by the data access address in the data fingerprint library, directly overwrite and write the data module to respond to the write request;
[0223] In response to detecting that the data module corresponding to the write request is pointed to by the data access address in the data fingerprint library, a new data module is written to respond to the write request, and at the same time, the count of the number of writes and the number of repeated occurrences of the data module is adjusted.
[0224] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0225] S1: creating a deduplication storage device group, adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein one storage device includes at least one data module;
[0226] S2: According to the first data processing strategy, determine and delete the first target data module in the first duplicate data module corresponding to the first target storage device;
[0227] S3: according to the second data processing strategy, determine and delete the second target data module in the second duplicate data modules corresponding to the multiple storage devices in the target deduplication storage device group;
[0228] S4: Responding to the data read and write request based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0229] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0230] Create a deduplication storage device group;
[0231] Adding a second target storage device in the data center to the deduplication storage device group, wherein the physical links between the second target storage device and other storage devices in the deduplication storage device group are interconnected;
[0232] Based on the physical link connectivity relationship between the storage devices, an association relationship is established between the second target storage device and other storage devices in the deduplication storage device group.
[0233] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0234] Testing the link connectivity between any two storage devices in the deduplication storage device group;
[0235] In response to detecting that the link connectivity test passes, adding the next storage device in the data center to the deduplication storage device group, and repeatedly performing the association relationship establishment and link connectivity test operations until the target number of storage devices in the data center are added;
[0236] Define the deduplication storage device group to which the target number of storage devices in the data center have been added as the target deduplication storage device group.
[0237] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0238] Based on the preset block size, the space occupied by the data in each storage device is divided into blocks to obtain multiple data modules.
[0239] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0240] Traversing the data modules in the first target storage device, and calculating and determining the data fingerprint of the data module;
[0241] In response to detecting the presence of a first duplicate data fingerprint, defining a data module corresponding to the first duplicate data fingerprint as a first duplicate data module;
[0242] The third target data module selected in the first duplicate data module is retained, and the first target data modules except the third target data module in the first duplicate data module are deleted.
[0243] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0244] Pointing the data access address of the first repeated data module to the third target data module, and obtaining first relevant information of the third target data module to generate a one-to-one mapping relationship, wherein the first relevant information at least includes a data fingerprint, a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0245] and / or, obtaining second relevant information of other data modules and generating a one-to-one mapping relationship, wherein the second relevant information at least includes a data fingerprint, a data access address, a number of repetitions, a number of updates, and a number of reads and writes;
[0246] The multiple mapping relationships are saved in a data fingerprint library of the first target storage device.
[0247] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0248] Record the number of repetitions corresponding to the first repetitive data module;
[0249] Based on the number of repetitions, the count of the number of repetitions corresponding to the third target data module in the data fingerprint library is adjusted.
[0250] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0251] Obtaining data fingerprint libraries of multiple storage devices in the target deduplication storage device group, and comparing data fingerprints in different data fingerprint libraries;
[0252] In response to detecting the presence of the second duplicate data fingerprint, comparing the number of the second duplicate data fingerprints with the preset data module storage amount;
[0253] In response to detecting that the number of the second duplicate data fingerprints is greater than the preset data module storage amount, obtaining a deduplication weight of each data module in the second duplicate data module corresponding to the second duplicate data fingerprint in a corresponding storage device;
[0254] Based on the deduplication weight of each data module in the corresponding storage device, determine and delete a second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint;
[0255] In response to detecting that the number of the second duplicate data fingerprints is less than or equal to the preset data module storage amount, the second duplicate data module corresponding to the second duplicate data fingerprint is retained.
[0256] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0257] Acquire third related information of each data module in the second repeated data module, wherein the third related information at least includes a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0258] Based on the third related information, a deduplication weight of each data module in a corresponding storage device is calculated and determined, wherein the calculation method includes:
[0259]
[0260] in, Indicates The deduplication weight of each storage device, , and Both represent weight coefficients, Indicates the number of repetitions. Indicates the number of read and write times. Indicates the number of updates. Indicates the number of storage devices;
[0261] In response to detecting that the deduplication weight has been calculated, the update times and the read / write times of the corresponding data module are cleared.
[0262] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0263] Based on the deduplication weights, sorting the data modules in the second duplicate data module in descending order to obtain a sorting result;
[0264] Based on the preset data module storage amount, the data modules in the sorting result are divided to determine the second target data module, and the deletion operation is performed, and the data access address corresponding to the reserved data module is adjusted.
[0265] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0266] In response to receiving a read request, detecting whether a data module corresponding to the read request in the target deduplication storage device group has been deduplicated;
[0267] In response to detecting that the data module corresponding to the read request in the target deduplication storage device group is not deduplicated, responding to the read request based on data in the local storage device;
[0268] In response to detecting that a data module corresponding to the read request in the target deduplication storage device group is deduplicated, adjusting the number of reads corresponding to the data module, and acquiring a data access address corresponding to the data fingerprint of the data module according to a data fingerprint library;
[0269] In response to detecting that the data access address is a local storage device, responding to the read request based on data in the local storage device;
[0270] In response to detecting that the data access address is not a local storage device, the read request is sent to the corresponding storage device, and data is obtained in response to the read request.
[0271] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0272] In response to receiving a write request, detecting whether a data module corresponding to the write request is pointed to by a data access address in a data fingerprint library;
[0273] In response to detecting that the data module corresponding to the write request is not pointed to by the data access address in the data fingerprint library, directly overwrite and write the data module to respond to the write request;
[0274] In response to detecting that the data module corresponding to the write request is pointed to by the data access address in the data fingerprint library, a new data module is written to respond to the write request, and at the same time, the count of the number of writes and the number of repeated occurrences of the data module is adjusted.
[0275] In one embodiment, a computer program product is provided, the computer program product comprising a computer program, the computer program when executed by a processor implements the following steps:
[0276] S1: creating a deduplication storage device group, adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein one storage device includes at least one data module;
[0277] S2: According to the first data processing strategy, determine and delete the first target data module in the first duplicate data module corresponding to the first target storage device;
[0278] S3: according to the second data processing strategy, determine and delete the second target data module in the second duplicate data modules corresponding to the multiple storage devices in the target deduplication storage device group;
[0279] S4: Responding to the data read and write request based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
[0280] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0281] Create a deduplication storage device group;
[0282] Adding a second target storage device in the data center to the deduplication storage device group, wherein the physical links between the second target storage device and other storage devices in the deduplication storage device group are interconnected;
[0283] Based on the physical link connectivity relationship between the storage devices, an association relationship is established between the second target storage device and other storage devices in the deduplication storage device group.
[0284] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0285] Testing the link connectivity between any two storage devices in the deduplication storage device group;
[0286] In response to detecting that the link connectivity test passes, adding the next storage device in the data center to the deduplication storage device group, and repeatedly performing the association relationship establishment and link connectivity test operations until the target number of storage devices in the data center are added;
[0287] Define the deduplication storage device group to which the target number of storage devices in the data center have been added as the target deduplication storage device group.
[0288] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0289] Based on the preset block size, the space occupied by the data in each storage device is divided into blocks to obtain multiple data modules.
[0290] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0291] Traversing the data modules in the first target storage device, and calculating and determining the data fingerprint of the data module;
[0292] In response to detecting the presence of a first duplicate data fingerprint, defining a data module corresponding to the first duplicate data fingerprint as a first duplicate data module;
[0293] The third target data module selected in the first duplicate data module is retained, and the first target data modules except the third target data module in the first duplicate data module are deleted.
[0294] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0295] Pointing the data access address of the first repeated data module to the third target data module, and obtaining first relevant information of the third target data module to generate a one-to-one mapping relationship, wherein the first relevant information at least includes a data fingerprint, a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0296] and / or, obtaining second relevant information of other data modules and generating a one-to-one mapping relationship, wherein the second relevant information at least includes a data fingerprint, a data access address, a number of repetitions, a number of updates, and a number of reads and writes;
[0297] The multiple mapping relationships are saved in a data fingerprint library of the first target storage device.
[0298] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0299] Record the number of repetitions corresponding to the first repetitive data module;
[0300] Based on the number of repetitions, the count of the number of repetitions corresponding to the third target data module in the data fingerprint library is adjusted.
[0301] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0302] Obtaining data fingerprint libraries of multiple storage devices in the target deduplication storage device group, and comparing data fingerprints in different data fingerprint libraries;
[0303] In response to detecting the presence of the second duplicate data fingerprint, comparing the number of the second duplicate data fingerprints with the preset data module storage amount;
[0304] In response to detecting that the number of the second duplicate data fingerprints is greater than the preset data module storage amount, obtaining a deduplication weight of each data module in the second duplicate data module corresponding to the second duplicate data fingerprint in a corresponding storage device;
[0305] Based on the deduplication weight of each data module in the corresponding storage device, determine and delete a second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint;
[0306] In response to detecting that the number of the second duplicate data fingerprints is less than or equal to the preset data module storage amount, the second duplicate data module corresponding to the second duplicate data fingerprint is retained.
[0307] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0308] Acquire third related information of each data module in the second repeated data module, wherein the third related information at least includes a number of repeated occurrences, a number of updates, and a number of reads and writes;
[0309] Based on the third related information, a deduplication weight of each data module in a corresponding storage device is calculated and determined, wherein the calculation method includes:
[0310]
[0311] in, Indicates The deduplication weight of each storage device, , and Both represent weight coefficients, Indicates the number of repetitions. Indicates the number of read and write times. Indicates the number of updates. Indicates the number of storage devices;
[0312] In response to detecting that the deduplication weight has been calculated, the update times and the read / write times of the corresponding data module are cleared.
[0313] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0314] Based on the deduplication weights, sorting the data modules in the second duplicate data module in descending order to obtain a sorting result;
[0315] Based on the preset data module storage amount, the data modules in the sorting result are divided to determine the second target data module, and the deletion operation is performed, and the data access address corresponding to the reserved data module is adjusted.
[0316] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0317] In response to receiving a read request, detecting whether a data module corresponding to the read request in the target deduplication storage device group has been deduplicated;
[0318] In response to detecting that the data module corresponding to the read request in the target deduplication storage device group is not deduplicated, responding to the read request based on data in the local storage device;
[0319] In response to detecting that a data module corresponding to the read request in the target deduplication storage device group is deduplicated, adjusting the number of reads corresponding to the data module, and acquiring a data access address corresponding to the data fingerprint of the data module according to a data fingerprint library;
[0320] In response to detecting that the data access address is a local storage device, responding to the read request based on data in the local storage device;
[0321] In response to detecting that the data access address is not a local storage device, the read request is sent to the corresponding storage device, and data is obtained in response to the read request.
[0322] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0323] In response to receiving a write request, detecting whether a data module corresponding to the write request is pointed to by a data access address in a data fingerprint library;
[0324] In response to detecting that the data module corresponding to the write request is not pointed to by the data access address in the data fingerprint library, directly overwrite and write the data module to respond to the write request;
[0325] In response to detecting that the data module corresponding to the write request is pointed to by the data access address in the data fingerprint library, a new data module is written to respond to the write request, and at the same time, the count of the number of writes and the number of repeated occurrences of the data module is adjusted.
[0326] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0327] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0328] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Creating a deduplication storage device group, adding a plurality of storage devices to the deduplication storage device group to generate a target deduplication storage device group, wherein a storage device includes at least one data module; According to the first data processing strategy, determining and deleting a first target data module in the first duplicate data module corresponding to the first target storage device; According to the second data processing strategy, determining and deleting a second target data module in the second duplicate data modules corresponding to a plurality of storage devices in the target deduplication storage device group includes: Obtain data fingerprint libraries of multiple storage devices in the target deduplication storage device group, and compare data fingerprints in different data fingerprint libraries; in response to detecting the presence of a second duplicate data fingerprint and the number of the second duplicate data fingerprints being greater than the preset data module storage amount, obtain the deduplication weight of each data module in the second duplicate data module corresponding to the second duplicate data fingerprint in the corresponding storage device, wherein the deduplication weight is calculated and determined based on the number of repeated occurrences, the number of updates, and the number of reads and writes of each data module; based on the deduplication weight of each data module in the corresponding storage device, determine and delete the second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint; The calculation method of the deduplication weight includes: Among them, Weight_i represents the deduplication weight of the i-th storage device, K1, K2 and K3 all represent weight coefficients, N dup (i) represents the number of repetitions, IOPS_r(i) represents the number of reads and writes, IOPS_w(i) represents the number of updates, and n represents the number of storage devices; Based on a preset time period, obtaining the number of deleted data modules of a target storage device in a target deduplication storage device group and a corresponding data module identifier, and defining the number as a first target number; Based on the number of occurrences of the target data module identifier, determine the number of deletions of the data module corresponding to the target data module identifier, and define the number of deletions as a second target number; According to the first target number and the second target number, weight information corresponding to the data module corresponding to the target data module identifier is determined, and the calculation method includes: Among them, Weight_e represents weight information, Q(i) represents the first target number, P(i) represents the second target number, K4 and K5 both represent weight coefficients; In response to detecting that the value of the weight information is greater than a preset threshold, when it is detected that the storage devices to which the storage modules corresponding to the duplicate data among the multiple storage devices belong include the target storage device, and the number of duplicate storage modules is greater than the preset data module storage capacity, the storage modules in the target storage device are directly included in the deduplication plan, and the deduplication operation is performed; The data read and write request is responded to based on the target deduplication storage device group after the first target data module and the second target data module have been deleted.
2. The data processing method according to claim 1, characterized in that: Creating a deduplication storage device group, and adding multiple storage devices to the deduplication storage device group to generate a target deduplication storage device group includes: Create a deduplication storage device group; Adding a second target storage device in the data center to the deduplication storage device group, wherein the physical links between the second target storage device and other storage devices in the deduplication storage device group are interconnected; Based on the physical link connectivity relationship between the storage devices, an association relationship is established between the second target storage device and other storage devices in the deduplication storage device group.
3. The data processing method according to claim 2, characterized in that: After establishing an association relationship between the second target storage device and other storage devices in the deduplication storage device group, the method further includes: Testing the link connectivity between any two storage devices in the deduplication storage device group; In response to detecting that the link connectivity test passes, adding the next storage device in the data center to the deduplication storage device group, and repeatedly performing the association relationship establishment and link connectivity test operations until the target number of storage devices in the data center are added; Define the deduplication storage device group to which the target number of storage devices in the data center have been added as the target deduplication storage device group.
4. The data processing method according to any one of claims 1 to 3, characterized in that: After generating the target deduplication storage device group, the method further includes: Based on the preset block size, the space occupied by the data in each storage device is divided into blocks to obtain multiple data modules.
5. The data processing method according to claim 1, characterized in that: According to the first data processing strategy, determining and deleting the first target data module in the first duplicate data module corresponding to the first target storage device includes: Traversing the data modules in the first target storage device, and calculating and determining the data fingerprint of the data module; In response to detecting the presence of a first duplicate data fingerprint, defining a data module corresponding to the first duplicate data fingerprint as a first duplicate data module; The third target data module selected in the first duplicate data module is retained, and the first target data modules except the third target data module in the first duplicate data module are deleted.
6. The data processing method according to claim 5, characterized in that: The method further comprises: Pointing the data access address of the first repeated data module to the third target data module, and obtaining first relevant information of the third target data module to generate a one-to-one mapping relationship, wherein the first relevant information at least includes a data fingerprint, a number of repeated occurrences, a number of updates, and a number of reads and writes; and / or, obtaining second relevant information of other data modules and generating a one-to-one mapping relationship, wherein the second relevant information at least includes a data fingerprint, a data access address, a number of repetitions, a number of updates, and a number of reads and writes; The multiple mapping relationships are saved in a data fingerprint library of the first target storage device.
7. The data processing method according to claim 5, characterized in that: The method further comprises: Record the number of repetitions corresponding to the first repetitive data module; Based on the number of repetitions, the count of the number of repetitions corresponding to the third target data module in the data fingerprint library is adjusted.
8. The data processing method according to claim 1, characterized in that: According to the second data processing strategy, determining and deleting the second target data module in the second duplicate data modules corresponding to the plurality of storage devices in the target deduplication storage device group comprises: Obtaining data fingerprint libraries of multiple storage devices in the target deduplication storage device group, and comparing data fingerprints in different data fingerprint libraries; In response to detecting the presence of the second duplicate data fingerprint, comparing the number of the second duplicate data fingerprints with the preset data module storage amount; In response to detecting that the number of the second duplicate data fingerprints is less than or equal to the preset data module storage amount, the second duplicate data module corresponding to the second duplicate data fingerprint is retained.
9. The data processing method according to claim 8, characterized in that: The method for determining the deduplication weight of each data module in the corresponding storage device includes: Acquire third related information of each data module in the second repeated data module, wherein the third related information at least includes a number of repeated occurrences, a number of updates, and a number of reads and writes; Based on the third relevant information, calculating and determining the deduplication weight of each data module in the corresponding storage device; In response to detecting that the deduplication weight has been calculated, the update times and the read / write times of the corresponding data module are cleared.
10. The data processing method according to claim 8, characterized in that: Based on the deduplication weight of each data module in the corresponding storage device, determining and deleting the second target data module in the second duplicate data module corresponding to the second duplicate data fingerprint includes: Based on the deduplication weights, sorting the data modules in the second duplicate data module in descending order to obtain a sorting result; Based on the preset data module storage amount, the data modules in the sorting result are divided to determine the second target data module, and the deletion operation is performed, and the data access address corresponding to the reserved data module is adjusted.
11. The data processing method according to claim 1, characterized in that: Based on the target deduplication storage device group after the first target data module and the second target data module have been deleted, responding to the data read and write request includes: In response to receiving a read request, detecting whether a data module corresponding to the read request in the target deduplication storage device group has been deduplicated; In response to detecting that the data module corresponding to the read request in the target deduplication storage device group is not deduplicated, responding to the read request based on data in the local storage device; In response to detecting that a data module corresponding to the read request in the target deduplication storage device group is deduplicated, adjusting the number of reads corresponding to the data module, and acquiring a data access address corresponding to the data fingerprint of the data module according to a data fingerprint library; In response to detecting that the data access address is a local storage device, responding to the read request based on data in the local storage device; In response to detecting that the data access address is not a local storage device, the read request is sent to the corresponding storage device, and data is obtained in response to the read request.
12. The data processing method according to claim 1, characterized in that: Based on the target deduplication storage device group after the first target data module and the second target data module have been deleted, responding to the data read and write request includes: In response to receiving a write request, detecting whether a data module corresponding to the write request is pointed to by a data access address in a data fingerprint library; In response to detecting that the data module corresponding to the write request is not pointed to by the data access address in the data fingerprint library, directly overwrite and write the data module to respond to the write request; In response to detecting that the data module corresponding to the write request is pointed to by the data access address in the data fingerprint library, a new data module is written to respond to the write request, and at the same time, the count of the number of writes and the number of repeated occurrences of the data module is adjusted.
13. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Data processing method and device, computer equipment and storage medium
CN114442961A