Testing Method, Device, Equipment, and Medium for Deduplication Fingerprint Threshold Based on Snapshot
By writing data in the data volume and taking ROW snapshots, recording the number of writes and analyzing the metadata mapping relationship, the problem of being unable to accurately test the redeletion data threshold in the prior art is solved, and efficient redeletion fingerprint threshold testing is achieved.
Patent Information
- Application Number
- CN202211033617.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-26
AI Technical Summary
In the prior art, when testing ROW snapshots and deleting volumes, it is impossible to accurately test the processing situation when the deleting data reaches the threshold. It usually requires a lot of repeated tests to increase the probability, and it is impossible to ensure that the threshold processing process will definitely be triggered.
By writing data in the data volume and taking ROW snapshots, recording the number of writes, judging the consistency of the redeleted data file, analyzing the metadata mapping relationship, ensuring that the processing process is accurately triggered when the redeleted fingerprint threshold is reached.
It realizes switching and processing of accurately testing the redeletion fingerprint threshold in ROW snapshots, avoiding large amounts of data writing, and improving the accuracy and efficiency of the test.
Smart Images

Figure CN115470040B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data deduplication, and particularly to a method, device, equipment, and medium for testing a deduplication fingerprint threshold based on snapshots. Background Art
[0002] Snapshot, that is, snapshot technology, is widely used in backup and disaster recovery. A snapshot is a complete and available copy of a specific data set, which contains the static image of the source data at the copy point; a snapshot can be a copy or replication of data reproduction. According to the definition of SNIA, there are two types of snapshots: full - volume snapshots and incremental snapshots, and different snapshot technologies are used for each. COW is called copy - on - write or copy - before - write. After creating a snapshot, if the data on the source volume changes, the snapshot system will first copy the original data to the corresponding data block on the snapshot volume, and then rewrite the source volume. ROW is called redirect - on - write, which is a concept opposite to COW. After creating a snapshot, the snapshot system redirects the write request to the data volume to the storage space reserved for the snapshot, and directly writes the new data to the snapshot volume. When the upper - layer service reads the source volume, the data before creating the snapshot is read from the source volume, and the data generated after creating the snapshot is read from the snapshot volume. It can avoid the performance loss caused by two write operations.
[0003] For writing data to a deduplicated volume, a logical address (LBA), a physical address (PBA), and a fingerprint lock (HBA) are recorded for each data write. The essence of deduplication technology lies in calculating fingerprints for the written data, and only retaining one copy of the data (PBA) and the corresponding metadata (that is, the metadata stores the mapping relationship L->P from the logical address LBA to the physical address PBA when the data is written) and the fingerprint lock (HBA) data. Duplicate data can then be deleted, and the metadata (L - P) is recorded. Generally, there is an upper limit on how much duplicate data can be deleted, that is, there is a maximum specification for the deduplication rate. For example, the deduplication fingerprint lock threshold of the Inspur MCS system is set to 32. After exceeding 32 duplicate data, new fingerprint data will be recorded. When the deduplication function is enabled in the source volume and then ROW snapshot processing is performed, data deduplication processing, metadata and fingerprint lock data, and metadata reading and writing in the ROW snapshot are involved.
[0004] Currently, for testing ROW snapshots and deduplicated volumes, generally io read - write tools such as vdbench are used to adjust the deduplication rate parameters for reading and writing. To improve the accuracy of testing, generally, a large amount of data is written, the write duration is extended, and the amount of written data is increased to increase the triggering probability during writing. When the deduplication threshold is in the scenario of ROW snapshots, during the test process, the processing situation when the deduplicated data reaches the threshold cannot be accurately tested. Only by increasing the test intensity and performing a large number of repeated tests can the probability be increased, and it cannot be guaranteed that the processing process can be tested for sure. Summary of the Invention
[0005] Currently, for the tests on ROW snapshots and deduplicated volumes, generally, io read and write tools such as vdbench are used to adjust the deduplication rate parameter for reading and writing. To improve the accuracy of the test, generally, a large amount of data is written, the writing duration is extended, and the amount of written data is increased to increase the triggering probability during writing. When the deduplication threshold is in the scenario of ROW snapshots, during the test process, the processing situation when the deduplicated data reaches the threshold cannot be accurately tested. Only by increasing the test intensity and conducting a large number of repeated tests to increase the probability can we not guarantee that the processing process will surely be tested. The present invention provides a method, device, equipment, and medium for testing the deduplication fingerprint threshold based on snapshots.
[0006] In a first aspect, the technical solution of the present invention provides a method for testing the deduplication fingerprint threshold based on snapshots, including the following steps:
[0007] Preset a deduplicated data file for the written data volume;
[0008] Copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the number of writes;
[0009] According to the ROW snapshot, when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, determine whether the recorded number of writes reaches the deduplication fingerprint threshold;
[0010] If not, execute the steps: copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the number of writes;
[0011] If so, copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume. According to the ROW snapshot, when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, parse the metadata in the ROW snapshot;
[0012] According to the parsing result, when there are two physical addresses in the metadata, output that the test passes, that is, the processing process after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal.
[0013] In this application, data (deduplicated data file) is written to the data volume, a ROW snapshot is taken, the metadata mapping is copied once, and writing duplicate data triggers deduplication. When the number of times of writing duplicate data reaches the deduplication fingerprint threshold, continue to write duplicate data to trigger the switching process of the deduplication fingerprint threshold.
[0014] Further, the step of determining whether the recorded number of writes reaches the deduplication fingerprint threshold according to the ROW snapshot when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file includes:
[0015] Read the deduplicated data file written to the data volume according to the metadata address mapping relationship in the ROW snapshot;
[0016] Perform consistency verification on the obtained deduplicated data file and the preset deduplicated data file;
[0017] When the verification is consistent, determine whether the recorded write count reaches the deduplication fingerprint threshold.
[0018] Furthermore, copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume. According to the ROW snapshot, when the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, in the step of parsing the metadata in the ROW snapshot, the steps of parsing the metadata in the ROW snapshot include:
[0019] Obtain the mapping relationship between the logical address and the physical address in the metadata;
[0020] Determine whether the relationship between the logical address and the physical address is many-to-one;
[0021] If so, there is one physical address in the metadata;
[0022] If not, there are two physical addresses in the metadata.
[0023] Furthermore, the steps of performing consistency verification on the obtained deduplicated data file and the preset deduplicated data file further include:
[0024] When the verification is inconsistent, output a write exception prompt message.
[0025] When taking a ROW snapshot of the data volume, cooperate with writing data and the snapshot to test the deduplication fingerprint threshold of the data volume. There is no need to continuously write a large amount of data to increase the triggering probability. This method can accurately test the switching and processing mechanism of the deduplication fingerprint threshold to determine whether the deduplication fingerprint processing mechanism in the ROW snapshot can handle it normally.
[0026] In a second aspect, the technical solution of the present invention further provides a test device for the deduplication fingerprint threshold based on snapshots, including a preset module, a write processing module, a first judgment module, an analysis module, and a test result output module;
[0027] The preset module is used to preset the deduplicated data file written to the data volume;
[0028] The write processing module is used to copy the deduplicated data file and write it to the data volume, take a ROW snapshot of the data volume when writing, and record the write count;
[0029] The first judgment module is used to judge whether the write count reaches the deduplication fingerprint threshold when the deduplicated data file written to the data volume is consistent with the preset deduplicated data file according to the ROW snapshot; trigger the write processing module;
[0030] The parsing module is used to, when the first judgment module judges that the write count reaches the deduplication fingerprint threshold and triggers the write processing module, after the write processing module finishes execution, parse the metadata in the ROW snapshot when the deduplicated data file written to the data volume is consistent with the preset deduplicated data file;
[0031] The test result output module is used to output that the test passes, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal, when it is judged according to the parsing result that there are two physical addresses in the metadata.
[0032] In this application, data (deduplicated data file) is written to the data volume, a ROW snapshot is taken, the metadata mapping is copied once, and writing duplicate data triggers deduplication. When the number of times of writing duplicate data reaches the deduplication fingerprint threshold, continue to write duplicate data to trigger the switching process of the deduplication fingerprint threshold.
[0033] Further, the device further includes a consistency verification module, which specifically includes a reading unit and a verification unit;
[0034] The reading unit reads the deduplicated data file written to the data volume according to the metadata address mapping relationship in the ROW snapshot;
[0035] The verification unit is used to perform consistency verification on the obtained deduplicated data file and the preset deduplicated data file;
[0036] The first judgment module is specifically used to judge whether the write count reaches the deduplication fingerprint threshold when the verification unit outputs consistent verification.
[0037] Further, the parsing module includes an acquisition unit, an address judgment unit, and a parsing result output unit;
[0038] The acquisition unit is used to acquire the mapping relationship between the logical address and the physical address in the metadata;
[0039] The address judgment unit judges whether the logical address and the physical address are in a many-to-one relationship and triggers the parsing result output unit;
[0040] The parsing result output unit is used to output that there is one physical address in the metadata or there are two physical addresses in the metadata according to the judgment result of the address judgment unit.
[0041] Further, the test result output module is also used to output a write exception prompt message when the consistency verification module outputs inconsistent verification.
[0042] When taking a ROW snapshot of a data volume, by writing data and snapshots in cooperation to test the deduplication fingerprint threshold of the data volume, there is no need to continuously write a large amount of data to increase the triggering probability. This method can accurately test the switching and processing mechanism of the deduplication fingerprint threshold to determine whether the deduplication fingerprint processing mechanism can handle it normally in the ROW snapshot.
[0043] In a third aspect, the technical solution of the present invention further provides an electronic device, and the electronic device includes:
[0044] At least one processor; and,
[0045] A memory communicatively connected to the at least one processor; wherein,
[0046] The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the method for testing the deduplication fingerprint threshold based on snapshots as described in the first aspect.
[0047] In a fourth aspect, the technical solution of the present invention provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method for testing the deduplication fingerprint threshold based on snapshots as described in the first aspect.
[0048] From the above technical solutions, it can be seen that the present invention has the following advantages: The test method used in the present invention is for testing the deduplication volume, verifying whether the data can be guaranteed to be consistent after data deduplication, and verifying whether the processing flow after the deduplication fingerprint reaches the threshold is accurate and normal. Different from the conventional test method of continuously writing a large amount of data for probabilistic testing, this method can accurately test the processing flow when the deduplication reaches the threshold, thereby verifying the processing flow of the ROW snapshot and the deduplicated data, and improving the quality and stability of the product.
[0049] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect.
[0050] It can be seen that compared with the prior art, the present invention has prominent substantive features and significant progress, and the beneficial effects of its implementation are also obvious. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a schematic flow chart of the method according to an embodiment of the present invention.
[0053] Figure 2 It is a schematic flow chart of the method according to another embodiment of the present invention.
[0054] Figure 3 It is a schematic block diagram of the device according to an embodiment of the present invention. Detailed implementation manners
[0055] For the data writing of the deduplicated volume, one piece of data writes and records a logical address (LBA), a physical address (PBA), and a fingerprint lock (HBA). The essence of the deduplication technology lies in calculating the fingerprint for the written data, and only retaining one piece of data (PBA) and the corresponding metadata (that is, the metadata stores the mapping relationship L->P from the logical address LBA to the physical address PBA when the data is written) and the fingerprint lock (HBA) data. The duplicate data can be deleted, and the metadata (L-P) is recorded. Generally, there is an upper limit on how much duplicate data is deleted, that is, there is a maximum specification for the deduplication rate. For example, the deduplication fingerprint lock threshold of the Inspur MCS system is set to 32. After exceeding 32 duplicate data, new fingerprint data will be recorded. When the deduplication function is enabled in the source volume and then the ROW snapshot processing is performed, the deduplication processing of the data, the metadata and the fingerprint lock data, and the reading and writing of the metadata in the ROW snapshot are involved.
[0056] Currently, for the tests of ROW snapshots and deduplicated volumes, generally, io read and write tools such as vdbench are used to adjust the deduplication rate parameter for reading and writing. To improve the accuracy of the test, generally, a large amount of data is written, the writing duration is lengthened, and the amount of written data is increased to increase the triggering probability during writing. When the deduplication threshold is in the scenario of ROW snapshots, during the test process, the processing situation when the deduplicated data reaches the threshold cannot be accurately tested. Only by increasing the test intensity and performing a large number of repeated tests can the probability be increased, and it cannot be guaranteed that the processing process can be surely tested. In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0057] As Figure 1 shown, an embodiment of the present invention provides a method for testing the deduplication fingerprint threshold based on snapshots, including the following steps:
[0058] Step 1: Preset the deduplicated data file of the written data volume;
[0059] Step 2: Copy and deduplicate the data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the write count.
[0060] Step 3: According to the ROW snapshot, when the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, determine whether the recorded write count reaches the deduplication fingerprint threshold.
[0061] If not, execute Step 2.
[0062] If so, execute Step 4.
[0063] Step 4: Copy and deduplicate the data file and write it to the data volume. When writing, take a ROW snapshot of the data volume. According to the ROW snapshot, when the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, parse the metadata in the ROW snapshot.
[0064] Step 5: According to the parsing result, when there are two physical addresses in the metadata, output that the test passes, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal.
[0065] In this application, data (deduplicated data file) is written to the data volume, a ROW snapshot is taken, the metadata mapping is copied once, and writing duplicate data triggers deduplication. When the number of times of writing duplicate data reaches the deduplication fingerprint threshold, continue to write duplicate data to trigger the switching process of the deduplication fingerprint threshold.
[0066] As Figure 2 shown, an embodiment of the present invention provides a method for testing the deduplication fingerprint threshold based on a snapshot, including the following steps:
[0067] S1: Preset the deduplicated data file to be written to the data volume.
[0068] S2: Copy and deduplicate the data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the write count.
[0069] S3: Read the deduplicated data file written to the data volume according to the metadata address mapping relationship in the ROW snapshot.
[0070] S4: Perform a consistency check on the obtained deduplicated data file and the preset deduplicated data file.
[0071] When the check is consistent, execute Step S5; when the check is inconsistent, execute Step S8.
[0072] S5: Determine whether the recorded write count reaches the deduplication fingerprint threshold.
[0073] If not, execute Step S2.
[0074] If so, execute step S6;
[0075] S6: Copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume. When it is determined according to the ROW snapshot that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, parse the metadata in the ROW snapshot. In this step, the step of parsing the metadata in the ROW snapshot includes: obtaining the mapping relationship between the logical address and the physical address in the metadata; determining whether the relationship between the logical address and the physical address is one-to-many; if so, there is one physical address in the metadata; if not, there are two physical addresses in the metadata;
[0076] S7: Determine according to the parsing result that when there are two physical addresses in the metadata, output that the test passes, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal.
[0077] S8: Output a write exception prompt message.
[0078] When taking a ROW snapshot of the data volume, the deduplication fingerprint threshold of the data volume is tested by cooperating with the written data and the snapshot. There is no need to continuously write a large amount of data to increase the triggering probability. This method can accurately test the switching of the deduplication fingerprint threshold and the processing mechanism to determine whether the deduplication fingerprint processing mechanism in the ROW snapshot can be processed normally.
[0079] Deduplication technology is generally divided into source-side deduplication and destination-side deduplication. Source-side deduplication first calculates the fingerprint of the data to be transmitted at the client and discovers and eliminates duplicate content by comparing the fingerprint with the server, and only sends non-duplicate data content to the server, so as to achieve the goal of saving network bandwidth and storage resources at the same time. Destination-side deduplication directly transmits the client's data to the server and detects and eliminates duplicate content inside the server. Both deployment methods can provide storage space efficiency. The main difference is that source-side deduplication exchanges the consumption of client computing resources for the improvement of network transmission efficiency.
[0080] This article involves the deduplication technology which belongs to the host-side deduplication. Based on the Inspur MCS storage software system, the online deduplication is a pool-based duplicate data detection and reduction technology, which is implemented based on software and does not require the assistance of a hardware acceleration card. Before the data is written to the hard disk, the number of writes to the SSD disk and the amount of written data are reduced by detecting and deleting duplicate data in real time, thereby saving storage space, reducing the wear of the SSD disk, and extending the service life of the SSD disk. The embodiment of the present invention also provides a test method for the deduplication fingerprint threshold in the ROW snapshot. When taking the ROW snapshot of the deduplication volume, the fingerprint threshold of the deduplication volume is tested by cooperating with the written data and the snapshot, and it is not necessary to continuously write a large amount of data to increase the triggering probability. This method can accurately test the threshold switching and processing mechanism of the deduplication fingerprint to determine whether the deduplication fingerprint processing mechanism can be normally processed in the ROW snapshot. In the Inspur MCS system, only under the all-flash stack can the ROW snapshot and the deduplication function be supported. The all-flash stack writes data based on the log structure method. Thus, it can be realized that the received random data is converted into sequential arrangement when writing to the disk, and then written to the disk after filling up the RAID full stripe size. After the storage pool is created, the storage space is managed by the LSA volume (logstructure). The data is sent down and saved in the LSA volume, and the metadata stores the mapping relationship from the corresponding thin volume address (LBA) of the data to the LSA volume address (PBA). The specific implementation process is as follows:
[0081] The specific implementation process is as follows:
[0082] 1. Prepare the deduplication data of the data volume to be written, such as the file file1, and copy file1 multiple times;
[0083] 2. After writing the file file1 to the data volume for the first time, take a ROW snapshot once, and copy the metadata once at this time;
[0084] 3. Write the copied file of file1 to the data volume, and then take a ROW snapshot again. At this time, the copied metadata is the metadata of the first deduplication, and the deduplication rate is 2:1 at this time;
[0085] 4. Write the copied file of file1 to the data volume for the second time, and then take a ROW snapshot again. At this time, the copied metadata is the metadata of the second deduplication, and the deduplication rate is 3:1 at this time;
[0086] 5. And so on. After writing the copied file of file1 for the 32nd time, the deduplication fingerprint reaches the threshold;
[0087] 6. After reaching the threshold, write the copied file of file1 for the 33rd time, and at this time, the switching of the deduplication fingerprint threshold will be triggered.
[0088] 7. After reaching the deduplication fingerprint threshold, write the deduplication data, and at this time, a new fingerprint will be recorded;
[0089] 8. After each data write, a ROW snapshot is taken of the source volume, and the metadata content is copied once. Through the rollback of the ROW snapshot, data consistency verification is performed to determine whether the deduplicated data is correct. After the 33rd write of duplicate data, a ROW snapshot is taken, and then the consistency verification of the snapshot data is performed to verify whether the switching process of the fingerprint threshold is accurate.
[0090] As Figure 3 shown, an embodiment of the present invention further provides a test device for the deduplication fingerprint threshold based on snapshots, including a preset module, a write processing module, a first judgment module, a parsing module, and a test result output module;
[0091] The preset module is used to preset the deduplicated data file of the data volume to be written.
[0092] The write processing module is used to copy the deduplicated data file and write it into the data volume. When writing, a ROW snapshot is taken of the data volume and the write count is recorded.
[0093] The first judgment module is used to judge whether the recorded write count reaches the deduplication fingerprint threshold when the deduplicated data file written into the data volume is consistent with the preset deduplicated data file according to the ROW snapshot; trigger the write processing module.
[0094] The parsing module is used to, when the first judgment module judges that the write count reaches the deduplication fingerprint threshold and triggers the write processing module, and after the write processing module finishes execution, parse the metadata in the ROW snapshot when the deduplicated data file written into the data volume is consistent with the preset deduplicated data file.
[0095] The test result output module is used to output a test pass according to the parsing result when there are two physical addresses in the metadata, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal.
[0096] In this application, data (deduplicated data file) is written into the data volume, a ROW snapshot is taken, the metadata mapping is copied once, writing duplicate data triggers deduplication, and when the number of times of writing duplicate data reaches the deduplication fingerprint threshold, continue to write duplicate data to trigger the switching process of the deduplication fingerprint threshold.
[0097] The device further includes a consistency verification module, which specifically includes a reading unit and a verification unit;
[0098] The reading unit reads the deduplicated data file written into the data volume according to the metadata address mapping relationship in the ROW snapshot.
[0099] The verification unit is used to perform consistency verification on the obtained deduplicated data file and the preset deduplicated data file.
[0100] The first judgment module is specifically configured to judge whether the recorded write count reaches the deduplication fingerprint threshold when the verification unit outputs consistent verification.
[0101] The parsing module includes an acquisition unit, an address judgment unit, and a parsing result output unit;
[0102] The acquisition unit is used to acquire the mapping relationship between the logical address and the physical address in the metadata;
[0103] The address judgment unit judges whether the logical address and the physical address are in a many-to-one relationship, and triggers the parsing result output unit;
[0104] The parsing result output unit is used to output that there is one physical address in the metadata according to the judgment result of the address judgment unit; or there are two physical addresses in the metadata.
[0105] The test result output module is further configured to output a write exception prompt message when the consistency verification module outputs inconsistent verification.
[0106] When making a ROW snapshot of the data volume, the deduplication fingerprint threshold of the data volume is tested by cooperating with writing data and the snapshot. There is no need to continuously write a large amount of data to increase the triggering probability. This method can accurately test the switching and processing mechanism of the deduplication fingerprint threshold to judge whether the deduplication fingerprint processing mechanism can be normally processed in the ROW snapshot.
[0107] The embodiment of the present invention further provides an electronic device, which includes: a processor, a communication interface, a memory, and a bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the bus. The bus can be used for information transmission between the electronic device and the sensor. The processor can call the logical instructions in the memory to execute the following method: Step 1: Preset the deduplicated data file written to the data volume; Step 2: Copy the deduplicated data file and write it to the data volume. When writing, make a ROW snapshot of the data volume and record the write count; Step 3: According to the ROW snapshot, judge whether the recorded write count reaches the deduplication fingerprint threshold when the deduplicated data file written to the data volume is consistent with the preset deduplicated data file; if not, execute Step 2; if so, execute Step 4; Step 4: Copy the deduplicated data file and write it to the data volume. When writing, make a ROW snapshot of the data volume. According to the ROW snapshot, when the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, parse the metadata in the ROW snapshot; Step 5: According to the parsing result, judge that when there are two physical addresses in the metadata, output that the test passes, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal.
[0108] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0109] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method provided by the above method embodiment, for example, including: S1: Preset a deduplicated data file for writing to a data volume; S2: Copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the write count; S3: Read the deduplicated data file written to the data volume according to the metadata address mapping relationship in the ROW snapshot; S4: Perform a consistency check on the obtained deduplicated data file and the preset deduplicated data file; when the check is consistent, execute step S5; when the check is inconsistent, execute step S8; S5: Determine whether the recorded write count reaches the deduplication fingerprint threshold; if not, execute step S2; if so, execute step S6; S6: Copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume, and when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file according to the ROW snapshot, parse the metadata in the ROW snapshot; in this step, the step of parsing the metadata in the ROW snapshot includes: obtaining the mapping relationship between the logical address and the physical address in the metadata; determining whether the logical address and the physical address are in a many-to-one relationship; if so, there is one physical address in the metadata; if not, there are two physical addresses in the metadata; S7: According to the parsing result, when there are two physical addresses in the metadata, output a test passed, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal. S8: Output a write exception prompt message.
[0110] Although the present invention has been described in detail by reference to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for testing the deduplication fingerprint threshold based on snapshots, characterized in that Including the following steps: Preset the deduplicated data file to be written to the data volume; Copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the write count; According to the ROW snapshot, when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, determine whether the recorded write count reaches the deduplication fingerprint threshold; Specifically include: Read the deduplicated data file written to the data volume according to the metadata address mapping relationship in the ROW snapshot; Perform a consistency check on the obtained deduplicated data file and the preset deduplicated data file; When the check is consistent, determine whether the recorded write count reaches the deduplication fingerprint threshold; If not, execute the steps: Copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the write count; If so, copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume. According to the ROW snapshot, when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, parse the metadata in the ROW snapshot; Among them, the steps of parsing the metadata in the ROW snapshot include: Obtain the mapping relationship between the logical address and the physical address in the metadata; Determine whether the logical address and the physical address are in a many-to-one relationship; If so, it is determined that there is one physical address in the metadata; If not, it is determined that there are two physical addresses in the metadata; According to the parsing result, when there are two physical addresses in the metadata, output that the test passes, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal.
2. The test method for the deduplication fingerprint threshold based on snapshots according to claim 1, wherein The steps of performing a consistency check on the obtained deduplicated data file and the preset deduplicated data file further include: When the check is inconsistent, output a write exception prompt message.
3. A test device for snapshot-based deduplication fingerprint threshold, characterized in that Including a preset module, a write processing module, a first judgment module, a parsing module, and a test result output module; The preset module is used to preset the deduplicated data file to be written to the data volume; The write processing module is used to copy the deduplicated data file and write it to the data volume. When writing, take a ROW snapshot of the data volume and record the write count; The first judgment module is used to, according to the ROW snapshot, when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, determine whether the recorded write count reaches the deduplication fingerprint threshold; Trigger the write processing module; The parsing module is used to, when the first judgment module determines that the write count reaches the deduplication fingerprint threshold and triggers the write processing module, after the write processing module finishes execution, according to the ROW snapshot, when it is determined that the deduplicated data file written to the data volume is consistent with the preset deduplicated data file, parse the metadata in the ROW snapshot; The test result output module is used to, according to the parsing result, when there are two physical addresses in the metadata, output that the test passes, that is, the processing flow after the deduplication fingerprint reaches the deduplication fingerprint threshold is accurate and normal; The device further includes a consistency check module, specifically including a reading unit and a verification unit; The reading unit reads the deduplicated data file written to the data volume according to the metadata address mapping relationship in the ROW snapshot; The verification unit is used to perform a consistency check on the obtained deduplicated data file and the preset deduplicated data file; The first judgment module is specifically configured to judge whether the recorded write count reaches the deduplication fingerprint threshold when the verification unit outputs consistent verification results; The parsing module includes an acquisition unit, an address judgment unit, and a parsing result output unit; The acquisition unit is used to acquire the mapping relationship between the logical address and the physical address in the metadata; The address judgment unit judges whether the logical address and the physical address are in a many-to-one relationship, and triggers the parsing result output unit; The parsing result output unit is used to output that there is one physical address in the metadata or there are two physical addresses in the metadata according to the judgment result of the address judgment unit.
4. The test device for the deduplication fingerprint threshold based on snapshots according to claim 3, characterized in that The test result output module is further configured to output a write exception prompt message when the consistency verification module outputs inconsistent verification results.
5. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the snapshot-based deduplication fingerprint threshold test method according to claim 1 or 2.
6. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the snapshot-based deduplication fingerprint threshold test method according to claim 1 or 2.
Citation Information
Patent Citations
ROW snapshot method, system and device and computer readable storage medium
CN110781133A
Data deduplication method and system, storage medium and equipment
CN113535708A