Data Management Method, Device, Electronic Device, Medium and Product

By using generation identification in the address mapping relationship to distinguish the modified version of the data, only the address mapping relationship corresponding to the modified data is stored, the problem of excessive memory space occupied in data management is solved and storage efficiency is improved.

CN119883141BActive Publication Date: 2025-06-13INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510369502.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-13
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

As the amount of data increases and/or the number of modifications increases, the address mapping relationship occupies too much memory space, resulting in a decrease in data management efficiency.

Method used

By adding generation identifiers to the address mapping relationship, different data modification versions are distinguished, and only the address mapping relationship corresponding to the modified data is stored, rather than the full address mapping relationship.

Benefits of technology

It effectively reduces the memory space usage, increases the number of storage in the effective mapping relationship, and solves the problem of excessive memory space usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883141B_ABST
    Figure CN119883141B_ABST
Patent Text Reader

Abstract

The present application discloses a data management method, apparatus, electronic device, medium and product, relating to the technical field of data processing, including: adding a generation identifier to the address mapping relationship. Since the generation identifier added to the address mapping relationship can distinguish different data modification versions, it is possible to store only the address mapping relationship corresponding to the modified data instead of storing the full amount of address mapping relationships, thereby solving the technical problem that the address mapping relationship occupies too much memory space and achieving the technical effect of increasing the storage quantity of effective mapping relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a data management method, apparatus, electronic device, medium, and product. Background Art

[0002] With the improvement of data processing technology and the expansion of data application scenarios, the current data volume is in a state of rapid growth. In this context, the complexity of data has increased significantly, bringing greater challenges to data management.

[0003] In the related art, after modifying data, the modified data is stored at a new physical address, a mapping relationship between the logical address and the physical address is established, and different historical versions of data can be rolled back through the mapping relationship. However, with the increase in data volume and / or the number of modification times, the number of mapping relationships will increase, which will in turn cause the mapping relationships to occupy too much memory space. Summary of the Invention

[0004] This application provides a data management method, apparatus, electronic device, medium, and product to at least solve the problem of excessive memory space occupation in the related art.

[0005] This application provides a data management method, including: receiving a data modification request, where the data modification request includes original data, an original logical address, an original physical address, and modification information; according to the data modification request, performing a modification process on the original data through the modification information to obtain modified data, determining a target physical address, and storing the modified data at the target physical address; determining a target generation identifier corresponding to the modified data, and creating a target address mapping relationship according to the target generation identifier, the original logical address, and the target physical address, where the target generation identifier represents the generation of the data modification request; creating a target snapshot according to the target generation identifier and the target address mapping relationship, and the target snapshot is used to restore the original data.

[0006] This application also provides a data management apparatus, including: a receiving module, configured to receive a data modification request, where the data modification request includes original data, an original logical address, an original physical address, and modification information; a modification module, configured to perform a modification process on the original data through the modification information according to the data modification request to obtain modified data, determine a target physical address, and store the modified data at the target physical address; a creating module, configured to determine a target generation identifier corresponding to the modified data, and create a target address mapping relationship according to the target generation identifier, the original logical address, and the target physical address, where the target generation identifier represents the generation of the data modification request; a generating module, configured to create a target snapshot according to the target generation identifier and the target address mapping relationship, and the target snapshot is used to restore the original data.

[0007] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above data management methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above data management methods when executed by a processor.

[0009] The present application also provides a computer program product including a computer program, and the computer program implements the steps of any of the above data management methods when executed by a processor.

[0010] Through the present application, since the generation identifier added to the address mapping relationship can distinguish different data modification versions, it is possible to store only the address mapping relationship corresponding to the modified data instead of storing the full amount of address mapping relationships, thereby solving the technical problem of excessive memory space occupied by the address mapping relationship and achieving the technical effect of increasing the storage quantity of effective mapping relationships. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 Schematic diagram of an application scenario of a data management method provided by an embodiment of the present application;

[0013] Figure 2 Schematic diagram of a process of a data management method provided by an embodiment of the present application;

[0014] Figure 3 Schematic diagram of data modification provided by an embodiment of the present application;

[0015] Figure 4 Schematic diagram of a process of a data management method provided by an embodiment of the present application;

[0016] Figure 5 Schematic diagram of a duration interval matrix table provided by an embodiment of the present application;

[0017] Figure 6 Schematic diagram of an address mapping relationship provided by an embodiment of the present application;

[0018] Figure 7 Schematic diagram of a process of a data management method provided by an embodiment of the present application;

[0019] Figure 8 Schematic diagram of data rollback provided by an embodiment of the present application;

[0020] Figure 9 Schematic diagram of data rollback provided by an embodiment of the present application;

[0021] Figure 10 Schematic structural diagram of a data management device provided by an embodiment of the present application;

[0022] Figure 11 Schematic structural diagram of a data management device provided by an embodiment of the present application;

[0023] Figure 12 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0025] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0026] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, processing, transmission, provision, disclosure, application and other processing of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0027] In the related art, for data management in a storage system, after modifying some data in a source volume, a full-scale address mapping relationship corresponding to the data in the source volume is generated and stored for data rollback. As the amount of data and / or the number of modification times increase, the number of stored address mapping relationships increases rapidly, resulting in excessive memory space occupation. To avoid excessive stored mapping relationships affecting normal data reading and writing, it is necessary to clean up the address mapping relationships in a timely manner, resulting in a limitation on the number of stored address mapping relationships at the same time. For example, even if only 1% of the data in the source volume is modified, due to the inability to accurately locate the modified position in the related art, it is still necessary to store the address mapping relationships corresponding to all the data, resulting in excessive memory space occupation.

[0028] This application generates an address mapping relationship according to the physical address involved in the data modification, and only stores the address mapping relationship corresponding to the modified data, and distinguishes the address mapping relationships through generation identifiers. During data rollback, different data modification records can be accurately distinguished through the generation identifiers, so as to accurately locate the data that needs to be rolled back, without storing the full-scale address mapping relationships in the memory, thereby solving the problem of excessive memory space occupation in the related art.

[0029] To enable those skilled in the art of this technical field to better understand the solution of this application, the following further elaborates on this application in combination with the accompanying drawings and specific implementation manners.

[0030] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the data management method depends, the specific application environment architecture or specific hardware architecture is described herein. Refer to Figure 1 , Figure 1 FIG. is a schematic diagram of the application scenario of the data management method. The original data is modified to obtain modified data, and the address mapping relationship corresponding to the original data is stored. The modified data is used for current use, and the address mapping relationship is used to roll back the original data.

[0031] Combined with the scenario example, in scenarios such as the system developing new functions, fixing system vulnerabilities, or correcting system errors, historical data needs to be modified to obtain modified data. During the process of using the modified data, if problems such as data corruption, system failures, or compatibility occur, the data can be rolled back to the historical data to solve the problems.

[0032] Figure 2 FIG. is a schematic flowchart of the data management method provided by the embodiment of this application. As Figure 2 shown, the embodiment of this application provides a data management method, and the method is described in detail as follows:

[0033] S201. Receive a data modification request, where the data modification request includes the original data, the original logical address, the original physical address, and the modification information.

[0034] Among them, the original data is the data before modification. The modification information is the specific modification requirement, such as deleting the specified data or adding the specified data, etc. The original logical address is the address for accessing the original data, and the original physical address is the address for storing the original data.

[0035] Exemplarily, the logical address is the virtual address for accessing data, providing a unified access interface, shielding the underlying physical storage details. The application can read and write data through the logical address without being aware of the actual storage location of the data. The logical address simplifies the file system design. The physical address is the actual location of the data on the physical storage medium. The logical address and the physical address are associated through an address mapping relationship, and through the address mapping relationship, the physical address corresponding to each logical address can be accurately determined, so as to access the data stored under the physical address through the logical address.

[0036] S202. According to the data modification request, modify the original data through the modification information to obtain the modified data, determine the target physical address, and store the modified data at the target physical address.

[0037] Exemplarily, modify the original data according to the specific modification requirement in the modification information to obtain the modified data. If the original data is overwritten by the modified data, that is, the original data is deleted, the original data cannot be rolled back when the original data needs to be used subsequently. Saving both the original data and the modified data can achieve subsequent data rollback.

[0038] Next, Figure 3 the data modification will be described.

[0039] Figure 3 is a schematic diagram of the data modification provided by the embodiment of the present application. As Figure 3 shown, the original data is stored at the original physical address, and the original data under the original physical address can be accessed through the original logical address. The original data is modified to obtain the modified data, and the modified data is stored at the target physical address without changing the original logical address. After the modification process is completed, the modified data under the target physical address is accessed through the original logical address, while the original data is stored under the original physical address, and the original data under the original physical address can only be accessed when performing data rollback.

[0040] S203. Determine the target generation identifier corresponding to the modified data, and create a target address mapping relationship according to the target generation identifier, the original logical address, and the target physical address. The target generation identifier represents the generation of the data modification request.

[0041] Among them, the generation identifier represents the generation corresponding to each data modification. The generation identifier is used to distinguish different data modification versions, and each generation identifier corresponds to a data version.

[0042] Optionally, by maintaining a generation identifier table that includes the generation identifier corresponding to each data modification, the corresponding modification record can be quickly queried. When rolling back data each time, the corresponding modification record is queried from the generation identifier table according to the specified generation identifier.

[0043] Exemplarily, the target address mapping relationship records the updated address mapping relationship. Refer to Figure 3 , the address mapping relationship before modifying the data is the mapping relationship between the original logical address and the original physical address, and the address mapping relationship after modifying the data is the mapping relationship between the original logical address and the target physical address. Adding the target generation identifier to the mapping relationship between the original logical address and the target physical address can accurately distinguish the target address mapping relationship from other address mapping relationships.

[0044] Combined with a scenario example, when the data is initially established, the generation identifier is set to 000. When the data is modified for the first time, the corresponding generation identifier is set to 001. When the data is modified for the second time, the corresponding generation identifier is set to 002, and so on. The generation of the corresponding data modification can be accurately determined through the generation identifier.

[0045] Exemplarily, the target address mapping relationship is generated only based on the changed physical address and is distinguished by the generation identifier, avoiding excessive occupation of memory space caused by full-scale updating of the address mapping relationship.

[0046] S204. Create a target snapshot according to the target generation identifier and the target address mapping relationship.

[0047] Among them, the target snapshot is used to restore the original data.

[0048] Among them, a snapshot is a mirror of the data state at a certain point in time in the storage system. Its core principle is based on version control of metadata rather than full-scale data copying. After the snapshot is created, the data state at a certain point in time is "frozen", and subsequent modifications do not affect the content of the snapshot.

[0049] Exemplarily, the data state in this application is the address mapping relationship. The target snapshot is used to record the target generation identifier and the target address mapping relationship. The address mapping relationship can be quickly queried through the snapshot, so as to quickly roll back the data. By adding the target generation identifier information to the snapshot, multiple snapshots can be accurately distinguished.

[0050] Combined with the scenario example, at time T3, a target snapshot is generated to record the target address mapping relationship at time T3. Any subsequent data modifications after time T3 will not change the target snapshot. New snapshots are generated based on subsequent new data modifications after time T3, and the new snapshots coexist with the target snapshot. Multiple snapshots can be rolled back to their corresponding address mapping relationships respectively.

[0051] Based on the above embodiments, by adding a target generation identifier to the snapshot, multiple snapshots can be accurately distinguished, so as to accurately determine the specified snapshot when rolling back data.

[0052] The data management method provided by the embodiments of the present application receives a data modification request, where the data modification request includes original data, an original logical address, an original physical address, and modification information; according to the data modification request, the original data is modified through the modification information to obtain modified data, a target physical address is determined, and the modified data is stored at the target physical address; a target generation identifier corresponding to the modified data is determined, and according to the target generation identifier, the original logical address, and the target physical address, a target address mapping relationship is created, where the target generation identifier represents the generation of the data modification request; according to the target generation identifier and the target address mapping relationship, a target snapshot is created, and the target snapshot is used to restore the original data. In the above solution, since the generation identifiers added in the address mapping can distinguish different data modification versions, only the address mapping relationships corresponding to the modified data need to be stored instead of storing all the address mapping relationships, thus solving the technical problem that the address mapping relationships occupy too much memory space and achieving the technical effect of increasing the storage quantity of effective mapping relationships.

[0053] Based on any of the above embodiments, below, in combination with Figure 4 , the detailed process of data management will be described.

[0054] Figure 4 It is a schematic flowchart of a data management method provided by the embodiments of the present application. As Figure 4 shown, the method includes:

[0055] S401. Receive a data modification request, where the data modification request includes original data, an original logical address, an original physical address, and modification information.

[0056] It should be noted that the execution process of S401 refers to S201, which will not be elaborated here.

[0057] S402. Determine the data size of the modified data.

[0058] Exemplarily, clarifying the data size is used to accurately determine the number of physical blocks required to store the modified data, so as to accurately determine the target physical address.

[0059] Illustrated with a scenario example, for instance, if the data size is 12 KB and the physical block size is 4 KB, then 3 physical blocks are required.

[0060] S403. Determine the free block bitmap. The free block bitmap includes multiple free states corresponding to multiple blocks, and the free states are free or not free.

[0061] Exemplarily, the free block bitmap is a binary metadata structure that records the free state of each physical block in the storage pool. The free state being not free indicates that the corresponding physical block has been occupied, and the free state being free indicates that the corresponding physical block can be allocated.

[0062] Optionally, after obtaining the free block bitmap, perform a version check on the free block bitmap to check whether the free block bitmap is the latest version. If it is not the latest version, update the free block bitmap to improve the accuracy of allocating physical blocks.

[0063] A feasible implementation method can determine the target physical address through the following method, including: scanning the free block bitmap according to the data size to determine multiple free blocks, and the free states of the multiple free blocks are all free; determining a block linked list, where the block linked list includes multiple physical addresses corresponding to multiple blocks; and determining multiple target physical addresses corresponding to the multiple free blocks according to the block linked list.

[0064] Exemplarily, the total space corresponding to the multiple free blocks should be sufficient to store the modified data corresponding to the data size. It can be understood that allocating free blocks as needed according to the data size can avoid wasting free blocks and ensure that the free blocks are sufficient to store the modified data.

[0065] Optionally, determine the allocation method of the multiple free blocks according to the application scenario of the modified data. For example, for modified data mainly for sequential read and write, determine multiple consecutive free blocks to reduce the seek time for accessing the modified data, thereby improving the access speed. For modified data with frequent random read and write, determine multiple discrete free blocks, which can improve the storage utilization rate.

[0066] Illustrated with a scenario example, the free block is not occupied. By scanning to obtain the free block, it can avoid the problem of storage failure caused by storing the modified data in an occupied block.

[0067] Exemplarily, the block linked list is a data structure that records the physical addresses of multiple physical blocks and is used for quickly allocating and recycling physical blocks. Query through the block linked list to obtain multiple target physical addresses corresponding to multiple free blocks.

[0068] In this feasible implementation method, the block linked list pre-records the physical addresses, and the target physical address can be quickly obtained through the block linked list, thereby improving the efficiency of data management.

[0069] A feasible implementation method can store the modified data into the target physical address through the following steps: generating an encryption key corresponding to the modified data; encrypting the modified data with the encryption key to obtain encrypted data; storing the encrypted data into the target physical address.

[0070] Optionally, the encryption key is generated by a random generator or derived from a user password.

[0071] Optionally, the encryption key and the encrypted data are stored separately to avoid the risk of single-point leakage and improve data security.

[0072] Optionally, the encryption key is split into multiple shards and stored in multiple distributed nodes through a distributed storage method, which can avoid the risk of single-point leakage and improve data security.

[0073] Optionally, the encryption process is performed using a symmetric encryption algorithm or an asymmetric encryption algorithm to obtain encrypted data.

[0074] Optionally, each free block is encrypted independently to avoid the problem of excessive performance overhead caused by full-disk encryption.

[0075] Optionally, if an interruption occurs during the encryption process, such as a terminal caused by a hardware failure or other problems, roll back to the state before encryption and re-encrypt.

[0076] Combined with a scenario example, the encrypted data obtained by encrypting the modified data is stored, and the user scope that can access the data can be determined.

[0077] In this feasible implementation method, even if the encrypted data is stored in the target physical address and is obtained by unauthorized personnel, since the data is already encrypted, direct access to the data can be avoided, thereby improving data security.

[0078] S404. Determine the global generation identifier table, which includes multiple allocated historical generation identifiers.

[0079] Exemplarily, the allocated historical generation identifier is a used generation identifier, and each historical generation identifier corresponds to a data modification operation. The global generation identifier table is a central metadata table in the storage system that records all allocated generation identifiers.

[0080] Optionally, the global generation identifier table can be stored in a single-point centralized manner or in a cross-node distributed manner. The global generation identifier table maintains data consistency across nodes and cycles.

[0081] Exemplarily, the historical generation identifier is unique, that is, any two historical generation identifiers in the global generation identifier table are different.

[0082] As illustrated by a scenario example, the storage system includes multiple logical addresses. When data under each logical address is modified, a corresponding generation identifier is allocated. The global generation identifier table stores the historical generation identifiers allocated to each logical address to avoid conflicts with any historical generation identifier when allocating generation identifiers.

[0083] S405. Determine a target generation identifier based on the historical generation identifiers, where the target generation identifier is different from any historical generation identifier.

[0084] Optionally, the target generation identifier is allocated in a monotonically increasing and globally unique manner. For example, if the largest generation identifier value among multiple historical generation identifiers is 019, then the target generation identifier is 020.

[0085] Exemplarily, each generation identifier is ensured to be a unique generation identifier through a globally unique manner.

[0086] Optionally, the global generation counter is incremented through an atomic operation to obtain a unique and incrementing target generation identifier, avoiding the generation of multiple identical generation identifiers at the same time, thereby improving the accuracy of data management.

[0087] As illustrated by a scenario example, the target generation identifier determined based on the historical generation identifiers can avoid conflicts between the target generation identifier and any historical generation identifier, thereby ensuring the uniqueness of each generation identifier to improve the reliability of data management.

[0088] Based on the above embodiments, the target generation identifier is determined through the global generation identifier table to ensure the uniqueness of the target generation identifier, thereby accurately distinguishing the target address mapping relationship from other address mapping relationships through the target generation identifier, avoiding conflicts, and further improving the reliability of data management.

[0089] A feasible implementation manner is that the data management method further includes: performing an update process on the global generation identifier table through the target generation identifier to obtain an updated generation identifier table; determining multiple distributed storage nodes; and sending the updated generation identifier table to the multiple distributed storage nodes.

[0090] Exemplarily, each distributed storage node holds the same generation identifier table to avoid conflicts of newly generated generation identifiers caused by version divergence. Moreover, each distributed storage node holding the same generation identifier table can be repaired in case of a single point of failure.

[0091] As illustrated by a scenario example, by sending the updated generation identifier table to multiple distributed storage nodes, any one of the distributed storage nodes can allocate generation identifiers according to the updated generation identifier table, thereby avoiding conflicts of generation identifiers. Correspondingly, after any one of the distributed storage nodes updates the generation identifier table, it is synchronized to other distributed storage nodes.

[0092] In this feasible implementation, by sending an updated generation identifier table to the distributed storage nodes, the uniqueness of the generation identifiers across nodes can be ensured, thereby improving the reliability of data management.

[0093] In a feasible implementation, the data management method further includes: determining multiple sub-generation identifiers corresponding to the original logical address, where the multiple sub-generation identifiers include a target generation identifier; arranging the multiple sub-generation identifiers in descending order according to a linked structure to obtain a sub-generation identifier list; and generating a sub-bitmap index based on the sub-generation identifier list, where the sub-bitmap index includes multiple bits, and each bit corresponds to a sub-generation identifier.

[0094] Exemplarily, the multiple sub-generation identifiers are the generation identifiers corresponding to multiple data modification records corresponding to the original logical address.

[0095] Optionally, the multiple sub-generation identifiers are determined from the global generation identifier table.

[0096] Exemplarily, arranging the multiple sub-generation identifiers in descending order according to a linked structure can reduce the complexity of traversal, thereby quickly locating a specified sub-generation identifier.

[0097] Optionally, the sub-generation identifier list is a singly or doubly linked list, and the head node points to the latest version. The head node pointing to the latest version can improve the traversal efficiency.

[0098] Optionally, the sub-generation identifier list is mapped to a binary bitmap, and each bit indicates whether a sub-generation identifier exists. For example, assuming that the global generation range is 001 - 005, and the sub-generation identifiers are [001, 003, 005], that is, the data modification records corresponding to 001, 003, and 005 are the data modification records of the data under the original logical address, while the data modification records corresponding to 002 and 004 are the data modification records of the data under other logical addresses, and the bitmap is 10101 (5 bits) represented in binary. It can be understood that by compressing to a binary representation, the memory occupancy can be reduced.

[0099] Combined with a scenario example, by separately managing the generation identifier list related to each logical address, during data rollback, according to the logical address specified for data rollback, the specified historical data can be quickly located from the generation identifier list corresponding to the specified logical address, thereby improving the efficiency of data rollback.

[0100] In this feasible implementation, through the collaborative design of linked descending order arrangement and sub-bitmap index, an efficient, lightweight, and scalable solution is provided for the multi-version management of logical addresses, which can improve the efficiency of data rollback.

[0101] S406. Obtain the current metadata writing speed, the current processor utilization rate, the current memory utilization rate, and the duration interval matrix table.

[0102] Exemplarily, the current metadata writing speed, the current processor utilization rate, and the current memory utilization rate are all current performance parameters. The metadata writing speed represents the update frequency of metadata per unit duration. The processor utilization rate is used to characterize the load of computing resources. The memory utilization rate is used to characterize the load of memory resources. The duration interval matrix table is used to map a matching duration interval according to the current performance. The current performance is sufficient to support generating a snapshot according to the matching duration interval.

[0103] Optionally, the duration interval table is pre-calculated and generated.

[0104] Optionally, recalculate regularly to update the duration interval table.

[0105] In the related art, the snapshot generation frequency is a fixed value. The snapshot generation frequency is small, and the mapping relationships covered by the snapshots are few. When it is necessary to roll back to any historical data, there may be no corresponding snapshot, resulting in the problem of rollback failure.

[0106] Combined with a scenario example, obtain multiple current performance parameters. According to the multiple current performance parameters, the maximum snapshot generation frequency supported by the current performance can be determined. According to the supported maximum snapshot generation frequency, the current actual snapshot generation frequency is determined. It can be understood that the current actual snapshot generation frequency is less than or equal to the supported maximum snapshot generation frequency, which can ensure that generating snapshots does not occupy too many system resources, so that the system can run normally while generating snapshots, thereby improving the reliability of the system.

[0107] S407. Determine the target duration interval according to the current metadata writing speed, the current processor utilization rate, the current memory utilization rate, and the duration interval matrix table.

[0108] Combined with a scenario example, the shorter the duration interval, the greater the snapshot generation frequency, and the more snapshots need to be generated within a period, which generates greater pressure on system resources. The target duration interval determined according to the real-time performance is the shortest duration interval supported by the system resources. When the actual duration interval for generating snapshots is greater than the target duration interval, the generation of snapshots will not affect the normal operation of the system.

[0109] Optionally, normalize the current metadata writing speed, the current processor utilization rate, and the current memory utilization rate to obtain the corresponding normalized metrics respectively, and match and process the target duration interval from the duration interval matrix table through the normalized metrics. It can be understood that normalization can eliminate the dimension difference, thereby improving the accuracy of matching.

[0110] Optionally, corresponding weights are determined for each performance parameter, and each performance parameter is corrected according to the weights to obtain a correction result. The target time interval is matched from the time interval matrix table according to the correction result. It can be understood that through the correction process, the contribution degree of each performance parameter can be accurately reflected, thereby improving the accuracy of matching.

[0111] Optionally, multiple parameter ranges are set for each performance parameter in the time interval matrix table, and the corresponding target time interval is determined according to which parameter range the performance parameter falls into.

[0112] Next, Figure 5 the time interval matrix table will be described.

[0113] Figure 5 is a schematic diagram of the time interval matrix table provided by the embodiment of the present application. As Figure 5 shown, for example, if the current metadata writing speed is in range B, the current processor utilization rate is in range E, and the current memory utilization rate is in range H, then the target time interval is determined to be interval B.

[0114] S408. At every target time interval, according to the target generation identifier and the target address mapping relationship, a target snapshot is created.

[0115] Exemplarily, the target address mapping relationship is the mapping relationship corresponding to the modified data. Recording the target address mapping relationship through the target snapshot can reduce the memory occupancy of the target snapshot.

[0116] Optionally, a time stamp for generating the target snapshot is added to the target snapshot. The time stamp is the moment when the target snapshot is generated.

[0117] Combined with the scenario example, by adding time stamp information, the snapshot at a specified moment can be quickly located according to the time stamp during data rollback, improving the positioning speed and thus the efficiency of data rollback.

[0118] Based on the above embodiments, creating a target snapshot according to the target time interval can increase the snapshot generation frequency without affecting the normal operation of the system, thereby covering more historical data and enabling data rollback for more historical data.

[0119] A feasible implementation method can determine the target duration interval through the following steps: matching the matching duration interval from the duration interval matrix table according to the current metadata writing speed, the current processor utilization rate, and the current memory utilization rate; performing calculation processing based on the current metadata writing speed, the current processor utilization rate, and the current memory utilization rate to obtain the theoretical metadata writing speed; if the historical metadata writing speeds of a continuous preset number of times are all less than the theoretical metadata writing speed, then determine the sum of the matching duration interval and the preset value as the target duration interval; if the historical metadata writing speeds of a continuous preset number of times are all greater than the theoretical metadata writing speed, then determine the difference between the matching duration interval and the preset value as the target duration interval; if there are no continuous preset number of historical metadata writing speeds that are all less than the theoretical metadata writing speed and there are no continuous preset number of historical metadata writing speeds that are all greater than the theoretical metadata writing speed, then determine the matching duration interval as the target duration interval.

[0120] Exemplarily, the current metadata writing speed is the data obtained, and the theoretical metadata writing speed is the calculated subsequent calculation data.

[0121] Exemplarily, dynamically correct the target duration interval. For example, if the calculated theoretical metadata writing speed is M, and then if the actual metadata writing speeds of a continuous preset number of times are all less than the theoretical metadata writing speed M, then increase the snapshot duration interval by a preset duration; if the metadata writing speeds of a continuous preset number of times are all greater than the theoretical metadata writing speed M, then reduce the snapshot duration interval by a preset duration, and so on. The same applies to the processor utilization rate and the memory utilization rate.

[0122] In this feasible implementation method, by dynamically adjusting the snapshot frequency, the snapshot frequency can be made to conform to the real-time resources of the system, thereby avoiding increasing the pressure on the system.

[0123] A feasible implementation method can create a target snapshot through the following steps: determining the coding length according to multiple historical generation identifiers; determining the blank fields according to the coding length and adding the blank fields to the target address mapping relationship; adding the target generation identifier to the blank fields to obtain the target snapshot.

[0124] In the related technology, during the process of generating a snapshot, the data reading and writing are set to a silent state, that is, the data reading and writing are locked until the snapshot generation is completed, so as to avoid conflicts between the data reading and writing during the snapshot generation process and the snapshot.

[0125] Exemplarily, generate the target snapshot through the generation identifier. Different snapshots can be distinguished through the generation identifier. Direct the newly generated data write operations during the snapshot generation process to the new physical address.

[0126] Illustrated with a scenario example, for instance, the target generation identifier corresponding to the current target snapshot is 020, and the corresponding physical address is the target physical address. During the process of generating the target snapshot, if a new data write operation occurs, the newly generated data write operation is directed to a physical address different from the target physical address and is identified by the generation identifier 021, thereby accurately distinguishing the data modified by the new data write operation from the current data.

[0127] Exemplarily, the encoding length is the number of encoding bits, and the encoding length determined by the historical generation identifier is sufficient to hold the target generation identifier to ensure that the target generation identifier can be correctly displayed. For example, the largest generation identifier among multiple historical generation identifiers can be 019 or 019653. The number of bits occupied by 019 and 019653 is different, and the encoding length determined according to multiple historical generation identifiers should be sufficient to display the target generation identifier.

[0128] Next, Figure 6 the address mapping relationship will be described.

[0129] Figure 6 is a schematic diagram of the address mapping relationship provided by the embodiments of the present application. As Figure 6 shown, the target address mapping relationship includes a logical address field and a physical address field. The logical address field stores the original logical address, and the physical address field stores the target physical address. A blank field is added to the target address mapping relationship, and the length of the blank field is the encoding length. The blank field is used to store the target generation identifier. After adding the target generation identifier to the blank field, a target snapshot is created according to the target address mapping relationship.

[0130] In this feasible implementation, the target snapshot created through the target generation identifier and the target address mapping relationship can be distinguished from other snapshots by the target generation identifier, and the data storage address can also be obtained through the target address mapping relationship, thereby achieving accurate data rollback.

[0131] A feasible implementation, the data management method further includes: determining the previous generation identifier of the target generation identifier from multiple historical generation identifiers; generating a parent version pointer according to the previous generation identifier; and adding the parent version pointer to the blank field.

[0132] Exemplarily, the parent version pointer is a shortcut link pointing to the previous generation identifier. A parent-child dependency relationship is constructed through the parent version pointer to form a global version chain.

[0133] Exemplarily, if the generation identifier is an incrementing serial number, the previous generation identifier is the generation identifier with the largest serial number among multiple historical generation identifiers. If the generation identifier is a timestamp, the previous generation identifier is the generation identifier closest to the current moment among multiple historical generation identifiers.

[0134] Exemplarily, chain jumps can be achieved through the parent version pointer. That is, each snapshot is a node on the chain, and the parent pointer is the chain connecting the nodes. Through the parent version pointer, the previous generation identifier can be quickly located, so as to quickly perform data rollback.

[0135] Exemplarily, the parent version pointer only stores the inheritance relationship between data versions, rather than specific data, which can save memory space.

[0136] Combined with the scenario example, through the parent version pointer, the data inheritance relationship can be clarified. Specifically, it can be quickly determined on the basis of which data version any data version is modified, so as to quickly locate the historical data.

[0137] In this feasible implementation method, a global version chain is established, and any historical version can be restored through pointer jumps, so as to reduce the time-consuming of data rollback and improve the efficiency of data rollback.

[0138] Figure 7 It is a schematic flowchart of the data management method provided by the embodiment of the present application. As Figure 7 shown, the embodiment of the present application provides a data management method, and the method is described in detail as follows:

[0139] S701. Receive a data rollback request, where the data rollback request includes a specified generation identifier and a source volume identifier.

[0140] Exemplarily, the specified generation identifier is the generation identifier corresponding to the data that needs to be rolled back indicated by the data rollback request. The source volume identifier is the identifier of the source volume corresponding to the data that needs to be rolled back.

[0141] Optionally, the relationship between the volume and the volume identifier is one-to-one, and the unique corresponding source volume can be determined through the source volume identifier.

[0142] Combined with the scenario example, the source volume corresponds to a logical address and multiple physical addresses, and different versions of data are stored under the multiple physical addresses. In the current address mapping relationship, the physical address pointed to by the logical address is the data of the latest version. The data rollback request is used to request to obtain the historical data stored under other physical addresses.

[0143] S702. According to the data rollback request, determine the source volume corresponding to the source volume identifier, and obtain multiple historical snapshots of the source volume.

[0144] Exemplarily, multiple historical snapshots are all snapshots corresponding to the physical addresses of the source volume.

[0145] Next, the data rollback will be described in combination with Figure 8 for data rollback.

[0146] Figure 8Schematic diagram of data rollback provided by an embodiment of the present application. As Figure 8 shown, the data currently displayed in the source volume is the data stored at physical address D, and the data stored at physical addresses A, B, and C are all historical data. The address mapping relationships between the logical address and physical addresses A, B, and C are all recorded through corresponding historical snapshots, and each historical snapshot corresponds to an address mapping relationship. For example, historical snapshot B records the address mapping relationship between the logical address and physical address B. Among them, except for the physical address where the data created for the first time is stored, the data stored in other physical addresses is only the data that has actually been modified.

[0147] S703. Perform data rollback processing according to the specified generation identifier and multiple historical snapshots to obtain a target volume, where the target volume includes rollback data.

[0148] Exemplarily, each historical snapshot corresponds to a generation identifier, and the rollback data is the data stored at the physical address corresponding to the historical snapshot corresponding to the specified generation identifier.

[0149] A feasible implementation manner may perform data rollback processing through the following method: create a first clone volume and a second clone volume of the source volume, where the second clone volume is used to store the newly generated input / output data after receiving a data rollback request; obtain rollback data according to the specified generation identifier and multiple historical snapshots; store the rollback data in the first clone volume to obtain a third clone volume; and merge the second clone volume and the third clone volume to obtain a target volume.

[0150] Exemplarily, the first clone volume is created to be the same as the source volume and is used to load the rollback data to generate the intermediate state after rollback (i.e., the third clone volume). The second clone volume is created as an empty volume, and the second clone volume receives the new data generated during the execution of the data rollback operation.

[0151] Exemplarily, after merging the second clone volume and the third clone volume, the obtained target volume includes rollback data and newly generated data.

[0152] Combined with the scenario example, by separating the first clone volume and the second clone volume, it is possible to avoid the conflict between data read / write operations and data rollback during the execution of the data rollback operation.

[0153] Next, Figure 9 data rollback will be described.

[0154] Figure 9 Schematic diagram of data rollback provided by an embodiment of the present application. As Figure 9As shown, generation identifier 000 indicates the initial creation of data, corresponding to the initial data A under each logical address. Each time data is modified, a corresponding generation identifier is assigned, such as 001, 002, 003, and 004, and a corresponding snapshot is generated for the incremental data of each modification. The generation identifier of the current version of the data is 004. If the data rollback request indicates rolling back to the data with generation identifier 002, the data under the corresponding physical address is obtained according to snapshot 2 corresponding to generation identifier 002 and stored in the first cloned volume to obtain the third cloned volume. The third cloned volume includes the modified data (i.e., data B) corresponding to generation identifier 001 incremented on the basis of the initial data and the data (i.e., data C) modified by generation identifier 002.

[0155] In this feasible implementation, by separating the first cloned volume and the second cloned volume, there is no need to lock data read and write operations during the data rollback operation, and there is no need to interrupt data read and write operations, ensuring the continuity of the service.

[0156] A feasible implementation can obtain rollback data through the following method, including: creating a logical view according to multiple historical snapshots, where the logical view includes multiple address mapping relationships corresponding to multiple historical snapshots; performing matching processing according to the specified generation identifier and the logical view to obtain the specified physical address; and obtaining the rollback data from the physical address.

[0157] Exemplarily, the logical view is a global metadata index layer that integrates the address mapping relationships of multiple historical snapshots to obtain a versioned address space topology.

[0158] Optionally, the process of creating a logical view may include scanning multiple historical snapshots, extracting the address mapping relationships recorded in each historical snapshot, sorting the multiple address mapping relationships according to the multiple generation identifiers corresponding to the multiple historical snapshots to obtain a multi-version address chain, and the multi-version address chain forms a logical view.

[0159] Exemplarily, the physical address corresponding to the specified generation identifier in the logical view is determined as the specified physical address.

[0160] Optionally, the logical view is stored according to the data type in the logical view. For example, the address mapping relationships with high-frequency access in the logical view are stored in memory, and the address mapping relationships with low-frequency access in the logical view are stored on disk to improve the access speed of the address mapping relationships with high-frequency access.

[0161] In this feasible implementation, the logical view is established only according to the historical snapshots corresponding to the source volume. Compared with establishing a logical view for all historical snapshots in the system, the execution complexity can be reduced, thereby improving the efficiency of data rollback.

[0162] A feasible implementation method can perform matching processing through the following method to obtain a specified physical address, including: performing matching processing through multiple threads according to a specified generation identifier and a logical view to obtain a matching result, where the matching result includes a specified physical address or a miss; if the matching result is a miss, then determine the physical address corresponding to the generation identifier adjacent to the specified generation identifier in the logical view as the specified physical address.

[0163] Exemplarily, performing matching processing through multiple threads simultaneously can improve the matching efficiency compared to performing matching processing with a single thread.

[0164] Optionally, tasks are allocated to multiple threads according to the loads corresponding to the multiple threads to achieve reasonable allocation. Specifically, a range of generation identifiers matching the load is allocated to each thread. For example, a matching task for the range of generation identifiers 0000 - 0113 is allocated to thread A, and since the load of thread B is lower, a matching task for the range of generation identifiers 0114 - 0506 is allocated to thread B.

[0165] Optionally, each thread performs matching through the binary search method to locate the specified generation identifier.

[0166] Combined with a scenario example, if the matching result is a miss, it indicates that the specified generation identifier is not included in the logical view, which may be caused by problems such as data corruption. At this time, the data in the physical address corresponding to the generation identifier closest to the specified generation identifier is closest to the data that the data rollback request hopes to roll back. For example, if the specified generation identifier is 075, the logical view does not include the generation identifier 075 but includes the generation identifier 060 and the generation identifier 080, then the physical address corresponding to the generation identifier 080 is determined as the specified physical address.

[0167] Optionally, if a specified physical address is matched through an adjacent generation identifier in the case where the matching result is a miss, a corresponding prompt message is generated so that the user can clarify the version of the rollback data according to the prompt message.

[0168] In this feasible implementation method, matching through adjacent generation identifiers can handle extreme scenarios and improve the fault tolerance of data rollback.

[0169] A feasible implementation method, the data management method further includes: obtaining the transaction log of the source volume, where the transaction log includes the metadata state and physical block reference count of the source volume before performing data rollback; performing integrity verification processing on the second cloned volume and the third cloned volume to obtain a verification result, where the verification result is verification passed or verification failed; if the verification result is verification failed, then perform recovery processing on the source volume through the transaction log to obtain a recovery volume; if the verification result is verification passed, then perform merging processing on the second cloned volume and the third cloned volume to obtain a target volume.

[0170] Exemplarily, the metadata status includes generation identification, snapshot, and address mapping relationship. The physical block reference count is the number of times each physical block is referenced by the snapshot (for example, any physical block is referenced by 3 snapshots).

[0171] Optionally, before performing data rollback, a transaction log is generated, and the transaction log reflects the state of the source volume before data rollback.

[0172] Optionally, check whether the address mapping relationships in the second clone volume and the third clone volume are self-consistent (for example, whether there are duplicate mapping relationships) for verification processing. Check whether the physical block reference counts in the second clone volume and the third clone volume match the physical block reference counts in the transaction log for verification processing.

[0173] Combined with the scenario example, if the verification passes, it indicates that both the second clone volume and the third clone volume are valid, and then the target volume is generated. If the verification passes, it indicates that the second clone volume and the third clone volume are invalid, and an exception may occur, so roll back to the source volume to avoid the impact of abnormal data on user usage.

[0174] In this feasible implementation, through verification processing, data rollback can be aborted when abnormal data appears, avoiding the impact of abnormal data on user usage, thereby improving the reliability of data rollback.

[0175] A feasible implementation, the data management method further includes: determining the used space of multiple historical snapshots and multiple timestamps corresponding to the multiple historical snapshots in the source volume; if the used space is greater than or equal to the preset space size, then delete the historical snapshots from the source volume according to the multiple timestamps until the used space is less than the preset space size, and release the corresponding physical addresses.

[0176] Exemplarily, the used space is the memory space occupied by multiple historical snapshots. The timestamp corresponding to each historical snapshot is determined according to the moment when the historical snapshot is generated. The preset space size is the upper limit of the memory space size that the preset historical snapshots can occupy. Exceeding the preset space size may affect the normal operation of the system.

[0177] Optionally, when deleting historical snapshots, preferentially delete the historical snapshot corresponding to the timestamp farthest from the current moment.

[0178] Combined with the scenario example, the farther the historical snapshot is from the current moment, the lower the corresponding data version and the lower the value of rollback, so it is preferentially deleted.

[0179] Optionally, generate the priority corresponding to each snapshot according to the data modification information corresponding to each snapshot, and preferentially delete the snapshots with lower priority.

[0180] In this feasible implementation, by periodically recycling snapshots, it is possible to avoid excessive space occupation caused by unlimited increase of snapshots, thereby improving the fluency of system operation.

[0181] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation.

[0182] Figure 10 It is a schematic structural diagram of the data management device provided by the embodiment of the present application. As Figure 10 shown, the embodiment of the present application also provides a data management device. The data management device 100 may include: a receiving module 101, a modifying module 102, a creating module 103, and a generating module 104, where

[0183] The receiving module 101 is configured to receive a data modification request, and the data modification request includes original data, an original logical address, an original physical address, and modification information.

[0184] The modifying module 102 is configured to, according to the data modification request, perform a modification process on the original data through the modification information to obtain modified data, determine a target physical address, and store the modified data at the target physical address.

[0185] The creating module 103 is configured to determine a target generation identifier corresponding to the modified data, and create a target address mapping relationship according to the target generation identifier, the original logical address, and the target physical address, where the target generation identifier represents the generation of the data modification request.

[0186] The generating module 104 is configured to create a target snapshot according to the target generation identifier and the target address mapping relationship, and the target snapshot is used to restore the original data.

[0187] Optionally, the receiving module 101 may execute Figure 2 S201 in the embodiment.

[0188] Optionally, the modifying module 102 may execute Figure 2 S202 in the embodiment.

[0189] Optionally, the creating module 103 may execute Figure 2 S203 in the embodiment.

[0190] Optionally, the generating module 104 may execute Figure 2 S204 in the embodiment.

[0191] It should be noted that the data management device shown in the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and their implementation principles and beneficial effects are similar, which will not be elaborated here.

[0192] In a possible implementation manner, the modification module 102 is specifically configured to:

[0193] Determine the data size of the modified data;

[0194] Determine the free block bitmap, where the free block bitmap includes multiple free states corresponding to multiple blocks, and the free state is free or not free;

[0195] Determine the target physical address according to the free block bitmap and the data size.

[0196] In a possible implementation manner, the modification module 102 is specifically configured to:

[0197] Perform a scanning process on the free block bitmap according to the data size to determine multiple free blocks, and the free states of the multiple free blocks are all free;

[0198] Determine the block linked list, where the block linked list includes multiple physical addresses corresponding to multiple blocks;

[0199] Determine the multiple target physical addresses corresponding to the multiple free blocks according to the block linked list.

[0200] In a possible implementation manner, the creation module 103 is specifically configured to:

[0201] Determine the global generation identifier table, where the global generation identifier table includes multiple allocated historical generation identifiers;

[0202] Determine the target generation identifier according to the historical generation identifier, and the target generation identifier is different from any historical generation identifier.

[0203] In a possible implementation manner, the generation module 104 is specifically configured to:

[0204] Obtain the current metadata writing speed, the current processor usage rate, the current memory usage rate, and the duration interval matrix table;

[0205] Determine the target duration interval according to the current metadata writing speed, the current processor usage rate, the current memory usage rate, and the duration interval matrix table;

[0206] Create a target snapshot every target duration interval according to the target generation identifier and the target address mapping relationship.

[0207] In a possible implementation manner, the generation module 104 is specifically configured to:

[0208] Match the matching duration interval from the duration interval matrix table according to the current metadata writing speed, the current processor utilization rate, and the current memory utilization rate;

[0209] Perform calculation processing according to the current metadata writing speed, the current processor utilization rate, and the current memory utilization rate to obtain the theoretical metadata writing speed;

[0210] If the historical metadata writing speeds for a continuous preset number of times are all less than the theoretical metadata writing speed, then determine the sum of the matching duration interval and the preset value as the target duration interval;

[0211] If the historical metadata writing speeds for a continuous preset number of times are all greater than the theoretical metadata writing speed, then determine the difference between the matching duration interval and the preset value as the target duration interval;

[0212] If there are no continuous preset number of historical metadata writing speeds that are all less than the theoretical metadata writing speed, and there are no continuous preset number of historical metadata writing speeds that are all greater than the theoretical metadata writing speed, then determine the matching duration interval as the target duration interval.

[0213] Figure 11 It is a schematic structural diagram of a data management device provided by an embodiment of the present application. In Figure 10 Based on the shown embodiment, as Figure 11 shown, the data management device 110 further includes: an encryption module 105, a distributed module 106, a bitmap module 107, a first addition module 108, a second addition module 109, a rollback module 111, a verification module 112, and a recycling module 113, where,

[0214] The encryption module 105 is used for:

[0215] Generate an encryption key corresponding to the modified data;

[0216] Perform encryption processing on the modified data through the encryption key to obtain encrypted data;

[0217] Store the encrypted data in the target physical address.

[0218] The distributed module 106 is used for:

[0219] Update the global generation identifier table through the target generation identifier to obtain an updated generation identifier table;

[0220] Determine multiple distributed storage nodes;

[0221] Send the updated generation identifier table to multiple distributed storage nodes.

[0222] The bitmap module 107 is used for:

[0223] Determine multiple sub-generation identifiers corresponding to the original logical address, where the multiple sub-generation identifiers include the target generation identifier;

[0224] Arrange the multiple sub-generation identifiers in descending order according to a chain structure to obtain a list of sub-generation identifiers;

[0225] Generate a sub-bitmap index according to the list of sub-generation identifiers, where the sub-bitmap index includes multiple bits, and each bit corresponds to a sub-generation identifier.

[0226] The first addition module 108 is used for:

[0227] Determine the coding length according to multiple historical generation identifiers;

[0228] Determine a blank field according to the coding length and add the blank field to the target address mapping relationship;

[0229] Add the target generation identifier to the blank field to obtain a target snapshot.

[0230] The second addition module 109 is used for:

[0231] Determine the previous generation identifier of the target generation identifier from multiple historical generation identifiers;

[0232] Generate a parent version pointer according to the previous generation identifier;

[0233] Add the parent version pointer to the blank field.

[0234] The rollback module 111 is used for:

[0235] Receive a data rollback request, where the data rollback request includes a specified generation identifier and a source volume identifier;

[0236] Determine the source volume corresponding to the source volume identifier according to the data rollback request, and obtain multiple historical snapshots of the source volume;

[0237] Perform data rollback processing according to the specified generation identifier and multiple historical snapshots to obtain a target volume, where the target volume includes rollback data.

[0238] In a possible implementation manner, the rollback module 111 is specifically used for:

[0239] Create a first clone volume and a second clone volume of the source volume, where the second clone volume is used to store the input / output data newly generated after receiving the data rollback request;

[0240] Obtain rollback data according to the specified generation identifier and multiple historical snapshots;

[0241] Store the rollback data in the first clone volume to obtain a third clone volume;

[0242] Merge the second cloned volume and the third cloned volume to obtain a target volume.

[0243] In a possible implementation manner, the rollback module 111 is specifically configured to:

[0244] Create a logical view based on multiple historical snapshots, where the logical view includes multiple address mapping relationships corresponding to the multiple historical snapshots;

[0245] Perform matching processing based on a specified generation identifier and the logical view to obtain a specified physical address;

[0246] Obtain rollback data from the physical address.

[0247] In a possible implementation manner, the rollback module 111 is specifically configured to:

[0248] Perform matching processing based on a specified generation identifier and the logical view through multiple threads to obtain a matching result, where the matching result includes a specified physical address or a miss;

[0249] If the matching result is a miss, then determine the physical address corresponding to the generation identifier adjacent to the specified generation identifier in the logical view as the specified physical address.

[0250] The verification module 112 is configured to: obtain the transaction log of the source volume, where the transaction log includes the metadata status and the physical block reference count of the source volume before performing data rollback; perform integrity verification processing on the second cloned volume and the third cloned volume to obtain a verification result, where the verification result is verification passed or verification failed; if the verification result is verification failed, then perform recovery processing on the source volume through the transaction log to obtain a recovered volume; if the verification result is verification passed, then merge the second cloned volume and the third cloned volume to obtain a target volume.

[0251] The recycling module 113 is configured to: determine the used space of multiple historical snapshots in the source volume, and multiple timestamps corresponding to the multiple historical snapshots; if the used space is greater than or equal to a preset space size, then delete historical snapshots from the source volume according to the multiple timestamps until the used space is less than the preset space size, and release the corresponding physical address.

[0252] For the description of the features in the embodiments corresponding to the data management device, reference can be made to the relevant description in the embodiments corresponding to the data management method, which will not be elaborated here one by one.

[0253] Figure 12 This is a schematic structural diagram of the electronic device provided by this application. As Figure 12As shown in the figure, the electronic device 120 provided in this embodiment includes: at least one processor 1201 and a memory 1202. Optionally, the electronic device 120 further includes a communication component 1203. Among them, the processor 1201, the memory 1202, and the communication component 1203 are connected through a bus.

[0254] In the specific implementation process, at least one processor 1201 executes the computer-executable instructions stored in the memory 1202, so that at least one processor 1201 executes the above data management method embodiment.

[0255] For the specific implementation process of the processor 1201, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0256] In the above embodiment, it should be understood that the processor may be a central processing unit (Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0257] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0258] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0259] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is set to execute the steps in any of the above data management method embodiments when running.

[0260] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0261] The embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data management method.

[0262] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data management method.

[0263] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0264] The above has introduced in detail a data management device provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A data management method, characterized in that: include: receiving a data modification request, wherein the data modification request includes original data, an original logical address, an original physical address, and modification information; According to the data modification request, modify the original data using the modification information to obtain modified data, determine a target physical address, and store the modified data at the target physical address; Determine a target generation identifier corresponding to the modified data, and create a target address mapping relationship according to the target generation identifier, the original logical address, and the target physical address, wherein the target generation identifier indicates the generation of the data modification request; Creating a target snapshot according to the target generation identifier and the target address mapping relationship, wherein the target snapshot is used to restore the original data; It also includes: arranging the multiple child-generation identifiers corresponding to the original logical address in descending order according to a chain structure to obtain a child-generation identifier list, wherein the multiple child-generation identifiers include the target generation identifier; generating a sub-bitmap index based on the child-generation identifier list, wherein the sub-bitmap index includes multiple bits, each bit corresponding to a child-generation identifier.

2. The method according to claim 1, characterized in that Determining the target physical address includes: Determining the data size of the modified data; Determine a free block bitmap, the free block bitmap comprising a plurality of free states corresponding to a plurality of blocks, the free state being free or not free; A target physical address is determined according to the free block bitmap and the data size.

3. The method according to claim 2, characterized in that Determining a target physical address according to the free block bitmap and the data size includes: Scanning the free block bitmap according to the data size to determine a plurality of free blocks, wherein the free blocks are all in free status; Determine a block linked list, wherein the block linked list includes a plurality of physical addresses corresponding to a plurality of blocks; According to the block linked list, a plurality of target physical addresses corresponding to the plurality of free blocks are determined.

4. The method according to any one of claims 1 to 3, characterized in that Storing the modified data to the target physical address includes: generating an encryption key corresponding to the modified data; Encrypting the modified data using the encryption key to obtain encrypted data; The encrypted data is stored to the target physical address.

5. The method according to claim 1, characterized in that Determining a target generation identifier corresponding to the modified data includes: Determine a global generation identification table, wherein the global generation identification table includes a plurality of allocated historical generation identifications; The target generation identifier is determined according to the historical generation identifier, and the target generation identifier is different from any historical generation identifier.

6. The method according to claim 5, characterized in that The method further comprises: The global generation identification table is updated by using the target generation identification to obtain an updated generation identification table; Determine multiple distributed storage nodes; The update generation identification table is sent to the multiple distributed storage nodes.

7. The method according to claim 5, characterized in that The method further comprises: Determine multiple child sub-identifiers corresponding to the original logical address.

8. The method according to claim 1, characterized in that Creating a target snapshot according to the target generation identifier and the target address mapping relationship includes: Get the current metadata writing speed, current processor usage, current memory usage, and time interval matrix table; Determine a target duration interval according to the current metadata writing speed, the current processor usage rate, the current memory usage rate, and the duration interval matrix table; At each target time interval, a target snapshot is created according to the mapping relationship between the target generation identifier and the target address.

9. The method according to claim 8, characterized in that Determining a target duration interval according to the current metadata writing speed, the current processor usage rate, the current memory usage rate, and the duration interval matrix table includes: According to the current metadata writing speed, the current processor usage rate, and the current memory usage rate, matching duration intervals are obtained from the duration interval matrix table; Calculating and processing according to the current metadata writing speed, the current processor usage rate, and the current memory usage rate to obtain a theoretical metadata writing speed; If a preset number of consecutive historical metadata writing speeds are all lower than the theoretical metadata writing speed, the sum of the matching time interval and the preset value is determined as the target time interval; If a preset number of consecutive historical metadata writing speeds are all greater than the theoretical metadata writing speed, the difference between the matching time interval and the preset value is determined as the target time interval; If there are not a continuous preset number of historical metadata writing speeds that are all lower than the theoretical metadata writing speed, and there are not a continuous preset number of historical metadata writing speeds that are all higher than the theoretical metadata writing speed, then the matching duration interval is determined to be the target duration interval.

10. The method according to claim 8 or 9, characterized in that: Creating a target snapshot according to the target generation identifier and the target address mapping relationship includes: Determine the coding length according to multiple historical generation identifiers; Determine a blank field according to the encoding length, and add the blank field to the target address mapping relationship; The target generation identifier is added to the blank field to obtain the target snapshot.

11. The method according to claim 10, characterized in that The method further comprises: Determining a previous generation identifier of the target generation identifier from the multiple historical generation identifiers; Generate a parent version pointer according to the previous generation identifier; Add the parent version pointer in the blank field.

12. The method according to claim 1, characterized in that The method further comprises: receiving a data rollback request, wherein the data rollback request includes a specified generation identifier and a source volume identifier; Determine, according to the data rollback request, a source volume corresponding to the source volume identifier, and obtain multiple historical snapshots of the source volume; Data rollback processing is performed according to the specified generation identifier and the multiple historical snapshots to obtain a target volume, where the target volume includes rollback data.

13. The method according to claim 12, characterized in that Performing data rollback processing according to the multiple historical snapshots to obtain a target volume includes: creating a first clone volume and a second clone volume of the source volume, wherein the second clone volume is used to store input and output data newly generated after receiving the data rollback request; Acquire rollback data according to the specified generation identifier and the multiple historical snapshots; The rollback data is stored in the first clone volume to obtain a third clone volume; The second clone volume and the third clone volume are merged to obtain the target volume.

14. The method according to claim 13, characterized in that Acquiring rollback data according to the specified generation identifier and the multiple historical snapshots includes: Creating a logical view according to the multiple historical snapshots, the logical view including multiple address mapping relationships corresponding to the multiple historical snapshots; Perform matching processing according to the specified generation identifier and the logical view to obtain a specified physical address; The rollback data is obtained from the physical address.

15. The method according to claim 14, characterized in that Performing matching processing according to the specified generation identifier and the logical view to obtain a specified physical address includes: According to the specified generation identifier and the logical view, matching processing is performed through multiple threads to obtain a matching result, wherein the matching result includes a specified physical address or a miss; If the matching result is a miss, the physical address corresponding to the generation identifier adjacent to the designated generation identifier in the logical view is determined as the designated physical address.

16. The method according to claim 13, characterized in that The method further comprises: Acquire a transaction log of the source volume, wherein the transaction log includes a metadata state and a physical block reference count of the source volume before performing data rollback; Performing integrity verification on the second clone volume and the third clone volume to obtain a verification result, where the verification result is verification passed or verification failed; If the verification result is that the verification fails, the source volume is restored through the transaction log to obtain a recovery volume; If the verification result is verification passed, the second clone volume and the third clone volume are merged to obtain the target volume.

17. The method according to any one of claims 12 to 16, characterized in that: The method further comprises: Determining, in the source volume, used spaces of the plurality of historical snapshots and a plurality of timestamps corresponding to the plurality of historical snapshots; If the used space is greater than or equal to a preset space size, historical snapshots are deleted from the source volume according to the multiple timestamps until the used space is less than the preset space size, and corresponding physical addresses are released.

18. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data management method according to any one of claims 1 to 17 when executing the computer program.

19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data management method according to any one of claims 1 to 17.

20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data management method according to any one of claims 1 to 17 are implemented.

Citation Information

Patent Citations

  • Data processing method and device

    CN103729301A

  • Data storage method and device, storage medium and computer program product

    CN118672516A