Snapshot data statistical method, apparatus and device, and readable storage medium
By building a snapshot status mapping table in the storage system, recording the relationship between the snapshot volume and the source data and calculating the storage capacity, the problem of inaccurate capacity statistics in COW snapshot technology is solved, the accuracy of storage resource management and operation and maintenance efficiency are improved, and the storage cost of enterprises is reduced.
Patent Information
- Application Number
- CN202510397726.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, COW snapshot technology lacks an effective snapshot capacity statistics mechanism in the storage system, which makes it impossible for users to accurately obtain snapshot capacity occupancy in real time, affecting storage resource management and operation and maintenance efficiency.
By constructing a snapshot status mapping table, recording the relationship between the snapshot volume and the source data, and using a preset algorithm to calculate the storage capacity occupied by the snapshot volume, achieving accurate measurement.
It realizes accurate measurement of snapshot capacity, improves the capacity management accuracy and operation and maintenance efficiency of the storage system, and reduces the storage costs of enterprises.
Smart Images

Figure CN120255812A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of communication technologies, and in particular, to a snapshot data statistics method, apparatus, device, and readable storage medium. Background Art
[0002] A snapshot is a fully available copy of a specified data set (such as a storage volume). Snapshot technology can capture the state of a storage volume at a specific moment and save it so that it can be restored to that moment's state when needed later. This technology provides strong support for data backup, recovery, and disaster recovery, ensuring that data can be quickly restored in the event of accidental damage or loss, thus guaranteeing business continuity.
[0003] The Copy On Write (COW) technology has been widely used in some read-intensive storage systems, especially in the field of file storage. The core principle of the COW technology is that when the data content of the source data volume is updated, whether it is a deletion operation or a modification operation, the system does not directly write in the original address space. Instead, the old data in the source data volume is read out and written into a new address space. At the same time, the system updates the data address mapping table and writes the new data content into the original address space of the source data volume. This mechanism effectively avoids the risks that may be brought by directly modifying the original data and ensures the consistency and integrity of the snapshot data. However, there are also some problems in the actual application process of the COW technology.
[0004] In the actual production environment, after users create a snapshot using the COW technology, with the continuous progress of the business, there will be a large number of write operations. These write operations will cause the storage capacity occupancy of the COW area, that is, the new address space opened by the COW, to gradually increase. The size of this part of the capacity occupancy mainly depends on two factors: one is the size of the file system data block. For example, after creating a snapshot, if only 1KB of data is modified, but the block size of the file system is 4KB, then the entire 4KB block will be copied to the COW area. Even if the actual modified data volume is small, it will occupy a large storage space; the other is the amount of data actually modified or deleted by the user after the snapshot is established. As the business continues to develop, the number of data blocks protected by the snapshot that are deleted and modified is increasing, and the snapshot capacity will also increase accordingly.
[0005] For users of storage systems, during the operation and maintenance of storage systems, they are very concerned about the actual occupied size of this part of the capacity. However, in current file systems that adopt the COW snapshot technology, there is a lack of an effective snapshot capacity statistics mechanism. Users cannot directly obtain real-time data on snapshot capacity and can only indirectly infer the occupancy of snapshot capacity through some external means. For example, users can roughly estimate the occupancy of snapshot capacity by comparing the changes in data before and after deleting a snapshot. However, this method is not only inaccurate but also cumbersome to operate and cannot meet the user's need for precise management of snapshot capacity in actual use. Summary of the Invention
[0006] In view of this, this specification provides a snapshot data statistics method, apparatus, electronic device, and readable storage medium to improve the problem of difficultly knowing the occupied size of the storage capacity of the snapshot volume.
[0007] Specific technical solutions are as follows:
[0008] This specification provides a snapshot data statistics method applied to a storage device. The method includes: in response to an event of snapshot volume generation, recording corresponding status information of the snapshot volume in a snapshot status mapping table according to snapshot data units included in the snapshot volume; the snapshot volume is generated based on a change in source data, the snapshot volume records at least one snapshot data unit, and the data of the snapshot data unit is the data of the corresponding data unit in the source data before the source data changed this time; the snapshot status mapping table corresponds to the source data and records the association relationship between the snapshot volume and each data unit of the source data, and the status information is used to reflect the existence status of whether there is a corresponding snapshot data unit in the snapshot volume for the data unit of the source data; in response to a snapshot data statistics signaling, according to the existence status reflected by the status information, obtaining storage capacity information occupied by the corresponding snapshot volume with a preset algorithm, and the preset algorithm includes accumulating the storage capacity occupied by each data unit of the source data for which there is a corresponding snapshot data unit in the snapshot volume to be statistically analyzed.
[0009] As a technical solution, in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, the corresponding status information of the snapshot volume is recorded in the snapshot status mapping table, including: according to the snapshot data units included in the snapshot volume, for each data unit of the source data corresponding to the status information associated with the snapshot volume in the snapshot status mapping table, a status value is recorded. If there is a corresponding snapshot data unit, the status value is recorded as 1, and the remaining status values are recorded as 0; in response to a snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: according to the status information of the snapshot volume in the snapshot status mapping table, the storage capacity value occupied by each data unit of the source data is multiplied by the status value corresponding to each data unit of the snapshot volume respectively, and the products are accumulated, and the accumulated result is used as the storage capacity information occupied by the snapshot volume.
[0010] As a technical solution, the snapshot status mapping table includes total status information, and the total status information includes count values with an initial value of 0 respectively recorded for each data unit of the source data; in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, the count value of the data unit of the corresponding source data in the total status information is incremented by 1; in response to a snapshot data total signaling, according to the total status information of the snapshot volume in the snapshot status mapping table, the storage capacity value occupied by each data unit of the source data is multiplied by the count value respectively recorded for each data unit, and the products are accumulated, and the accumulated result is used as the total storage capacity information of all snapshot volumes corresponding to the source data.
[0011] As a technical solution, in response to a snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: in response to a snapshot data statistics signaling triggered by an event of generating a snapshot volume, the storage capacity information of the generated snapshot volume is obtained and recorded.
[0012] As a technical solution, in response to a snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: in response to a snapshot data statistics signaling issued by a user, the snapshot volume information to be statistically analyzed indicated by the snapshot data statistics signaling is parsed, and the storage capacity information occupied by the corresponding snapshot volume is obtained and fed back.
[0013] As a technical solution, the snapshot status mapping table includes a snapshot serial number, and the snapshot serial number is used to identify and distinguish each separately generated and independently stored snapshot volume.
[0014] As a technical solution, the storage device records a number of snapshot state mapping tables respectively associated with different source data; in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, record the corresponding state information of the snapshot volume in the snapshot state mapping table, including: in response to the event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume and the source data associated with the snapshot volume, record the corresponding state information of the snapshot volume in the snapshot state mapping table associated with the source data.
[0015] This specification also provides a snapshot data statistics device, the device includes: a first module, configured to, in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, record the corresponding state information of the snapshot volume in the snapshot state mapping table; the snapshot volume is generated based on a change in the source data, the snapshot volume records at least one snapshot data unit, and the data of the snapshot data unit is the data of the corresponding data unit in the source data before the source data changes this time; the snapshot state mapping table corresponds to the source data and records the association relationship between the snapshot volume and each data unit of the source data, and the state information is used to reflect the existence state of the data unit of the source data having a corresponding snapshot data unit in the snapshot volume; a second module, configured to, in response to a snapshot data statistics signaling, according to the existence state reflected by the state information, obtain the storage capacity information occupied by the corresponding snapshot volume by using a preset algorithm, and the preset algorithm includes accumulating the storage capacity occupied by each data unit of the source data having a corresponding snapshot data unit in the snapshot volume to be statistically analyzed.
[0016] As a technical solution, in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, recording the corresponding state information of the snapshot volume in the snapshot state mapping table includes: according to the snapshot data units included in the snapshot volume, the state information associated with the snapshot volume in the snapshot state mapping table respectively records state values corresponding to each data unit of the source data, if there is a corresponding snapshot data unit, the state value is recorded as 1, and the remaining state values are recorded as 0; in response to a snapshot data statistics signaling, according to the existence state reflected by the state information, obtaining the storage capacity information occupied by the corresponding snapshot volume by using a preset algorithm includes: according to the state information of the snapshot volume in the snapshot state mapping table, multiplying the storage capacity value occupied by each data unit of the source data by the state value corresponding to each data unit for the snapshot volume respectively, and accumulating each product and using the accumulated result as the storage capacity information occupied by the snapshot volume.
[0017] As a technical solution, the snapshot status mapping table includes total status information, and the total status information includes count values with an initial value of 0 respectively recorded for each data unit of the source data; the first module is further configured to, in response to an event of generating a snapshot volume, increment by 1 the count value of the data unit of the corresponding source data in the total status information according to the snapshot data units included in the snapshot volume; the second module is further configured to, in response to a snapshot data total signaling, respectively multiply the storage capacity value occupied by each data unit of the source data by the count value recorded for each data unit according to the total status information of the snapshot volume in the snapshot status mapping table, accumulate the products, and use the accumulated result as the total storage capacity information of all snapshot volumes corresponding to the source data.
[0018] As a technical solution, the obtaining the storage capacity information occupied by a corresponding snapshot volume according to the existence status reflected by the status information in response to a snapshot data statistics signaling includes: obtaining and recording the storage capacity information of the generated snapshot volume in response to a snapshot data statistics signaling triggered by an event of generating a snapshot volume.
[0019] As a technical solution, the obtaining the storage capacity information occupied by a corresponding snapshot volume according to the existence status reflected by the status information in response to a snapshot data statistics signaling includes: parsing the information of the snapshot volume to be statistically analyzed indicated by the snapshot data statistics signaling in response to a snapshot data statistics signaling issued by a user, and obtaining and feeding back the storage capacity information of the corresponding snapshot volume according to the information of the snapshot volume to be statistically analyzed.
[0020] As a technical solution, the snapshot status mapping table includes a snapshot serial number, and the snapshot serial number is used to identify and distinguish each separately generated and independently stored snapshot volume.
[0021] As a technical solution, the storage device records a plurality of snapshot status mapping tables respectively associated with different source data; the recording the corresponding status information of the snapshot volume in the snapshot status mapping table according to the snapshot data units included in the snapshot volume in response to an event of generating a snapshot volume includes: recording the corresponding status information of the snapshot volume in the snapshot status mapping table associated with the source data according to the snapshot data units included in the snapshot volume and the source data associated with the snapshot volume in response to an event of generating a snapshot volume.
[0022] This specification also provides an electronic device, including a processor and a readable storage medium, where the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the foregoing snapshot data statistics method.
[0023] This specification also provides a readable storage medium storing machine-executable instructions, which, when called and executed by a processor, cause the processor to implement the aforementioned snapshot data statistics method.
[0024] The above technical solutions provided in this specification at least bring the following beneficial effects:
[0025] By constructing a snapshot state mapping table to track the data unit association relationship in real time, and realizing accurate measurement of the snapshot capacity based on the existence status flag, the measurement deviation between the data block specification and the actual modification amount in the traditional COW mechanism is eliminated, and the capacity statistics accuracy is improved to the data unit level; and through the dynamic mapping of the status information, it supports real-time calculation of the storage occupancy of a single snapshot volume or a multi-level snapshot superposition scenario, enabling the operation and maintenance personnel to accurately identify the actual capacity distribution of each snapshot; at the same time, this statistical mechanism provides data support for storage resource recovery, snapshot dependency analysis and fragmentation reorganization, effectively improving the capacity management accuracy and operation and maintenance efficiency of the storage system, and reducing the enterprise storage cost. Description of the Drawings
[0026] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments of this specification or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings of the embodiments of this specification.
[0027] Figure 1 is a flowchart of the snapshot data statistics method in an embodiment of this specification;
[0028] Figure 2 is a structural diagram of the snapshot data statistics device in an embodiment of this specification;
[0029] Figure 3 is a hardware structural diagram of the electronic device in an embodiment of this specification.
[0030] Figure 4 is a schematic diagram of the snapshot state mapping table in an embodiment of this specification.
[0031] Reference numerals: the first module 21, the second module 22. Detailed Embodiments
[0032] The terms used in the embodiments of this specification are for the purpose of describing specific embodiments only and do not limit this specification. The singular forms "a", "the", and "said" used in this specification and the claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more of the associated listed items.
[0033] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" used may be interpreted as "when" or "while" or "in response to a determination".
[0034] The operating principle of the COW snapshot mechanism is as follows: when a data modification or deletion operation occurs on the source data volume, the system reads the original data block from the source address space and writes it into a newly allocated storage area (i.e., the COW area), then writes the new data into the source address space, and updates the address mapping table to reflect the change in data location. This process involves two write operations (migrating the old data to the COW area and writing the new data to the source address) and one read operation, achieving an instant snapshot function while ensuring data consistency.
[0035] The COW snapshot technology has significant capacity management defects. In the scenario of continuous business operation, frequent modifications to the source data volume will cause the storage occupancy of the COW area to continue to grow, and its growth rate is affected by both the file system data block specification and the actual data change amount. For example, when the file system uses a 4KB fixed block size, even if only 1KB of data is modified, the system still needs to migrate the entire 4KB data block to the COW area. This design feature makes the snapshot capacity form a non-linear relationship with the amount of data actually modified by the user, and the existing technology system lacks an accurate measurement mechanism for snapshot capacity. Operation and maintenance personnel can only indirectly estimate the capacity consumption by comparing the storage space changes before and after snapshot deletion. This method not only has poor real-time performance but also causes error accumulation in the scenario of multiple snapshots superimposed, seriously affecting the effectiveness of storage resource planning.
[0036] The above technical solutions have defects. First, enterprise users cannot grasp the resource occupancy of each snapshot in real time, and it is difficult to clean up historical snapshots with low value but high occupancy in a timely manner, resulting in waste of storage resources. Second, there are blind spots in multi-level snapshot management. When the system creates snapshots at multiple time points, the COW regions associated with each snapshot may contain data blocks of different versions, but the existing mapping table structure only records the association relationship between the latest snapshot and the source address, and cannot trace the data block distribution of historical snapshots. Moreover, the performance optimization of the storage system lacks data support. The random write characteristics of the COW region may cause storage fragmentation problems, and the lack of a capacity statistics mechanism makes it impossible to formulate a defragmentation strategy based on quantification. These problems are particularly prominent in long-running business systems and seriously restrict the operation and maintenance efficiency and economy of the storage system.
[0037] In view of this, this specification provides a snapshot data statistics method, device, electronic device, and readable storage medium to at least improve one of the above technical problems.
[0038] The specific technical solution is as follows.
[0039] In one implementation, this specification provides a snapshot data statistics method applied to a storage device. The method includes: in response to an event of snapshot volume generation, recording the corresponding status information of the snapshot volume in a snapshot status mapping table according to the snapshot data units included in the snapshot volume; the snapshot volume is generated based on a change in source data, and the snapshot volume records at least one snapshot data unit, and the data of the snapshot data unit is the data of the corresponding data unit in the source data before the source data changes this time; the snapshot status mapping table corresponds to the source data and records the association relationship between the snapshot volume and each data unit of the source data, and the status information is used to reflect the existence status of the data unit of the source data having a corresponding snapshot data unit in the snapshot volume; in response to a snapshot data statistics signaling, obtaining the storage capacity information occupied by the corresponding snapshot volume according to the existence status reflected by the status information, and the preset algorithm includes accumulating the storage capacity occupied by each data unit of the source data having a corresponding snapshot data unit in the snapshot volume to be statistically analyzed.
[0040] Specifically, as Figure 1 , it includes the following steps:
[0041] Step S11, in response to an event of snapshot volume generation, recording the corresponding status information of the snapshot volume in a snapshot status mapping table according to the snapshot data units included in the snapshot volume.
[0042] When the source data changes (such as being modified or deleted), according to the COW snapshot mechanism, a snapshot volume will be automatically generated. This snapshot volume contains at least one snapshot data unit, and each snapshot data unit records the data of the corresponding data unit in the source data before the change. To effectively manage the snapshot data units, the system maintains a snapshot status mapping table. This mapping table not only corresponds to the source data but also records the association relationships between the snapshot volume and each data unit of the source data. In this way, whenever a new snapshot volume is generated, the system will update the corresponding status information in the mapping table to reflect the existence status of the corresponding snapshot data unit for the data unit of the source data in the snapshot volume.
[0043] Suppose there is a file A with a size of 10MB, and two snapshots have been generated: Snapshot 1 and Snapshot 2. If, after creating Snapshot 2, the user modifies file A and only changes 1KB of the data. Since the block size of the file system is 4KB, the COW mechanism will copy and save this part of the data to the snapshot volume. At this time, the snapshot status mapping table will record this change, indicating which part of the data in file A has been saved by Snapshot 2. This recording method enables the system to accurately track which parts of the data have been protected by the snapshots and the space occupied by each of these snapshots even if file A is modified multiple times subsequently.
[0044] Step S12, in response to the snapshot data statistics signaling, obtain the storage capacity information occupied by the corresponding snapshot volume according to the existence status reflected by the status information with a preset algorithm.
[0045] When the system receives the snapshot data statistics signaling, it calculates the storage capacity occupied by the corresponding snapshot volume according to the status information previously recorded in the snapshot status mapping table. Specifically, the core of the preset algorithm is to accumulate the storage capacity occupied by the data units of the source data for which there are corresponding snapshot data units in the snapshot volume to be statistically analyzed. For each data unit protected by the snapshot, the system accumulates according to its actual size, thereby obtaining the total storage capacity of the entire snapshot volume. This method not only takes into account the situation of a single snapshot volume but also can handle complex situations in a multi-level snapshot environment. For example, in the previous example, if the user wants to know how much storage space Snapshot 1 and Snapshot 2 each occupy, the system can accurately calculate the storage capacity occupied by these two snapshots respectively by analyzing the information in the snapshot status mapping table.
[0046] To improve the statistical efficiency and accuracy, an incremental update strategy can be adopted, that is, only update the relevant entries when the source data changes each time, rather than reconstructing the entire mapping table every time.
[0047] In one embodiment, when a snapshot volume is generated, the system responds to this event and records corresponding status information in the snapshot status mapping table according to the snapshot data units included in the snapshot volume. The snapshot volume is generated based on changes in the source data and records at least one snapshot data unit, such as one, three, ten, one hundred, etc. The data of these snapshot data units is the data of the corresponding data units in the source data before the source data changes. The snapshot status mapping table corresponds to the source data and records the association relationship between the snapshot volume and each data unit of the source data. The status information is used to reflect the existence status of whether there is a corresponding snapshot data unit for the data unit of the source data in the snapshot volume.
[0048] When receiving the snapshot data statistics signaling, the system obtains the storage capacity information occupied by the corresponding snapshot volume through a preset algorithm according to the existence status reflected by the status information. The preset algorithm includes accumulating the storage capacity occupied by each data unit of the source data for which there is a corresponding snapshot data unit in the snapshot volume to be statistically analyzed.
[0049] In a storage system, there is a source data volume. This source data volume is divided into multiple data units, and each data unit has a unique address identifier. When a user modifies some data units in the source data volume, the system triggers the snapshot mechanism. For example, if the user modifies the data unit with the address 0x1 in the source data volume, a corresponding snapshot volume is created.
[0050] The snapshot volume will contain a snapshot data unit, and the data of this snapshot data unit is the data of the data unit with the address 0x1 in the source data before the modification. At the same time, the system records the status information of this snapshot volume in the snapshot status mapping table. The status information will indicate that there is a corresponding snapshot data unit in the snapshot volume for the data unit with the address 0x1 in the source data.
[0051] In the snapshot status mapping table, each column corresponds to a data unit of the source data, and each row corresponds to a snapshot volume. The status information in the table can be a simple flag bit, such as 0 or 1, where 0 indicates that there is no corresponding snapshot data unit in the snapshot volume for the data unit of the source data, and 1 indicates that there is a corresponding snapshot data unit. In this way, the system can quickly query the status of any data unit of the source data in each snapshot volume.
[0052] When a snapshot data statistics signaling is received, it is considered that the storage capacity occupied by the snapshot volume needs to be counted. At this time, the system will calculate the storage capacity occupied by the snapshot volume through a preset algorithm according to the status information in the snapshot status mapping table. Specifically, the system will traverse the snapshot status mapping table to find all data units of the source data with the status information of 1. For each such data unit, the system will obtain the storage capacity it occupies and accumulate them. The final accumulated value is the storage capacity occupied by the snapshot volume.
[0053] Suppose there are 5 data units in the source data volume, located at addresses 0x1 to 0x5 respectively. The size of each data unit is 4KB. The user has performed the following operations on the data in the source data volume: First, modify the data unit at address 0x1; then, modify the data unit at address 0x3; finally, modify the data unit at address 0x5. Each modification will trigger the snapshot mechanism to generate a snapshot volume. Therefore, the system will generate three snapshot volumes, corresponding to these three modification operations respectively.
[0054] In the snapshot status mapping table, the system will record the status information of each snapshot volume. For the first snapshot volume, the status information will indicate that there is a corresponding snapshot data unit for the data unit at address 0x1; for the second snapshot volume, the status information will indicate that there is a corresponding snapshot data unit for the data unit at address 0x3; for the third snapshot volume, the status information will indicate that there is a corresponding snapshot data unit for the data unit at address 0x5. There are no corresponding snapshot data units for data units at other addresses in these snapshot volumes, so their status information is 0.
[0055] When it is necessary to count the storage capacity occupied by the first snapshot volume, the system will receive the snapshot data statistics signaling. The system will find the data unit at address 0x1 according to the status information in the snapshot status mapping table. Since the size of each data unit is 4KB, the storage capacity occupied by the first snapshot volume is 4KB. Similarly, when counting the storage capacity occupied by the second snapshot volume, the system will find the data unit at address 0x3, and its occupied storage capacity is also 4KB. For the third snapshot volume, the system will find the data unit at address 0x5, and its occupied storage capacity is also 4KB.
[0056] In this way, the method of the present invention can accurately count the storage capacity occupied by each snapshot volume. This method is not only applicable to the statistics of a single snapshot volume, but can also be extended to the statistics of multiple snapshot volumes. For example, if it is necessary to count the total storage capacity occupied by all snapshot volumes, the system can simply add up the storage capacities occupied by all snapshot volumes. In addition, this method can also be used to analyze the trend of the storage capacity change of the snapshot volume. For example, by regularly counting the storage capacity occupied by the snapshot volume, the system can generate a storage capacity change curve to help users better understand the storage requirements of the snapshot volume.
[0057] When the storage capacity occupied by the snapshot volume reaches a certain threshold, the system can automatically trigger a warning to remind the user to optimize the storage capacity. In addition, the system can also provide the user with optimization suggestions for the storage capacity according to the storage capacity information of the snapshot volume. For example, if a certain snapshot volume occupies a large storage capacity but has a low importance, the system can suggest that the user delete the snapshot volume to free up storage space.
[0058] In a multi-level snapshot, a source data volume may have multiple snapshot volumes, and there may be dependencies between these snapshot volumes. For example, one snapshot volume may be generated based on another snapshot volume. In this case, the technical solution of this embodiment can still effectively count the storage capacity occupied by each snapshot volume. The system will respectively count the storage capacity occupied by each snapshot volume according to the status information in the snapshot status mapping table. At the same time, the system can also consider the dependencies between the snapshot volumes to calculate the storage capacity more accurately.
[0059] For example, there is an initial snapshot volume A for a source data volume. Subsequently, the user makes some modifications to the source data volume and generates a new snapshot volume B. Snapshot volume B is generated based on snapshot volume A. In this case, there is a dependency between snapshot volume A and snapshot volume B. When counting the storage capacity occupied by snapshot volume B, it is necessary to consider the snapshot data units that already exist in snapshot volume A. If snapshot volume A already contains the snapshot data units of a certain data unit of the source data, then when counting the storage capacity occupied by snapshot volume B, the storage capacity occupied by this data unit is not recalculated. This can avoid double counting of the storage capacity and improve the accuracy of the statistics.
[0060] The above embodiment provides an efficient and accurate solution for the statistics of snapshot data by introducing a snapshot status mapping table and a preset algorithm. This technical solution can not only help users better understand the storage requirements of snapshot volumes, but also optimize the use of storage resources, improve the performance and reliability of the storage system.
[0061] In one embodiment, in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, the corresponding status information of the snapshot volume is recorded in the snapshot status mapping table, including: according to the snapshot data units included in the snapshot volume, for each data unit of the source data, the status value corresponding to the status information associated with the snapshot volume in the snapshot status mapping table is recorded. If there is a corresponding snapshot data unit, the status value is recorded as 1, and the remaining status values are recorded as 0; in response to the snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: according to the status information of the snapshot volume in the snapshot status mapping table, multiplying the storage capacity value occupied by each data unit of the source data by the status value corresponding to each data unit for the snapshot volume respectively, and accumulating the products, and taking the accumulated result as the storage capacity information occupied by the snapshot volume.
[0062] In one embodiment, the snapshot status mapping table includes total status information, and the total status information includes count values with an initial value of 0 respectively recorded for each data unit of the source data; in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, the count value of the data unit of the corresponding source data in the total status information is incremented by 1; in response to the snapshot data total signaling, according to the total status information of the snapshot volume in the snapshot status mapping table, multiplying the storage capacity value occupied by each data unit of the source data by the count value respectively recorded for each data unit, and accumulating the products, and taking the accumulated result as the total storage capacity information of all snapshot volumes corresponding to the source data.
[0063] In the above embodiment, a dynamic accumulation counting system is introduced to achieve accurate measurement of the total storage capacity consumption in the multi-snapshot superposition scenario. The core improvement of this solution is to upgrade the binary status flag to an accumulative numerical flag, and the frequency of snapshot reference to the data unit is reflected by numerical superposition, so as to construct a capacity statistical model for the global snapshot set. During specific implementation, the storage system creates an independent total status information record item for each data unit (Data Unit, DU) of the source data volume during the initialization phase. This record item includes a counter field (Counter Field, CF) with an initial value of 0, and the counter field records the count value. When a snapshot generation event occurs, the system traverses all snapshot data units (Snapshot Data Unit, SDU) included in the current snapshot volume, and performs an atomic increment operation on the CF of the source DU corresponding to each SDU. The final total storage capacity is calculated by traversing the CF values of all source DUs, multiplying them by the standard capacity of the DUs, and then accumulating the results.
[0064] Specifically, assume that the source data volume contains 5 DUs each with a size of 4KB, and their logical addresses are 0x1 to 0x5 respectively. In the initial state, the CF value of each DU in the total status information table is 0. When the user first creates snapshot seq1, assume that source DU 0x4 is modified, triggering the COW mechanism to generate the corresponding SDU. At this time, the system increments the CF value of 0x4 in the total status information table from 0 to 1. If snapshot seq2 is created subsequently and source DU 0x4 is modified again to generate a new SDU, the system continues to increment the CF value of 0x4 from 1 to 2. When the administrator triggers the snapshot data total signaling, the system executes the statistical algorithm: traverses the CF values of all DUs, calculates 4KB × CF value for 0x1 (CF = 0), 0x2 (CF = 0), 0x3 (CF = 0), 0x4 (CF = 2), 0x5 (CF = 0) respectively, and finally obtains a total storage capacity of 8KB (2 × 4KB for 0x4). This calculation method accurately reflects the repeated occupation of the same DU by two snapshots, avoiding the calculation redundancy caused by traversing the status bits of multiple snapshots in the traditional method.
[0065] In the above embodiment, by replacing discrete marking with numerical accumulation, the capacity statistics are extended from the single snapshot dimension to the global dimension, enabling the administrator to obtain the total capacity consumption of historical snapshots without querying each snapshot one by one; at the same time, the calculation complexity of the above embodiment is significantly reduced, and the total capacity calculation only requires one full table traversal instead of multiple snapshot traversals, optimizing the time complexity from O(N×M) to O(M) (M is the total number of DUs) in a system with N snapshots; in addition, the above embodiment supports storage hot spot analysis, and the source DUs with high-frequency modifications can be intuitively presented by high CF values, providing data support for storage optimization. For example, the CF value of a DU in the log of an e-commerce database reaches 1000 times, indicating that this area has retained old data for 1000 historical snapshots, prompting the administrator to adopt a higher redundancy strategy for this DU or migrate it to a high-performance storage medium.
[0066] Furthermore, the system adopts a hierarchical storage architecture to manage the total status information table: the CF values of active DUs are stored in the cache area in memory, and continuous modifications are batch processed through the Write Coalescing technology; the CF values of inactive DUs are persistently stored in the metadata area of the NVMe SSD, and a B+ tree index is used to achieve fast retrieval. For large-scale storage systems, a distributed counting mechanism is implemented: the source DUs are scattered to multiple storage nodes according to the hash algorithm, each node maintains local CF values, and during statistics, the local products are calculated in parallel through the MapReduce framework and then globally reduced. This design enables the system to maintain a sub-second total capacity response speed even under the EB-level data scale.
[0067] The exception handling mechanism ensures the accuracy of counting: when the snapshot rollback operation causes the SDU to become invalid, the system does not simply decrement the CF value, but maintains an independent Invalidation Counter (IC). For example, when snapshot seq2 is rolled back to seq1, the 0x4 SDU corresponding to seq2 is marked as invalid, but the CF value of 0x4 in the total status information table remains 2 unchanged, while the IC value increases by 1. When performing physical space recycling, the system determines the actual valid reference count based on the (CF - IC) value, avoiding counting errors caused by snapshot state changes. This design ensures the statistical logic consistency while meeting the accuracy requirements for storage space recycling.
[0068] In one implementation, in response to the snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: obtaining and recording the storage capacity information of the generated snapshot volume in response to the snapshot data statistics signaling triggered by an event generated according to the snapshot volume.
[0069] In one implementation, in response to the snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: in response to the snapshot data statistics signaling issued by the user, parsing the information of the snapshot volume to be counted indicated by the snapshot data statistics signaling, and obtaining and feeding back the storage capacity information occupied by the corresponding snapshot volume according to the information of the snapshot volume to be counted.
[0070] In one implementation, the snapshot state mapping table includes snapshot sequence numbers, which are used to identify and distinguish each separately generated and independently stored snapshot volume.
[0071] In one implementation, the storage device records several snapshot state mapping tables respectively associated with different source data; in response to an event generated by the snapshot volume, according to the snapshot data unit included in the snapshot volume, recording the corresponding status information of the snapshot volume in the snapshot state mapping table, including: in response to an event generated by the snapshot volume, according to the snapshot data unit included in the snapshot volume and the source data associated with the snapshot volume, recording the corresponding status information of the snapshot volume in the snapshot state mapping table associated with the source data.
[0072] In one implementation, a mapping table between the snapshot sequence number and the source data address is newly defined in the original COW snapshot mechanism. A method for calculating the snapshot capacity through the mapping table between the snapshot sequence number and the source data address is defined:
[0073] Data 1, 2, 3, 4, and 5 in the source data volume are stored in the address space from 0x1 to 0x5. At this time, a snapshot of the source data volume is created with a snapshot sequence number of seq1. After the snapshot is created, the user newly writes data 6 and modifies the data 4 at 0x4. At this time, the COW mechanism is triggered, the data 4 is copied to a new address space in the COW area, and an address mapping table is established.
[0074] This embodiment defines a mapping table between the snapshot sequence number and the source data address. For example, Figure 4 , in this table, each row of seq1 represents the number of times the address space in the source data volume has been modified or deleted after the snapshot seq1 is created. The default value of the status row is 0, and when the value of seq1 is greater than or equal to 1, the status is set to 1.
[0075] There is a mapping relationship between address 0x4 and 1x4, and their storage space occupancy is also the same, which depends on how many data blocks this address space occupies and the data block size defined by the file system (for example, the default data block size of the file system is 4KB. Assuming that address 0x4 spans two data blocks, then the occupied space of this address is 8KB, even if the actual used space of the data itself may only be 7KB). That is to say, to calculate the capacity size of the snapshot volume, only the total size of all the data blocks modified on the source data volume needs to be known.
[0076] In the actual production environment, for important data, there are often multiple snapshots at different times. This algorithm is still applicable to the scenario of multi-level snapshots. In the previous example, if data 6 is written to the address space 0x4 after the seq1 snapshot, the data in the source data volume at this time is 1, 2, 3, 6, 5; at this time, snapshot 2 is created to form a multi-level snapshot. After snapshot 2 is created, data 8 is written to 0x4 again.
[0077] The count corresponding to the 0x4 address space of the source data volume corresponding to seq1 is incremented by 1, representing the number of modifications and deletions. Based on the previous assumption, the capacity of the snapshot with the sequence number seq1 is still 8KB, and the capacity of the snapshot with the sequence number seq2 is 8KB.
[0078] Simply put, after creating a snapshot, whenever the storage system receives a request to modify or delete the source data address space, the count corresponding to the seq sequence number in the newly defined mapping table between the snapshot sequence number and the source data volume is incremented. Finally, the status of the capacity occupancy is generated according to the count. After multiplying the status by the capacity information of the corresponding address space and summing, the capacity occupancy of a certain seq snapshot is obtained.
[0079] In one embodiment, such as Figure 2, this specification also provides a snapshot data statistics device, which includes: a first module for, in response to an event of generating a snapshot volume, recording the corresponding status information of the snapshot volume in a snapshot status mapping table according to the snapshot data units included in the snapshot volume; the snapshot volume is generated based on a change in source data, the snapshot volume records at least one snapshot data unit, and the data of the snapshot data unit is the data of the corresponding data unit in the source data before the source data changes this time; the snapshot status mapping table corresponds to the source data and records the association relationship between the snapshot volume and each data unit of the source data, and the status information is used to reflect the existence status of the data unit of the source data having a corresponding snapshot data unit in the snapshot volume; a second module for, in response to a snapshot data statistics signaling, obtaining the storage capacity information occupied by the corresponding snapshot volume according to the existence status reflected by the status information, and the preset algorithm includes accumulating the storage capacity occupied by each data unit of the source data having a corresponding snapshot data unit in the snapshot volume to be statistically analyzed.
[0080] In one implementation, the step of, in response to an event of generating a snapshot volume, recording the corresponding status information of the snapshot volume in a snapshot status mapping table according to the snapshot data units included in the snapshot volume includes: according to the snapshot data units included in the snapshot volume, the status information associated with the snapshot volume in the snapshot status mapping table records status values corresponding to each data unit of the source data respectively. If there is a corresponding snapshot data unit, the status value is recorded as 1, and the remaining status values are recorded as 0; the step of, in response to a snapshot data statistics signaling, obtaining the storage capacity information occupied by the corresponding snapshot volume according to the existence status reflected by the status information includes: according to the status information of the snapshot volume in the snapshot status mapping table, multiplying the storage capacity value occupied by each data unit of the source data by the status value corresponding to each data unit for the snapshot volume respectively, accumulating each product, and using the accumulated result as the storage capacity information occupied by the snapshot volume.
[0081] In one implementation, the snapshot status mapping table includes total status information, and the total status information includes count values with an initial value of 0 recorded corresponding to each data unit of the source data respectively; the first module is further configured to, in response to an event of generating a snapshot volume, increment the count value of the data unit of the source data corresponding to the total status information according to the snapshot data units included in the snapshot volume; the second module is further configured to, in response to a snapshot data total statistics signaling, according to the total status information of the snapshot volume in the snapshot status mapping table, multiplying the storage capacity value occupied by each data unit of the source data by the count value recorded for each data unit respectively, accumulating each product, and using the accumulated result as the total storage capacity information of all snapshot volumes corresponding to the source data.
[0082] In one embodiment, in response to the snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: obtaining and recording the storage capacity information of the generated snapshot volume in response to the snapshot data statistics signaling triggered by an event generated according to the snapshot volume.
[0083] In one embodiment, in response to the snapshot data statistics signaling, according to the existence status reflected by the status information, the storage capacity information occupied by the corresponding snapshot volume is obtained by a preset algorithm, including: in response to the snapshot data statistics signaling issued by the user, parsing the information of the snapshot volume to be statistically analyzed indicated by the snapshot data statistics signaling, and obtaining and feeding back the storage capacity information occupied by the corresponding snapshot volume according to the information of the snapshot volume to be statistically analyzed.
[0084] In one embodiment, the snapshot status mapping table includes a snapshot sequence number, and the snapshot sequence number is used to identify and distinguish each separately generated and independently stored snapshot volume.
[0085] In one embodiment, the storage device records several snapshot status mapping tables respectively associated with different source data; in response to an event of generating a snapshot volume, according to the snapshot data unit included in the snapshot volume, the corresponding status information of the snapshot volume is recorded in the snapshot status mapping table, including: in response to an event of generating a snapshot volume, according to the snapshot data unit included in the snapshot volume and the source data associated with the snapshot volume, the corresponding status information of the snapshot volume is recorded in the snapshot status mapping table associated with the source data.
[0086] In one embodiment, the snapshot data statistics device implements a data statistics method, which realizes the accurate statistics of the capacity occupation of snapshot volumes in the storage device by constructing a measurement system combining dynamic mapping and status marking. The core of this method lies in establishing a snapshot status mapping table with multi-dimensional associations, and through a status tracking mechanism at the data unit level, converting the traditional fuzzy statistics based on the overall utilization rate of the storage pool into a precise calculation based on the snapshot granularity.
[0087] During the initialization phase of the storage device, a basic metadata structure needs to be created for the source data volume. This structure includes a source data address space partitioning module, a snapshot metadata management module, and a status flag persistence module. Among them, the source data address space partitioning module divides the source data volume into several data units (DUs) of a fixed size. Each DU corresponds to the smallest management unit of physical storage. In a typical implementation, 4KB is used as the standard size of the DU, but it can be adjusted to any value within the range of 512B to 64KB according to the physical characteristics of the storage medium. This partitioning operation is implemented through the address translation layer of the storage controller. Each DU is assigned a unique logical unit address (LUA), and the corresponding relationship between the LUA and the physical block address (PBA) is established in the physical address mapping table of the storage medium.
[0088] When a data change event occurs in the source data volume, the write request processing flow of the storage device triggers the snapshot generation mechanism. Taking the Nth modification of the source data volume as an example, assume that the user performs a data overwrite operation on the DU with the LUA of 0x1000. The storage controller first detects the current active snapshot status: if there is an existing but unfinished snapshot volume in the system, the data capture of the existing snapshot is preferentially executed; if it is in a snapshot-free state, a new snapshot volume is created according to the preset policy or external instruction. The core operations of snapshot generation are divided into three stages: the first stage is data pre-reading, where the controller reads the original data of the DU to be modified from the source address space, and this data is the historical version that will be overwritten; the second stage is the allocation of snapshot data units (SDUs). The storage system allocates idle physical blocks from the reserved snapshot storage pool, writes the read original data into this physical block to form an SDU corresponding to the source DU, and records the association relationship between the physical address of the SDU and the logical address of the source DU; the third stage is the update of the status flag. A new status record entry is created for the current snapshot volume in the snapshot status mapping table (SSMT). This record entry includes the snapshot volume identifier (Snapshot ID, SID), the logical address of the source DU, the physical address of the SDU, and the status flag bit. Among them, the status flag bit uses a binary marking mechanism. When there is a corresponding SDU for a certain source DU, it is marked as "1", otherwise it is "0". This marking system can be extended to a multi-bit pattern in specific implementations to support the coexistence management of multiple versions of snapshots. For example, when using 3 bits, the status information of 8 different snapshot versions can be represented simultaneously.
[0089] The construction and maintenance of the snapshot status mapping table constitute the key technical features of this method. At the physical implementation level, the SSMT adopts a distributed storage architecture, and each source DU corresponds to an independent status record linked list. Taking the source DU 0x1000 as an example, in its initial state, the SSMT record is empty. When the first snapshot is generated, the system creates the first status record node, and the node data structure includes SID = 001, SDU_PBA = 0x5000 (assuming the allocated physical address), and status bit = 1. When this DU is modified again in subsequent snapshot cycles, the system will create a new status record node and append it to the end of the linked list. For example, when SID = 002, a new node is generated, with SDU_PBA = 0x6000 and status bit = 1. This linked list structure enables the system to trace the historical versions of the same source DU in different snapshot cycles. At the same time, through the time series marking of the status bits, it can accurately reflect the occupancy of each snapshot volume on this DU. It should be noted that when the source DU is not modified in a certain snapshot cycle, the system only sets the corresponding status bit to 0 in the SSMT record of this snapshot, without the need to allocate physical storage space. This design significantly reduces the metadata storage overhead.
[0090] When the snapshot data statistics signaling is triggered, the storage device executes the preset algorithm of the capacity calculation engine. This algorithm performs multi-dimensional statistics based on the status flag bits of the SSMT. The specific implementation process includes the following steps: First, the statistics engine receives the statistics request parameters, including the SID list of the target snapshot volume (supporting joint statistics of single or multiple snapshots), the statistics granularity (DU level or file object level), and the statistics mode (real-time calculation or offline batch processing). Second, the system traverses the status record linked lists of all source DUs in the SSMT and scans the status bits for each target SID. For example, when calculating the capacity of the snapshot volume with SID = 001, the algorithm will retrieve the SSMT linked lists of all source DUs and check whether the status bit corresponding to SID = 001 is 1. For the qualified record nodes, the system accumulates the standard capacity (i.e., 4KB) of the corresponding DUs. In the specific optimization implementation, the bitmap acceleration technology can be adopted: The system maintains an independent bitmap for each snapshot volume, where each bit corresponds to the status flag of a source DU. When statistics are needed, directly calculate the number of set bits in the bitmap, and then multiply by the DU standard capacity to obtain the accurate statistical value. This method reduces the time complexity from O(n) to O(1), significantly improving the statistical efficiency of large-scale storage systems.
[0091] This method demonstrates unique advantages in the scenario of multi-level snapshot superposition. Suppose the system creates three snapshot volumes with SID = 001, 002, and 003 successively, and the source DU 0x1000 is modified at SID = 001 and 003, but not touched at SID = 002. When the user requests to count the independent capacity of SID = 003, the algorithm locates the status bit of this DU at SID = 003 as 1 through the SSMT linked list, and the subsequent snapshots (none with a higher SID in this example) do not overwrite this record, so 4KB is included in the statistical result. If the user needs to count the aggregated capacity of SID = 001 - 003, the algorithm will check the status bits of the three snapshots respectively, and finds that both SID = 001 and 003 are 1, while SID = 002 is 0, so the statistical value is 8KB. This statistical mechanism accurate to the DU level completely solves the problem of duplicate capacity calculation caused by data block reuse in traditional snapshot technologies. In addition, in the snapshot deletion scenario, this method realizes precise recovery of storage space through cascaded updates of status tags. For example, when deleting the SID = 002 snapshot, the system traverses its bitmap, and for each bit set to 1, checks whether the status bit of a subsequent snapshot with a higher SID is 1. If a certain DU has a snapshot record only at SID = 002, the physical space of the corresponding SDU is released; if this DU also has a record at SID = 003, the SDU is retained to ensure data integrity. This recovery strategy aware of dependencies avoids the risk of data loss caused by broken snapshot chains in traditional methods.
[0092] In a specific implementation case, suppose an enterprise storage system uses 4KB DU division, and the source data volume contains 1 million DUs (about 4GB in capacity). When the user randomly modifies 10% of the DUs (i.e., 100,000) and generates a snapshot, the traditional COW mechanism will migrate the entire 4KB blocks of all modified DUs, resulting in a snapshot capacity of 400MB. However, the actual data change amount may be only 1KB modified for each DU, and the theoretical minimum snapshot capacity should be 100MB. This method can accurately count the actual number of covered DUs (100,000) through the status tags of SSMT and accurately give a capacity report of 400MB. When the system accumulates 100 snapshots after running for three months, the administrator can obtain the accurate capacity of each snapshot volume in real time through this method. For example, it is found that the SID = 057 snapshot occupies 85GB due to a large number of full-DU overwriting operations, while the SID = 092 snapshot only modifies a small number of DUs and occupies 3.2GB, so as to formulate a differentiated snapshot retention strategy.
[0093] The extended implementation of this method also includes multiple optimization sub-modules: at the storage medium layer, the Allocate-On-Write strategy is adopted, and the physical space of the SDU is actually allocated only when the status flag is set to 1, avoiding storage waste caused by pre-allocation; at the metadata cache layer, a cache pool for the hot area of the SSMT is established, and the frequently modified source DU status records are stored in the high-speed storage medium to improve the statistical response speed; in the distributed storage scenario, a cross-node status flag synchronization protocol is designed to ensure the global consistency of snapshot capacity statistics in the cluster environment.
[0094] This technical solution constructs a full-link measurement system from data change capture, status flag update to accurate capacity calculation. Its innovation is reflected in three aspects: First, it breaks through the limitation of traditional snapshot technology with physical block migration as the statistical benchmark and establishes a direct mapping relationship between logical data units and snapshot status; Second, it invents a spatio-temporal encoding method for status flag bits, which not only records the storage occupancy of the current snapshot but also retains the associated information of historical versions; Finally, it designs a distributed statistical algorithm for massive snapshots, and realizes the non-intrusive real-time measurement of the storage system through bitmap acceleration and cache optimization. The organic combination of these technical points makes the snapshot capacity management of storage devices change from experience-driven to data-driven, providing a reliable infrastructure support for the resource scheduling of intelligent storage systems.
[0095] In one implementation, this specification provides an electronic device, including a processor and a readable storage medium, where the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the aforementioned snapshot data statistical method. In terms of the hardware level, the schematic diagram of the hardware architecture can be seen Figure 3 as shown.
[0096] In one implementation, this specification provides a readable storage medium, where the readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the aforementioned snapshot data statistical method.
[0097] Here, the readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0098] The systems, devices, modules, or units described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, laptop computer, cellular phone, camera phone, smart phone, personal digital assistant, media player, navigation device, email transceiver device, game console, tablet computer, wearable device, or a combination of any several of these devices.
[0099] For the convenience of description, when describing the above devices, they are described as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0100] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, system, or computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0101] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0102] Moreover, these computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.
[0104] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (which may include, but are not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0105] The above is only the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification should be included within the scope of the claims of this specification.
Claims
1. A snapshot data statistics method, characterized in that, Applied to a storage device, the method includes: In response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, record the corresponding status information of the snapshot volume in the snapshot status mapping table; The snapshot volume is generated based on a change in source data. The snapshot volume records at least one snapshot data unit, and the data of the snapshot data unit is the data of the corresponding data unit in the source data before the change of the source data this time; The snapshot status mapping table corresponds to the source data and records the association relationship between the snapshot volume and each data unit of the source data. The status information is used to reflect the existence status of the data unit of the source data having a corresponding snapshot data unit in the snapshot volume; In response to a snapshot data statistics signaling, according to the existence status reflected by the status information, obtain the storage capacity information occupied by the corresponding snapshot volume with a preset algorithm. The preset algorithm includes accumulating the storage capacity occupied by each data unit of the source data having a corresponding snapshot data unit in the snapshot volume to be statistically analyzed.
2. The method according to claim 1, wherein: The step of, in response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, record the corresponding status information of the snapshot volume in the snapshot status mapping table includes: According to the snapshot data units included in the snapshot volume, for the status information associated with the snapshot volume in the snapshot status mapping table, record a status value for each data unit of the source data. If there is a corresponding snapshot data unit, the status value is recorded as 1, and the remaining status values are recorded as 0; The step of, in response to a snapshot data statistics signaling, according to the existence status reflected by the status information, obtain the storage capacity information occupied by the corresponding snapshot volume with a preset algorithm includes: According to the status information of the snapshot volume in the snapshot status mapping table, multiply the storage capacity value of each data unit of the source data by the status value corresponding to each data unit for the snapshot volume respectively, accumulate the products, and use the accumulated result as the storage capacity information occupied by the snapshot volume.
3. The method according to claim 1, characterized in that, The snapshot status mapping table includes total status information, and the total status information includes a count value with an initial value of 0 recorded for each data unit of the source data; In response to an event of generating a snapshot volume, according to the snapshot data units included in the snapshot volume, increment the count value of the corresponding data unit of the source data in the total status information by 1; In response to a snapshot data total statistics signaling, according to the total status information of the snapshot volume in the snapshot status mapping table, multiply the storage capacity value of each data unit of the source data by the count value recorded for each data unit respectively, accumulate the products, and use the accumulated result as the total storage capacity information of all snapshot volumes corresponding to the source data.
4. The method according to claim 1, wherein The step of, in response to a snapshot data statistics signaling, according to the existence status reflected by the status information, obtain the storage capacity information occupied by the corresponding snapshot volume with a preset algorithm includes: In response to a snapshot data statistics signaling triggered by an event of generating a snapshot volume, obtain and record the storage capacity information of the generated snapshot volume.
5. The method according to claim 1, wherein The step of, in response to a snapshot data statistics signaling, according to the existence status reflected by the status information, obtain the storage capacity information occupied by the corresponding snapshot volume with a preset algorithm includes: In response to the snapshot data statistics signaling issued by the user, parse the snapshot volume information to be statistically analyzed indicated by the snapshot data statistics signaling, obtain the storage capacity information occupied by the corresponding snapshot volume according to the snapshot volume information to be statistically analyzed, and feed it back.
6. The method according to claim 1, characterized in that The snapshot status mapping table includes a snapshot sequence number, which is used to identify and distinguish each separately generated and independently stored snapshot volume.
7. The method according to claim 1, wherein The storage device records several snapshot status mapping tables respectively associated with different source data; In response to the event of snapshot volume generation, according to the snapshot data units included in the snapshot volume, record the corresponding status information of the snapshot volume in the snapshot status mapping table, including: In response to the event of snapshot volume generation, according to the snapshot data units included in the snapshot volume and the source data associated with the snapshot volume, record the corresponding status information of the snapshot volume in the snapshot status mapping table associated with the source data.
8. A snapshot data statistics device, characterized in that, Applied to a storage device, the apparatus includes: A first module, configured to, in response to the event of snapshot volume generation, record the corresponding status information of the snapshot volume in the snapshot status mapping table according to the snapshot data units included in the snapshot volume; The snapshot volume is generated based on a change in the source data, and the snapshot volume records at least one snapshot data unit, and the data of the snapshot data unit is the data of the corresponding data unit in the source data before the change of the source data this time; The snapshot status mapping table corresponds to the source data and records the association relationship between the snapshot volume and each data unit of the source data. The status information is used to reflect the existence status of the data unit of the source data having a corresponding snapshot data unit in the snapshot volume; A second module, configured to, in response to the snapshot data statistics signaling, obtain the storage capacity information occupied by the corresponding snapshot volume according to the existence status reflected by the status information, where the preset algorithm includes accumulating the storage capacity occupied by each data unit of the source data having a corresponding snapshot data unit in the snapshot volume to be statistically analyzed.
9. The apparatus according to claim 8, wherein In response to the event of snapshot volume generation, according to the snapshot data units included in the snapshot volume, record the corresponding status information of the snapshot volume in the snapshot status mapping table, including: According to the snapshot data units included in the snapshot volume, the status information associated with the snapshot volume in the snapshot status mapping table respectively records status values corresponding to each data unit of the source data. If there is a corresponding snapshot data unit, the status value is recorded as 1, and the remaining status values are recorded as 0; In response to the snapshot data statistics signaling, according to the existence status reflected by the status information, obtain the storage capacity information occupied by the corresponding snapshot volume, including: According to the status information of the snapshot volume in the snapshot status mapping table, multiply the storage capacity value of each data unit of the source data by the status value corresponding to each data unit for the snapshot volume respectively, accumulate the products, and use the accumulated result as the storage capacity information occupied by the snapshot volume.
10. The device according to claim 8, characterized in that, The snapshot status mapping table includes total status information, and the total status information includes count values with an initial value of 0 respectively recorded for each data unit of the source data; The first module is further configured to, in response to an event of generating a snapshot volume, increment by 1 the count value of the data unit of the corresponding source data in the total status information according to the snapshot data units included in the snapshot volume; The second module is further configured to, in response to a snapshot data total signaling, according to the total status information of the snapshot volume in the snapshot status mapping table, multiply the storage capacity value occupied by each data unit of the source data by the count value respectively recorded by each data unit, accumulate each product, and use the accumulated result as the total storage capacity information of all snapshot volumes corresponding to the source data.
11. The device according to claim 8, characterized in that, The obtaining, in response to a snapshot data statistics signaling, of the storage capacity information occupied by a corresponding snapshot volume according to the existence status reflected by the status information includes: Obtaining and recording the storage capacity information of the generated snapshot volume in response to a snapshot data statistics signaling triggered by an event of generating a snapshot volume.
12. The device according to claim 8, characterized in that The obtaining, in response to a snapshot data statistics signaling, of the storage capacity information occupied by a corresponding snapshot volume according to the existence status reflected by the status information includes: In response to a snapshot data statistics signaling issued by a user, parsing the information of the snapshot volume to be statistically analyzed indicated by the snapshot data statistics signaling, obtaining the storage capacity information occupied by the corresponding snapshot volume according to the information of the snapshot volume to be statistically analyzed, and feeding it back.
13. The device according to claim 8, characterized in that, The snapshot status mapping table includes a snapshot serial number, and the snapshot serial number is used to identify and distinguish each separately generated and independently stored snapshot volume.
14. The device according to claim 8, characterized in that, The storage device records a plurality of snapshot status mapping tables respectively associated with different source data; The recording, in response to an event of generating a snapshot volume, of the corresponding status information of the snapshot volume in the snapshot status mapping table according to the snapshot data units included in the snapshot volume includes: The recording, in response to an event of generating a snapshot volume, of the corresponding status information of the snapshot volume in the snapshot status mapping table associated with the source data according to the snapshot data units included in the snapshot volume and the source data associated with the snapshot volume.
15. An electronic device, characterized in that, Including: A processor and a readable storage medium, where the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1-7.
16. A readable storage medium, characterized in that, The readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the method according to any one of claims 1-7.
Citation Information
Cited By
Data protection method, memory, electronic equipment and storage medium
CN121680752A