Efficient retention locking of large namespace copies on deduplicated file systems
By generating flat structure horizontal files on the backup storage device and applying retention locks, the problem of high computing resources consumption in large namespace backups is solved, and efficient backup and incremental backup operations are achieved.
Patent Information
- Application Number
- CN202510083703.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-25
AI Technical Summary
When the prior art backups of large namespaces, it is necessary to copy and apply the locking operation, resulting in high consumption of computing resources and pause of applications, making it difficult to efficiently manage the locking operation of a large number of files.
By generating working frozen copies of the active namespace and copying them to the backup storage device, a flat structure of horizontal files is formed, and the application retains locking is reduced, reducing computing resource consumption.
Efficient initial backup and frequent incremental backups are realized, reducing the demand for computing resources, and reducing the time and resource consumption of backup operations.
Smart Images

Figure CN120371555A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the data backup process. More specifically, at least some embodiments of the present invention relate to systems, hardware, software, computer-readable media, and methods for reserving and locking large namespaces on backup storage devices. Background Art
[0002] Consider an application with a large namespace that may have millions of files that need to be protected with reservation locks. Since these files are modifiable, the system must create copies of these files and reserve and lock these files. One way to perform reservation locking is to copy one version of these files to another namespace and reserve and lock that namespace or all individual files. However, this process requires reading data, writing the data to another physical namespace, and enumerating the files at the target location for the same reservation locking, all of which may require a large amount of computing resources. In addition, the application must be paused while these files are being copied. In addition, reservation locking will involve enumerating the entire namespace. Therefore, a more efficient method for reserving and locking large file namespaces is needed. Summary of the Invention
[0003] Embodiments of the present invention generally relate to the data backup process. More specifically, at least some embodiments of the present invention relate to systems, hardware, software, computer-readable media, and methods for reserving and locking large namespaces on backup storage devices.
[0004] An example method includes generating a first timepoint copy of a first working frozen copy of an active namespace in a shared protection namespace, the first working frozen copy including a first file associated with the active namespace stacked in a first level file. Applying a reservation lock to the first level file. Generating a second timepoint copy of a second working frozen copy of the active namespace, the second working frozen copy including a second file associated with the active namespace stacked in a second level file. Applying a reservation lock to the second level file. The second level file is generated by copying files from the first level file into the second level file and copying files from the second working frozen copy of the active namespace into the second level file.
[0005] Accordingly, the embodiments disclosed herein advantageously provide a novel system and method that are capable of performing an initial backup in an efficient manner and applying a reservation lock to the initial backup, thereby reducing the amount of computing resources required. In addition, the novel system and method also allow for frequent incremental backups and apply a reservation lock to the frequent incremental backups in an initial manner. Embodiments of the novel system and method will now be explained. Brief Description of the Drawings
[0006] To describe the manner in which at least some of the advantages and features of the present invention can be obtained, embodiments of the present invention will be described in more detail with reference to specific embodiments of the present invention shown in the accompanying drawings. It is to be understood that these drawings only depict typical embodiments of the present invention and should not be regarded as limiting its scope, and the embodiments of the present invention will be described and explained with additional features and details by using the drawings.
[0007] Figures 1A - 1C Aspects of a backup computing system according to embodiments disclosed herein are disclosed;
[0008] Figures 2A - 2C Aspects of a backup computing system according to embodiments disclosed herein are disclosed;
[0009] Figures 3A - 3C Aspects of a backup computing system according to embodiments disclosed herein are disclosed;
[0010] Figure 4A and Figure 4B Aspects of a backup computing system according to embodiments disclosed herein are disclosed;
[0011] Figure 5 Aspects of a method according to embodiments disclosed herein are disclosed; and
[0012] Figure 6 Aspects of a computing device, system, or entity according to embodiments disclosed herein are disclosed. Detailed Description
[0013] Embodiments of the present invention, such as the examples disclosed herein, may be beneficial in various aspects. For example, and as will be apparent from the present invention, one or more embodiments of the present invention may provide one or more advantageous and unexpected effects in any combination, some examples of which are described below. It should be noted that these effects are neither intended nor should they be construed as limiting the scope of the claimed invention in any way. It should also be noted that nothing in this document should be construed as constituting a necessary or indispensable element of any invention or embodiment. Rather, the various aspects of the disclosed embodiments may be combined in various ways to define further embodiments. For example, any element(s) of any embodiment may be combined with any element(s) of any other embodiment to define a further embodiment. Such further embodiments are considered to be within the scope of the present invention. Similarly, any embodiment included within the scope of the present invention should not be construed as solving or being limited to solving any particular problem(s). Nor should any such embodiment be construed as achieving or being limited to achieving any particular technical effect(s) or solution(s). Finally, it is not required that any embodiment achieve any of the advantageous and unexpected effects disclosed herein.
[0014] It should be noted that embodiments of the present invention, whether claimed or not, cannot be actually performed or otherwise carried out in the human mind. Accordingly, nothing in this document should be construed as teaching or suggesting that any aspect of any embodiment of the present invention can or will be actually performed or otherwise carried out in the human mind. Additionally, unless explicitly indicated otherwise herein, the disclosed methods, processes, and operations are contemplated to be implemented by a computing system that may include hardware and / or software. That is, these methods, processes, and operations are defined as computer-implemented.
[0015] The following is a discussion of various aspects of an example operating environment for various embodiments of the present invention. This discussion is not intended to limit the scope and applicability of the present invention in any way.
[0016] Generally, embodiments of the present invention may be implemented in conjunction with systems, software, and components that alone and / or together effectuate data protection operations and / or cause the implementation of data protection operations, which may include, but are not limited to, data replication operations, IO replication operations, data read / write / delete operations, data deduplication operations, data backup operations, data recovery operations, data cloning operations, data archiving operations, and disaster recovery operations. More generally, the scope of the present invention includes any operating environment in which the disclosed concepts may be useful.
[0017] At least some embodiments of the present invention support the implementation of the disclosed functionality in existing backup platforms, examples of which include Dell EMC NetWorker and Avamar platforms and related backup software, and storage environments, such as Dell EMC PowerProtect DataDomain storage environments. However, in general, the scope of the present invention is not limited to any particular data backup platform or data storage environment.
[0018] New and / or modified data collected and / or generated in connection with some embodiments may be stored in a data protection environment, which may take the form of a public or private cloud storage environment, an on-premises storage environment, and a hybrid storage environment including both public and private elements. Any of these example storage environments may be partially or fully virtualized. The storage environment may include or consist of a data center operable to serve read, write, delete, backup, restore, and / or clone operations initiated by one or more clients or other elements of the operating environment. In cases where the backup includes multiple sets of data with respective different characteristics, the data may be assigned and stored to respective different targets in the storage environment, where each target corresponds to a set of data having one or more specific characteristics.
[0019] Example cloud computing environments may or may not be public and include storage environments that may provide data protection functionality for one or more clients. Another example of a cloud computing environment is an environment that may perform processing, data protection, and other services for one or more clients. Some example cloud computing environments that may be used in connection with embodiments of the present invention include, but are not limited to, Microsoft Azure, Amazon AWS, Dell EMC cloud storage services, and Google Cloud. However, more generally, the scope of the present invention is not limited to the use of any particular type or implementation of cloud computing environment.
[0020] In addition to cloud environments, the operating environment may also include one or more clients capable of collecting, modifying, and creating data. Thus, a particular client may use one or more instances of each of one or more applications, or otherwise be associated therewith, which perform such operations on the data. Such clients may include physical machines or virtual machines (VMs).
[0021] In particular, devices in the operating environment can take the form of software, physical machines, or virtual machines (VMs), or any combination thereof, although no particular device implementation or configuration is required for any implementation. Similarly, components of the data protection system (such as databases, storage servers, storage volumes (LUNs), storage disks, replication services, backup servers, recovery servers, backup clients, and recovery clients) can likewise take the form of software, physical machines, or virtual machines (VMs), although no particular component implementation is required for any implementation. In the case of using VMs, a hypervisor or other virtual machine monitor (VMM) can be used to create and control the VMs. The term virtual machine (VM) includes, but is not limited to, any virtualization, emulation, or other representation of one or more computing system elements (such as computing system hardware). A VM can be based on one or more computer architectures and provide the functionality of a physical computer. The implementation of a VM can include or at least involve the use of hardware and / or software. The image of a VM can, for example, take the form of a.VMX file and one or more.VMDK files (VM hard disks).
[0022] As used herein, the term "data" is intended to have a broad scope. Thus, by way of example and not limitation, the term includes, for example, data segments, data chunks, data blocks, atomic data, emails, any type of object, any type of file (including media files, word processing files, spreadsheet files, and database files), as well as contact information, directories, subdirectories, volumes, and any combination of one or more of the foregoing.
[0023] Example embodiments of the present invention are applicable to any system capable of storing and processing various types of objects in analog, digital, or other forms. Although these terms may be used by way of example as documents, files, segments, blocks, or objects, the principles of the present invention are not limited to any particular form of representing and storing data or other information. Instead, these principles apply equally to any object capable of representing information.
[0024] As used herein, the term "backup" is intended to have a broad scope. Thus, example backups that can be used in conjunction with embodiments of the present invention include, but are not limited to, full backups, partial backups, clones, snapshots, and incremental or differential backups.
[0025] Now, with particular attention Figure 1A , an embodiment of the backup computing system 100 will be described. As Figure 1AAs shown, backup computing system 100 includes main memory 110 and backup memory 120 for backing up the large file namespace 130 for backup activities. The large file namespace 130 (also referred to herein as "large namespace 130") is an active production system that may have millions of files required by applications executing on the system. The large namespace 130 is hosted on the main memory 110, which is typically a high-performance main storage system. To protect the files in the large namespace 130, a frozen copy 155 of the large namespace 30 can be hosted on the same main memory 110 or a similar main storage system, although this is not shown in Figure 1A . These frozen copies 155 are point-in-time snapshots of the active large namespace 130.
[0026] In operation, backup control system 140 copies the frozen copies 155 of the large namespace 130 into the shared protection namespace 150 hosted on the backup memory 120. That is, the system uses the backup memory 120 as the namespace (i.e., the shared protection namespace 150) where the frozen copies 155 can be directly stored. In the event of a disaster or data corruption, the copies hosted in the shared protection namespace 150 can be restored back to the main memory 110. The backup and recovery control paths are coordinated by the backup control system 140. In some embodiments, the frozen copies 155 can be incremental in nature, so the backup memory 120 can include an aggregation of different frozen copies. Once the files are no longer needed, i.e., the files associated with the frozen copies 155 are no longer needed, the frozen copies 155 can be deleted from the shared protection namespace 150.
[0027] Part of the responsibility of the backup memory 120 is to retain the frozen copies 155 for a certain predetermined expiration period. One way to retain the frozen copies 155 is to create a point-in-time (PIT) copy of the shared protection namespace 150. For example, as Figure 1B shown, a PIT copy 156 of the shared protection namespace 150 is created. The PIT copy 156 is a snapshot of the shared protection namespace 150 that includes a set of working copies stored in the shared protection namespace 150 at the time the PIT copy 156 is created. Then, when the predetermined expiration period expires, the PIT copy 156 can be deleted from the shared protection namespace 150.
[0028] Copy-on-write protected namespace 150 requires reading and writing files, including both metadata and data. In some backup memories 120, an advanced file system such as Data Domain Filesystem (DDFS) is implemented. Such a file system has special methods such as fastcopy, which creates a copy of the copy-on-write protected namespace 150 by only copying the metadata of the file to a new inode.
[0029] Therefore, creating such a copy of the copy-on-write protected namespace 150 can be expensive. Even using a super-efficient method such as fastcopy, creating a copy still takes a finite amount of time. In addition, for a large copy-on-write protected namespace 150, creating a copy is very difficult for the file system. For example, if there are one million files in the working freeze set and the protection memory hosts 14 copies, there are 14 million files in the system.
[0030] Figure 1C An embodiment in which the computing system 100 is running is shown. The figure depicts an application view, where there is a real-time data set 160 and a freeze-time data set 170 including snapshots taken at different time points at any given point in time. In this embodiment, the freeze-time data set 170 retains the last three snapshots at any given point in time. Thus, at time T4, the active state is captured and the active state at time T1 is discarded, such that the freeze set at time T4 includes the freeze sets at times T2, T3, and T4. Figure 1C Also shown are the freeze-time snapshots of four generations 180 of the freeze-time data set 170. Thus, at time T7, there are four freeze-time snapshots, the freeze-time data set 170, and the active data set 160.
[0031] In Figure 1C it is assumed that there is a namespace with one million files and three copies of the namespace (e.g., 160, 170, and at least a portion of 180). If a moderate change rate of 5% is assumed, this means that each of the three copies increases by 50,000 files. This means that in the working freeze set with three time-point sets, there will be 1.1 million files. If there are freeze-time snapshots of four generations 180, there will be 4.4 million files to be protected.
[0032] Therefore, Figure 1CDuplicate spread may be shown. For large namespaces, this may create significant namespace management issues. In particular, in some cases, it is desirable to apply a hold lock to each file in the namespace. A hold lock is a mechanism that ensures that each PIT copy of the shared protection namespace 150 cannot be changed in any way during the duration of the hold lock. In some embodiments, the backup memory 120 includes a hold lock (RL) engine 125 capable of applying a hold lock to each individual file in the PIT copy. However, when the PIT copy includes millions of individual files, applying a hold lock to each individual file may consume a large amount of computing system resources and may also take unnecessary time.
[0033] Figures 2A - 2C A backup computing system 200 is shown, which may be an embodiment of the backup computing system 100. Thus, the backup computing system 200 includes all of the elements previously described with respect to the backup computing system 100, and not all of these elements need to be shown or described with respect to Figures 2A - 2C shown or described.
[0034] Figure 2A A backup operation is shown. During the backup operation, the backup control system 140 copies a frozen working copy 220 of the large namespace 130 into the shared protection namespace 210 of the backup memory 120, which may correspond to the shared protection namespace 150. As shown, the frozen working copy 220 includes six individual files: File A, File B, File C, File D, File E, and File F. It should be understood that including only six files in the frozen working copy 220 is for ease of explanation only, and the frozen working copy 220 may include any number of additional files, typically millions of files.
[0035] The backup control system 140 also instructs the shared protection namespace 210 to generate a PIT copy 230 of the frozen working copy 220. As previously mentioned, applying a hold lock to each individual file of the frozen working copy 220 is burdensome for the backup computing system 200, especially when the total number of files reaches millions. Therefore, while generating the PIT copy 230, the backup control system 140 instructs the shared protection namespace 210 to stack each individual file in the frozen working copy 220 into at least one horizontal file with a flat structure. As Figure 2AAs shown, individual files A - F are extracted from the working freeze copy 220 and stacked into a horizontal file 232. In one embodiment, the backup memory 120 is capable of implementing a virtual synthesis process that generates the horizontal file 232 as a VS file, which will be a synthetic complete copy of the individual files A - F in the working freeze copy 220. The advantage of using the virtual synthesis process is that the horizontal file 232 does not involve data reading. Instead, this can be considered a scatter - gather method of creating a new object using pointers from other objects, so it is more efficient than traditional tar / zip methods.
[0036] However, it should be noted that the horizontal file 232 does not need to be a single file and can instead include multiple files. Generally, the concept of stacking the individual files of the working freeze copy 220 into a horizontal file means going from a large number of files to a much smaller number of files. Thus, in Figure 2A the example, six individual files are stacked into a single horizontal file 232. However, if the working freeze copy 220 includes millions of files, then these millions of files may be stacked into a few hundred or fewer horizontal files. Thus, in the embodiments and claims disclosed herein, the individual files of the working freeze copy are stacked into at least one horizontal file, but can be stacked into more than one horizontal file as long as the total number is only a small fraction of the total individual files.
[0037] Figure 2B An embodiment 260 of the horizontal file 232 is shown. As shown, in embodiment 260, the horizontal file 232 is generated by alternating the metadata and data of each individual file A - F. This results in a single horizontal file. Metadata can describe file permissions, size, and directory structure, as well as its permissions. By reading the metadata, the corresponding directory structure can be restored.
[0038] Figure 2B An embodiment 270 is also shown when the individual files A - F of the working freeze copy 220 are stacked into more than one horizontal file, where the number of these more horizontal files is still less than the number of individual files. As shown, the metadata of each individual file A - F is stacked in the horizontal file 272, and the data of each individual file A - F is stacked in the horizontal file 274. Thus, embodiments 260 and 270 show that there can be any number of alternative ways to generate a horizontal file. It should be noted that the discussion of generating the horizontal file 232 using the virtual synthesis process and the discussion of embodiments 260 and 270 will also apply to any horizontal file discussed with respect to other embodiments discussed in more detail below.
[0039] Once the horizontal file 232 has been generated, the reservation lock engine 125 of the backup memory 120 is able to apply a reservation lock to the horizontal file 232. As shown, the PIT copy 230 includes the horizontally file 232 that has been reservation-locked. Since the horizontal file 232 is a single file (or a small number of files as described above), the amount of computing resources and time required to apply a reservation lock to the horizontal file 232 is greatly reduced compared to applying a reservation lock to each individual file.
[0040] Figure 2C A subsequent backup operation is shown after changes have been made to the large namespace 130 during subsequent operations of an application using the main memory 110. During the subsequent backup operation, the backup control system 140 copies a frozen working copy 240 of the large namespace 130 into the shared protected namespace 210. As shown, the frozen working copy 240 includes seven individual files: File A, File B’, File C, File D, File E, File F, and File G. Thus, the changes made to the large namespace 130 during the subsequent operation include updating File B to File B’ and adding File G, while no changes were made to Files A, C, D, E, and F.
[0041] The backup control system 140 instructs the shared protected namespace 210 to generate a PIT copy 250 of the frozen working copy 240, and while generating the PIT copy 250, stacks each individual file in the frozen working copy 240 into at least one horizontal file having a flat structure. As Figure 2C shown, the individual files A, B’, C, D, E, F, and G are extracted from the frozen working copy 240 and stacked into the horizontal file 252 generated in the manner described previously.
[0042] Once the horizontal file 252 has been generated, the reservation lock engine 125 of the backup memory 120 is able to apply a reservation lock to the horizontal file 252. As shown, the PIT copy 250 includes the horizontally file 252 that has been reservation-locked. Since the horizontal file 252 is a single file (or a small number of files as described above), the amount of computing resources and time required to apply a reservation lock to the horizontal file 252 is greatly reduced compared to applying a reservation lock to each individual file.
[0043] For Figure 2C all backup operations after the backup operation described in, this process will be repeated. That is, the frozen working copy will be copied into the shared protected namespace 210. Then, when generating the PIT copy of the frozen working copy, at least one horizontal file will be generated for the individual files of the frozen working copy. A reservation lock will be applied to the at least one horizontal file.
[0044] Figures 3A - 3CIllustrated is a backup computing system 300, which may be an embodiment of backup computing systems 100 and 200. Thus, backup computing system 300 includes all of the elements previously described with respect to backup computing systems 100 and 200, and these elements need not all be Figures 3A - 3C shown or described.
[0045] Figure 3A Illustrated is a backup operation. During the backup operation, backup control system 140 copies a frozen working copy 320 of large namespace 130 into a shared protected namespace 310 of backup memory 120, which shared protected namespace 310 may correspond to shared protected namespace 150. As shown, frozen working copy 320 includes six individual files: File A, File B, File C, File D, File E, and File F. It should be understood that including only six files in frozen working copy 320 is for ease of explanation only, and frozen working copy 320 may include any number of additional files, typically on the order of millions of files.
[0046] Backup control system 140 instructs shared protected namespace 310 to generate a PIT copy 330 of frozen working copy 320, and while generating PIT copy 330, stacks each individual file in frozen working copy 320 into at least one horizontal file having a flat structure. As Figure 3A shown, individual files A - F are extracted from frozen working copy 320 and stacked into horizontal file 332 generated in the manner previously described.
[0047] Once horizontal file 332 has been generated, retention lock engine 125 of backup memory 120 is able to apply a retention lock to horizontal file 332. As shown, PIT copy 330 includes horizontal file 332 that has been retention - locked. Since horizontal file 332 is a single file (or a small number of files as previously described), the amount of computing resources and time required to apply a retention lock to horizontal file 332 is greatly reduced compared to applying a retention lock to each individual file.
[0048] Figure 3B Illustrated is a subsequent backup operation after changes have been made to large namespace 130 during subsequent operations of an application utilizing main memory 110. During the subsequent backup operation, backup control system 140 copies a frozen working copy 340 of large namespace 130 into shared protected namespace 310. As shown, frozen working copy 340 includes seven individual files: File A, File B', File C, File D, File E, File F, and File G. Thus, the changes made to large namespace 130 during the subsequent operation include updating File B to File B' and adding File G, while no changes were made to Files A, C, D, E, and F.
[0049] The backup control system 140 instructs the shared protection namespace 310 to generate a point-in-time (PIT) copy 350 of the working freeze copy 340, and while generating the PIT copy 350, stacks each individual file in the working freeze copy 340 into at least one horizontal file having a flat structure. In Figure 3B an implementation, to generate at least one horizontal file, the shared protection namespace 310 performs a metadata difference operation (also referred to as a "Diff" operation) 360, which will now be explained.
[0050] The purpose of the metadata difference operation 360 is to identify the subset of files included in both the PIT copy 330 and the working freeze copy 340, and to identify the subset of files included only in the working freeze copy 340. However, since the horizontal file 332 of the PIT copy 330 is a single file (or a small number of files) - which is the purpose of generating a horizontal file, and the working freeze copy 340 includes multiple individual files (seven in the present invention), it is not possible to directly compare the horizontal file 332 and the working freeze copy 340.
[0051] As previously mentioned, the individual files A - F of the working freeze copy 320 include both metadata files and data files. During the generation of the horizontal file 332, a metadata index file 126 is also created in the backup memory 120 in which the metadata of the individual files A - F is stored. Then, the metadata index file 126 is used in the metadata difference operation 360. Since the number of files in the index file 126 and the working freeze copy 340 is closer, it is possible to compare the metadata in both.
[0052] During the metadata difference operation 360, the shared namespace 310 traverses and compares the metadata of the individual files A - F stored in the metadata index file 126 with the metadata of the individual files A, B’, C - F, and G of the working freeze copy 340. Thus, when comparing the metadata in the metadata index file 126 with the metadata of the files in the working freeze copy 340, as shown at 370, it is found that files A, C, D, E, and F are included in both the PIT copy 330 and the working freeze copy 330, while files B’ and G are new and are therefore included only in the working freeze copy 340. In other words, the presence of the same metadata for files A, C, D, E, and F in both the metadata index file 126 and the files in the working freeze copy 340 indicates that these files are included in both the PIT copy 330 and the working freeze copy 340. Since the metadata index file 126 lacks the metadata for files B’ and G in the working freeze copy 340, it is inferred that these are new / updated files that do not exist in the PIT copy 330.
[0053] Then, the shared protection namespace 310 stacks the individual files A, B', C, D, E, F, and G into a horizontal file 352. As Figure 3B shown, the files A, C, D, E, and F included in both the PIT copy 330 and the working freeze copy 340 are extracted from the horizontal file 332 into the horizontal file 352. In some embodiments, the metadata index file 126 is used to determine which files in the horizontal file 332 are to be extracted into the horizontal file 352 in embodiments where the backup system cannot directly determine the content of the horizontal file 332. Only the files B' and G that are included only in the working freeze copy 340 are extracted from the working freeze copy 340 into the horizontal file 352. Since only the changed files are extracted from the working freeze copy 340, the number of files that need to be extracted from the working freeze copy 340 is greatly reduced.
[0054] In addition, extracting the files A, C, D, E, and F included in both the PIT copy 330 and the working freeze copy 340 from the horizontal file 332 into the horizontal file 352 can save the number of operations required to extract these files. For example, when the metadata difference operation 360 determines that the file A is to be extracted from the horizontal file 332 into the horizontal file 352, the operation to extract the file A is not automatically performed. Instead, the operation to extract the file A is put on hold. Next, when the metadata difference operation 360 determines that the file B' is to be extracted from the working freeze copy 340 into the horizontal file 352, it is determined that the file B' comes from a different location than the file A, so the operation to extract the file A from the horizontal file 332 into the horizontal file 352 is performed. However, the operation to extract the file B' is not performed, but the operation to extract the file B' is put on hold.
[0055] Next, when the metadata difference operation 360 determines that the file C is to be extracted from the horizontal file 332 into the horizontal file 352, it is determined that the file B' comes from a different location than the file C, so the operation to extract the file B' from the working freeze copy 340 into the horizontal file 352 is performed. However, the operation to extract the file C is not performed, but the operation to extract the file C is put on hold.
[0056] Next, when the metadata difference operation 360 determines that the file D is to be extracted from the horizontal file 332 into the horizontal file 352, it is determined that the file D is adjacent to the file C and thus comes from the same location as the file C. Therefore, the operation to extract the file D is not performed, but the operation to extract the file D is put on hold.
[0057] The metadata difference operation 360 then determines that the files E and F are also to be extracted from the horizontal file 332 into the horizontal file 352. It is also determined that the files E and F are adjacent to the files C and D. Therefore, the operations to extract the files E and F are not performed, but the operations to extract the files E and F are put on hold.
[0058] Finally, when the metadata difference operation 360 determines that file G is to be extracted from the working frozen copy 340 to the horizontal file 352, it is determined that files C - F are from a different location than file G, so an operation to extract files C - F from the horizontal file 332 to the horizontal file 352 is performed. However, the operation to extract file G is not performed and is instead shelved. Since file G is the last file, the operation to extract file G from the working frozen copy 340 to the horizontal file 352 will be performed.
[0059] The above process shows that the embodiments disclosed herein reduce the number of operations required to extract files to the horizontal file 352. For example, in this embodiment, seven files are extracted to the horizontal file 352. However, only four operations, instead of seven operations, are required to extract all seven files. This is because files C - F are adjacent. That is, since only two files, namely file B' and G, have changed, there are several adjacent files that have not changed. As discussed, only one operation is required for adjacent unchanged files. Thus, in the embodiments disclosed herein, the number of operations required to extract files from the horizontal file 332 and the frozen working copy 340 to the horizontal file 352 is at most 2*N, where N is the number of changed files.
[0060] Although for ease of explanation the above process is shown as using a small number of files, in actual operation, millions of files will be used. For embodiments with frequent backups, there will be a large number of adjacent files that do not change. This will greatly reduce the number of operations required to extract files from the horizontal file 332 to the horizontal file 352. For example, assume there are one million files, but only 1000 files have changed since the last backup. In this case, at most only 2000 operations (i.e., 2*N, where N is 1000 changed files) are required to extract files from the horizontal file 332 to the horizontal file 352.
[0061] Figure 3C An embodiment 380 of generating the horizontal file 352 in the described manner is shown. As Figure 3CAs shown, the horizontal file 332 is a file with alternating metadata and data of files A - F. Additionally, for illustrative purposes only, the files in the working freeze copy 340 are also shown as having alternating metadata and data. As further shown in the figure, the metadata and data of files A and F are extracted from the horizontal file 332 to the horizontal file 352. Although not shown, the ellipsis indicates that the metadata and data of files C - E are also extracted from the horizontal file 332 to the horizontal file 352. In some embodiments, this can be achieved by extracting pointers to the metadata and data in the horizontal file 332 from the horizontal file 332 to the horizontal file 352. It should be noted that since file B is not included in both the horizontal file 332 and the working freeze copy 340, the metadata and data of file B are not extracted from the horizontal file 332 to the horizontal file 352. The metadata and data of file B' and file G are extracted from the working freeze copy 340 to the horizontal file 352 because these two files are only included in the working freeze copy 340.
[0062] Once the horizontal file 352 has been generated, the reservation lock engine 125 of the backup storage 120 can apply a reservation lock to the horizontal file 352. As shown, the PIT copy 350 includes the horizontal file 352 that has been reservation - locked. Since the horizontal file 352 is a single file (or a small number of files as described above), the amount of computing resources and time required to apply a reservation lock to the horizontal file 352 is greatly reduced compared to applying a reservation lock to each individual file.
[0063] For Figure 3B all backup operations after the backup operation described above, this process will be repeated. That is, the working freeze copy will be copied to the shared protection namespace 310. Then, at least one horizontal file will be generated for the multiple individual files of the working freeze copy, where the files included in the horizontal file of the latest PIT copy and the working freeze copy are extracted from the horizontal file of the latest PIT copy, and the files only included in the working freeze copy are extracted from the working freeze copy. A reservation lock will be applied to the at least one horizontal file.
[0064] When the PIT copy 330 and the PIT copy 350 are later copied to another backup storage server, Figures 3A - 3C the embodiments provide a non - restrictive advantage. Since most of the files in the horizontal file 352 of the PIT copy 350 are the same as the files in the horizontal file 332 of the PIT copy 330, this can be utilized to increase the copying speed because for the files that are the same as those in the horizontal file 332 of the PIT copy 330, only the data in the horizontal file 322 of the PIT copy 300 needs to be copied because pointers to these files can be included in the horizontal file 352 of the PIT copy 350.
[0065] Figure 4A and Figure 4B illustrates a backup computing system 400, which may be an implementation of backup computing systems 100, 200, and 300. Accordingly, backup computing system 400 includes all of the elements previously described with respect to backup computing systems 100, 200, and 300, and these elements need not all be shown or described with respect to Figure 4A and Figure 4B shown or described.
[0066] Figure 4A illustrates a backup operation. During the backup operation, backup control system 140 copies a frozen working copy 420 of large namespace 130 into a shared protected namespace 410 of backup memory 120, which may correspond to shared protected namespace 150. As shown, frozen working copy 420 includes six individual files: File A, File B, File C, File D, File E, and File F. It should be understood that including only six files in frozen working copy 420 is for ease of explanation only, and frozen working copy 420 may include any number of additional files, typically on the order of millions of files.
[0067] It should be understood that while the backup operation is in progress, changes made to large namespace 130 cannot be written to frozen working copy 420. In other words, the system cannot update frozen working copy 420 until a PIT copy 440 is generated. This may cause application write stalls. As explained below, Figure 4A and Figure 4B implementations of [and] advantageously address this issue by generating snapshots 430 and 460. Since snapshots 430 and 460 can be used in the backup operation as described above, frozen working copy 420 can be updated almost immediately because the backup operation no longer needs it, thereby lifting the stall on the frozen working copy portion of the shared protected namespace.
[0068] Accordingly, in this implementation, backup control system 140 instructs shared protected namespace 410 to take a snapshot 430 of shared protected namespace 410 when frozen working copy 420 is copied into shared protected namespace 410. Note that the frozen working copy included in snapshot 430 is labeled frozen working copy 420A. This is to indicate that the frozen working copy 420 in namespace 410 can now be freely changed as described above, and frozen working copy 420A is the version of frozen working copy 420 that existed when snapshot 430 was generated.
[0069] Then, the backup control system 140 instructs the shared protection namespace 410 to generate a point-in-time (PIT) copy 440 of the working frozen copy 420, and while generating the PIT copy 440, stacks each individual file in the working frozen copy 420 into at least one horizontal file with a flat structure. As Figure 4A shown, different from extracting files A - F from the working frozen copy 420 in the shared protection namespace 410 as in the Figure 3A embodiment, the individual files A - F are extracted from the working frozen copy 420A included in the snapshot 430. The individual files A - F are stacked into the horizontal file 442 generated in the previously described manner.
[0070] Once the horizontal file 442 has been generated, the reservation lock engine 125 of the backup memory 120 can apply a reservation lock to the horizontal file 442. As shown, the PIT copy 440 includes the horizontal file 442 that has been reservation - locked. Since the horizontal file 442 is a single file (or a small number of files as described above), the amount of computing resources and time required to apply a reservation lock to the horizontal file 442 is greatly reduced compared to applying a reservation lock to each individual file. Once the PIT copy 470 has been generated, the snapshot 430 can be released.
[0071] Figure 4B Illustrated is a subsequent backup operation after changes are made to the large namespace 130 during subsequent operations of an application using the main memory 110. During the subsequent backup operation, the backup control system 140 copies the working frozen copy 450 of the large namespace 130 into the shared protection namespace 410. As shown, the working frozen copy 450 includes seven individual files: File A, File B’, File C, File D, File E, File F, and File G. Thus, the changes made to the large namespace 130 during the subsequent operation include updating File B to File B’ and adding File G, while no changes are made to Files A, C, D, E, and F.
[0072] Then, the backup control system 140 instructs the shared protection namespace 410 to take a snapshot 460 of the shared protection namespace 410 when the working frozen copy 450 is copied into the shared protection namespace 410. Thus, as Figure 4B shown, the snapshot 460 includes the working frozen copy 450A and the PIT copy 440. It should be noted that the working frozen copy included in the snapshot 460 is labeled as the working frozen copy 450A. This is to indicate that the working frozen copy 450 in the namespace 410 can now be freely changed as described above, and the working frozen copy 450A is the version of the working frozen copy 450 that existed when the snapshot 460 was generated.
[0073] The backup control system 140 instructs the shared protection namespace 410 to generate a PIT copy 470 of the working freeze copy 450A, and while generating the PIT copy 470, stacks each individual file in the working freeze copy 450 into at least one horizontal file having a flat structure. In Figure 4B an embodiment, to generate the at least one horizontal file, the shared protection namespace 410 performs a metadata difference operation (also referred to as a "Diff" operation) 480, which will now be explained.
[0074] The purpose of the metadata difference operation 480 is to identify a subset of files that are included in both the PIT copy 440 and the working freeze copy 450A included in the snapshot 460, and to identify a subset of files that are only included in the working freeze copy 450A included in the snapshot 460. However, since the horizontal file 442 of the PIT copy 440 is a single file (or a small number of files) - which is the purpose of generating the horizontal file, and the working freeze copy 450A includes multiple individual files (seven in the present invention), a direct comparison between the horizontal file 442 and the working freeze copy 450A cannot be made.
[0075] As previously mentioned, the individual files A - F of the working freeze copy 420A include both metadata files and data files. During the generation of the horizontal file 442, a metadata index file 126 is also created in the backup memory 120 in which the metadata of the individual files A - F is stored. Then, the metadata index file 126 is used in the metadata difference operation 480. Since the number of files in the index file 126 and the working freeze copy 450A is closer, the metadata in both can be compared.
[0076] During the metadata differential operation 480, the shared namespace 410 traverses and compares the metadata of individual files A - F stored in the metadata index file 126 with the metadata of individual files A, B’, C - F, and G of the working frozen copy 450A. Thus, when comparing the metadata in the metadata index file 126 with the metadata of the files in the working frozen copy 450A, as shown at 485, it is found that files A, C, D, E, and F are included in both the PIT copy 440 and the working frozen copy 450A included in the snapshot 460, while files B’ and G are new and are thus only included in the working frozen copy 450A included in the snapshot 460. In other words, the presence of the same metadata for files A, C, D, E, and F in both the metadata index file 126 and the files in the working frozen copy 450A indicates that these files are included in both the PIT copy 440 and the working frozen copy 450A. Since the metadata index file 126 lacks the metadata for files B’ and G in the working frozen copy 450A, it is inferred that these are new / updated files that do not exist in the PIT copy 440.
[0077] Then, the shared protected namespace 410 stacks the individual files A, B’, C, D, E, F, and G into a horizontal file 472. As Figure 4B shown, files A, C, D, E, and F that are included in both the PIT copy 440 and the working frozen copy 450A included in the snapshot 460 are extracted from the horizontal file 442 to the horizontal file 472. In some embodiments, the metadata index file 126 is used to determine the files in the horizontal file 442 to be extracted to the horizontal file 472 in embodiments where the backup system cannot directly determine the content of the horizontal file 442. Only files B’ and G that are included only in the working frozen copy 450A included in the snapshot 460 are extracted from the working frozen copy 450A to the horizontal file 472. Since only the changed files are extracted from the working frozen copy 450A included in the snapshot 460, the number of files that need to be extracted from the working frozen copy 450A is greatly reduced.
[0078] In addition, extracting files A, C, D, E, and F, which are included in both the PIT copy 440 and the working freeze copy 450A, from the horizontal file 442 to the horizontal file 472 can save the number of operations required to extract the files. For example, when the metadata difference operation 480 determines that file A is to be extracted from the horizontal file 442 to the horizontal file 472, the operation of extracting file A is not automatically executed. Instead, the operation of extracting file A is put on hold. Next, when the metadata difference operation 480 determines that file B' is to be extracted from the working freeze copy 450A to the horizontal file 472, it is determined that file B' comes from a different location than file A, so the operation of extracting file A from the horizontal file 442 to the horizontal file 472 is executed. However, the operation of extracting file B' is not executed, but rather the operation of extracting file B' is put on hold.
[0079] Next, when the metadata difference operation 480 determines that file C is to be extracted from the horizontal file 442 to the horizontal file 472, it is determined that file B' comes from a different location than file C, so the operation of extracting file B' from the working freeze copy 450A to the horizontal file 472 is executed. However, the operation of extracting file C is not executed, but rather the operation of extracting file C is put on hold.
[0080] Next, when the metadata difference operation 480 determines that file D is to be extracted from the horizontal file 442 to the horizontal file 472, it is determined that file D is adjacent to file C and thus comes from the same location as file C. Therefore, the operation of extracting file D is not executed, but rather the operation of extracting file D is put on hold.
[0081] The metadata difference operation 480 will then determine that files E and F are also to be extracted from the horizontal file 442 to the horizontal file 472. It will also be determined that files E and F are adjacent to files C and D. Therefore, the operations of extracting files E and F are not executed, but rather the operations of extracting files E and F are put on hold.
[0082] Finally, when the metadata difference operation 480 determines that file G is to be extracted from the working freeze copy 450A to the horizontal file 472, it is determined that files C - F come from a different location than file G, so the operation of extracting files C - F from the horizontal file 442 to the horizontal file 472 is executed. However, the operation of extracting file G is not executed, but rather the operation of extracting file G is put on hold. Since file G is the last file, the operation of extracting file G from the working freeze copy 450A to the horizontal file 472 will be executed.
[0083] The above process shows that the embodiments disclosed herein reduce the number of operations required to extract files to the horizontal file 472. For example, in this embodiment, seven files are extracted to the horizontal file 472. However, only four operations, instead of seven operations, are required to extract all seven files. This is because files C-F are adjacent. That is, since only two files, namely file B' and G, have changed, there are several unchanged files that are adjacent to each other. As discussed, only one operation is required for unchanged files that are adjacent to each other. Therefore, in the embodiments disclosed herein, the number of operations required to extract files from the horizontal file 442 and the frozen working copy 450A to the horizontal file 472 is at most 2*N, where N is the number of changed files.
[0084] It should be noted that Figure 4B As shown, files A, C, D, E, and F that are included in both the PIT copy 440 and the working frozen copy 450A included in the snapshot 460 are all extracted from the horizontal file 442 of the PIT copy 440 included in the shared protected namespace 410. However, this does not have to be the case. As described above, the snapshot 460 includes a copy in the PIT copy 440 that includes the horizontal file 442. Therefore, as shown in 490, files A, C, D, E, and F that are included in both the PIT copy 440 and the working frozen copy 450A included in the snapshot 460 can alternatively be extracted from the horizontal file 442 in the snapshot 460 to the horizontal file 472.
[0085] Once the horizontal file 472 has been generated, the retention lock engine 125 of the backup memory 120 can apply a retention lock to the horizontal file 472. As shown, the PIT copy 470 includes the horizontal file 472 that has been retention-locked. Since the horizontal file 472 is a single file (or a small number of files as described above), the amount of computing resources and time required to apply a retention lock to the horizontal file 472 is greatly reduced compared to applying a retention lock to each individual file. Then, once the PIT copy 470 has been generated, the snapshot 460 can be released because the snapshot 460 is no longer needed. This in turn helps to reduce the resources of the backup memory 120 required.
[0086] For Figure 4BAll backup operations after the backup operation described in will repeat this process. That is, the working freeze copy will be copied to the shared protection namespace 410, and a snapshot of the working freeze copy will be taken. Then, at least one horizontal file will be generated for the individual files of the working freeze copy included in the snapshot, where the files included in the horizontal file of the latest PIT copy and the files included in the working freeze copy included in the snapshot are extracted from the horizontal file in the latest PIT copy, and the files only included in the working freeze copy included in the snapshot are extracted from the working freeze copy included in the snapshot. A retention lock will be applied to the at least one horizontal file. It should be understood that Figure 4A and Figure 4B The implementation of combines the non-limiting advantages of the previously discussed implementations with respect to Figures 3A - 3C discussed.
[0087] Regarding the disclosed methods, including Figure 5 the example methods of, it should be noted that any operation of any of these methods can be performed in response to, due to, and / or based on the execution of any (one or more) previous operations. Accordingly, for example, the execution of one or more operations can be predictive or triggering of the subsequent execution of one or more additional operations. Thus, for example, the various operations that can constitute a method can be related to each other or otherwise related to each other by the relationships of the examples just mentioned. Finally, although not necessary, in some implementations, the individual operations that constitute the various example methods disclosed herein are performed in the specific order described in these examples. In other implementations, the individual operations that constitute the disclosed methods can be performed in an order different from the described specific order.
[0088] Attention is now turned to Figure 5 , and example method 500 is disclosed. Method 500 will be described in conjunction with one or more of the previously described figures, although method 500 is not limited to any particular implementation.
[0089] Method 500 includes generating a first point-in-time copy of the active namespace in the shared protection namespace of a backup storage device, the first working freeze copy including a first plurality of files associated with the active namespace, the first plurality of files being stacked in at least one first horizontal file in the first point-in-time copy (510). For example, as previously described, a first PIT copy 230, 330, or 440 is generated in the shared protection namespace 210, 310, or 410 of the backup memory 120. Each first PIT copy includes a horizontal file 232, 332, or 442, and the horizontal file 232, 332, or 442 includes files extracted from the first working freeze copy 220, 320, or 420.
[0090] Method 500 includes applying a retention lock (520) to at least one first-level file of a first point-in-time copy. For example, as previously described, the retention lock engine 125 of the backup memory 120 applies a retention lock to each first-level file 232, 332, or 442.
[0091] Method 500 includes generating a second point-in-time copy of a second working frozen copy of the active namespace in a shared protected namespace of the backup storage device, the second working frozen copy including a second plurality of files associated with the active namespace, the second plurality of files being stacked in at least one second-level file in the second point-in-time copy (530). For example, as previously described, a second PIT copy 250, 350, or 470 is generated in the shared protected namespace 210, 310, or 410 of the backup memory 120. Each second PIT copy includes a level file 252, 352, or 472, and the level files 252, 352, or 472 include files extracted from the second working frozen copy 240, 340, or 450A.
[0092] Method 500 includes applying a retention lock (540) to at least one second-level file of the second point-in-time copy. For example, as previously described, the retention lock engine 125 of the backup memory 120 applies a retention lock to each second-level file 252, 352, or 472.
[0093] Method 500 includes: generating at least one second-level file of the second point-in-time copy includes copying one or more files from at least one first-level file of the first point-in-time copy to at least one second-level file of the second point-in-time copy, and copying one or more files from the second working frozen copy of the active namespace to at least one second-level file (550). For example, as previously described, according to the embodiments disclosed herein, each second-level file 252, 352, or 472 is generated to include files extracted from the first-level files 232, 332, or 442 and the second working frozen copy 240, 340, or 450A.
[0094] The following are some other exemplary embodiments of the present invention. These embodiments are presented by way of example only and are not intended to limit the scope of the present invention in any way.
[0095] Embodiment 1. A method includes: generating a first point-in-time copy of a first working frozen copy of an active namespace in a shared protection namespace of a backup storage device, the first working frozen copy including a first plurality of files associated with the active namespace, the first plurality of files being stacked in at least one first horizontal file in the first point-in-time copy; applying a retention lock to the at least one first horizontal file of the first point-in-time copy; generating a second point-in-time copy of a second working frozen copy of the active namespace in the shared protection namespace of the backup storage device, the second working frozen copy including a second plurality of files associated with the active namespace, the second plurality of files being stacked in at least one second horizontal file in the second point-in-time copy; and applying a retention lock to the at least one second horizontal file of the second point-in-time copy, wherein generating the at least one second horizontal file of the second point-in-time copy includes copying one or more files of the first plurality of files from the at least one first horizontal file of the first point-in-time copy to the at least one second horizontal file of the second point-in-time copy and copying one or more files from the second working frozen copy of the active namespace to the at least one second horizontal file.
[0096] Embodiment 2. The method according to Embodiment 1, wherein copying one or more files of the first plurality of files from the at least one first horizontal file of the first point-in-time copy to the at least one second horizontal file of the second point-in-time copy and copying one or more files from the second working frozen copy of the active namespace to the at least one second horizontal file includes: determining a first file subset of the first plurality of files, the first file subset including files that are in the at least one first horizontal file of the first point-in-time copy and are also included in the second plurality of files included in the second working frozen copy of the active namespace; determining a second file subset that includes only files included in the second plurality of files included in the second working frozen copy of the active namespace; copying the first file subset from the at least one first horizontal file of the first point-in-time copy to the at least one second horizontal file of the second point-in-time copy; and copying the second file subset from the second working frozen copy of the active namespace to the at least one second horizontal file of the second point-in-time copy.
[0097] Embodiment 3. The method according to Embodiments 1-2, wherein the determination of the first file subset and the second file subset is based on a metadata difference operation between the second working frozen copy of the active namespace and the at least one first horizontal file of the first point-in-time copy.
[0098] Embodiment 4. The method according to Embodiments 1-3, wherein copying one or more files from the one or more files in the first plurality of files from the at least one first-level file of the first time-point copy to the at least one second-level file of the second time-point copy and copying one or more files from the second working freeze copy of the active namespace to the at least one second-level file includes: generating a snapshot of the shared protected namespace, the snapshot including the second working freeze copy of the active namespace and the first time-point copy of the first working freeze copy of the active namespace; determining a first file subset of the first plurality of files, the first file subset being included in the at least one first-level file of the first time-point copy and also being included in the second plurality of files included in the snapshot of the second working freeze copy of the active namespace; determining a second file subset, the second file subset including only the second plurality of files included in the snapshot of the second working freeze copy of the active namespace; copying the first file subset from the at least one first-level file of the first time-point copy to the at least one second-level file of the second time-point copy; and copying the second file subset from the snapshot of the second working freeze copy of the active namespace to the at least one second-level file of the second time-point copy.
[0099] Embodiment 5. The method according to Embodiments 1-4, wherein the first file subset copied from the at least one first-level file of the first time-point copy to the at least one second-level file of the second time-point copy is copied from the at least one first-level file of the first time-point copy of the first working freeze copy of the active namespace included in the snapshot.
[0100] Embodiment 6. The method according to Embodiments 1-5, wherein the determination of the first file subset and the second file subset is based on a metadata difference operation between the second working freeze copy of the active namespace included in the snapshot and the at least one first-level file of the first time-point copy.
[0101] Embodiment 7. The method according to Embodiments 1-6, wherein after copying the second file subset to the at least one second-level file of the second time-point copy, the snapshot is released.
[0102] Embodiment 8. The method according to Embodiments 1-7, wherein the number of operations required to copy one or more files from among the first plurality of files from the at least one first-level file of the first time-point copy to the at least one second-level file of the second time-point copy and to copy one or more files from the second working freeze copy of the active namespace to the at least one second-level file is at most twice the number of files that have changed between the first working freeze copy and the second working freeze copy.
[0103] Embodiment 9. The method according to Embodiments 1-8, wherein adjacent files in the at least one first-level file of the first time-point copy are copied to the at least one second-level file in the second time-point copy in the same operation.
[0104] Embodiment 10. The method according to Embodiments 1-9, wherein the first plurality of files associated with the active namespace are stacked in the at least one first-level file in the first time-point copy by alternating metadata and file data in the at least one first-level file.
[0105] Embodiment 11. A system comprising hardware and / or software operable to perform any one or any part of any of the operations, methods, or processes disclosed herein.
[0106] Embodiment 12. A non-transitory storage medium storing instructions executable by one or more hardware processors to perform operations including any one or more of Embodiments 1-10.
[0107] The embodiments disclosed herein may include the use of a special-purpose or general-purpose computer including various computer hardware or software modules, as discussed in more detail below. The computer may include a processor and a computer storage medium carrying instructions that, when executed by the processor and / or cause the instructions to be executed by the processor, perform any one or more of the methods disclosed herein or any part of any of the disclosed methods.
[0108] As described above, embodiments within the scope of the present invention also include a computer storage medium that is a physical medium for carrying or having stored thereon computer-executable instructions or data structures. Such a computer storage medium can be any available physical medium accessible by a general-purpose or special-purpose computer.
[0109] By way of example and not limitation, such a computer storage medium can include hardware memory such as a solid state drive / device (SSD), RAM, ROM, EEPROM, CD-ROM, flash memory, phase change memory (“PCM”), or other optical disk memory, magnetic disk memory, or other magnetic storage devices, or any other hardware storage device that can be used to store program code in the form of computer-executable instructions or data structures, which can be accessed and executed by a general or special purpose computer system to implement the functions disclosed herein. Combinations of the above should also be included within the scope of computer storage media. These media are also examples of non-transitory storage media, and non-transitory storage media also include cloud-based storage systems and architectures, although the scope of the present invention is not limited to these examples of non-transitory storage media.
[0110] Computer-executable instructions include, for example, instructions and data that, when executed, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a particular function or group of functions. Thus, some embodiments of the present invention may, for example, be downloaded to one or more systems or devices from a website, grid topology, or other source. Similarly, the scope of the present invention includes any hardware system or device that includes an instance of an application that includes the disclosed executable instructions.
[0111] Although the subject matter has been described in language specific to structural features and / or methodological steps, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or steps described above. Rather, the specific features and steps disclosed herein are disclosed as example forms of implementing the claims.
[0112] As used herein, the term “module” or “component” can refer to a software object or routine executing on a computing system. The different components, modules, engines, and services described herein can be implemented as objects or processes executing on a computing system, such as as separate threads. Although the systems and methods described herein can be implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated. In the present invention, a “computing entity” can be any computing system as defined hereinbefore, or any module or combination of modules executing on a computing system.
[0113] In at least some instances, a hardware processor is provided that is operable to execute executable instructions for implementing a method or process, such as the methods and processes disclosed herein. The hardware processor may or may not include other hardware elements, such as elements of the computing devices and systems disclosed herein.
[0114] In terms of computing environments, embodiments of the present invention may be executed in a client-server environment (whether a network environment or a local environment), or any other suitable environment. Suitable operating environments for at least some embodiments of the present invention include cloud computing environments, in which one or more of the client, server, or other machines may reside in and operate within the cloud environment.
[0115] Now briefly refer to Figure 6 , any one or more entities disclosed or implied by the previously described figures and / or elsewhere herein may take the form of a physical computing device, or include a physical computing device, or be implemented on a physical computing device, or be hosted by a physical computing device, one example of which is denoted by 600. Similarly, if any of the above elements includes or consists of a virtual machine (VM), then the VM may constitute Figure 6 virtualization of any combination of the physical components disclosed in
[0116] In Figure 6 's example, the physical computing device 600 includes a memory component 602, which may include one, some, or all of random access memory (RAM), non-volatile memory (NVM) 604 (such as NVRAM), read-only memory (ROM), and persistent memory, one or more hardware processors 606, a non-transitory storage medium 608, a UI device 610, and a data memory 612. One or more of the memory components 602 of the physical computing device 600 may take the form of solid-state device (SSD) memory. Similarly, one or more application programs 614 may be provided, which include instructions executable by one or more hardware processors 606 to perform any operation or part thereof disclosed herein.
[0117] Such executable instructions may take various forms, for example, including instructions executable to perform any method or part thereof disclosed herein, and / or instructions executable by / at any storage site, whether enterprise local or cloud computing site, client, data center, data protection site including cloud storage sites, or backup server, to perform any function disclosed herein. Similarly, such instructions may be executable to perform any other operations and methods disclosed herein and any part thereof.
[0118] Without departing from the spirit or essential characteristics of the present invention, the present invention may be embodied in other specific forms. The described embodiments are to be considered in all respects only as illustrative and not restrictive. Thus, the scope of the present invention is indicated by the appended claims rather than by the foregoing description. All changes within the meaning and scope of the claims should be included within their scope.
Claims
1. A method, comprising: generating a first point-in-time copy of a first working frozen copy of an active namespace in a shared protected namespace of a backup storage device, the first working frozen copy including a first plurality of files associated with the active namespace, the first plurality of files being stacked in at least one first-level file in the first point-in-time copy; applying a retention lock to the at least one first-level file of the first point-in-time copy; generating a second point-in-time copy of a second working frozen copy of the active namespace in the shared protected namespace of the backup storage device, the second working frozen copy including a second plurality of files associated with the active namespace, the second plurality of files being stacked in at least one second-level file in the second point-in-time copy; and applying a retention lock to the at least one second-level file of the second point-in-time copy, wherein generating the at least one second-level file of the second point-in-time copy includes copying one or more files of the first plurality of files from the at least one first-level file of the first point-in-time copy to the at least one second-level file of the second point-in-time copy and copying one or more files from the second working frozen copy of the active namespace to the at least one second-level file.
2. The method according to claim 1, wherein, Copying one or more files of the first plurality of files from the at least one first-level file of the first point-in-time copy to the at least one second-level file of the second point-in-time copy and copying one or more files from the second working frozen copy of the active namespace to the at least one second-level file includes: determining a first file subset of the first plurality of files, the first file subset including files that are in the at least one first-level file of the first point-in-time copy and are also included in the second plurality of files included in the second working frozen copy of the active namespace; determining a second file subset that includes only files that are included in the second plurality of files included in the second working frozen copy of the active namespace; copying the first file subset from the at least one first-level file of the first point-in-time copy to the at least one second-level file of the second point-in-time copy; and copying the second file subset from the second working frozen copy of the active namespace to the at least one second-level file of the second point-in-time copy.
3. The method according to claim 2, wherein The determination of the first file subset and the second file subset is based on a metadata difference operation between the second working frozen copy of the active namespace and the at least one first-level file of the first point-in-time copy.
4. The method according to claim 1, wherein Copying one or more files from the one or more files in the first plurality of files from the at least one first level file of the first point-in-time copy to the at least one second level file of the second point-in-time copy and copying one or more files from the second working frozen copy of the active namespace to the at least one second level file includes: Generating a snapshot of the shared protected namespace, the snapshot including the second working frozen copy of the active namespace and the first point-in-time copy of the first working frozen copy of the active namespace; Determining a first file subset of the first plurality of files, the first file subset being included in the at least one first level file of the first point-in-time copy and also being included in the second plurality of files included in the snapshot of the second working frozen copy of the active namespace; Determining a second file subset that is only included in the second plurality of files included in the snapshot of the second working frozen copy of the active namespace; Copying the first file subset from the at least one first level file of the first point-in-time copy to the at least one second level file of the second point-in-time copy; and Copying the second file subset from the snapshot of the second working frozen copy of the active namespace to the at least one second level file of the second point-in-time copy.
5. The method according to claim 4, wherein, The first file subset copied from the at least one first level file of the first point-in-time copy to the at least one second level file of the second point-in-time copy is copied from the at least one first level file of the first point-in-time copy of the first working frozen copy of the active namespace included in the snapshot.
6. The method according to claim 4, wherein, The determination of the first file subset and the second file subset is based on a metadata difference operation between the second working frozen copy of the active namespace included in the snapshot and the at least one first level file of the first point-in-time copy.
7. The method according to claim 4, wherein, After copying the second file subset to the at least one second level file of the second point-in-time copy, releasing the snapshot.
8. The method according to claim 1, wherein The number of operations required to copy one or more files from the at least one first level file of the first point-in-time copy to the at least one second level file of the second point-in-time copy and to copy one or more files from the second working frozen copy of the active namespace to the at least one second level file is at most twice the number of files that have changed between the first working frozen copy and the second working frozen copy.
9. The method according to claim 8, wherein Copying adjacent files in the at least one first level file of the first point-in-time copy to the at least one second level file of the second point-in-time copy in the same operation.
10. The method according to claim 1, wherein, Stacking the first plurality of files associated with the active namespace in the at least one first-level file in the first timepoint copy by alternating metadata and file data in the at least one first-level file.
11. A non-transitory storage medium storing instructions executable by one or more hardware processors to perform operations, the operations including the operations in the method according to any one of claims 1 to 10.