File system change block tracking for data platforms
By utilizing volume-level CBT and file mapping information in the application system, the problem of unreliable file-level backup during server restart is solved, and a more efficient backup process is achieved, reducing the consumption of computing resources.
Patent Information
- Application Number
- CN202410731774.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-06-06
- Publication Date
- 2025-05-02
AI Technical Summary
During server restarts of the application system, it is difficult to predict uninstallation of the file system, resulting in unreliable execution of file-level backups, consuming a large amount of computing resources for a comprehensive scan to determine file changes.
By leveraging volume-level CBT, the agent can more reliably track all block changes of the volume and combine it with file mapping information to identify block changes that store specific file data, thereby starting subsequent file-level backups.
It realizes more reliable file-level backup during server restart, reduces the consumption of computing resources, improves backup efficiency and timeliness of data access requests.
Smart Images

Figure CN119917459A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a data platform for a computing system. Background Art
[0002] Data platforms that support computing applications rely on primary storage systems to support latency-sensitive applications. However, because primary storage is often more difficult or expensive to scale, secondary storage systems are often relied upon to support secondary use cases such as backup and archiving. Summary of the invention
[0003] Aspects of the present disclosure include techniques for file system (FS) changed block tracking (CBT) for a data platform. The data platform is typically integrated with an application system to back up data stored to a storage device of the application system. In some cases, the data platform or an operator may install an agent on the application system, wherein the agent is responsible for managing the backup of data stored to the storage device and coordinating other operations associated with the data platform. The host system may not natively support volume-level CBT, in which blocks associated with volumes defined within a storage device are tracked to facilitate subsequent backups, wherein only changed blocks of the volume are backed up, rather than all blocks forming a particular volume. CBT can generally facilitate more efficient backups, thereby reducing the consumption of computing resources compared to a full backup of an entire volume.
[0004] In some cases, FS CBT can be used to facilitate backing up only specific files stored to a given volume. That is, FS CBT can allow an agent to identify changes to blocks storing file data associated with a specific file (which can be user-defined), eliminating the need to back up blocks for the entire volume when volume-level CBT is not enabled. Alternatively, volume-level CBT can be used with FS CBT, where FS CBT can facilitate file-level backups at a different frequency than volume-level CBT or according to various other different backup parameters. Therefore, using both volume-level CBT and FS CBT can allow for general volume-level backups of the entire volume at a longer frequency (e.g., daily, weekly, monthly, etc.) and file-level backups at a shorter frequency (e.g., hourly, daily, weekly, etc.). Therefore, users can configure volume-level backups and file-level backups to accommodate many different contexts and data management goals.
[0005] However, performing FS CBT to achieve file-level backups with sufficient reliability may become problematic given that many file systems may be unmounted unpredictably during a reboot of an application system's server (e.g., to accommodate performance degradation of the application system, installation of patches or other updates, power outages, planned maintenance downtime, and the like). That is, during a reboot, the operating system executed by the application system's server may unmount the file system while writes are still pending, which may prevent the agent from performing reliable FS CBT given that the file system is no longer available. The operating system may perform writes by issuing writes to the underlying volume internally, but the agent cannot accurately record file-level changes to blocks storing file data for a particular file. Without the ability to accurately track changes to blocks, the agent can only perform a full scan of the volume after recovery from the reboot to determine whether the file data for a particular file has changed, which may consume significant computing resources of the application system and the data protection system.
[0006] Given that a volume is not unmounted until very late in the restart process (e.g., until after all pending writes have been successfully executed), various aspects of the technology described in the present disclosure can enable an agent to perform FS CBT (also referred to as file system level incremental backup) in a more reliable manner using volume-level CBT. Volume-level CBT can therefore more reliably track all changes to the blocks forming the volume. The agent can interface with the file system to identify file mapping information that identifies which blocks in the volume store file data associated with a particular file undergoing a file level backup. The agent can then identify FS CBT information based on the intersection of the volume-level CBT and the file mapping information, the FS CBT information identifying whether at least one block of the blocks forming the volume that stores file data associated with a particular file has changed. The agent can then initiate a subsequent file level backup based on the file system change block information.
[0007] The technology can provide one or more technical advantages for implementing practical applications. For example, the technology can implement a more reliable FS CBT that can adapt to restarts without having to expend a large amount of computing resources to reconstruct the FS CBT information using a comprehensive scan of each block that forms a file relative to each block of the previous backup of the file. Therefore, various aspects of the technology can improve the operation of the application system itself in terms of reducing the consumption of computing resources (e.g., processor cycles, memory, memory bus bandwidth, and associated power) and network resources (e.g., network bandwidth) compared to standard FSCBT. Such a reduction in computing resources can also improve efficiency because restarting may not take a long time, and data access requests may also be improved in terms of timeliness (considering that memory bus bandwidth is limited, and performing a comprehensive scan of the volume may consume most or even all of the memory bus bandwidth).
[0008] In one example, the present disclosure describes a method, the method comprising: obtaining, by a processing device of a computing device, volume changed block tracking information identifying one or more blocks of a plurality of blocks forming a volume of a storage device, the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determining, by the processing device, file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with a file; determining, by the processing device and based on the volume changed block tracking information and the file mapping information, file system changed block information indicating whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed; and initiating, by the processing device and based on the file system changed block information, a subsequent backup of at least a portion of the file data associated with the file.
[0009] In another example, the present disclosure describes a computing device, the computing device including: a storage device having a plurality of blocks forming a volume; and a processing device capable of accessing the storage device and configured to: obtain volume changed block tracking information, the volume changed block tracking information identifying one or more blocks of the plurality of blocks forming the volume, the one or more blocks storing update data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determine file mapping information, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with a file; determine file system changed block information based on the volume changed block tracking information and the file mapping information, the file system changed block information identifying whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed; and initiate a subsequent backup of at least a portion of the file data associated with the file based on the file system changed block information.
[0010] In another example, the present disclosure describes a computer-readable storage medium that includes instructions that, when executed, configure one or more processors to: obtain volume changed block tracking information, the volume changed block tracking information identifying one or more blocks of a plurality of blocks forming a volume, the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determine file mapping information, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with a file; determine file system changed block information based on the volume changed block tracking information and the file mapping information, the file system changed block information identifying whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed; and initiate a subsequent backup of at least a portion of the file data associated with the file based on the file system changed block information. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1A to Figure 1B is a block diagram illustrating an example system that performs file system block change tracking in accordance with one or more aspects of the present disclosure.
[0012] Figure 2 is a block diagram illustrating an example system in accordance with techniques of this disclosure.
[0013] Figure 3 is a diagram showing various aspects of the technology described in the present disclosure. Figure 1A to Figure 2 A conceptual diagram of example operations of a changed block tracking (CBT) driver in performing file system CBT in a manner that utilizes volume CBT is shown in the example of FIG.
[0014] Figure 4 It also shows various aspects of the technology described in this disclosure. Figure 1A to Figure 2 A conceptual diagram of example operations of a changed block tracking (CBT) driver in performing file system CBT in a manner that utilizes volume CBT is shown in the example of FIG.
[0015] Figure 5 It also shows various aspects of the technology described in this disclosure. Figure 1A to Figure 2 A conceptual diagram of example operations of a changed block tracking (CBT) driver in performing file system CBT in a manner that utilizes volume CBT is shown in the example of FIG.
[0016] Figure 6 is a diagram showing various aspects of the CBT technology described in the present disclosure Figure 1A to Figure 2 A flow chart of example operations of a digital management module is shown in an example.
[0017] Like reference numerals refer to like elements throughout the text and drawings. DETAILED DESCRIPTION
[0018] Figure 1A to Figure 1B is a block diagram illustrating an example system for performing file system block change tracking according to one or more examples of the present disclosure. Figure 1A In the example of , system 100 includes application system 102. Application system 102 can represent a collection of hardware devices, software components, and / or data storage, which can be used to implement one or more applications or services provided to one or more mobile devices 108 and one or more client devices 109 via network 113. Application system 102 can include one or more physical or virtual computing devices that execute workloads 174 for the applications or services. Workloads 174 can include one or more virtual machines, containers, Kubernete pods each including one or more containers, bare metal processes, and / or other types of workloads.
[0019] exist Figure 1A , application system 102 includes application servers 170A–170M (collectively, “application servers 170”) connected via a network to a database server 172 that implements a database. Other examples of application system 102 may include one or more load balancers, network servers, network devices such as switches or gateways, or other devices for implementing one or more applications or services and delivering the one or more applications and services to mobile devices 108 and client devices 109. Application system 102 may include one or more file servers. One or more file servers may implement a primary file system for application system 102. (In such cases, file system 153 may be a secondary file system that provides backup, archiving, and / or other services for the primary file system. References to a file system herein may include a primary file system or a secondary file system, e.g., a primary file system for application system 102 or a file system 153 operating as a primary file system or a secondary file system.)
[0020] The application system 102 can be located locally and / or in one or more data centers, each of which is part of a public, private, or hybrid cloud. The application or service can be a distributed application. The application or service can support enterprise software, financial software, office or other productivity software, data analysis software, end-user relationship management, network services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications or services. The application or service can be provided as software as a service (SaaS), platform as a service (PaaS), infrastructure as a service (IaaS), data storage as a service (DSaaS), or other types of services.
[0021] In some examples, application system 102 may represent an enterprise system that includes one or more workstations in the form of desktop computers, laptop computers, mobile devices, enterprise servers, network devices, and other hardware for supporting enterprise applications. Enterprise applications may include enterprise software, financial software, office or other productivity software, data analysis software, end-user relationship management, network services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications. Enterprise applications may be delivered as a service from an external cloud service provider or other provider, executed natively on application system 102, or both.
[0022] exist Figure 1A In the example of , the system 100 includes a data platform 150 that provides a file system 153 and archiving functions to the application system 102 using the storage system 105 and the separate storage system 115. The data platform 150 implements a distributed file system 153 and storage architecture to facilitate the application system 102 to access the file system data and to facilitate the transmission of data between the storage system 105 and the application system 102 via the network 111. In terms of a distributed file system, the data platform 150 enables the device of the application system 102 to access the file system data via the network 111 using a communication protocol, as if such file system data is stored locally (e.g., stored to the hard disk of the device of the application system 102). Example communication protocols for accessing files and objects include server message block (SMB), network file system (NFS), or AMAZON simple storage service (S3). The file system 153 can be the primary file system or a secondary file system of the application system 102.
[0023] The file system manager 152 may represent a collection of hardware devices and software components that implement the file system 153 for the data platform 150. Examples of file system functions provided by the file system manager 152 include storage space management, including deduplication, file naming, directory management, metadata management, partitioning, and access control. The file system manager 152 may execute a communication protocol to facilitate the application system 102 to access files and objects stored in the storage system 105 via the network 111.
[0024] The data platform 150 includes a storage system 105 having one or more storage devices 180A-180N (collectively referred to as "storage devices 180"). Storage devices 180 may represent one or more physical or virtual computing and / or storage devices that include or have access to storage media. Such storage media may include one or more of the following: flash drives, solid-state drives (SSDs), hard disk drives (HDDs), various electrically programmable memories (EPROMs) or electrically erasable programmable (EEPROM) memories, and / or other types of storage media used to support the data platform 150. Different storage devices of storage devices 180 may have different mixes of various types of storage media. Each of the storage devices 180 may include system memory. Each of the storage devices 180 may be a storage server, a network attached storage (NAS) device, or may represent disk storage of a computing device. The storage system 105 may be a redundant array of independent disks (RAID) system. In some examples, one or more of the storage devices 180 is both a computing device and a storage device that executes software for the data platform 150, such as the file system manager 152 and the archive manager 154 in the example of the system 100, and stores objects and metadata for the data platform 150 to a storage medium. In some examples, a separate computing device (not shown) may execute software for the data platform 150, such as the file system manager 152 and the archive manager 154 in the example of the system 100. Each of the storage devices 180 may be considered and referred to as a "storage node" or simply a "node". The storage device 180 may represent a virtual machine running on a supported hypervisor, a cloud virtual machine, a physical rack-mounted server, or a computing node installed in a converged platform.
[0025] In various examples, the data platform 150 can run on a physical system, virtually, or cloud-natively. For example, the data platform 150 can be deployed as a physical cluster, a virtual cluster, or a cloud-based cluster that runs in a private cloud, a hybrid private / public cloud, or a public cloud deployed by a cloud service provider. In some examples of the system 100, multiple instances of the data platform 150 can be deployed, and the file system 153 can be replicated in each instance. In some cases, the data platform 150 can be a computing cluster representing a single management domain. The number of storage devices 180 can be expanded to meet performance needs.
[0026] The data platform 150 may implement multiple storage domains and provide the storage domains to one or more tenants or separate workloads 174 that require different data policies. A storage domain may be a data policy domain that determines policies for deduplication, compression, encryption, tiering, and other operations performed on objects stored using the storage domain. In this way, the data platform 150 may provide users with the flexibility to select global data policies or workload-specific data policies. The data platform 150 may support partitioning.
[0027] A view may be a protocol export that resides within a storage domain. A view inherits the data policy of its storage domain, but additional data policies may be specified for the view. Views may be exported via SMB, NFS, S3, and / or another communication protocol. Policies that determine data processing and storage for the data platform 150 may be assigned at the view level. A protection policy may specify backup frequency and retention policies, which may include data lockout periods. Archives 142 or snapshots created in accordance with a protection policy inherit the data lockout and retention periods specified by the protection policy.
[0028] Each of network 113 and network 111 may be the Internet, or may include or represent any public or private communication network or other network. For example, network 113 may be a cellular, Near field communication (NFC), satellite, enterprise, service provider, and / or other types of networks capable of transmitting data between computing systems, servers, computing devices, and / or storage devices. One or more of such devices may use any suitable communication technology to transmit and receive data, communications, control signals, and / or other information across network 113 or network 111. Each of network 113 or network 111 may include one or more network hubs, network switches, network routers, satellite dishes, or any other network equipment. Such network devices or components may be operatively coupled to each other, thereby being used to exchange information between computers, devices, or other components (e.g., between one or more client devices or systems and one or more computers / servers / storage devices or systems). Figure 1A to Figure 1BEach of the devices or systems shown in the figure may be operatively coupled to the network 113 and / or the network 111 using one or more network links. The links coupling such devices or systems to the network 113 and / or the network 111 may be Ethernet, asynchronous transfer mode (ATM), or other types of network connections, and such connections may be wireless and / or wired connections. Figure 1A to Figure 1B One or more of the devices or systems shown in or on network 113 and / or network 111 may be remotely located relative to one or more of the other shown devices or systems.
[0029] The application system 102 can generate objects and other data using the file system 153 provided by the data platform 150, and the file system manager 152 can create, manage and store the objects and data in the storage system 105. For this purpose, the application system 102 can be alternatively referred to as a "source system", and the file system 153 used for the application system 102 can be alternatively referred to as a "source file system". The application system 102 can communicate directly with the storage system 105 via the network 111 to transfer objects for some purposes, and communicate with the file system manager 152 via the network 111 to indirectly obtain objects or metadata from the storage system 105 for some purposes. The file system manager 152 can generate metadata and store the metadata in the storage system 105. The collection of data stored in the storage system 105 and used to implement the file system 153 is referred to as file system data in this article. The file system data may include the aforementioned metadata and objects. The metadata may include file system objects, tables, trees or other data structures; metadata generated to support deduplication; or metadata used to support snapshots. The stored objects may include files, virtual machines, databases, applications, pods, containers, any workload 174, system images, directory information, or other types of objects used by the application system 102. Objects of different types and objects of the same type may be deduplicated relative to each other.
[0030] Data platform 150 includes archive manager 154 that provides archiving for file system data of file system 153. In the example of system 100, archive manager 154 archives file system data stored by storage system 105 to storage system 115 via network 111.
[0031] The storage system 115 includes one or more storage devices 140A-140X (collectively referred to as "storage devices 140"). Storage device 140 may represent one or more physical or virtual computing and / or storage devices that include or have access to storage media. Such storage media may include one or more of the following: flash drives, solid-state drives (SSDs), hard disk drives (HDDs), optical disks, various electrically programmable memories (EPROMs) or electrically erasable programmable (EEPROM) memories, and / or other types of storage media. Different storage devices of storage device 140 may have different mixes of various types of storage media. Each of storage devices 140 may include system memory. Each of storage devices 140 may be a storage server, a network attached storage (NAS) device, or may represent disk storage of a computing device. Storage system 115 may include a redundant array of independent disks (RAID) system. Storage system 115 is capable of storing much more data than storage system 105. Storage device 140 may also be configured for long-term storage of information, more suitable for archival purposes.
[0032] In some examples, storage systems 105 and / or 115 may be storage systems deployed and managed by a cloud storage provider and are referred to as "cloud storage systems". Example cloud storage providers include, for example, Amazon Web Services (AWS) and Amazon Web Services (AWS. TM ), MICROSOFT DROPBOX from DROPBOX TM 、ORACLE CLOUD TMand GOOGLE CLOUD PLATFORM (GCP) of GOOGLE Corporation. In some examples, storage system 115 is co-located with storage system 105 in a data center, locally, or in a private, public, or hybrid private / public cloud. Storage system 115 can be considered a "backup" or "secondary" storage system for primary storage system 105. Storage system 115 can be referred to as an "external target" for archive 142. In the case of being deployed and managed by a cloud storage provider, storage system 115 can be referred to as "cloud storage". Storage system 115 can include one or more interfaces for managing the transmission of data between storage system 105 and storage system 115 and / or between application system 102 and storage system 115. The data platform 150 supporting application system 102 can rely on primary storage system 105 to support latency-sensitive applications. However, because storage system 105 is generally more difficult to scale or more expensive to scale, data platform 150 can use secondary storage system 115 to support secondary use cases such as backup and archiving. In general, a file system backup may be a copy of a file system 153 used to support protection of the file system 153 for quick recovery (typically due to the loss of some data in the file system 153), and a file system archive ("archive") may be a copy of a file system 153 used to support long-term retention and viewing. A "copy" of a file system 153 may include such data as is necessary to restore or view the state of the file system 153 at the time of the backup or archive.
[0033] Archive manager 154 can archive the file system data of file system 153 at any time according to an archive policy, which specifies, for example, archive periodicity and timing (daily, weekly, etc.), which file system data will be archived, archive retention period, storage location, access control, etc. The initial archive of file system data can correspond to the state of the file system data at the initial archive time (archive creation time of the initial archive). According to the archive policy, the initial archive can include a complete archive of the file system data, or can include a less complete archive of the file system data. For example, the initial archive can include all objects of file system 153 or one or more selected objects of file system 153.
[0034] One or more subsequent incremental archives of file system 153 may correspond to respective states of file system 153 at respective subsequent archive creation times (i.e., after the archive creation time corresponding to the initial archive). Subsequent archives may include incremental archives of file system 153. Subsequent archives may correspond to incremental archives of one or more objects of file system 153. Some file system data of file system 153 stored on storage system 105 at the initial archive creation time may also be stored on storage system 105 at the subsequent archive creation time. Subsequent incremental archives may include data that was not previously archived to storage system 115. Archive manager 154 may deduplicate file system data included in subsequent archives against file system data included in one or more previous archives (including the initial archive) to reduce the amount of storage used. (References to "time" in this disclosure may refer to dates and / or times. Times may be associated with dates. For example, multiple archives may occur at different times on the same day.)
[0035] In system 100, archive manager 154 may archive file system data to storage system 115 as archive 142 using block files 162. Archive manager 154 may use any of archives 142 to later restore the file system (or portion thereof) to its state at the time of archive creation, or, for example, may use the archive to create or present a new file system (or "view") based on the archive. As described above, archive manager 154 may de-duplicate file system data included in subsequent archives against file system data included in one or more previous archives. For example, a second object of file system 153 and included in a second archive may be de-duplicated against a first object of file system 153 and included in an earlier first archive. Archive manager 154 may remove data blocks ("blocks") of the second object and generate metadata having references (e.g., pointers) to the stored blocks in blocks 164 in one of block files 162. The stored blocks in this example are instances of blocks stored for the first object.
[0036] Archive manager 154 may apply deduplication as part of the write process of writing (i.e., storing) objects of file system 153 to one of archives 142 in storage system 115. Deduplication may be implemented in a variety of ways. For example, the method may be fixed length or variable length, the block size of the file system may be fixed or variable, and the deduplication domain may be applied globally or per workload. Fixed length deduplication involves dividing a data stream at fixed intervals. Variable length deduplication involves dividing a data stream at variable intervals to improve the ability to match data regardless of which file system block size method is used. The algorithm is more complex than the fixed length deduplication algorithm, but may be more efficient for most situations and generally produces less metadata. Variable length deduplication may include variable length, sliding window deduplication. The length of any deduplication operation (whether fixed length or variable length) determines the size of the blocks that are deduplicated.
[0037] In some examples, for variable length deduplication, the block size may be within a fixed range. For example, archive manager 154 may calculate blocks with block sizes in the range of 16kB to 48kB. Archive manager 154 may avoid deduplication of objects smaller than 16kB. In some example implementations, when considering deduplication of data for an object, archive manager 154 compares a block identifier (ID) of the data (e.g., a hash value of the entire block) with an existing block ID of a stored block. If a match is found, archive manager 154 updates the metadata of the object to point to the matching, stored block. If no matching block is found, archive manager 154 writes the data of the object to memory as one of the blocks 164 of one of the block files 162. In addition, archive manager 154 stores the block ID in association with the newly stored block in the block metadata to allow future deduplication against the newly stored block. In general, for any of the archives 142, chunk metadata may be used to generate, view, retrieve, or restore objects stored as chunks 164 (and references thereto) within chunk files 162, and this is described in more detail below.
[0038] Each of the block files 162 includes a plurality of blocks 164. The block files 162 may be of a fixed size (e.g., 8MB) or a variable size. The block files 162 may be stored using a data structure provided by a cloud storage provider for the storage system 115. For example, each of the block files 162 may be one of the following: an S3 object in an AWS cloud storage bucket, an object in an Azure Blob storage, an object in an object storage for an ORACLE CLOUD, or other similar data structures used in another cloud storage provider storage system. Any of the block files 162 may be subject to a write-once-ready-many (WORM) lock having a WORM lock expiration time. The WORM lock for an S3 object is referred to as an "object lock," and the WORM lock for an object in an Azure Blob storage is referred to as "blob immutability."
[0039] The process of deduplicating multiple objects across multiple archives results in block files 162, each of which has multiple blocks 164 for multiple different objects associated with the multiple archives. In some examples, different archives 142 may have objects that are actually copies of the same data, e.g., objects that have not been modified with respect to the file system. Archived objects may be represented or "stored" as metadata with references to the blocks, which enables access to the archived objects. Thus, descriptions herein of an archive "storing," "having," or "including" an object include instances in which the archive does not store the object's data in its native form.
[0040] The initial archive and one or more subsequent incremental archives of archive 142 can each be associated with a corresponding retention period and, in some cases, with a data lock period for the archive. As described above, a data management policy (not shown) can specify a retention period for an archive and a data lock period for an archive. The retention period for an archive is the amount of time that blocks referenced by the archive's objects are stored in the archive and before the blocks can be removed from the storage area. The retention period for an archive begins when the archive is stored (archive creation time). Block files containing blocks referenced by the archive's objects and subject to the archive's retention period but not to the archive's data lock period can be modified at any time before the retention period expires. The nature of such modifications must be to be able to preserve the data referenced by the archive's objects.
[0041] Archived data stored in storage system 115 may be accessed (e.g., read or write) by a user or application associated with application system 102. A user or application may delete portions of the data due to a malicious attack (e.g., a virus, ransomware, etc.), a rogue or malicious administrator, and / or human error. A user's credentials may be compromised, and therefore, the archived data stored in storage system 115 may be subject to ransomware attacks. To reduce the possibility of accidental or malicious data deletion or corruption, a data lock with a data lock period may be applied to the archive.
[0042] As described above, the block file 162 may represent an object in an archival storage system (illustrated as “storage system 115”, which may also be referred to as “archival storage system 115”) that conforms to the underlying architecture of the archival storage system 115. The data platform 150 includes an archive manager 154 that supports archiving data in the form of block files 162, which interface with the archival storage system 115 to store the block files 162 after forming the block files 162 from one or more blocks 164 of data. The archive manager 154 may apply a process known as “de-duplication” to the blocks 164 to remove redundant blocks and generate metadata that links the redundant blocks to previously archived blocks 164, thereby reducing consumed storage (thereby reducing storage costs in terms of storage required to store the blocks). The archive manager 154 may use the metadata to aggregate the blocks 164 to form the block files 162 at the archival storage system 115.
[0043] like Figure 1A In the example of FIG. 1 , the application system 102 can execute the data management module 120. In this example, the application server 170M is shown as executing the data management module 120, but the database server 172 and / or the application server 170 (and Figure 1A Any one or more of the additional servers, computing devices, etc. not shown in the examples may execute a corresponding instance of the data management module 120.
[0044] The data management module 120 may represent software (or hardware or a combination of both) for managing data stored locally at the application system 102 (e.g., via the database server 172 and / or the application server 170), wherein the data management module 120 may interface with the data platform 150 to store the data remotely (e.g., backup) via the storage system 105 or archive the data to the storage system 115. The data management module 120 may include an agent 122 that may coordinate the backup and / or archiving of data stored locally at the application system 102 on behalf of the data platform 150. The agent 122 may interface with the data platform 150 via the network 111 to coordinate the delivery of the data to be backed up, provide metadata specific to the data to be backed up (e.g., data type, archive duration, archive periodicity, data volume, etc.), and otherwise help coordinate the operation of the data platform 150 with respect to the application system 102.
[0045] The data management module 120 may also include one or more changed block tracking (CBT) drivers 124 that intercept input / output control ("IOCTL") operations that are sent to the underlying storage device to perform data operations such as read, write, delete, copy (read first then write), etc. The CBT driver 124 can process these requests to identify whether the data stored locally on the application system 102 has been updated since a previous backup of the local data. The CBT driver 124 can track these "changes" in order to identify whether the local data needs to be included in a subsequent backup.
[0046] In other words, the data platform 150 can be integrated with the application system 102 to back up the storage devices of the servers of the application system 102 through the data management module 120. As described above, the agent 122 is responsible for managing the backup of data stored to the storage devices and coordinating other operations associated with the data platform 150. The application system 150 may not natively support volume-level CBT, in which blocks associated with volumes defined within the storage devices are tracked to facilitate subsequent backups, where only changed blocks of the volume are backed up, rather than all blocks forming a particular volume. CBT can generally facilitate more efficient backups, thereby reducing the consumption of computing resources (e.g., processing cycles, memory, memory bus bandwidth, and associated power) compared to a full backup of the entire volume.
[0047] In some cases, file system (FS) CBT can be used to facilitate backing up only specific files stored to a given volume. That is, FS CBT can allow agent 122 to identify changes to blocks storing file data associated with specific files (which can be user-defined) through CBT driver 124, thereby eliminating the need to back up blocks of the entire volume when volume-level CBT is not enabled. Alternatively, volume-level CBT can be used with FS CBT, where FS CBT can facilitate file-level backups at a different frequency than volume-level CBT or according to various other different backup parameters. Therefore, using both volume-level CBT and FS CBT can allow for general volume-level backups of the entire volume at a longer frequency (e.g., daily, weekly, monthly, etc.) and file-level backups at a shorter frequency (e.g., hourly, twice a day, weekly, etc.). Therefore, users can configure volume-level backups and file-level backups to accommodate many different contexts and data management goals.
[0048] However, given that it is difficult to predict the unmounting of many file systems during a restart of the application server 170M (e.g., to accommodate performance degradation of the application server 170M, installation of patches or other updates, power outages, planned maintenance downtime, etc.), performing FS CBT to achieve file-level backups with sufficient reliability may become problematic. That is, during a restart, the operating system executed by the application system 170M may unmount the file system while writes are still pending, which may prevent the CBT driver 124 (and the agent 122) from performing reliable FS CBT given that the file system is no longer available. The operating system can perform writes, but the CBT driver 124 cannot accurately record file-level changes to the blocks that store the file data of a particular file. If changes to the blocks cannot be accurately tracked, the agent 122 (via the CBT driver 124) can only perform a full scan of the volume after recovering from the restart to determine whether the file data of a particular file has changed, which may consume a large amount of computing resources of the application system and may result in a degraded user experience.
[0049] Considering that a volume is not unmounted until very late in the restart process (e.g., until after all pending writes are successfully executed), various aspects of the technology described in the present disclosure can enable a CBT driver 124 to perform FS CBT in a more reliable manner using volume-level CBT. Volume-level CBT can therefore more reliably track all changes to the blocks forming the volume in the form of volume CBT information 125 ("vol CBT info 125"). The CBT driver 124 can interface with the file system of the underlying operating system executed by the application server 170M to identify file mapping information 127 ("FMI 127"), which identifies which blocks in the volume store file data associated with a particular file undergoing file-level backup. The CBT driver 124 can then identify FS CBT information 129 ("FSCBT info 129") based on the intersection of the volume-level CBT information 125 and the file mapping information 127, which identifies whether at least one block of the blocks forming the volume storing file data associated with a particular file has changed. The agent 122 may then initiate subsequent file level backups based on the FS CBT information 129 .
[0050] The technology can provide one or more technical advantages for implementing practical applications. For example, the technology can implement a more reliable FS CBT that can adapt to restarts without having to expend a large amount of computing resources to reconstruct the FS CBT information 129 using a full scan of each block that forms the volume relative to each block of the previous backup of the volume. Therefore, various aspects of the technology can improve the operation of the application server 170M itself in terms of reducing the consumption of computing resources (e.g., processor cycles, memory, memory bus bandwidth, and associated power) compared to standard FS CBT. Such a reduction in computing resources can also improve the computing experience, because restarting may not take a long time, and data access requests may also be improved in terms of timeliness (considering that memory bus bandwidth is limited, and performing a full scan of the volume may consume most or even all of the memory bus bandwidth), which can also improve the user experience.
[0051] In this regard, the data management module 120 may obtain volume changed block tracking information 125, the volume changed block tracking information identifying one or more blocks of a plurality of blocks forming a volume stored to a storage device (coupled to the application server 170M), the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume. The data management module 120 may then identify a file stored to the volume for which changed block tracking is to be performed, and determine file mapping information 125, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with the file. The data management module 120 may then determine file system changed block information 129 based on the volume changed block tracking information 125 and the file mapping information 127, the file system changed block information identifying whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed. The data management module 120 may initiate a subsequent backup of at least a portion of the file data associated with the file based on the file system changed block information 129.
[0052] Figure 1B The system 190 is Figure 1A A variation of system 100 in which data platform 150 stores archive 142 using block files 162, which are stored to archive storage system 115, which resides locally or in other words, local to data platform 150. In some examples of system 190, storage system 115 enables a user or application to create, modify, or delete block files 162 via file system manager 152. In system 190, Figure 1B Storage system 105 is the local storage system that archive manager 154 initially uses to store and accumulate chunks before they are archived to storage system 115 .
[0053] Figure 2 is a block diagram illustrating an example of an application server in accordance with the techniques of this disclosure. Figure 2 The computing system 202 may represent an example of any of the application server 170 and / or the database server 172 and may be located at Figure 1A System 100 or Figure 1B The system 190 is described in the context of FIG. 1 .
[0054] exist Figure 2In the examples of the present disclosure, computing system 202 may be implemented as any suitable computing system, such as one or more server computers, workstations, mainframes, devices, cloud computing systems, and / or other computing systems that may be capable of performing operations and / or functions described according to one or more aspects of the present disclosure. In some examples, computing system 202 may represent a cloud computing system, a server farm, and / or a server cluster (or a portion thereof) that provides services to other devices or systems. In other examples, computing system 202 may represent or be implemented by one or more virtualized computing instances (e.g., virtual machines, containers) of a cloud computing system, a server farm, a data center, and / or a server cluster.
[0055] exist Figure 2 In the example of , computing system 202 may include one or more communication units 215, one or more input devices 217, one or more output devices 218, and one or more storage devices of a local storage system 205 ("storage system 205"). Local storage system 205 stores data management module 120 and file system (FS) manager module 220 ("FS manager module 220") and presents volumes 230A-230N ("volumes 230"). One or more of the devices, modules, storage areas, or other components of computing system 202 may be interconnected to enable inter-component communication (physically, communicatively, and / or operationally). In some examples, such connections may be provided through a communication channel (e.g., communication channel 212), which may represent one or more of a system bus, a network connection, an inter-process communication data structure, or any other method for transferring data.
[0056] The computing system 202 includes a processing device. Figure 2 In the example of FIG. 1 , the processing device includes one or more processors 213. The one or more processors 213 of the computing system 202 may implement a processor associated with or associated with the computing system 202. Figure 2 Functionality and / or execution associated with one or more modules shown and described below may be associated with computing system 202 or with Figure 2Instructions associated with one or more modules shown and described below. One or more processors 213 can be a processing circuit, can be a part of a processing circuit, and / or can include a processing circuit that performs operations according to one or more aspects of the present disclosure. Examples of processor 213 include a microprocessor, an application processor, a display controller, an auxiliary processor, one or more sensor hubs, and any other hardware configured to function as a processor, a processing unit, or a processing device. The computing system 202 can use a processing device (e.g., one or more processors 213) to perform operations according to one or more aspects of the present disclosure using software, hardware, firmware, or a mixture of hardware, software, and firmware resident in and / or executed at the computing system 202.
[0057] One or more communication units 215 of computing system 202 may communicate with devices external to computing system 202 by transmitting and / or receiving data, and in some aspects, may operate as both an input device and an output device. In some examples, communication unit 215 may communicate with other devices over a network. In other examples, communication unit 215 may send and / or receive radio signals over a radio network (such as a cellular radio network). In other examples, communication unit 215 of computing system 202 may transmit and / or receive satellite signals over a satellite network. Examples of communication unit 215 include a network interface card (e.g., such as an Ethernet card), an optical transceiver, a radio frequency transceiver, a GPS receiver, or any other type of device that can send and / or receive information. Other examples of communication unit 215 may include a device that is capable of transmitting and / or receiving information over a network. GPS, NFC, and cellular networks (e.g., 3G, 4G, 5G) and found in mobile devices devices that communicate with other devices such as wireless radios and Universal Serial Bus (USB) controllers. Such communications may comply with, implement, or follow appropriate protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), Ethernet, NFC or other technologies or protocols.
[0058] The one or more input devices 217 may represent any input device of the computing system 202 that is not otherwise separately described herein. The input device 217 may generate, receive, and / or process input. For example, the one or more input devices 217 may generate input or receive input from a network, a user input device, or any other type of device for detecting input from a human or a machine.
[0059] One or more output devices 218 may represent any output device of computing system 202 that is not otherwise described separately herein. Output device 218 may generate, present, and / or process output. For example, one or more output devices 218 may generate, present, and / or process output in any form. Output device 218 may include one or more USB interfaces, video and / or audio output interfaces, or any other type of device capable of generating tactile, audio, visual, video, electrical, or other output. Some devices may be used as both input devices and output devices. For example, a communication device may send data to other systems or devices and receive data from other systems or devices via a network.
[0060] One or more storage devices of the local storage system 205 within the computing system 202 can store information for processing during operation of the computing system 202, such as random access memory (RAM), flash memory, solid state drive (SSD), hard disk drive (HDD), etc. The storage device can store program instructions and / or data associated with one or more modules described according to one or more aspects of the present disclosure. One or more processors 213 and one or more storage devices can provide an operating environment or platform for such modules, which can be implemented as software, but in some examples can include any combination of hardware, firmware and software. One or more processors 213 can execute instructions, and one or more storage devices of the storage system 205 can store instructions and / or data of one or more modules. The combination of the processor 213 and the local storage system 205 can retrieve, store and / or execute instructions and / or data of one or more applications, modules or software. The processor 213 and / or the storage device of the local storage system 205 can also be operably coupled to one or more other software and / or hardware components, including but not limited to the computing system 202 and / or one or more components of one or more devices or systems shown as connected to the computing system 202.
[0061] The file system manager module 220 (also referred to as "file system manager 220" or "FS manager 220") may perform functions related to providing a file system (FS) 221 ("FS221"). The file system manager 220 may generate and manage file system metadata for constructing file system data of the file system 221, and store the file system metadata and the file system data 230 to the local storage system 205 (on one or more of the volumes 230). The file system metadata may include one or more trees that describe objects within the file system 221 and the hierarchical structure of the file system 221, and may be used to write or retrieve objects within the file system 221. The file system manager 221 may interact and / or operate in coordination with one or more modules of the computing system 202, including the data management module 120. Examples of the file system 221 may include a new technology file system (NFTS), a resilient file system (ReFS), a file allocation table (FAT), exFAT, or other suitable file systems.
[0062] The data management module 120 may perform archiving functions related to backing up the file system 221 (eg, part or all of the file system 221). The data management module 120 may generate one or more backups of the file system 221, which may be stored to the storage system 105 of the data platform 150.
[0063] Data management module 120 may include a local agent 122 (illustrated as “agent” 122) that executes locally within computing system 202. Agent 122 may operate as an agent for data platform 150 to coordinate archive 142 locally with respect to computing system 202, as well as perform various other operations to facilitate compact data transmission (e.g., perform deduplication, compression, encoding, etc.), secure transmission (e.g., encryption), etc. of updates to archive 142.
[0064] The data management module 120 may also include a CBT driver 124 configured to perform different levels of CBT (e.g., volume-level CBT and FS-level CBT). The CBT driver 124 may be a single driver 124 that is located at the back during the reboot process (in order to perform volume CBT). Therefore, the CBT driver 124 may implement volume CBT (or, in some cases, rely on the underlying storage system 205 itself, where some storage devices may implement volume CBT themselves to generate volume CBT information 125) so as to utilize the volume CBT information 125 to build a more reliable FS CBT information 129.
[0065] The agent 122 may interface with the data platform 150 to configure (or reconfigure) the CBT driver 124 when the data management module 120 is installed and / or executed. The agent 122 may initialize and configure (or reconfigure) the CBT driver 124 to obtain the volume CBT information 125. Figure 2 In the example of , it is assumed that the volume CBT information is applied to the volume 230N storing the file 231. As described above, the storage system 205 may include a storage device (e.g., a flash memory) storing the volume 230N, wherein the storage device natively supports volume CBT. The CBT driver 124 may be configured to retrieve the volume CBT information 125 from the storage device itself.
[0066] When the underlying storage device presenting the volume 230N does not natively support volume CBT, the CBT driver 124 may be configured to perform volume CBT in order to obtain volume CBT information 125. The volume CBT information 125 may include a volume CBT bitmap, wherein each bit in the bitmap identifies whether a corresponding block of the plurality of blocks forming the volume 230N (which may be referred to as a "sector" of the volume) has changed relative to a previous backup of one or more of the plurality of blocks forming the volume 230N. After performing a backup of any blocks that have changed relative to the previous backup, the CBT driver 124 may clear the bitmap to indicate that all blocks have not changed relative to the now previous backup.
[0067] In any event, to generate FS CBT information 129, CBT driver 124 may interface with FS manager 220 via application programming interface (API) 223 to obtain file mapping information (FMI) 127 (“FMI 127”). That is, FS manager 220 may expose or otherwise present API 223, which CBT driver 124 may be configured to call in order to determine FMI 127. Although described with respect to API 223, FS manager 220 may present various FS tools that are also capable of determining FMI 127.
[0068] FMI 127 may identify one or more blocks of the plurality of blocks forming volume 230N that store file data associated with file 231. The blocks storing file data associated with file 231 may be contiguous or fragmented. FMI 127 may specify whether a particular file, such as file 231, is subject to FS CBT within the volume block structure. In some cases, FMI 127 may include a file mapping bitmap that maps file data associated with file 231 to one or more blocks of the plurality of blocks forming volume 230N.
[0069] In some examples, since a file may have multiple different paths (e.g., direct paths, shortcuts, other hard / soft links, etc.), the CBT driver 124 may perform a translation on the path to access the file 231 to obtain a common file name. The CBT driver 124 may then interface with the FS manager 220 via the API 223, passing the common file name to the FS manager 220. The FS manager 220 may then output the FMI 127 (or some derivative thereof, such as file layout information, which the CBT driver 124 may process to obtain the FMI 127, e.g., in the bitmap format described above).
[0070] In some cases, the CBT driver 124 may also interface with the FS manager module 220 via the API 223 to provide the FMI 127 regarding the backup, thereby maintaining data consistency (so that when querying the file layout information (which may form part of the FMI 127), the file layout does not change, or when reading data from the file, the file content does not change, etc.). That is, considering that the current version of the file system 221 may have changed in terms of hierarchy, file location, fragmentation, etc., data consistency may not be maintained, and the CBT driver 124 may request the FS manager 220 to output the FMI 127 regarding the previous backup of FS221 via the API 223 (where FS221 may represent the previous backup of FS221).
[0071] The CBT driver 124 may determine the FS CBT information 129 based on the volume CBT information 125 and the FMI 127. As an example, the CBT driver 124 may intersect the volume CBT information 125 with the FMI 127 to derive the changed blocks of the file data associated with the file 231 (which is subject to FS CBT) stored in the volume 230N. The agent 122 may receive an indication via user input, the indication identifying the file 231 stored in the volume 230N for which changed block tracking is to be performed; and configure (or reconfigure) the CBT driver 124 to perform the above-mentioned FS CBT on the file 231 by utilizing the more reliable volume CBT information 125. At some time, the agent 122 may initiate a subsequent backup of at least a portion of the file data associated with the file 231 based on the FS CBT information 129.
[0072] Figure 3 is a diagram showing various aspects of the technology described in the present disclosure. Figure 1A to Figure 2 A conceptual diagram of example operations of a changed block tracking (CBT) driver in performing a file system CBT process in a manner that utilizes volume CBT is shown in the example of FIG. Figure 3In the example of , the CBT driver 124 can obtain or, in some cases, determine volume CBT information 125 and FMI 127 (each in the form of a bitmap, the length of which has been shortened to eight (8) bits to facilitate discussion herein, but each of which can be of any length).
[0073] The CBT driver 124 may then determine the FS CBT information 129 based on the volume CBT bitmap 125 (which is another way of referring to the volume CBT information 125) and the file mapping bitmap 127 (which is another way of referring to the FMI 127), the FS CBT information identifying the FS CBT information forming the volume 230N (shown in FIG. Figure 2 The CBT driver 124 may determine the FS CBT information 129 (again, in the form of a bitmap, and therefore may also be referred to as “FS CBT bitmap 129”) as the intersection (e.g., logical AND operation) of the volume CBT bitmap 125 and the file mapping bitmap 127.
[0074] like Figure 3 As shown in the example of , the volume CBT bitmap 125 has 8 bits, each of which corresponds to a different block of the volume 230N. The volume CBT bitmap 125 includes: a first bit (bit 0) having a value of zero (0), indicating that the first block of the volume 230N has not changed relative to the previous backup of the volume 230N; a second bit (bit 1) having a value of zero (0), indicating that the block directly adjacent to the block corresponding to bit 0 has not changed relative to the previous backup of the volume 230N; a third bit (bit 2) having a value of one (1), indicating that the block directly adjacent to the block corresponding to bit 1 has changed relative to the previous backup of the volume 230N, and so on.
[0075] The structure of the file mapping bitmap 127 is similar to the volume CBT bitmap 125, but each bit indicates whether a particular block of the volume 230N stores file data associated with the file 231. Figure 3 In the example of , the file mapping bitmap 127 may include a first bit (bit 0) having a value of one (1), indicating that the first block of the volume 230N stores file data associated with the file 231. The values of each of the second bit (bit 1), the fourth bit (bit 3), and the sixth bit (bit 5) of the file mapping bitmap 127 are one (1), thereby indicating that the second, fourth, and sixth blocks of the volume 230N each store file data associated with the file 231. The values of the third bit (bit 2), the fifth bit (bit 4), the seventh bit (bit 6), and the eighth bit (bit 7) of the file mapping bitmap 127 are all zero (0), indicating that the third, fifth, seventh, and eighth blocks of the volume 230N do not store file data associated with the file 231.
[0076] Thus, the CBT driver 124 can obtain the FS CBT bitmap 129 according to the bitwise logical AND of the volume CBT bitmap 125 and the file mapping bitmap 127 (wherein the bitwise logical AND can mean that each bit is individually performed logical AND between the volume CBT bitmap 125 and the file mapping bitmap 127 to generate each corresponding bit of the FS CBT bitmap 129). The result of the bitwise logical AND operation can generate the FS CBT bitmap 129, and the value of each bit of the FS CBT bitmap is one only when the values of the corresponding bits of the volume CBT bitmap 125 and the file mapping bitmap 127 are both one (1).
[0077] This bitwise logical AND operation can produce a FS CBT bitmap 129 with values of zero (0) for bits 0-2, bit 4, bit 6, and bit 7 and values of one (1) for bits 3 and 5. A value of zero (0) for each bit of the FS CBT bitmap 129 indicates that no file data is stored in the corresponding block in volume 230N, or the block may store file data associated with file 231, but the file data has not changed. A value of one (1) for each bit of the FS CBT bitmap 129 indicates that the corresponding block of volume 230N stores file data associated with file 231, and such file data has changed. In this regard, the CBT driver 124 can utilize the volume CBT bitmap 125 and the file mapping bitmap 127 to effectively recreate the FS CBT bitmap 129 in a manner that may be more reliable and efficient (even between reboots).
[0078] Figure 4 It also shows various aspects of the technology described in this disclosure. Figure 1A to Figure 2 A conceptual diagram of an example operation of a changed block tracking (CBT) driver in performing a file system CBT process in a manner utilizing volume CBT is shown in the example of FIG. Figure 3 Similar to the described example, the CBT driver 124 can obtain the volume CBT bitmap 125 and the file mapping bitmap 127 ′, but the file mapping bitmap 127 ′ is different from the file mapping bitmap 127 because the file mapping bitmap 127 ′ is defined at a different (lower) granularity than the file mapping bitmap 127 .
[0079] In general, the granularity of any given file mapping bitmap depends on the underlying file structure, which may define its block size as an integer multiple of the volume's underlying sector size. Figure 3 In the example of , this integer multiple is assumed to be one (1), so that each block size of the file system is actually equal to the underlying sector of the volume (where a sector refers to a block of the volume). Figure 4 In the example of , this integer multiple is assumed to be two (2), where each file system block (which may be called a "group") occupies two blocks (or, in other words, sectors) of the volume.
[0080] Considering that the volume CBT bitmap 125 has a higher granularity than the file mapping bitmap 127' (where the apostrophe (') indicates Figure 3 The CBT driver 124 may upsample (US) the file mapping bitmap 127' to generate an upsampled (US) file mapping bitmap 147 (also referred to as upsampled (US) file mapping information (FMI) 147 - USFMI 147) having the same granularity as the volume CBT bitmap 125. The CBT driver 124 may obtain the FS CBT bitmap 129 based on the bitwise logical AND of the volume CBT bitmap 125 and the US file mapping bitmap 147.
[0081] Figure 5 It also shows various aspects of the technology described in this disclosure. Figure 1A to Figure 2 A conceptual diagram of an example operation of a changed block tracking (CBT) driver in performing a file system CBT process in a manner utilizing volume CBT is shown in the example of FIG. Figure 4 Similar to the described example, the CBT driver 124 can obtain the volume CBT bitmap 125 ′ and the file mapping bitmap 127 , but the volume CBT bitmap 125 ′ is different from the volume CBT bitmap 125 because the volume CBT bitmap 125 ′ is defined at a different (lower) granularity than the volume CBT bitmap 125 .
[0082] Considering that the volume CBT bitmap 125' has a lower granularity than the file mapping bitmap 127, the CBT driver 124 may downsample (DS) the file mapping bitmap 127 to generate a downsampled (DS) file mapping bitmap 149 (also referred to as downsampled (DS) file mapping information (FMI) 149 - DSFMI 149) having the same granularity as the volume CBT bitmap 125'. The CBT driver 124 may obtain the FS CBT bitmap 129 based on the bitwise logical AND of the volume CBT bitmap 125 and the DS file mapping bitmap 149.
[0083] Figure 6 is a diagram showing various aspects of the CBT technology described in the present disclosure Figure 1A to Figure 2Flowchart of an example operation of a data management module shown in an example of . The data management module 120 may obtain volume changed block tracking information 125 (600), the volume changed block tracking information identifying one or more blocks of a plurality of blocks forming a volume stored on a storage device (coupled to an application server 170M), the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume. The data management module 120 may then identify a file stored on the volume for which changed block tracking is to be performed (602), and determine file mapping information 125 (604), the file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with the file. The data management module 120 may then determine file system changed block information 129 (606) based on the volume changed block tracking information 125 and the file mapping information 127, the file system changed block information identifying whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed. Data management module 120 may initiate a subsequent backup of at least a portion of the file data associated with the file based on file system changed block information 129 (608).
[0084] Although the techniques described in this disclosure are primarily described with respect to backup functions performed by agents of a data platform, similar techniques may be applied additionally or alternatively to backup, replication, cloning, or snapshot functions performed by a data platform. In such cases, archive 142 will be a backup, replica, clone, or snapshot, respectively.
[0085] Thus, once all open handles to a file are closed and the corresponding file object is destroyed, the CBT driver may lose track of the CBT information. For example, for SQL DB-level backups, the FSCBT driver loses track of changes when the DB is unloaded and more commonly between SQL server restarts. The CBT driver supports incremental backups between restarts. However, the following reasons may be more reasonable: 1) On systems where the user has both volume-level incremental backup use cases and file-level incremental backup use cases, IO performance will be lost due to having two different drives in the IO path. It may be advantageous to have one drive address both use cases. 2) Software Development Life Cycle (SDLC) cost of 2 drives relative to 1 drive. 3) Tracking file-level CBT for multiple files on a volume may be more expensive than tracking CBT information for the entire volume in terms of: a) the memory footprint required to track CBT information, because each file being tracked has additional metadata; and b) processing file metadata information (for example, the file name may need to be converted to a standardized path so that all access mechanisms to the path result in IO being tracked).
[0086] For each file-level backup:
[0087] 1) Send necessary IOCTLs (SnapshotBegin, GetBitmap, SnapshotComplete) to the volume CBT driver;
[0088] 2) When incremental file-level backup is required:
[0089] a) Get the changed blocks bitmap of the volume CBT driver.
[0090] b) On the snapshot volume, find the layout of the files. For example, on NTFS, this can be done using FSCTL_GET_RETRIEVAL_POINTERS and FSCTL_GET_RETRIEVAL_POINTER_BASE
[0091] c) Use the file layout information to create a bitmap.
[0092] d) Make the file layout bitmap have the same granularity as the volume CBT bitmap (by upsampling or downsampling the file layout bitmap). This can enable more optimizations in the process of finding the amount of data to be read. What is meant by “upsampling” and “downsampling” is as follows – the volume CBT driver tracks changes at 4K granularity / block size, and the file layout information is provided by the file system at 8K granularity. Subsequently, the file layout bitmap is converted to a bitmap that represents the data at 4K granularity. Similarly, if the file layout information is at 2K granularity, then that information is converted to a bitmap that represents the layout at 4K granularity.
[0093] e) Intersect these 2 bitmaps to find the changed blocks of the file.
[0094] 3) For data consistency, snapshots will be used (so that the file layout does not change when querying file layout information, or the file content does not change when reading data from the file, etc.).
[0095] In this regard, various aspects of the technology may implement the following examples.
[0096] Example 1. A method, the method comprising: obtaining, by processing circuitry of a computing device, volume changed block tracking information identifying one or more blocks of a plurality of blocks forming a volume of a storage device, the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determining, by the processing circuitry, file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with a file; determining, by the processing circuitry and based on the volume changed block tracking information and the file mapping information, file system changed block tracking information indicating whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed; and initiating, by the processing circuitry and based on the file system changed block tracking information, a subsequent backup of at least a portion of the file data associated with the file.
[0097] Example 2. The method of example 1, wherein obtaining the volume changed block tracking information comprises obtaining the volume changed block tracking information after restarting the computing device.
[0098] Example 3. The method of example 2, wherein the volume changed block tracking information is more resilient to reboots of the computing device in terms of accuracy and reliability than natively tracking the file system changed block information separate from the volume changed block tracking information.
[0099] Example 4. The method of any of Examples 1-3, further comprising identifying, by the processing circuit, the file stored to the volume for which changed block tracking is to be performed.
[0100] Example 5. The method of example 4, wherein identifying the file stored to the volume for which changed block tracking is to be performed comprises receiving an indication via a user interface, the indication identifying the file of the volume for which changed block tracking is to be performed.
[0101] Example 6. A method as described in any of Examples 1-5, wherein determining file mapping information includes interfacing with an application programming interface presented by a file system manager to obtain file mapping information, the file system manager managing a file system mapped to the volume stored on the storage device, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume that store the file data associated with the file.
[0102] Example 7. A method as described in Example 6, wherein the file has multiple paths for accessing the file, and wherein determining the file mapping information further includes: converting the multiple paths for accessing the file into a common file name; and interfacing with the application programming interface to pass the common file name to obtain the file mapping information.
[0103] Example 8. A method as described in any of Examples 1-7, wherein the volume changed block tracking information includes a volume changed block tracking bitmap, wherein the file mapping information includes a file mapping bitmap having a higher granularity than the volume changed block tracking bitmap, and wherein determining the file system changed block information includes downsampling the file mapping bitmap to have the same granularity as the volume changed block tracking bitmap.
[0104] Example 9. The method of any of Examples 1-8, wherein the volume changed block tracking information comprises a volume changed block tracking bitmap, wherein the file mapping information comprises a file mapping bitmap having a lower granularity than the volume changed block tracking bitmap, and
[0105] wherein determining the file system changed block information comprises upsampling the file mapping bitmap to have the same granularity as the volume changed block tracking bitmap. Example 10. The method of any one of Examples 1-9, wherein initiating the subsequent backup comprises: executing, by the processing circuit, a local agent installed on the computing device, the local agent interfacing with a remote data platform to initiate the subsequent backup of at least the portion of the file data associated with the file to the remote data platform.
[0106] Example 11. A method as described in any of Examples 1-10, wherein the file system changed block information includes a file system changed block bitmap, which identifies whether at least one of the one or more blocks storing file data associated with the file in the plurality of blocks forming the volume has been changed.
[0107] Example 12. A computing device, the computing device comprising: a storage device having a plurality of blocks forming a volume; and a processing circuit capable of accessing the storage device and configured to: obtain volume changed block tracking information, the volume changed block tracking information identifying one or more blocks of the plurality of blocks forming the volume, the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determine file mapping information, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with a file; determine file system changed block tracking information based on the volume changed block tracking information and the file mapping information, the file system changed block tracking information identifying whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed; and initiate a subsequent backup of at least a portion of the file data associated with the file based on the file system changed block tracking information.
[0108] Example 13. The computing device of Example 12, wherein to obtain the volume changed block tracking information, the processing circuit is further configured to obtain the volume changed block tracking information after restarting the computing device.
[0109] Example 14. The computing device of Example 13, wherein the volume changed block tracking information is more resilient to reboots of the computing device in terms of accuracy and reliability than natively tracking the file system changed block information separate from the volume changed block tracking information.
[0110] Example 15. The computing device of any of Examples 12-14, wherein the processing circuit is further configured to identify the file stored to the volume for which changed block tracking is to be performed.
[0111] Example 16. The computing device of Example 15, wherein to identify the file stored to the volume for which changed block tracking is to be performed, the processing circuit is configured to receive an indication via a user interface, the indication identifying the file stored to the volume for which changed block tracking is to be performed.
[0112] Example 17. A computing device as described in any of Examples 12-16, wherein to determine file mapping information, the processing circuit is configured to interface with an application programming interface presented by a file system manager to obtain file mapping information, the file system manager managing a file system mapped to the volume stored on the storage device, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume that store the file data associated with the file.
[0113] Example 18. A computing device as described in Example 17, wherein the file has multiple paths for accessing the file, and wherein in order to determine the file mapping information, the processing circuit is configured to: convert the multiple paths for accessing the file into a common file name; and interface with the application programming interface to pass the common file name to obtain the file mapping information.
[0114] Example 19. A computing device as described in any of Examples 12-18, wherein the volume changed block tracking information includes a volume changed block tracking bitmap, wherein the file mapping information includes a file mapping bitmap having a higher or lower granularity than the volume changed block tracking bitmap, wherein to determine the file system changed block information, the processing circuit is configured to downsample or upsample the file mapping bitmap to have the same granularity as the volume changed block tracking bitmap.
[0115] Example 20. A computer-readable storage medium comprising instructions that, when executed, configure processing circuitry of a computing system to: obtain volume changed block tracking information, the volume changed block tracking information identifying one or more blocks of a plurality of blocks forming a volume, the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determine file mapping information, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with a file; determine file system changed block tracking information based on the volume changed block tracking information and the file mapping information, the file system changed block tracking information identifying whether at least one of the one or more blocks of the plurality of blocks forming the volume storing file data associated with the file has changed; and initiate a subsequent backup of at least a portion of the file data associated with the file based on the file system changed block tracking information.
[0116] For the processes, devices and other examples or illustrations described herein included in any flow chart, certain operations, actions, steps or events included in any technology described herein may be performed in a different order, may be added, merged or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technology). In addition, in some examples, operations, actions, steps or events may be performed simultaneously, such as by multithreading, interrupt processing or multiple processors, rather than sequentially. In addition, even if not explicitly identified as automatically executed, certain operations, actions, steps or events may also be automatically executed. In addition, certain operations, actions, steps or events described as automatically executed may alternatively not be automatically executed, but, in some examples, such operations, actions, steps or events may be executed in response to an input or another event.
[0117] The detailed descriptions set forth herein in conjunction with the accompanying drawings are intended as descriptions of various configurations and are not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed descriptions include specific details for providing a comprehensive understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some cases, well-known structures and components are shown in block diagram form in order to avoid obscuring such concepts.
[0118] According to one or more aspects of the present disclosure, when the context does not specify otherwise, the term "or" may be interrupted to "and / or". In addition, although phrases such as "one or more" or "at least one" may be used in some cases; however, if the context does not specify otherwise, those cases where such language is not used may be interpreted as implying such meaning.
[0119] In one or more examples, the described functions may be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on and / or transmitted through a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include: computer-readable storage media, which corresponds to tangible media, such as data storage media; or communication media, which includes any media that facilitates the transfer of a computer program from one place to another (e.g., in accordance with a communication protocol). Thus, computer-readable media may generally correspond to (1) tangible computer-readable storage media, which is non-transitory, or (2) communication media, such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present disclosure. A computer program product may include a computer-readable medium.
[0120] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, flash memory, or any other medium that can be used to store the desired program code in the form of an instruction or data structure and can be accessed by a computer. In addition, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using coaxial cable, optical cable, twisted pair, digital subscriber line (DSL) or wireless technology such as infrared, radio and microwave, then coaxial cable, optical cable, twisted pair, DSL or wireless technology such as infrared, radio and microwave are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals or other transient media, but are instead directed to non-transient tangible storage media. Disks and optical disks such as those used include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks and blue-ray disks, where disks typically reproduce data magnetically, and optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0121] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" or "processing circuit" as used herein may each refer to any of the aforementioned structures or any other structure suitable for implementing the described techniques. In addition, in some examples, the described functionality may be provided in dedicated hardware and / or software modules. Furthermore, the described techniques may be fully implemented in one or more circuits or logic elements.
[0122] As used herein, a processing device may include processing circuitry as described above. In some examples, the processing device may include at least one processor and at least one memory having computer code, the computer code including a set of instructions that, when executed by the at least one processor, causes the at least one processor to perform any of the functions described herein. In some examples, the processing device may receive computer code including the set of instructions from at least one memory coupled to the processing device.
[0123] The technology of the present disclosure can be implemented in various devices or apparatuses, including wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs) or a set of ICs (e.g., chipsets). The present disclosure describes various components, modules, or units to emphasize the functional methods of devices configured to perform the disclosed technology, but does not necessarily require implementation by different hardware units. Instead, as described above, the various units can be combined in a hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) in combination with appropriate software and / or firmware.
Claims
1. A method, comprising: obtaining, by a processing device of a computing device, volume changed block tracking information, the volume changed block tracking information identifying one or more blocks of a plurality of blocks forming a volume of a storage device, the one or more blocks storing updated data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determining, by the processing device, file mapping information, the file mapping information identifying one or more blocks of the plurality of blocks forming the volume storing file data associated with a file; determining, by the processing device and based on the volume changed block tracking information and the file mapping information, file system changed block tracking information indicating whether at least one of the one or more blocks storing file data associated with the file in the plurality of blocks forming the volume has changed; as well as A subsequent backup of at least a portion of the file data associated with the file is initiated, by the processing device and based on the file system changed block tracking information.
2. The method of claim 1, wherein obtaining the volume changed block tracking information comprises obtaining the volume changed block tracking information after restarting the computing device.
3. The method of claim 1, further comprising identifying, by the processing device, the file stored to the volume for which changed block tracking is to be performed.
4. The method of claim 1 , wherein determining file mapping information comprises interfacing with an application programming interface presented by a file system manager that manages a file system mapped to the volume stored on the storage device to obtain file mapping information, the file mapping information identifying the one or more blocks of the plurality of blocks forming the volume that store the file data associated with the file.
5. The method according to claim 4, wherein the file has multiple paths to access the file, and Wherein determining the file mapping information further comprises: converting the plurality of paths for accessing the file into a common file name; as well as The application programming interface is interfaced to pass the generic file name to obtain the file mapping information.
6. The method according to claim 1, wherein the volume changed block tracking information comprises a volume changed block tracking bitmap, wherein the file mapping information comprises a file mapping bitmap having a higher granularity than the volume changed block tracking bitmap, and Wherein determining the file system changed block information comprises downsampling the file mapping bitmap to have the same granularity as the volume changed block tracking bitmap.
7. The method according to claim 1, wherein the volume changed block tracking information comprises a volume changed block tracking bitmap, wherein the file mapping information comprises a file mapping bitmap having a lower granularity than the volume changed block tracking bitmap, and Wherein determining the file system changed block information comprises upsampling the file mapping bitmap to have the same granularity as the volume changed block tracking bitmap.
8. A method as claimed in claim 1, wherein initiating the subsequent backup includes executing a local agent installed on the computing device through the processing device, and the local agent interfaces with a remote data platform to initiate the subsequent backup of at least the portion of the file data associated with the file to the remote data platform.
9. The method of claim 1, wherein the file system changed block information comprises a file system changed block bitmap that identifies whether at least one of the one or more blocks storing file data associated with the file in the plurality of blocks forming the volume has been changed.
10. A computing device, comprising: a storage device having a plurality of blocks forming a volume; as well as a processing device having access to the storage device and configured to: obtaining volume changed block tracking information, the volume changed block tracking information identifying one or more blocks of the plurality of blocks forming the volume, the one or more blocks storing update data that has changed relative to a previous backup of the one or more blocks of the plurality of blocks forming the volume; determining file mapping information that identifies one or more blocks of the plurality of blocks forming the volume that store file data associated with a file; determining file system changed block tracking information based on the volume changed block tracking information and the file mapping information, the file system changed block tracking information identifying whether at least one of the one or more blocks storing file data associated with the file in the plurality of blocks forming the volume has changed; and A subsequent backup of at least a portion of the file data associated with the file is initiated based on the file system changed block tracking information.
11. The computing device of claim 10, wherein in order to obtain the volume changed block tracking information, the processing device is further configured to obtain the volume changed block tracking information after restarting the computing device.
12. The computing device of claim 10, wherein the processing device is further configured to identify the file stored to the volume for which changed block tracking is to be performed.
13. The computing device of claim 10 , wherein to determine file mapping information, the processing device is configured to interface with an application programming interface presented by a file system manager that manages a file system mapped to the volume stored on the storage device to obtain file mapping information, the file mapping information identifying the one or more blocks of the plurality of blocks forming the volume that store the file data associated with the file.
14. A computing device as claimed in any one of claims 10 to 13, wherein the volume changed block tracking information comprises a volume changed block tracking bitmap, wherein the file mapping information comprises a file mapping bitmap having a higher or lower granularity than the volume changed block tracking bitmap, and To determine the file system changed block information, the processing circuit is configured to downsample or upsample the file mapping bitmap to have the same granularity as the volume changed block tracking bitmap.
15. A computer-readable storage medium comprising instructions which, when executed by one or more processors, configure the one or more processors to perform the method of any one of claims 1 to 9.