File system metadata deduplication
By deduplication of multiple file metadata structures in the storage cluster, and identifying and deleting duplicate metadata elements that do not have the highest reference count, the problem of waste of SSD storage space caused by duplicate metadata in the storage cluster is solved, and storage efficiency and processor performance are improved.
Patent Information
- Application Number
- CN202080032635.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-21
- Filing Date
- 2020-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-04-29
AI Technical Summary
There are duplicate file metadata storage problems in existing storage clusters, resulting in inefficient use of SSD storage space and the inability to effectively reduce the amount of metadata associated with multiple files stored.
By deduplication of multiple file metadata structures in the storage cluster, duplicate metadata elements that do not have the highest reference count, retain the metadata elements that have the highest reference count, and modify its parent node to reference the metadata element, reducing the number of writes.
It effectively reduces the SSD storage space of the storage cluster, improves storage efficiency, and reduces the number of times the processor writes to metadata repeatedly.
Smart Images

Figure CN113767378B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 840,614, filed Apr. 30, 2019, entitled FILE SYSTEM METADATA DEDUPLICATION (Attorney Docket No. COHEP042+), which is incorporated herein by reference in its entirety for all purposes. BACKGROUND OF THE INVENTION
[0004] Data associated with files and metadata of files are typically stored in a storage cluster. The data and metadata can be generated by the storage cluster and / or can be generated as part of a process for backing up files and / or restoring files. File operation requests such as reads, writes, deletes involve accessing data in the storage cluster and incur processor clock cycles as well as accesses to solid - state and / or hard disk memories. It is desirable to store data as efficiently as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.
[0006] Figure 1 is a block diagram showing a system for file system metadata deduplication according to some embodiments.
[0007] Figure 2A is a block diagram showing an embodiment of a tree - type data structure.
[0008] Figure 2B is a block diagram showing an embodiment of a cloned file system metadata snapshot tree.
[0009] Figure 2C is a block diagram showing an embodiment of modifying a file system metadata snapshot tree.
[0010] Figure 2D is a block diagram showing an embodiment of a modified snapshot tree.
[0011] Figure 3A is a block diagram showing an embodiment of a tree - type data structure.
[0012] Figure 3B is a block diagram showing an embodiment of adding a file metadata structure to a tree - type data structure.
[0013] Figure 3C is a block diagram showing an embodiment of modifying a file metadata structure.
[0014] Figure 3D is a block diagram showing an embodiment of a modified file metadata structure.
[0015] Figure 4A A block diagram showing an example of duplicate metadata according to some embodiments.
[0016] Figure 4B A block diagram showing an example of deduplicated metadata according to some embodiments.
[0017] Figure 4C A block diagram showing an example of duplicate metadata according to some embodiments.
[0018] Figure 4D A block diagram showing an example of deduplicated metadata according to some embodiments.
[0019] Figure 5 A flowchart showing a process for deduplicating metadata associated with multiple files according to some embodiments.
[0020] Figure 6 A flowchart showing a process for deduplicating metadata elements according to some embodiments.
[0021] Figure 7 A flowchart showing a process for deduplicating metadata associated with multiple files according to some embodiments.
[0022] Figure 8 A flowchart showing a process for deduplicating metadata elements according to some embodiments. Detailed Description
[0023] A storage cluster may be configured to store multiple files. The storage cluster may be configured to back up multiple files stored on a primary system. The storage cluster may be configured to store multiple files generated on or by the storage cluster (e.g., system-generated files, user-generated files, application-generated files, etc.). Regardless of the file source, the storage cluster may be configured to organize metadata associated with the files using a tree data structure.
[0024] The tree data structure of metadata associated with a file may be referred to as a "file metadata tree" or "file metadata structure". The tree data structure of metadata associated with a file may include a root node, one or more levels of one or more intermediate nodes, and a plurality of leaf nodes. Each node of the file metadata structure may be referred to as a "metadata element". The leaf nodes of the file metadata structure may store key-value pairs (KVPs). The value of a KVP may include a brick identifier associated with one or more data blocks of the file. A data brick may be associated with one or more block identifiers (e.g., SHA-1 hash values). A block file may be configured to store a plurality of data blocks. A block metadata table may store information associating a brick identifier with one or more block identifiers and one or more block file identifiers. A block file metadata table may associate a block file identifier with a block file storing a plurality of data blocks. The block metadata table and the block file metadata table may be used based on the brick identifier to locate the data blocks associated with the file corresponding to the file metadata structure.
[0025] A storage cluster may include a plurality of storage nodes. Each storage node may include a corresponding processor and a plurality of storage layers. For example, a first storage layer may include one or more solid state drives (SSDs), and a second storage layer may include one or more hard disk drives (HDDs) and solid state drives (SSDs). The storage associated with the first storage layer may be faster than the storage associated with one or more other storage layers. The metadata associated with a plurality of files may be stored in the first storage layer of the storage cluster, such as an SSD, while the data associated with a plurality of files, particularly the one or more data blocks mentioned above, may be stored in the first storage layer or the second storage layer (e.g., SSD or HDD) of the storage cluster. The storage cluster may receive one or more file operation requests (e.g., read, write, delete) for data associated with a plurality of files. The metadata associated with a plurality of files is stored in the first storage layer (SSD) of the storage cluster so that the processor of the storage cluster processing requests can quickly locate the data associated with the files.
[0026] Multiple files stored in a storage cluster may include duplicate data blocks, and the duplicate data blocks may be deduplicated to reduce the storage amount used to store the multiple files. This frees up storage space on HDDs and SSDs to store data associated with one or more other files. Although this enables data associated with and between files to be deduplicated, the storage cluster may separately store duplicate metadata associated with multiple files. For example, the storage cluster may store multiple file metadata structures that include corresponding leaf nodes storing the same value, e.g., because the leaf nodes point to the same data brick via the same brick identifier. This scenario is inefficient because multiple leaf nodes reference the same data brick. The storage amount available in the storage cluster, especially in SSDs, is limited, and efficient use of storage for both SSDs and HDDs is desired.
[0027] In embodiments disclosed herein, the storage amount used to store metadata associated with multiple files may be reduced by deduplicating the metadata associated with the multiple files. In embodiments described herein, one or more duplicate metadata elements (e.g., duplicate leaf nodes) are removed from the SSD. To identify potentially duplicate metadata elements, the bottom two levels of the multiple file metadata structures stored by the storage cluster may be scanned. In this way, those leaf nodes storing the same brick identifier may be identified. The bottom two levels of the file metadata structure correspond to the levels referred to herein as the leaf node level and the lowest intermediate node level (i.e., the node that references the leaf nodes included in the leaf node level). Leaf nodes storing the same brick identifier may be associated with different files and / or different versions of the same file.
[0028] It is possible to maintain a reference count for at least some of the leaf nodes, where the reference count indicates the number of other nodes that reference the leaf node, such as pointers to leaf nodes. Identify the leaf node that stores the same value (e.g., brick identifier) as one or more other leaf nodes and has the highest reference count. Among those leaf nodes that store the same value, select and retain the node with the highest reference count, and modify any one or more parent nodes (e.g., intermediate nodes included in the lowest intermediate node level) associated with one or more leaf nodes that do not have the highest reference count to reference the leaf node with the highest reference count. Then delete one or more leaf nodes that do not have the highest reference count. Retaining the leaf node with the highest reference count can reduce the number of writes required by the processors of the storage cluster to deduplicate metadata associated with multiple files. For example, two leaf nodes may store the same brick identifier. The first leaf node may have a reference count of 4, and the second leaf node may have a reference count of 1. Modifying the parent node of the second leaf node to reference the first leaf node will require a single write operation, while modifying the parent node of the first leaf node to reference the second leaf node will require four separate write operations.
[0029] In other embodiments, remove one or more duplicate portions of the file metadata structure. A leaf node associated with at least a portion of the file metadata structure may store a particular sequence of values (e.g., brick identifiers). For example, a first file metadata structure may include four leaf nodes, where the first leaf node stores the brick identifier "1", the second leaf node stores the brick identifier "2", the third leaf node stores the brick identifier "3", and the fourth leaf node stores the brick identifier "4". In this instance, the particular sequence of values for the first file metadata structure is "1234". The particular sequence of values may be referred to as a "tree fingerprint". A breadth-level search of the leaf node level of the file metadata structure may be performed to identify the tree fingerprint associated with the file metadata structure. In some embodiments, a portion of the tree fingerprint associated with a first file metadata structure is the same as a portion of the tree fingerprint associated with the same or one or more other file metadata structures.
[0030] For example, the tree fingerprint associated with the first file metadata structure may be "1234", the tree fingerprint associated with the second file metadata structure may be "5634", and the tree fingerprint associated with the third file metadata structure may be "8934". The tree fingerprints associated with the first, second, and third file metadata structures share a common value sequence (e.g., a brick identifier). In this instance, the common value sequence is "34". So-called common nodes associated with the common sequence can be identified. A common node may include a direct or indirect reference to a leaf node associated with the common sequence and does not include a direct or indirect reference to a leaf node not associated with the common sequence. If a node is one level higher than a leaf node associated with the common sequence, the node may include a direct reference to a leaf node associated with the common sequence. If a node is two or more levels higher than a leaf node associated with the common sequence, the node may include an indirect reference to a leaf node associated with the common sequence. For example, an intermediate node associated with a first-level intermediate node may include a reference to an intermediate node associated with a second-level intermediate node. The intermediate node associated with the second-level intermediate node may include a reference to a leaf node. The intermediate node associated with the first-level intermediate node includes an indirect reference to a leaf node.
[0031] A reference count can be maintained for each of the common nodes associated with the common sequence. The reference count indicates the number of other nodes that reference the common node, e.g., pointers including the common node. The common node with the highest reference count can be identified. Among those common nodes associated with the common sequence, the common node with the highest reference count is selected and retained, and any parent nodes of one or more other common nodes that do not have the highest reference count are modified to reference the common node with the highest reference count. One or more common nodes that do not have the highest reference count are deleted, and one or more nodes directly or indirectly referenced by one or more common nodes that do not have the highest reference count are deleted. Retaining the common node with the highest reference count can reduce the number of writes required by a processor of a storage cluster to deduplicate metadata associated with multiple files. For example, two common nodes may be associated with the same common sequence. The first common node may have a reference count of 4, and the second common node may have a reference count of 1. Modifying the parent node of the second common node to reference the first common node will require a single write operation, while modifying the parent node of the first common node to reference the second common node will require four separate write operations.
[0032] It can be used as a background process of the storage cluster to deduplicate metadata associated with multiple files stored in the storage cluster. The metadata deduplication process can be scheduled when the storage cluster has available resources, such as when the storage cluster is not performing a backup of the primary storage system. The metadata deduplication techniques described herein can allow the storage cluster to reclaim valuable SSD storage without burdening the system resources of the storage cluster.
[0033] Figure 1 FIG. 4 is a block diagram illustrating a system for file system metadata deduplication according to some embodiments. In the illustrated example, system 100 includes a primary system 102 and a storage cluster 112.
[0034] The primary system 102 is a computing system that stores file system data. The file system data can be stored in a storage volume 104. The file system data can be stored across one or more objects, one or more virtual machines, one / more physical entities, one or more file systems, one or more array backups, and / or one or more volumes of the primary system 102. The file system data can include one or more files (e.g., content files, text files). The primary system 102 can include one or more servers, one or more computing devices, one or more storage devices, and / or combinations thereof.
[0035] The file system data stored on the primary system 102 can include one or more data blocks. The primary system 102 can be configured to perform agent-based tracking, for example, by means of a change block tracker 105, which monitors one or more data blocks and stores an indication of when one of the one or more data blocks has been modified. The change block tracker 105 can receive one or more data blocks associated with one or more files in transit in one or more objects, one or more virtual machines, one / more physical entities, one or more file systems, one or more array backups, and / or one or more volumes of the primary system 102. The change block tracker can be configured to maintain a mapping of one or more changes to the file system data. The mapping can include one or more data blocks that have been changed, values associated with the one or more changed data blocks, and associated timestamps. If the primary system 102 performs a backup snapshot (full or incremental), the change block tracker 105 is configured to clear (e.g., empty) the mapping of one or more data blocks that have been modified. As an alternative to agent-based tracking, changes can be identified via, for example, APIs and appropriate function calls (so-called "agentless" tracking).
[0036] A backup snapshot can be initiated by the backup agent 106 or by means of an instruction received, for example, via an API; once initiated, the primary system 102 sends the file system data stored in the storage volume 104 to the storage cluster 112. The backup snapshot can be a full backup snapshot or an incremental backup snapshot. The backup agent 106 can receive a command to execute the backup snapshot from the storage cluster 112. The primary system 102 is coupled to the storage cluster 112 via the network connection 110. The connection 110 can be a wired connection or a wireless connection.
[0037] The storage cluster 112 is a storage system configured to extract and store the file system data received from the primary system 102 via the connection 110. The storage cluster 112 can include one or more storage nodes 111, 113, 115. Each storage node can include a corresponding processor and multiple storage layers. For example, as described above, the first storage layer can include one or more SSDs, and the second storage layer can include one or more HDDs and SSDs. The file system data included in the backup snapshot can be stored in one or more of the storage nodes 111, 113, 115. In some embodiments, one or more storage nodes store one or more copies of the file system data. In one instance, the storage cluster 112 includes one solid state drive SSD and three hard disk drives HDD. The storage cluster 112 can receive one or more file operation requests (e.g., read, write, delete) for data associated with multiple files. Metadata associated with the multiple files is stored in the SSD of the storage cluster so that the processor of the storage cluster processing the request can quickly locate the data associated with the file.
[0038] The storage cluster 112 can include a file system manager 117. The file system manager 117 can be configured to organize the file system data received in a backup snapshot from the primary system 102 into a tree - type data structure. An example of the tree - type data structure is a file system metadata snapshot tree (e.g., Cohesity ) This can be based on a B+ tree structure (or other types of tree structures in other embodiments). The tree data structure provides a view of the file system data corresponding to the backup snapshot. The view of the file system data corresponding to the backup snapshot may include a file system metadata snapshot tree and a plurality of file metadata structures. The file metadata structure may correspond to one of the files included in the backup snapshot. The file metadata structure stores the metadata associated with the file. The file system manager 117 may be configured to perform one or more modifications to the file system metadata snapshot tree and the file metadata structure as disclosed herein. The file system metadata snapshot tree and the file metadata structure may be stored in the metadata store 114. The metadata store 114 may store the view of the file system data corresponding to the backup snapshot. The metadata store 114 may also store data associated with content files smaller than a limit size (e.g., 256 kB). The metadata store 114 may span the SSDs of the storage nodes 111, 113, 115.
[0039] The tree data structure can be used to capture different versions of the backup snapshot. The tree data structure allows a series of file system metadata snapshot trees corresponding to different versions of the backup snapshot (i.e., different versions of the file system metadata snapshot tree) to be linked together (e.g., a "snapshot tree forest") by allowing nodes of the file system metadata snapshot tree of a later version to reference nodes of the file system metadata snapshot tree of a previous version. For example, the root node or an intermediate node of the second file system metadata snapshot tree corresponding to the second backup snapshot may reference an intermediate node or a leaf node of the first file system metadata snapshot tree corresponding to the first backup snapshot.
[0040] The file system metadata snapshot tree is a representation of a fully - integrated backup because it provides a complete view of one or more storage volumes at a particular moment in time. A fully - integrated backup is a backup that can be used without reconstruction. Other systems that do not maintain fully - integrated backups can reconstruct the backup by starting from or restoring a full backup and applying one or more changes associated with one or more incremental backups to the data associated with the full backup. In contrast, in the presence of a file system metadata snapshot tree, any file stored in a storage volume at a particular time and the content of the file (for which there is an associated backup) can be determined from the file system metadata snapshot tree, regardless of whether the associated backup snapshot is a full - backup snapshot or an incremental - backup snapshot. Creating an incremental - backup snapshot can include only copying data that has not been previously backed up on one or more storage volumes. The file system metadata snapshot tree corresponding to the incremental - backup snapshot provides a complete view of one or more storage volumes at a particular moment in time because the file system metadata snapshot tree includes references to the previously stored data of the storage volume. For example, the root node associated with the file system metadata snapshot tree can include one or more references to leaf nodes associated with one or more previous backup snapshots and one or more references to leaf nodes associated with the current backup snapshot. This provides a significant savings in the amount of time required to restore or recover a storage volume and / or database. In contrast, other recovery / restoration methods may require a large amount of time, storage, and computational resources to reconstruct a particular version of a volume or database from a full backup and a series of incremental backups.
[0041] The storage cluster 112 can store a set of one or more file system metadata snapshot trees. Each file system metadata snapshot tree can correspond to a particular moment in time associated with the state of the file system data of the primary system 102. The file system metadata snapshot tree can include a root node, one or more levels of one or more intermediate nodes associated with the root node, and one or more leaf nodes associated with the intermediate nodes of the lowest intermediate level. The root node of the file system metadata snapshot tree can include one or more pointers to one or more intermediate nodes. Each intermediate node can include one or more pointers to other nodes (e.g., lower - level intermediate nodes or leaf nodes).
[0042] A leaf node may store file system metadata, data associated with a file smaller than a limit size, an identifier of a data brick, a pointer to a data block stored on a storage cluster, or a pointer to another file metadata structure. For example, data associated with a file that is less than or equal to a limit size (e.g., 256 kB) may be stored in a leaf node of a file system metadata snapshot tree. Optionally, if the data associated with a file is greater than or equal to the limit size, the leaf node may include a pointer to the aforementioned data block stored by the storage cluster. As another alternative, an additional file metadata structure may be generated for a file that is greater than the limit size, in which case the leaf node includes a pointer to the additional file metadata structure. In these latter two instances, the leaf node may be referred to as an index node (inode).
[0043] A file metadata structure may include a root node, one or more levels of one or more intermediate nodes associated with the root node, and one or more leaf nodes associated with the intermediate nodes of the lowest intermediate level. The file metadata structure is configured to store metadata associated with a particular version of a file. The tree-like data structure associated with the file metadata structure allows a series of file metadata structures corresponding to different versions of a file to be linked together by allowing nodes of a later version of the file metadata structure to reference nodes of a previous version of the file metadata structure. For example, the root node or an intermediate node of a second file metadata structure corresponding to a second version of a file may reference an intermediate node or a leaf node of a first file metadata structure corresponding to a first version of the file. A file metadata structure may be associated with a plurality of block files. A block file may include a plurality of file segment data blocks. The storage cluster 112 may store a set of one or more file metadata structures. Each file metadata structure may correspond to a file. In other embodiments, the file metadata structure corresponds to a portion of a file.
[0044] As mentioned above, a leaf node of a file metadata structure may store values such as an identifier of a data brick associated with one or more data blocks. For example, also as described above, a file metadata structure may correspond to a file, and a leaf node of the file metadata structure may include a pointer or an identifier of a data brick associated with one or more data blocks of the file. A data brick may be associated with one or more data blocks. In some embodiments, the size of a brick is 256 kB. One or more data blocks may have a variable length within a particular range (e.g., 4 kB to 64 kB).
[0045] The locations of one or more data blocks associated with a data brick can be identified using one or more data structures (e.g., lists, tables, etc.) stored in the metadata store 114. In one embodiment, a first data structure (e.g., a block metadata table) stores information that associates a brick identifier with one or more block identifiers and one or more block file identifiers. A second data structure (e.g., a block file metadata table) associates a block file identifier with a block file that stores a plurality of data blocks. In some embodiments, the first data structure and the second data structure are combined into a single data structure. One or more data blocks associated with a data brick can be located based on the block metadata table and the block file metadata table. For example, a first data brick having a first brick identifier can be associated with a first block identifier, e.g., a SHA-1 hash value, which is a hash of the content of the associated block file. A block file having the identified block file identifier can include a plurality of data blocks, and the block file metadata table can be used to identify the locations of the plurality of data blocks. For example, the block file metadata table can include offset information for the plurality of data blocks within the block file.
[0046] In some embodiments, the storage cluster 112 is configured to back up a plurality of files stored on the primary system and generate a corresponding file metadata structure for each of the plurality of files. In other embodiments, the storage cluster 112 is configured to store a plurality of files generated by the storage cluster 112 and generate a corresponding file metadata structure for each of the plurality of files. Different versions of files included in the plurality of backups can be configured to share nodes. As discussed above, nodes associated with a later version of a file can reference nodes associated with a previous version of the file. In contrast, file metadata structures corresponding to files generated by the storage cluster 112 can be independent of each other, i.e., file metadata structures corresponding to files generated by the storage cluster 112 do not share nodes. In either case, the storage cluster 112 may store duplicate metadata elements (e.g., nodes).
[0047] As explained above, file metadata structures stored by the storage cluster 112 can be stored in the SSDs of the storage cluster 112, and the amount of storage available in the SSDs of the storage cluster is limited. Storing duplicate metadata elements is an inefficient use of the SSDs. For example, a plurality of file metadata structures can include leaf nodes that store the same value (e.g., referencing the same data brick). The amount of storage used by the SSDs of the storage cluster 112 to store metadata associated with a plurality of files can be reduced by deduplicating the metadata associated with the plurality of files.
[0048] In some embodiments, one or more duplicate metadata elements (e.g., duplicate leaf nodes) are removed from the SSD. To identify potentially duplicate metadata elements, the storage cluster 112 can partially identify one or more duplicate metadata elements by scanning the bottom two levels of the plurality of file metadata structures stored by the storage cluster 112. As explained above, the bottom two levels of the file metadata structure correspond to the leaf node level and the lowest intermediate node level (i.e., the nodes that reference the leaf nodes included in the leaf node level). Those metadata elements that store the same value (e.g., brick identifier) can be associated with different files and / or different versions of the same file. The storage cluster 112 can then identify the metadata elements that store the same value, that is, instances of the metadata elements. In some instances, the plurality of file metadata structures corresponding to the files included in the backup snapshot received from the primary system 102 include metadata elements that store the same value. In other instances, one or more file metadata structures corresponding to the files included in the backup snapshot received from the primary system 102, and one or more file metadata structures corresponding to the files generated by the storage cluster 112 include metadata elements that store the same value. In other instances, the plurality of file metadata structures corresponding to the files generated by the storage cluster 112 can include metadata elements that store the same value.
[0049] Maintain a reference count for each metadata element that stores the same value as one or more other metadata elements. The reference count indicates the number of other nodes that reference one or more metadata elements, such as pointers to leaf nodes. For example, the metadata elements of the file metadata structure corresponding to the files included in the backup snapshot can be referenced by the nodes included in the plurality of file metadata structures (e.g., different versions of the same file). In contrast, the metadata elements of the file metadata structure corresponding to the files generated by the storage cluster 112 may not be referenced by other file metadata structures because the file metadata structures corresponding to the files generated by the storage cluster 112 are stored independently (i.e., not linked together).
[0050] Metadata associated with multiple files can be partially deduplicated by identifying the metadata element (e.g., leaf node) that stores the same value (e.g., brick identifier) as one or more other metadata elements and has the highest reference count. For metadata elements that store the same value, the metadata element with the highest reference count is selected and retained, and any parent metadata elements associated with the metadata elements that do not have the highest reference count (e.g., nodes that include pointers to the metadata elements that do not have the highest reference count) are modified to reference the metadata element with the highest reference count. Then, the metadata elements that do not have the highest reference count are deleted. Retaining the metadata element with the highest reference count can reduce the number of writes required by the processors of storage cluster 112 to deduplicate the metadata associated with multiple files. In other embodiments, multiple metadata elements that store the same value have the same reference count, i.e., there is no single metadata element with the highest reference count. In this case, the metadata associated with multiple files can be deduplicated by selecting and retaining one of the metadata elements; then, the parent metadata elements associated with one or more of the unselected metadata elements can be modified to reference the selected metadata element, while deleting the one or more unselected metadata elements.
[0051] In other embodiments, metadata associated with multiple files is deduplicated by removing one or more duplicate portions of the file metadata structure from the SSDs of storage cluster 112. The metadata elements associated with at least a portion of the file metadata structure can store a particular sequence of values (e.g., brick identifier). For example, the file metadata structure can include four leaf nodes, where the first leaf node stores the brick identifier "1", the second leaf node stores the brick identifier "2", the third leaf node stores the brick identifier "3", and the fourth leaf node stores the brick identifier "4". In this instance, the particular sequence of values for the file metadata structure is "1234". The particular sequence of values can be referred to as a "tree fingerprint". Storage cluster 112 can perform a breadth-level search at the leaf node level for each of the multiple file metadata structures to identify the corresponding tree fingerprint associated with the file metadata structure. In some embodiments, a portion of the tree fingerprint associated with the first file metadata structure is the same as a portion of the tree fingerprint associated with the same or one or more other file metadata structures.
[0052] For example, the tree fingerprint associated with the first file metadata structure may be "1234", the tree fingerprint associated with the second file metadata structure may be "5634", and the tree fingerprint associated with the third file metadata structure may be "8934". The tree fingerprints associated with the first, second, and third file metadata structures share a common value sequence (e.g., a brick identifier). In this example, the common value sequence is "34". The so-called common metadata elements (also referred to as common nodes) associated with the common sequence can be identified. A common metadata element can be a node that includes a direct or indirect reference to a leaf node associated with the common sequence and does not include a direct or indirect reference to a leaf node not associated with the common sequence. If a node is one level higher than a leaf node associated with the common sequence, the node may include a direct reference to a leaf node associated with the common sequence. If a node is two or more levels higher than a leaf node associated with the common sequence, the node may include an indirect reference to a leaf node associated with the common sequence. A common metadata element can be an intermediate node.
[0053] A reference count associated with the corresponding common metadata element can be maintained. The reference count indicates the number of other metadata elements that reference the common metadata element, such as a pointer including the common node. The common metadata element with the highest reference count can be identified. Among those common metadata elements associated with the common sequence, the common metadata element with the highest reference count is selected and retained, and one or more corresponding parent metadata elements of one or more other common metadata elements are modified to reference the common metadata element with the highest reference count. One or more unselected common metadata elements are deleted, and the metadata elements directly or indirectly referenced by one or more unselected common metadata elements are deleted. Retaining the common metadata element with the highest reference count can reduce the number of writes required by the processors of the storage cluster to deduplicate the metadata associated with multiple files. In other embodiments, multiple common metadata elements have the same reference count, i.e., there is no single common metadata element with the highest reference count. In this case, the metadata associated with multiple files can be deduplicated by selecting and retaining one of the common metadata elements; then the parent metadata elements associated with one or more unselected common metadata elements can be modified to reference the selected common metadata element, while deleting the metadata elements referenced by one or more unselected common metadata elements.
[0054] In some embodiments, there is no single common metadata element associated with a common sequence of values because at least one of the common metadata elements of multiple file metadata structures includes a direct or indirect reference to a leaf node that is not associated with the common sequence. For example, a first file metadata structure may have a tree fingerprint of "12345678", and a second file metadata structure may have a tree fingerprint of "92345678". In this instance, the common sequence of values is "2345678". A first portion of the first file metadata structure may be associated with a leaf node having a value sequence of "1234", and a second portion of the first file metadata structure may be associated with a leaf node having a value sequence of "5678". The first portion of the first file metadata structure may be associated with a first intermediate node, and the second portion of the first file metadata structure may be associated with a second intermediate node, where the first intermediate node and the second intermediate node are associated with different branches of the first file metadata structure. A first portion of the second file metadata structure may be associated with a leaf node having a value sequence of "9234", and a second portion of the second file metadata structure may be associated with a leaf node having a value sequence of "5678". The first portion of the second file metadata structure may be associated with a first intermediate node, and the second portion of the second file metadata structure may be associated with a second intermediate node, where the first intermediate node and the second intermediate node are associated with different branches of the second file metadata structure.
[0055] For such cases, the common sequence of values may be divided into multiple portions, and the corresponding common metadata elements of the multiple portions may be determined. In this instance, the common metadata element for the value sequence "234" may be determined, and the common metadata element for the value sequence "5678" may be determined. The metadata associated with the file metadata structure may be deduplicated based on the determined common metadata elements. If the determined common metadata elements include one or more references to metadata elements that are not part of the common sequence of values, the metadata associated with the multiple files may be deduplicated at the leaf node level as described above.
[0056] Figures 2A to 2D Details of an exemplary tree data structure are provided, and the embodiments disclosed herein may be advantageously configured to operate under the tree data structure. Figure 2A is a block diagram showing an embodiment of a tree data structure that represents file system data stored on a storage cluster such as storage cluster 112. The file system data may include metadata of a distributed file system and may include information such as block identifiers, block offsets, file sizes, directory structures, file permissions, physical storage locations of files, etc. A file system manager such as file system manager 117 may generate the tree data structure 200.
[0057] The tree - type data structure 200 includes a file - system metadata snapshot tree, and the file - system metadata snapshot tree includes a root node 202, intermediate nodes 212, 214, and leaf nodes 222, 224, 226, 228, and 230. Although the tree - type data structure 200 includes an intermediate level between the root node 202 and the leaf nodes 222, 224, 226, 228, 230, any number of intermediate levels can be implemented. The tree - type data structure 200 can correspond to a backup snapshot of the file - system data at a specific time point t, for example, at time t = 1. The backup snapshot can be received from the primary system at the storage cluster. The file - system metadata snapshot tree, in combination with multiple file metadata structures, can provide a complete view of the primary system at a specific time point.
[0058] The root node is the starting point of the file - system metadata snapshot tree and can include pointers to one or more other nodes. An intermediate node is a node that is pointed to by another node (e.g., the root node, another intermediate node) and includes one or more pointers to one or more other nodes. A leaf node is a node at the bottom level of the file - system metadata snapshot tree. Each node of the tree structure includes a view identifier (e.g., tree ID) of the view associated with the node.
[0059] The leaf nodes can be configured to store key - value pairs of the file - system data. The data key k is a lookup value that can be used to access a specific leaf node. For example, "1" can be used as the data key to look up "Data 1" in leaf node 222. The data key k can correspond to the brick identifier (e.g., brick number) of a data brick. The data brick can be associated with one or more data blocks. In some embodiments, the leaf nodes are configured to store file - system metadata (e.g., block identifier (e.g., hash value, SHA - 1, etc.), file size, directory structure, file permissions, physical storage location of the file, etc.). The leaf node can store the data key k and a pointer to the location storing the value associated with the data key. In other embodiments, the leaf nodes are configured to store the actual data when the data associated with the file is less than or equal to a limit size (e.g., 256 kb). As mentioned above, in some instances, when the size of the file is greater than the limit size, the leaf node includes a pointer to the file metadata structure.
[0060] A root node or an intermediate node may include one or more node keys. The node keys may be integer values or non-integer values. Each node key indicates a division between branches of the node and indicates how to traverse the tree structure to find a leaf node, i.e., which pointer to follow. For example, the root node 202 may include the node key "3". The data key k of the key-value pair that is less than or equal to the node key is associated with the first branch of the node, and the data key k of the key-value pair that is greater than the node key is associated with the second branch of the node. In the above example, to find the leaf node storing the value associated with the data key "1", "2", or "3", the first branch of the root node 202 will be traversed to the intermediate node 212 because the data keys "1", "2", and "3" are less than or equal to the node key "3". To find the leaf node storing the value associated with the data key "4" or "5", the second branch of the root node 202 will be traversed to the intermediate node 214 because the data keys "4" and "5" are greater than the node key "3".
[0061] The data key k of the key-value pair is not limited to numerical values. In some embodiments, non-numerical data keys may be used for the data key-value pairs (e.g., "name", "age", etc.), and numerical values may be associated with the non-numerical data keys. For example, the data key "name" may correspond to the numerical key "3". The data keys that appear alphabetically before the word "name" or are the data key of the word "name" may be found by following the left branch associated with the node. The data keys that appear alphabetically after the word "name" may be found by following the right branch associated with the node. In some embodiments, a hash function may be associated with the non-numerical data keys. The hash function may determine which branch in the node the non-numerical data key is associated with.
[0062] In the illustrated example, the root node 202 includes pointers to the intermediate node 212 and the intermediate node 214. The root node 202 includes the node ID "R1" and the tree ID "1". The node ID identifies the name of the node. The tree ID identifies the view with which the node is associated. When changes are made to the data stored in the leaf node as described with respect to Figure 2B , Figure 2C and Figure 2D , the tree ID is used to determine whether a copy of the node will be made.
[0063] The root node 202 includes a node key that divides a set of pointers into two different subsets. Leaf nodes with data keys k less than or equal to the node key (e.g., "1-3") are associated with a first branch, and leaf nodes with data keys k greater than the node key (e.g., "4-5") are associated with a second branch. A leaf node with data keys "1", "2", or "3" can be found by traversing the tree data structure 200 from the root node 202 to the intermediate node 212, because the data keys have values less than or equal to the node key. A leaf node with data keys "4" or "5" can be found by traversing the tree data structure 200 from the root node 202 to the intermediate node 214, because the data keys have values greater than the node key.
[0064] The root node 202 includes a first set of pointers. The first set of pointers associated with data keys less than the node key (e.g., "1", "2", or "3") indicates that traversing the tree data structure 200 from the root node 202 to the intermediate node 212 leads to a leaf node with data keys "1", "2", or "3". The intermediate node 214 includes a second set of pointers. The second set of pointers associated with data keys greater than the node key indicates that traversing the tree data structure 200 from the root node 202 to the intermediate node 214 leads to a leaf node with data keys "4" or "5".
[0065] The intermediate node 212 includes pointers to the leaf node 222, the leaf node 224, and the leaf node 226. The intermediate node 212 includes a node ID "I1" and a tree ID "1". The intermediate node 212 includes a first node key "1" and a second node key "2". The data key k of the leaf node 222 is a value less than or equal to the first node key. The data key k of the leaf node 224 is a value greater than the first node key and less than or equal to the second node key. The data key k of the leaf node 226 is a value greater than the second node key. The pointer to the leaf node 222 indicates that traversing the tree data structure 200 from the intermediate node 212 to the leaf node 222 leads to a node with the data key "1". The pointer to the leaf node 224 indicates that traversing the tree data structure 200 from the intermediate node 212 to the leaf node 224 leads to a node with the data key "2". The pointer to the leaf node 226 indicates that traversing the tree data structure 200 from the intermediate node 212 to the leaf node 226 leads to a node with the data key "3".
[0066] The intermediate node 214 includes pointers to leaf node 228 and leaf node 230. The intermediate node 212 includes a node ID "I2" and a tree ID "1". The intermediate node 214 includes a node key "4". The data key k of leaf node 228 is a value less than or equal to the node key. The data key k of leaf node 230 is a value greater than the node key. The pointer of leaf node 228 indicates that traversing the tree data structure 200 from intermediate node 214 to leaf node 228 will lead to a node with data key "4". The pointer of leaf node 230 indicates that traversing the tree data structure 200 from intermediate node 214 to leaf node 230 will lead to a node with data key "5".
[0067] Leaf nodes 222, 224, 226, 228, 230 respectively include data key-value pairs "1: Data 1", "2: Data 2", "3: Data 3", "4: Data 4", "5: Data 5". Leaf nodes 222, 224, 226, 228, 230 respectively include node IDs "L1", "L2", "L3", "L4", "L5". Each of leaf nodes 222, 224, 226, 228, 230 includes a tree ID "1". To view the value associated with data key "1", traverse the tree data structure 200 from root node 202 to intermediate node 212 and then to leaf node 222. To view the value associated with data key "2", traverse the tree data structure 200 from root node 202 to intermediate node 212 and then to leaf node 224. To view the value associated with data key "3", traverse the tree data structure 200 from root node 202 to intermediate node 212 and then to leaf node 226. To view the value associated with data key "4", traverse the tree data structure 200 from root node 202 to intermediate node 214 and then to leaf node 228. To view the value associated with data key "5", traverse the tree data structure 200 from root node 202 to intermediate node 214 and then to leaf node 230. In some instances, leaf nodes 222, 224, 226, 228, 230 are configured to store metadata associated with a file. In other instances, leaf nodes 222, 224, 226, 228, 230 are configured to store pointers to file metadata structures.
[0068] Figure 2BA block diagram showing an implementation of a cloned file system metadata snapshot tree. The file system metadata snapshot tree can be cloned when added to a tree data structure. In some implementations, the tree data structure 250 can be created by a storage system such as the storage cluster 112. The file system data of a primary system such as the primary system 102 can be backed up to a storage cluster such as the storage cluster 112. Subsequent backup snapshots can correspond to full backup snapshots or incremental backup snapshots. The way the file system data corresponding to the subsequent backup snapshots is stored in the storage cluster can be represented by the tree data structure. The tree data structure corresponding to the subsequent backup snapshot is created by cloning the file system metadata snapshot tree associated with the last backup snapshot.
[0069] In the example shown, the tree data structure 250 includes root nodes 202, 204, intermediate nodes 212, 214, and leaf nodes 222, 224, 226, 228, and 230. The tree data structure 250 can be a snapshot of the file system data at a particular point in time, such as t = 2. The tree data structure can be used to capture different versions of the file system data at different times. The tree data structure can allow a series of backup snapshot versions (i.e., file system metadata snapshot trees) to be linked together by allowing nodes of a later version of the file system metadata snapshot tree to reference nodes of a previous version of the file system metadata snapshot tree. For example, the file system metadata snapshot tree with root node 204 is linked to the file system metadata snapshot tree with root node 202. Each time a backup snapshot is executed, a new root node can be created, and the new root node includes the same set of pointers included in the previous root node, i.e., the new root node of the file system metadata snapshot tree can be linked to one or more intermediate nodes associated with the previous file system metadata snapshot tree. The new root node also includes a different node ID and a different tree ID. The tree ID is a view identifier associated with the view of the primary system corresponding to a particular moment.
[0070] In some implementations, the root node is associated with the current view of the file system data. The current view can still accept one or more changes to the data. The tree ID of the root node indicates the backup snapshot with which the root node is associated. For example, the root node 202 with tree ID "1" is associated with the first backup snapshot, and the root node 204 with tree ID "2" is associated with the second backup snapshot. In the example shown, the root node 204 is associated with the current view of the file system data.
[0071] In other implementations, the root node is associated with a snapshot view of the file system data. The snapshot view can represent the state of the file system data at a particular past moment and has not been updated. In the example shown, the root node 202 is associated with a snapshot view of the file system data.
[0072] In the illustrated example, root node 204 is a clone (e.g., copy) of root node 202. Similar to root node 202, root node 204 includes the same pointers as root node 202. Root node 204 includes a first set of pointers to intermediate nodes 212. Root node 204 includes node ID "R2" and tree ID "2".
[0073] Figure 2C is a block diagram showing an implementation for modifying a file system metadata snapshot tree. In the illustrated example, the tree - type data structure 255 can be modified by a file system manager such as file system manager 117. The file system metadata snapshot tree with root node 204 can be the current view of the file system data at time t = 2. The current view represents the state of the file system data that is the most recent and capable of receiving one or more modifications to the snapshot tree corresponding to the file system data. Since the snapshot represents a "frozen" perspective of the file system data in time, one or more copies of one or more nodes affected by changes to the file system data are made.
[0074] In the illustrated example, the value "data4" has been modified to "data4'". In some implementations, the value of a key - value pair has been modified. For example, the value "data4" can be a pointer to a file metadata structure corresponding to a first version of a file, and the value "data4'" can be a pointer to a file metadata structure corresponding to a second version of the file. In other implementations, the value of the key - pair is data associated with a content file that is less than or equal to a limit size. In other implementations, the value of the key - value pair points to a different file metadata structure. The different file metadata structure can be a modified version of the file metadata structure previously pointed to by the leaf node.
[0075] To modify the file system metadata snapshot tree, the file system manager may start at the root node 204, as the root node is the root node associated with the file system metadata snapshot tree at time t = 2 (i.e., the root node associated with the last backup snapshot). The value "Data 4" is associated with the data key "4". The file system manager may traverse the tree - type data structure 255 from the root node 204 until it reaches the target node, which in this instance is the leaf node 228. The file system manager may compare the tree ID at each intermediate node and leaf node with the tree ID of the root node. If the node's tree ID matches the tree ID of the root node, the file system manager may proceed to the next node. If the node's tree ID does not match the tree ID of the root node, a shadow copy of the node with the mismatched tree ID may be made. A shadow copy is a copy of a node and includes the same pointers as the copied node, but also includes a different node ID and tree ID. For example, to reach the leaf node with the data key "4", the file system manager starts at the root node 204 and proceeds to the intermediate node 214. The file system manager compares the tree ID of the intermediate node 214 with the tree ID of the root node 204, determines that the tree ID of the intermediate node 214 does not match the tree ID of the root node 204, and creates a copy of the intermediate node 214. The intermediate node copy 216 includes the same set of pointers as the intermediate node 214, but also includes the tree ID "2" that matches the tree ID of the root node 204. The file system manager may update the pointer of the root node 204 to point to the intermediate node 216 instead of the intermediate node 214. The file system manager may traverse the tree - type data structure 255 from the intermediate node 216 to the leaf node 228, determine that the tree ID of the leaf node 228 does not match the tree ID of the root node 204, and create a copy of the leaf node 228. The leaf node copy 232 stores the modified value "Data 4'" and includes the same tree ID as the root node 204. The file system manager may update the pointer of the intermediate node 216 to point to the leaf node 232 instead of the leaf node 228.
[0076] In some embodiments, the leaf node 232 stores the value of the key - value pair that has been modified. In other embodiments, the leaf node 232 stores the modified data associated with a file that is less than or equal to a limit size. In other embodiments, the leaf node 232 stores a pointer to a file metadata structure corresponding to a file such as a virtual machine container file.
[0077] Figure 2D is a block diagram showing an embodiment of a modified file system metadata snapshot tree. Figure 2D The tree - type data structure 255 shown illustrates the result of the modifications made to the snapshot tree as described with respect to Figure 2C what has been described.
[0078] Figure 3Ais a block diagram showing an implementation of a tree data structure, which closely corresponds to Figure 2A .
[0079] Leaf nodes of a file system metadata snapshot tree (such as the leaf nodes of tree data structures 200, 250, 255) may include pointers to tree data structures such as tree data structure 300 corresponding to files.
[0080] A tree data structure corresponding to a content file at a specific point in time (e.g., a specific version) may include a root node, one or more levels of one or more intermediate nodes, and one or more leaf nodes. In some implementations, the tree data structure corresponding to a content file includes a root node and one or more leaf nodes, without any intermediate nodes. Tree data structure 300 may be a snapshot of a content file at a specific time point t, e.g., at time t = 1.
[0081] In the illustrated example, tree data structure 300 includes file root node 302, file intermediate nodes 312, 314, and file leaf nodes 322, 324, 326, 328, 330. Although tree data structure 300 includes one intermediate level between root node 302 and leaf nodes 322, 324, 326, 328, 330, any number of intermediate levels may be implemented. Similar to the file system metadata snapshot tree described above, each node includes a "node ID" that identifies the node, and a "tree ID" that identifies the snapshot / view with which the node is associated.
[0082] In the illustrated example, root node 302 includes pointers to intermediate node 312 and intermediate node 314. Root node 202 includes node ID "FR1" and tree ID "1".
[0083] In the illustrated example, intermediate node 312 includes pointers to leaf node 322, leaf node 324, and leaf node 326. Intermediate node 312 includes node ID "FI1" and tree ID "1". Intermediate node 312 includes a first node key and a second node key. The data key k of leaf node 322 is a value less than or equal to the first node key. The data key of leaf node 324 is a value greater than the first node key and less than or equal to the second node key. The data key of leaf node 326 is a value greater than the second node key. The pointer of leaf node 322 indicates that traversing tree data structure 300 from intermediate node 312 to leaf node 322 leads to a node with data key "1". The pointer of leaf node 324 indicates that traversing tree data structure 300 from intermediate node 312 to leaf node 324 leads to a node with data key "2". The pointer of leaf node 326 indicates that traversing tree data structure 300 from intermediate node 312 to leaf node 326 leads to a node with data key "3".
[0084] In the illustrated example, the intermediate node 314 includes pointers to leaf node 328 and leaf node 330. The intermediate node 314 includes a node ID "FI2" and a tree ID "1". The intermediate node 314 includes a node key. The data key k of leaf node 328 is a value less than or equal to the node key. The data key of leaf node 330 is a value greater than the node key. The pointer of leaf node 328 indicates that traversing the tree data structure 300 from intermediate node 314 to leaf node 328 leads to a node with a data key of "4". The pointer of leaf node 330 indicates that traversing the tree data structure 300 from intermediate node 314 to leaf node 330 leads to a node with a data key of "5".
[0085] Leaf node 322 includes the data key-value pair "1: Brick 1". "Brick 1" is a brick identifier that identifies a data brick associated with one or more data blocks of a content file corresponding to the tree data structure 300. Leaf node 322 includes a node ID "FL1" and a tree ID "1". To view the value associated with the data key "1", traverse the tree data structure 300 from root node 302 to intermediate node 312 and then to leaf node 322.
[0086] Leaf node 324 includes the data key-value pair "2: Brick 2". "Brick 2" may be associated with one or more data blocks associated with the content file. Leaf node 324 includes a node ID "FL2" and a tree ID "1". To view the value associated with the data key "2", traverse the tree data structure 300 from root node 302 to intermediate node 312 and then to leaf node 324.
[0087] Leaf node 326 includes the data key-value pair "3: Brick 3". "Brick 3" may be associated with one or more data blocks associated with the content file. Leaf node 326 includes a node ID "FL3" and a tree ID "1". To view the value associated with the data key "3", traverse the tree data structure 300 from root node 302 to intermediate node 312 and then to leaf node 326.
[0088] Leaf node 328 includes the data key-value pair "4: Brick 4". "Brick 4" may be associated with one or more data blocks associated with the content file. Leaf node 328 includes a node ID "FL4" and a tree ID "1". To view the value associated with the data key "4", traverse the tree data structure 300 from root node 302 to intermediate node 314 and then to leaf node 328.
[0089] Leaf node 330 includes the data key-value pair "5: Brick 5". "Brick 5" may be associated with one or more data blocks associated with the content file. Leaf node 330 includes a node ID "FL5" and a tree ID "1". To view the value associated with the data key "5", traverse the tree data structure 300 from root node 302 to intermediate node 314 and then to leaf node 330.
[0090] The content file may include a plurality of data bricks and one or more block files. The data bricks may be associated with one or more block identifiers (e.g., SHA-1 hash values). In the illustrated example, the leaf nodes 322, 324, 326, 328, 330 each store a corresponding brick identifier. The block metadata table may store information that associates brick identifiers with one or more block identifiers and one or more block file identifiers corresponding to the one or more block identifiers. The block file metadata table may associate block file identifiers with block files that store a plurality of data bricks. The block metadata table and the block file metadata table may be used based on the brick identifiers to locate the data bricks associated with the file corresponding to the file metadata structure.
[0091] Figure 3B FIG. is a block diagram showing an embodiment of adding a file metadata structure to a tree data structure. In some embodiments, the tree data structure 350 may be created by a storage system such as the storage cluster 104. The tree data structure corresponding to a file may be used to capture different versions of the file at different times. When a backup snapshot is received, the root node of the file metadata structure may be linked to one or more intermediate nodes associated with the previous file metadata structure. This may occur when the data associated with the file is included in two backup snapshots.
[0092] In the illustrated example, the tree data structure 350 includes: a first file metadata structure, the first file metadata structure including: a root node 302, intermediate nodes 312, 314, and leaf nodes 322, 324, 326, 328, and 330; and a second file metadata structure, the second file metadata structure including a root node 304, intermediate nodes 312, 314, and leaf nodes 322, 324, 326, 328, and 330. The second file metadata structure may correspond to a version of the file at a particular point in time, e.g., at time t = 2. The first file metadata structure may correspond to a first version of the content file, and the second file metadata structure may correspond to a second version of the content file.
[0093] To create a snapshot of the file data at time t = 2, a new root node is created. The new root node may be a clone of the original node and includes the same set of pointers as the original node, but includes a different node ID and a different tree ID. In the illustrated example, the root node 304 includes a set of pointers to the intermediate nodes 312, 314, which are the intermediate nodes associated with the previous snapshot. In the illustrated example, the root node 304 is a copy of the root node 302. Similar to the root node 302, the root node 304 includes the same pointers as the root node 302. The root node 304 includes the node ID "FR2" and the tree ID "2".
[0094] Figure 3C A block diagram showing an implementation for modifying a file metadata structure. In the illustrated example, the tree data structure 380 may be modified by a file system manager such as the file system manager 117. The file metadata structure having the root node 304 may be the current view of the file data at a time, e.g., at time t = 2.
[0095] In some implementations, the file data of a content file may be modified such that one of the data blocks is replaced by another data block. When the data block of the file data associated with the previous backup snapshot is replaced by a new data block, the data brick associated with the new data block may be different. The leaf nodes of the file metadata structure may be configured to store the brick identifier of the brick associated with the new data block. To represent such a modification to the file data, a corresponding modification is made to the current view of the file metadata structure. The replaced data block of the file data has a corresponding leaf node in the previous file metadata structure. A new leaf node corresponding to the new data block may be created in the current view of the file metadata structure as described herein. The new leaf node may include an identifier associated with the current view. The new leaf node may also store the block identifier associated with the modified data block.
[0096] In the illustrated example, the data block corresponding to "brick 4" of the content file has been modified to the data block associated with "brick 6". At t = 2, the file system manager starts at the root node 304 because the root node is the root node associated with the file metadata structure at time t = 2. The value "brick 4" is associated with the data key "4". The file system manager can traverse the tree data structure 380 from the root node 304 until it reaches the target node, which is the leaf node 328 in this example. The file system manager can compare the tree ID at each intermediate node and leaf node with the tree ID of the root node. If the tree ID of the node matches the tree ID of the root node, the file system manager can proceed to the next node. If the tree ID of the node does not match the tree ID of the root node, a shadow copy of the node with the mismatched tree ID can be made. For example, to reach the leaf node with the data key "4", the file system manager can start at the root node 304 and proceed to the intermediate node 314. The file system manager can compare the tree ID of the intermediate node 314 with the tree ID of the root node 304, determine that the tree ID of the intermediate node 314 does not match the tree ID of the root node 304, and create a copy of the intermediate node 314. The intermediate node copy 316 can include the same set of pointers as the intermediate node 314, but also includes the tree ID "2" that matches the tree ID of the root node 304. The file system manager can update the pointer of the root node 304 to point to the intermediate node 316 instead of the intermediate node 314. The file system manager can traverse the tree data structure 380 from the intermediate node 316 to the leaf node 328, determine that the tree ID of the leaf node 328 does not match the tree ID of the root node 304, and create a copy of the leaf node 328. The leaf node 332 is a copy of the leaf node 328, but stores the brick identifier "brick 6" and includes the same tree ID as the root node 304. The file system manager updates the pointer of the intermediate node 316 to point to the leaf node 332 instead of the leaf node 328.
[0097] Figure 3D is a block diagram showing an embodiment of a modified file metadata structure. Figure 3D The illustrated tree data structure 380 shows the result of the modification made to the tree data structure 380 as described with respect to Figure 3C Each leaf node may have an associated reference count. The reference count may indicate the number of other nodes that reference the leaf node, such as pointers that include the leaf node. In the illustrated example, the leaf node 322 has a reference count of "1", the leaf node 324 has a reference count of "1", the leaf node 326 has a reference count of "1", the leaf node 328 has a reference count of "1", the leaf node 330 has a reference count of "2", and the leaf node 332 has a reference count of "1".
[0098] It is possible that multiple file metadata structures include leaf nodes that store the same brick identifier. For example, multiple file metadata structures corresponding to files included in a backup snapshot may include leaf nodes that store the same brick identifier. In another instance, one or more file metadata structures corresponding to files included in a backup snapshot, and one or more file metadata structures corresponding to files generated by a storage cluster may include leaf nodes that store the same brick identifier. In yet another instance, multiple file metadata structures corresponding to files generated by a storage cluster may include leaf nodes that store the same brick identifier.
[0099] To reduce the amount of storage used by the SSDs to store metadata associated with multiple files, such metadata is deduplicated and duplicate leaf nodes are removed from the SSDs. To identify potentially duplicate metadata elements, the bottom two levels of the multiple file metadata structures stored by the storage cluster are scanned. Multiple leaf nodes that store the same brick identifier are identified. The bottom two levels of the file metadata structure correspond to the leaf node level and the lowest intermediate node level (i.e., the node that references the leaf nodes included in the leaf node level). Leaf nodes that store the same brick identifier may be associated with different files and / or different versions of the same file.
[0100] A reference count for the leaf nodes is maintained, which indicates the number of other nodes that reference the leaf node, such as pointers to the leaf node. The leaf node that stores the same brick identifier as one or more other leaf nodes and has the highest reference count is identified. For leaf nodes that store the same brick identifier, the leaf node with the highest reference count is selected and retained, and any one or more parent nodes associated with one or more leaf nodes that do not have the highest reference count are modified to reference the leaf node with the highest reference. One or more leaf nodes that do not have the highest reference count are deleted. Retaining the leaf node with the highest reference count reduces the number of writes required by the processors of the storage cluster to deduplicate the metadata associated with multiple files.
[0101] In addition to the file metadata structure shown in the tree data structure 380, the storage system may store one or more other independent file metadata structures with leaf nodes storing the value "brick 5". An independent file metadata structure is a file metadata structure that does not include one or more references to one or more other file metadata structures. For example, an independent file metadata structure may correspond to a file generated by the storage system. For an independent file metadata structure, the reference count for the leaf node storing the value "brick 5" will be "1". When the storage system performs a file system metadata deduplication process, for an independent file metadata structure, the parent node of the leaf node storing the value "brick 5" may be modified to reference leaf node 330 because leaf node 330 has a higher reference count and the leaf node storing the value "brick 5" may be deleted. As explained above, retaining the leaf node with the highest reference count (leaf node 330 in this example) reduces the number of writes required by the processor of the storage system to deduplicate the metadata associated with multiple files.
[0102] Figure 4A is a block diagram showing an example of duplicate metadata according to some embodiments. In the example shown, the file metadata structure 400 includes a file metadata structure corresponding to a first file F1, a second file metadata structure corresponding to a second file F2, and a third file metadata structure corresponding to a third file F3.
[0103] The first file metadata structure corresponding to the first file F1 includes a root node 401, intermediate nodes 402, 403, and leaf nodes 404, 405, 406, 407. Leaf node 404 stores the value "brick 1", leaf node 405 stores the value "brick 2", leaf node 406 stores the value "brick 3", and leaf node 407 stores the value "brick 4".
[0104] The second file metadata structure corresponding to the second file F2 includes a root node 411, intermediate nodes 412, 413, and leaf nodes 414, 415, 416, 417. Leaf node 414 stores the value "brick 5", leaf node 415 stores the value "brick 6", leaf node 416 stores the value "brick 7", and leaf node 417 stores the value "brick 4".
[0105] The third file metadata structure corresponding to the third file F3 includes a root node 421, intermediate nodes 422, 423, and leaf nodes 424, 425, 426, 427. Leaf node 424 stores the value "brick 8", leaf node 425 stores the value "brick 9", leaf node 426 stores the value "brick 10", and leaf node 427 stores the value "brick 4".
[0106] In the illustrated example, the file metadata structures corresponding to files F1, F2, F3 each include leaf nodes storing the value "brick 4". Although leaf nodes 407, 417, 427 have different node IDs and different tree IDs, they are duplicate leaf nodes because they store the same value. Storing multiple leaf nodes that reference the same data brick is an inefficient use of storage. The storage amount required to store the file metadata structures corresponding to files F1, F2, F3 can be reduced by deduplicating the metadata associated with files F1, F2, F3.
[0107] Figure 4B FIG. is a block diagram illustrating an example of deduplicated metadata according to some embodiments. In the illustrated example, file metadata structure 430 includes a file metadata structure corresponding to a first file F1, a second file metadata structure corresponding to a second file F2, and a third file metadata structure corresponding to a third file F3.
[0108] In Figure 4A , the file metadata structures corresponding to files F1, F2, F3 each previously included leaf nodes referencing the value "brick 4". Metadata associated with multiple files can be deduplicated by scanning the bottom two levels of the file metadata structure to identify any leaf nodes referencing the same value (e.g., brick identifier) and the parent nodes of the leaf nodes referencing the same value.
[0109] In some embodiments, metadata associated with multiple files is partially deduplicated by identifying instances of leaf nodes storing the same value as one or more other leaf nodes. The instance with the highest reference count is selected and retained, and the parent nodes associated with leaf nodes not having the highest reference count are modified to reference the leaf node with the highest reference count. Instances of leaf nodes not having the highest reference count are deleted.
[0110] In other instances, multiple instances storing the same value have the same reference count, i.e., there is no single leaf node with the highest reference count. For these instances, metadata associated with multiple files is deduplicated by selecting and retaining one of the instances, and modifying the parent nodes associated with one or more of the unselected instances to reference the selected instance of the leaf node. The unselected instances are deleted.
[0111] In the illustrated example, leaf nodes (or leaf node instances) 407, 417, 427 have the same reference count. Leaf node 407 is selected as the retained leaf node, and the parent nodes of leaf nodes 417 (i.e., intermediate node 413) and 427 (i.e., intermediate node 427) are modified to reference leaf node 407 instead of leaf nodes 417 and 427 respectively. Leaf nodes 417, 427 are deleted after their parent leaf nodes are modified to reference the selected node.
[0112] Leaf node 407 may have an associated tree ID. The tree ID indicates the file metadata structure that owns the node, i.e., the original file with which the node is associated. The tree ID of leaf node 407 may be the same as the tree ID of root node 401. In this case, the value stored by leaf node 407 can be modified without creating a new leaf node. If the tree ID of leaf node 407 is different from the tree ID of root node 401, the value stored by leaf node 407 can be modified by creating a new leaf node (such as Figure 3C shown).
[0113] Problems arise when nodes are shared between file metadata structures because the values stored by the leaf nodes are now associated with multiple file metadata structures. Changes to a leaf node associated with a first file metadata structure corresponding to a first file may not apply to one or more other file metadata structures of shared leaf nodes corresponding to one or more other files. In the illustrated example, if leaf node 407 and root node 401 have the same tree ID, the value associated with leaf node 407 can be modified without creating a new leaf node. However, modifications to the file metadata structure corresponding to file F1 also cause changes to the file metadata structures corresponding to files F2 and F3. To prevent this from happening, if a file that owns a leaf node (e.g., the tree ID of the leaf node matches the tree ID of the root node) is modified such that the leaf node stores a different value and the leaf node is shared by multiple file metadata structures, then, for example, as Figure 3C shown, the file metadata structure is modified as if the tree ID of the leaf node does not match the tree ID of the root node. A new leaf node can be generated, and the new leaf node can store the modified value, but the tree ID of the node matches the tree ID of the root node. The value stored by a leaf node shared between file metadata structures can be modified to include a value that indicates that a new leaf node needs to be created to modify the value stored by the leaf node. This can prevent changes to the first file from also applying to one or more other files of the shared leaf node.
[0114] Figure 4CA block diagram showing an example of duplicate metadata according to another example, where the file metadata structure 450 includes a file metadata structure corresponding to the first file F1, a second file metadata structure corresponding to the second file F2, and a third file metadata structure corresponding to the third file F3.
[0115] The first file metadata structure corresponding to the first file F1 includes a root node 451, intermediate nodes 452, 453, and leaf nodes 454, 455, 456, 457. The leaf node 454 stores the value "brick 1", the leaf node 455 stores the value "brick 2", the leaf node 456 stores the value "brick 3", and the leaf node 457 stores the value "brick 4".
[0116] The second file metadata structure corresponding to the second file F2 includes a root node 461, intermediate nodes 462, 463, and leaf nodes 464, 465, 466, 467. The leaf node 464 stores the value "brick 5", the leaf node 465 stores the value "brick 6", the leaf node 466 stores the value "brick 3", and the leaf node 467 stores the value "brick 4".
[0117] The third file metadata structure corresponding to the third file F3 includes a root node 471, intermediate nodes 472, 473, and leaf nodes 474, 475, 476, 477. The leaf node 474 stores the value "brick 8", the leaf node 475 stores the value "brick 9", the leaf node 476 stores the value "brick 3", and the leaf node 477 stores the value "brick 4".
[0118] The metadata associated with multiple files can be deduplicated by removing duplicate portions of the file metadata structure. Leaf nodes associated with a portion of the file metadata structure can store a specific sequence of values (e.g., brick identifiers). In Figure 4C the example, the file metadata structure corresponding to the first file F1 includes four leaf nodes, where the first leaf node stores the brick identifier "1", the second leaf node stores the brick identifier "2", the third leaf node stores the brick identifier "3", and the fourth leaf node stores the brick identifier "4". In this example, the specific sequence of values for the file metadata structure corresponding to the first file is "1234". The specific sequence of values for the file metadata structure corresponding to the second file F2 is "5634". The specific sequence of values for the file metadata structure corresponding to the third file is "8934". The specific sequence of values can be referred to as a "tree fingerprint".
[0119] A breadth-level search at the leaf node level can be performed for each of the multiple file metadata structures to identify the corresponding tree fingerprint associated with the file metadata structure. In some embodiments, a portion of the tree fingerprint associated with the first file metadata structure is the same as a portion of the tree fingerprint associated with the same or one or more other file metadata structures. In Figure 4CIn it, the part of the first file metadata structure including leaf nodes 456 and 457 is the same as the part of the second file metadata structure including leaf nodes 466 and 467, and is also the same as the part of the third file metadata structure including leaf nodes 476 and 477.
[0120] Figure 4D is a block diagram showing an example of deduplicated metadata for an example of Figure 4C In the illustrated example, the tree fingerprints associated with the first, second, and third file metadata structures share a common value sequence. Specifically, the common value sequence is "34". A common metadata element can be a node that includes a direct or indirect reference to a leaf node associated with the common sequence and does not include a direct or indirect reference to a leaf node not associated with the common sequence. If a node is one level higher than a leaf node associated with the common sequence, the node can include a direct reference to a leaf node associated with the common sequence. If a node is two or more levels higher than a leaf node associated with the common sequence, the node can include an indirect reference to a leaf node associated with the common sequence.
[0121] In this example, the common node included in the common sequence corresponding to the leaf node of the first file metadata structure for the first file F1 is the intermediate node 453. The common node included in the common sequence corresponding to the leaf node of the second file metadata structure for the second file F2 is the intermediate node 463. The common node included in the common sequence corresponding to the leaf node of the third file metadata structure for the third file F3 is the intermediate node 473.
[0122] Determine the reference count associated with the corresponding common node and identify the common node with the highest reference. In some embodiments, for multiple common nodes associated with the common sequence, select and retain the common node with the highest reference count. Modify one or more corresponding parent nodes of the other common nodes to reference the common node with the highest reference count, delete the unselected common nodes, and delete the nodes directly or indirectly referenced by the unselected common nodes. Retaining the common node with the highest number of references can reduce the number of writes required by the processor of the storage cluster to deduplicate the metadata associated with multiple files.
[0123] In other examples, multiple common nodes have the same reference count, that is, there is no single common node with the highest reference count. In this case, the metadata associated with multiple files is deduplicated by: selecting and retaining one of the common nodes, modifying the parent nodes associated with one or more unselected common nodes to reference the selected common node and deleting the nodes referenced by one or more unselected common nodes.
[0124] In this example, the intermediate nodes 453, 463, 473 have the same reference count. The intermediate node 453 is selected as the retained leaf node, and the parent nodes of the intermediate node 463 (i.e., the root node 461) and the leaf node 473 (i.e., the root node 471) are modified to reference the intermediate node 453 instead of referencing the intermediate node 463 and the intermediate node 473 respectively. After the root nodes 461, 471 are modified to reference the selected nodes, the intermediate nodes 463, 473 and the leaf nodes 466, 467, 476, 477 are deleted.
[0125] Figure 5 is a flowchart showing a process for deduplicating metadata associated with multiple files according to some embodiments. In the illustrated example, the process 500 may be implemented by a storage cluster such as the storage cluster 112.
[0126] At 502, multiple file metadata structures of the file system are analyzed. As would be understood from the foregoing example, the storage cluster may store multiple files and use a tree data structure to organize the metadata associated with the multiple files. Some of the files stored by the storage cluster may correspond to files backed up from the primary system to the storage cluster. Other files stored by the storage cluster may correspond to files generated by the storage cluster.
[0127] The metadata associated with a file may be organized using a tree data structure and is referred to as a "file metadata tree" or "file metadata structure". Each of the file metadata structures includes a corresponding plurality of metadata elements. For example, a file metadata structure includes a root node, one or more levels of one or more intermediate nodes, and a plurality of leaf nodes. The leaf nodes of the file metadata structure store KVPs. The value of the KVP may be a brick identifier associated with one or more data blocks of the file. The metadata associated with the multiple files stored by the storage cluster may be stored in the SSD of the storage cluster.
[0128] At 504, metadata elements that are repeated among the multiple file metadata structures are identified. For example, if multiple leaf nodes store the same value, the storage cluster may store duplicate metadata, and / or the storage cluster may identify multiple file metadata structures that include metadata elements (e.g., leaf nodes) storing the same value. A reference count is determined for each metadata element that stores a value identical to one or more other metadata elements. The reference count indicates the number of other metadata elements that reference the metadata element, such as pointers including leaf nodes. Metadata elements that store a value identical to one or more other metadata elements and have the highest reference count are identified.
[0129] At 506, the identified metadata elements are de-duplicated by modifying at least one of the multiple file metadata structures to reference the same instance of the metadata element that is referenced by another file metadata structure of the multiple file metadata structures.
[0130] For instances of metadata elements that store the same value, select and retain the instance with the highest reference count, and modify the parent metadata element associated with the instance(s) that do not have the highest reference count to reference the instance of the metadata element that has the highest reference count. Then delete the instance(s) of the metadata element that do not have the highest reference count. Retaining the instance with the highest reference count reduces the number of writes required by the processors of the storage cluster to de-duplicate the metadata associated with multiple files.
[0131] For instances of metadata elements that have the same reference count, i.e., when no single instance of a metadata element has the highest reference count, de-duplicate the metadata associated with multiple files by selecting and retaining one of the instances, and modifying the parent metadata element associated with the one or more unselected instances to reference the selected instance of the metadata element. Delete the one or more unselected instances of the metadata element.
[0132] Figure 6 is a flow chart illustrating a process for de-duplicating metadata elements according to some embodiments. In the illustrated example, process 600 may be implemented by a storage cluster such as storage cluster 112. In some embodiments, process 600 is implemented to perform some or all of step 506 of process 500.
[0133] At 602, identify duplicate metadata elements with the highest reference count. Multiple metadata elements may store the same value, e.g., a brick identifier. A reference count may be determined for each metadata element that stores the same value. The reference count may indicate the number of other metadata elements (e.g., nodes) that reference the metadata element that stores the same value as one or more other metadata elements. Identify the metadata element that stores the same value as one or more other metadata elements and has the highest reference count.
[0134] The identified metadata elements may be associated with files backed up from a primary system to the storage cluster. In other embodiments, the duplicate metadata elements are associated with files generated by the storage cluster. Regardless of the file source, the storage cluster may be configured to organize the metadata associated with files using a tree data structure.
[0135] If the identified metadata element is associated with a file backed up from the primary system to the storage cluster, the file metadata structure including the identified metadata element may be associated with one or more other file metadata structures, i.e., one or more nodes of one or more other file metadata structures may include one or more references to one or more nodes of the file metadata structure including the identified metadata element. Each version of a file backed up from the primary system to the storage cluster may have an associated file metadata structure. When multiple versions of a file are backed up from the primary system to the storage cluster, the reference count may be greater than 1.
[0136] In other embodiments, the file metadata structure is associated with a file generated by the storage cluster. The storage cluster may be used to store multiple versions of a file generated by the storage cluster. A file metadata structure may be generated for each version of a file generated by a user associated with the storage cluster; however, each file metadata structure may be independent of one or more other file metadata structures corresponding to different versions of a file generated by the storage cluster. Duplicate metadata elements included in a file metadata structure corresponding to a file generated by the storage cluster have a reference count of 1 because the file metadata structure is independent of one or more other file metadata structures corresponding to different versions of a file generated by the storage cluster, i.e., nodes of one of the other file metadata structures do not include references to nodes of the file metadata structure including the duplicate metadata element.
[0137] At 604, one or more parent metadata elements associated with duplicate metadata elements that do not have the highest reference count are modified. The bottom two levels of multiple file metadata structures may be scanned to identify duplicate metadata elements and the corresponding parent metadata elements of the duplicate metadata elements, i.e., the nodes including direct references to the duplicate metadata elements. One or more corresponding parent metadata elements associated with one or more duplicate metadata elements that do not have the highest reference count may be modified to reference the duplicate metadata element having the highest reference count. For example, a parent metadata element may be modified to include a pointer to the duplicate metadata element having the highest reference count.
[0138] At 606, one or more duplicate metadata elements that do not have the highest reference count are deleted. Retaining the metadata element having the highest reference count may reduce the number of write operations required by a processor of the storage cluster to deduplicate metadata associated with multiple files. For example, two leaf nodes may store the same brick identifier. The first leaf node may have a reference count of 4, and the second leaf node may have a reference count of 1. Modifying the parent node of the second leaf node to reference the first leaf node would require a single write operation, while modifying the parent node of the first leaf node to reference the second leaf node would require four separate write operations.
[0139] Figure 7 FIG. Figure 7 is a flowchart showing a process for deduplicating metadata associated with multiple files according to some embodiments. In the illustrated example, process 700 may be implemented by a storage cluster such as storage cluster 112.
[0140] At 702, multiple file metadata structures of a file system are analyzed. A storage cluster may store multiple files and use a tree data structure to organize metadata associated with the multiple files. Some of the files stored by the storage cluster may correspond to files backed up from a primary system to the storage cluster. Some of the files stored by the storage cluster may correspond to files generated by the storage cluster.
[0141] Metadata associated with a file may be organized using a tree data structure and is referred to as a “file metadata tree” or “file metadata structure”. The file metadata structure may include a root node, one or more levels of one or more intermediate nodes, and multiple leaf nodes. Each node of the metadata structure may be referred to as a “metadata element”. The leaf nodes of the file metadata structure may store KVPs. The value of a KVP may be a brick identifier associated with one or more data blocks of a file. Metadata associated with multiple files stored by the storage cluster may be stored in the SSD of the storage cluster.
[0142] The multiple files stored in the storage cluster may include duplicate data blocks. The duplicate data blocks may be deduplicated to reduce the storage amount used to store the multiple files. This frees up storage space on the HDD and SSD to store data associated with one or more other files. Although data associated with files may be deduplicated, the storage cluster may store duplicate metadata associated with multiple files. For example, the storage cluster may store multiple file metadata structures that include corresponding metadata elements storing the same value (e.g., the same brick identifier). The storage amount used by the SSD to store metadata associated with multiple files may be reduced by deduplicating the metadata associated with the multiple files. The bottom two levels of the multiple file metadata structures stored by the storage cluster may be scanned to determine whether the storage cluster stores duplicate metadata. If multiple metadata elements store the same value, the storage cluster may store duplicate metadata. The storage cluster may identify multiple file metadata structures that include metadata elements storing the same value.
[0143] At 704, portions of file metadata structures that are repeated among the identified multiple file metadata structures are identified. Storage cluster 112 may perform a breadth-level search at the leaf node level for each of the multiple file metadata structures to identify the corresponding tree fingerprints associated with the file metadata structures. In some embodiments, a portion of the tree fingerprint associated with a first file metadata structure is the same as a portion of the tree fingerprint associated with the same one or more other file metadata structures. For example, the tree fingerprint associated with a first file metadata structure may be "1234", the tree fingerprint associated with a second file metadata structure may be "5634", and the tree fingerprint associated with a third file metadata structure may be "8934". The tree fingerprints associated with the first, second, and third file metadata structures share a common sequence of values (e.g., brick identifiers). In this example, the common sequence of the brick identifiers is "34". Common metadata elements (also referred to as common nodes) associated with the common sequence may be identified. A common metadata element may be a node that includes a direct or indirect reference to a leaf node associated with the common sequence and does not include a direct or indirect reference to a leaf node not associated with the common sequence. If a node is one level higher than a leaf node associated with the common sequence, the node may include a direct reference to the leaf node associated with the common sequence. If a node is two or more levels higher than a leaf node associated with the common sequence, the node may include an indirect reference to the leaf node associated with the common sequence.
[0144] At 706, the identified portions of the file metadata structures are de-duplicated by modifying at least one of the multiple file metadata structures to reference the same instance of the identified portion that is referenced by another file metadata structure of the multiple file metadata structures. When the identified portion is defined by a common sequence of values, such as the brick identifier "34" in the above example, the so-called common metadata elements (also referred to as common nodes) associated with the common sequence represent the identified portion.
[0145] A reference count associated with the corresponding common node may be determined. The common node with the highest reference count is identified, selected, and retained, and any parent metadata elements of the other common nodes are modified to reference the common node with the highest reference count. The unselected common nodes are deleted, and the nodes directly or indirectly referenced by the unselected common nodes are deleted. Retaining the common node with the highest number of references may reduce the number of writes required by the processors of the storage cluster to de-duplicate the metadata associated with multiple files.
[0146] In other instances, multiple common nodes have the same reference count, i.e., there is no single common node with the highest reference count. Metadata associated with multiple files is partially deduplicated by selecting and retaining one of the common nodes, while modifying the parent nodes associated with one or more unselected common nodes to reference the selected common node, and deleting the nodes referenced by one or more unselected common nodes.
[0147] In some embodiments, there is no single common node associated with a common sequence of values. For example, a first file metadata structure may have a tree fingerprint of "12345678", and a second file metadata structure may have a tree fingerprint of "92345678". In this instance, the common sequence of values is "2345678". A first portion of the first file metadata structure may be associated with a metadata element having a sequence of values "1234", and a second portion of the first file metadata structure may be associated with a metadata element having a sequence of values "5678". The first portion of the first file metadata structure may be associated with a first intermediate node, and the second portion of the first file metadata structure may be associated with a second intermediate node, where the first intermediate node and the second intermediate node are associated with different branches of the first file metadata structure. A first portion of the second file metadata structure may be associated with a metadata element having a sequence of values "9234", and a second portion of the second file metadata structure may be associated with a metadata element having a sequence of values "5678". The first portion of the second file metadata structure may be associated with a first intermediate node, and the second portion of the second file metadata structure may be associated with a second intermediate node, where the first intermediate node and the second intermediate node are associated with different branches of the second file metadata structure.
[0148] In this case, the common sequence of brick identifiers can be divided into multiple portions, and the corresponding nodes for the multiple portions can be determined. For example, the common node for the sequence of values "234" can be determined, and the common node for the sequence of values "5678" can be determined. Metadata associated with the file metadata structure can be deduplicated based on the determined common nodes. If the determined common nodes include one or more references to one or more nodes that are not part of the common sequence of values, then the metadata associated with the multiple files can be deduplicated at the leaf node level as described above.
[0149] Figure 8 is a flowchart illustrating a process for deduplicating metadata elements according to some embodiments. In the illustrated instance, process 800 may be implemented by a storage cluster such as storage cluster 112. In some embodiments, process 800 is implemented to perform some or all of the steps of step 706 of process 700.
[0150] As explained above, multiple file metadata structures may share a common sequence of values. A so-called common metadata element or common node is a member that has a common sequence of values. A common node may include a direct or indirect reference to a leaf node associated with the common sequence and does not include a node that has a direct or indirect reference to a leaf node not associated with the common sequence. At 802, the common metadata element (common node) with the highest reference count is identified. The reference count indicates the number of other nodes that reference the common node, such as pointers that include the common node.
[0151] The common sequence may be associated with a file backed up from a primary system to a storage cluster. In other instances, the common sequence is associated with a file generated by the storage cluster. Regardless of the file source, the storage cluster may be configured to use a tree data structure to organize the metadata associated with the file.
[0152] If the common sequence is associated with a file backed up from a primary system to a storage cluster, the file metadata structure that includes the common sequence may be associated with one or more other file metadata structures, i.e., one or more nodes of one or more other file metadata structures may include one or more references to one or more nodes of the file metadata structure that includes the common sequence. Each version of a file backed up from a primary system to a storage cluster may have an associated file metadata structure. If a common metadata element is associated with multiple versions of a file backed up from a primary system to a storage cluster, the common metadata element that indirectly or directly references the common sequence may have a reference count greater than 1.
[0153] In other instances, a file metadata structure is associated with a file generated by the storage cluster. The storage cluster may be used to store multiple versions of a file generated by the storage cluster. A file metadata structure may be generated for each version of a file generated by a user associated with the storage cluster; however, each file metadata structure may be independent of one or more other file metadata structures corresponding to different versions of a file generated by the storage cluster. A common metadata element that directly or indirectly references the common sequence and is included in a file metadata structure corresponding to a file generated by the storage cluster has a reference count of 1 because the file metadata structure is independent of one or more other file metadata structures corresponding to different versions of a file generated by the storage cluster, i.e., a node in one of the other file metadata structures does not include a reference to a node of the file metadata structure that includes the common sequence.
[0154] At 804, one or more parent metadata elements associated with one or more common metadata elements that do not have the highest reference count are modified. Multiple file metadata structures including the common sequence can be scanned to identify corresponding common metadata elements that indirectly or directly reference the common sequence. The corresponding parent metadata elements of the common metadata elements can be determined, i.e., the nodes that include direct references to the common metadata elements. One or more corresponding parent metadata elements associated with one or more common metadata elements that do not have the highest reference count can be modified to reference the common metadata element that has the highest reference count. For example, a parent metadata element can be modified to include a pointer to the common metadata element that has the highest reference count.
[0155] At 806, one or more common metadata elements that do not have the highest reference count are deleted. Retaining the common metadata element that has the highest reference count can reduce the number of writes required by the processors of the storage cluster to deduplicate the metadata associated with multiple files. For example, two metadata elements can have the same common sequence and can be referred to as instances of a common metadata element. A first instance of the metadata element can have a reference count of 4, and a second instance of the metadata element can have a reference count of 1. Then the parent metadata element of the second instance of the metadata element is modified to reference the first instance of the metadata element. This is efficient because it only requires a single write operation, while modifying the parent metadata element of the first instance of the metadata element to reference the second instance of the metadata element would require four separate write operations. If the metadata elements referenced by the second common metadata element are also referenced by the retained first instance of the common metadata element, these metadata elements can also be deleted.
[0156] In the above embodiments, the deduplication of the metadata associated with multiple files has been described with respect to the values maintained by the respective leaf nodes of the tree structure, particularly with respect to individual leaf nodes of the tree structure ( Figure 5 and Figure 6 ), multiple leaf nodes of the tree structure ( Figure 7 and Figure 8 ) and intermediate nodes within the tree structure. It will be appreciated that the methods described herein can also be used to deduplicate the metadata among such different tree structures by deduplicating values, such as identifiers of one or more data bricks associated with one or more data blocks maintained by one or more leaf nodes of different tree structures.
[0157] The present invention may be implemented in many ways, including as a process; an apparatus; a system; a combination of articles; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor. In this specification, these implementations, or any other form that the present invention may take, may be referred to as techniques. Generally, within the scope of the present invention, the order of steps of the disclosed processes may be changed. Unless otherwise stated, components such as a processor or a memory described as being configured to perform a task may be implemented as a general component temporarily configured to perform the task at a given time, or as a special component manufactured to perform the task. As used herein, the term 'processor' refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0158] Together with the accompanying Figure 1 drawings that illustrate the principles of the invention, provide a detailed description of one or more embodiments of the present invention. The present invention is described in connection with such embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims and the present invention encompasses many alternatives, modifications, and equivalent forms. Numerous specific details are set forth in the specification in order to provide a thorough understanding of the present invention. These details are provided for purposes of example and the present invention may be practiced according to the claims without some or all of these specific details. For the sake of clarity, technical material known in the relevant art to the present invention has not been described in detail so as not to unnecessarily obscure the present invention.
[0159] Although the foregoing embodiments have been described in some detail for purposes of clarity of understanding, the present invention is not limited to the details provided. There are many alternative ways of implementing the present invention. The disclosed embodiments are illustrative and not restrictive.
Claims
1. A method for data deduplication, the method comprising: Analyzing a file metadata structure of a file system, wherein each of the file metadata structures includes a plurality of metadata elements; Identifying metadata elements that are repeated among the analyzed file metadata structures; and Deduplicating the identified metadata elements that are repeated among the file metadata structures by: Determining a reference count associated with each of two or more instances of the identified metadata element, each reference count indicating the number of one or more other metadata elements that reference the corresponding instance; Identifying one instance among two or more instances having the highest reference count; and Modifying the parent metadata element of each of the two or more instances that do not have the highest reference count to reference the instance identified as having the highest reference count.
2. The method according to claim 1, wherein, Deduplicating the identified metadata element includes deleting the instances that do not have the highest reference count.
3. The method according to claim 2, wherein Deleting the instances that do not have the highest reference count from a solid state disk of a storage cluster.
4. The method according to claim 1, wherein The determined reference counts of each of the two or more instances are the same, and wherein identifying one instance among two or more instances having the highest reference count includes selecting one of the two or more instances of the identified metadata element.
5. The method according to claim 4, wherein, Deduplicating two or more instances of the identified metadata element further includes modifying the parent metadata element of each unselected instance of the identified metadata element to reference the selected instance of the identified metadata element.
6. The method according to claim 1, wherein, Analyzing the file metadata structure of the file system includes scanning a bottom level of the file metadata structure and levels above the bottom level of the file metadata structure.
7. The method according to claim 1, wherein At least some of the plurality of metadata elements are configured to store values, and the file metadata structure includes at least one metadata element that stores the values.
8. The method according to claim 1, wherein At least one of the file metadata structures corresponds to a file generated by a storage cluster and / or wherein at least one of the file metadata structures corresponds to a file backed up from a primary system to the storage cluster.
9. The method according to claim 1, wherein Deduplicating the identified metadata element as a background process of the storage cluster.
10. The method according to claim 1, wherein, The storage cluster includes a plurality of storage nodes, and metadata associated with a plurality of files is stored across the plurality of storage nodes.
11. The method according to claim 10, wherein, Each storage node has a processor and a plurality of storage layers.
12. The method according to claim 11, wherein, The plurality of storage layers includes: a first storage layer that includes solid state disks; and a second storage layer that includes one or more corresponding hard disk drives.
13. The method according to claim 12, wherein, The metadata associated with the plurality of files is stored in corresponding solid state disk drives of the storage cluster.
14. A system for data deduplication, the system comprising: A memory and at least one processor, the memory including a set of instructions that, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 13.
15. A computer storage medium, comprising a program which, when executed by at least one processor comprised by a computer including the computer storage medium, causes the at least one processor to execute the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Systems and methods for protecting deduplicated data
US9235588B1
Data storage system, process, and computer program for de-duplication of distributed data in a scalable cluster system
WO2018075042A1