Metadata storage method and device, computer equipment and storage medium
By employing an unordered index structure and linked list-based metadata storage method in the distributed file system, the problems of metadata operation latency and poor throughput performance are solved, fast querying and decoupling of directory semantics are achieved, and the scalability and metadata operation efficiency of the system are improved.
Patent Information
- Application Number
- CN202511021795.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing distributed file systems suffer from problems in metadata management, such as poor metadata operation latency and throughput, high cost of maintaining ordered indexes, and the need for complex logging or transaction mechanisms for frequent updates. They are particularly inadequate in high-speed network and persistent memory environments.
A storage method that uses an unordered index structure and linked lists to link metadata is adopted. Metadata is stored through a hash table, and the relationships are maintained within the metadata to achieve fast metadata retrieval and decoupling of directory semantics.
It reduces the complexity of metadata operations, improves the speed of metadata queries and the scalability of the system, reduces the overhead of metadata updates, and optimizes directory traversal performance.
Smart Images

Figure CN120909997A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a metadata storage method and device, computer equipment and storage medium. BACKGROUND
[0002] A distributed file system refers to a file system that stores data on multiple network nodes and coordinates management, which is widely used in industry for data management and sharing. The distributed file system is usually divided into two parts of metadata management and data management, wherein the metadata management is responsible for maintaining the structural information of the file system, such as file name, directory structure, permission information, etc., and the data management is responsible for the storage of actual file content. The current data center and high-performance computing cluster widely adopt distributed file system to manage large-scale data, wherein the metadata management of the file is one of the key performance bottlenecks.
[0003] The existing distributed file system generally adopts an ordered key-value storage system (such as RocksDB) to store and manage metadata, and relies on the range query operation provided by the ordered index to realize the directory traversal semantics in the file system. In this way, the directory semantics and the storage index of the metadata are tightly coupled, and as the scale of the file system metadata grows, it is difficult to maintain the ordered index, and the delay and throughput performance of the metadata operation is poor, and frequent metadata updates also require complex and high-overhead log recording or transaction mechanism. SUMMARY
[0004] Therefore, the present application provides a metadata storage method and device, computer equipment and storage medium.
[0005] Specifically, the present application is realized by the following technical solutions:
[0006] In a first aspect, the present application provides a metadata storage method applied to a distributed file system, the method comprising:
[0007] For any metadata of the file system,
[0008] storing the metadata in a metadata server, wherein the metadata server is configured to store the metadata in the file system based on an unordered index structure;
[0009] determine a second node associated with the first node in the hierarchical directory tree based on the position information of the first node in the hierarchical directory tree, wherein the hierarchical directory tree is used to describe a logical structure of storing data of the file system in an ordered key-value based manner, and the second node is associated with the first node including that the second node is a child node or a sibling node of the first node;
[0010] obtain storage position information of the associated metadata corresponding to the second node in the metadata server; and
[0011] based on the association relationship between the first node and the second node, link the storage position information of the associated metadata corresponding to the second node in the metadata server in the metadata.
[0012] Optionally, the unordered index structure includes a hash table.
[0013] The metadata is stored in the metadata server, including:
[0014] based on the data type of the metadata, determine the key value corresponding to the metadata;
[0015] perform a hash operation on the key value corresponding to the metadata to determine the storage position of the metadata in the hash table;
[0016] store the metadata according to the storage position.
[0017] Optionally, the metadata includes file metadata and directory metadata, and the directory metadata includes directory access metadata and directory content metadata.
[0018] The key corresponding to the file metadata is the name of the file metadata and the identifier of the parent node of the corresponding node of the file metadata.
[0019] The key corresponding to the directory access metadata is the name of the directory access metadata and the identifier of the parent node of the corresponding node of the directory access metadata.
[0020] The key corresponding to the directory content metadata is the identifier of the corresponding node of the directory content metadata.
[0021] Optionally, the metadata is stored according to the storage position, including:
[0022] store the metadata according to the preset memory layout at the corresponding storage position;
[0023] The data fields related to the same metadata operation are arranged in the same aligned persistent memory block in the preset memory layout; and the data in the same aligned persistent memory block is processed in the same CPU atomic operation.
[0024] Optionally, the data fields of the metadata include a file content modification time and a file state change time.
[0025] The method further includes:
[0026] In response to receiving a target metadata operation request of any metadata, if the metadata update field corresponding to the target metadata operation request includes the file content modification time and the file state change time, only the file content modification time of the any metadata is updated.
[0027] After receiving an acquisition request for the file state change time of the any metadata, the maximum value between the updated file content modification time and the file state change time recorded in the any metadata is taken as the file state change time corresponding to the acquisition request.
[0028] Optionally, the metadata of the file system includes file metadata and directory metadata, the file metadata corresponds to a first memory layout, the directory metadata corresponds to a second memory layout, and the first memory layout and the second memory layout are different.
[0029] In the first memory layout, the file content modification time and the file byte number are arranged in the same aligned persistent memory block; the mode, the user identifier of the file, the file state change time, and the low two sections of the group identifier of the file are arranged in the same aligned persistent memory block; and the high two sections of the group identifier of the file, the state information of the file, and the creation time are arranged in the same aligned persistent memory block.
[0030] In the second memory layout, the file content modification time and the head pointer pointing to the child node linked list are arranged in the same aligned persistent memory block; the mode, the user identifier of the file, the file state change time, and the low two sections of the group identifier of the file are arranged in the same aligned persistent memory block; and the high two sections of the group identifier of the file, the state information of the file, and the creation time are arranged in the same aligned persistent memory block.
[0031] Optionally, the method further includes:
[0032] The directory traversal request sent by the client is received, and a directory traversal result is returned to the client based on at least one of the following manners:
[0033] While reading the linked list node, the read part of the result is returned to the client.
[0034] splitting the linked list into multiple segments, reading the multiple segments of the linked list in parallel through multiple threads, returning the merged directory traversal result to the client after merging the read results.
[0035] determining whether the directory traversal result corresponding to the directory traversal request is cached in the dynamic random access memory (DRAM), reading the directory traversal result corresponding to the directory traversal request from the DRAM if the directory traversal result is cached in the DRAM, and returning the read directory traversal result to the client; wherein the DRAM caches historical directory traversal results within a preset time range before the current time.
[0036] returning the read directory traversal result to the client after compression.
[0037] In a second aspect, an embodiment of the present application provides a metadata storage device, and the device comprises:
[0038] a storage module configured to store any metadata of the file system in a metadata server, wherein the metadata server is configured to store the metadata in the file system based on an unordered index structure.
[0039] a determination module configured to determine a second node associated with a first node of a hierarchical directory tree based on position information of the first node, wherein the hierarchical directory tree is configured to describe a logical structure of storing data of the file system based on an ordered key value, and the second node is associated with the first node, including that the second node is a child node or a sibling node of the first node.
[0040] a processing module configured to obtain storage location information of associated metadata of the second node in the metadata server, and link the storage location information of the associated metadata of the second node in the metadata server in the metadata based on an association relationship between the first node and the second node.
[0041] The present application provides a computer readable storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to implement the metadata storage method.
[0042] The present application provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the metadata storage method when executing the program.
[0043] In the metadata storage method provided in this application, for any metadata in the file system, the metadata in the file system can be stored based on an unordered index structure. Under the unordered index structure, even if the size of the metadata increases, fast querying of the metadata can be achieved, reducing the complexity of metadata operations. In addition, within the metadata, the storage location information of the associated metadata of the second node associated with the first node of the metadata is linked in the form of a linked list. In this way, when determining the semantics of the directory, it can be done by traversing the linked list within the directory, thereby achieving the maintenance of the logical relationship between the metadata. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating a metadata storage method according to an exemplary embodiment of this application;
[0045] Figure 2 This is a schematic diagram of a hierarchical directory tree in a metadata storage method according to an exemplary embodiment of this application;
[0046] Figure 3 This is a schematic diagram of the hash mapping process in a metadata storage method according to an exemplary embodiment of this application;
[0047] Figure 4 This is a schematic diagram of the first memory layout in a metadata storage method according to an exemplary embodiment of this application;
[0048] Figure 5 This is a schematic diagram of the second memory layout in a metadata storage method according to an exemplary embodiment of this application;
[0049] Figure 6 This is a schematic diagram illustrating the file creation process in a metadata storage method according to an exemplary embodiment of this application;
[0050] Figure 7 This is a schematic diagram illustrating the optimization process of directory traversal in a metadata storage method according to an exemplary embodiment of this application;
[0051] Figure 8 This is a structural diagram of a metadata storage device illustrated in an exemplary embodiment of this application;
[0052] Figure 9 This is a schematic diagram of the structure of a computer device provided in this application. Detailed Implementation
[0053] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The description herein refers to the accompanying drawings, which show by way of example specific embodiments. In the following description, like reference numerals refer to like elements, unless the context clearly dictates otherwise. The following description of exemplary embodiments is not representative of all possible embodiments consistent with the present application. Rather, it is merely an example of apparatus and methods consistent with some aspects of the present application as detailed in the appended claims.
[0054] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0055] It is to be understood that the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. It is to be understood that the term "and / or" as used herein encompasses all possible combinations of one or more of the associated listed items and can be abbreviated as "or".
[0056] The existing distributed file system generally uses an ordered key-value storage system (such as RocksDB) to store and manage metadata, and relies on the range query operation provided by the ordered index to realize the directory traversal semantics in the file system. In this way, the directory semantics and the storage index of the metadata are tightly coupled, so that only an ordered key-value index structure can be used to manage the metadata, and only the range query operation provided by the ordered index can be relied on to realize the directory traversal operation of the file system when determining the directory semantics. This makes the prior art have the following problems under the high-speed network and persistent memory storage:
[0057] (1) The delay and throughput performance of metadata operations are poor, especially in the high-speed network and persistent memory hardware environment.
[0058] For example, a query for a certain file metadata may need to go through layer upon layer of directory queries to finally query the specific file metadata, and the layer upon layer of directory queries results in a large amount of time delay.
[0059] (2) As the size of the file system metadata grows, the maintenance cost of the ordered index increases significantly, and the system scalability is limited.
[0060] Specifically, since the ordered key-value index needs to arrange all the key values of the metadata in order, when the metadata scale is large, the sorting of the key values will consume a large maintenance cost.
[0061] (3) Frequent metadata updates require complex and costly logging or transaction mechanisms.
[0062] Based on this, the present application provides a metadata storage method, device, computer equipment and storage medium. For any metadata of a file system, the metadata in the file system can be stored based on an unordered index structure. Under the unordered index structure, even if the metadata scale is enhanced, the metadata can be quickly queried, and the complexity of metadata operation is reduced. In addition, in the metadata, the storage location information of the associated metadata of the second node associated with the first node of the metadata is linked by a chain table. When determining the directory semantics, the directory chain table can be traversed to complete the determination, thereby maintaining the logical relationship between the metadata.
[0063] The metadata storage method provided by the present application will be described below in combination with specific embodiments. The method is applied to a distributed file system. Referring to FIG. 1, a flowchart of a metadata storage method provided by the present application is shown, which includes the following steps: Figure 1
[0064] S101, for any metadata of the file system, the metadata is stored in a metadata server, wherein the metadata server is used to store the metadata in the file system based on an unordered index structure.
[0065] S102, based on the location information of the first node corresponding to the metadata on a hierarchical directory tree, a second node associated with the first node in the hierarchical directory tree is determined, wherein the hierarchical directory tree is used to describe a logical structure for storing data of the file system based on an ordered key value, and the second node is associated with the first node, including that the second node is a child node or a sibling node of the first node.
[0066] S103, the storage location information of the associated metadata of the second node in the metadata server is obtained, and based on the association relationship between the first node and the second node, the storage location information of the associated metadata of the second node in the metadata server is linked in the metadata.
[0067] The following is a detailed description of the above steps.
[0068] For S101,
[0069] The method provided in the present application is executed by a metadata server, and all metadata of a file system can be distributed to multiple metadata servers, and the metadata storage can be performed in each metadata server according to the method provided in the present application.
[0070] In a possible implementation, the hierarchical directory tree of the file system can be scattered / split, each node of the hierarchical directory tree can be metadata, and the scattered metadata can include multiple file metadata and directory metadata, the file metadata corresponds to a file node of the hierarchical directory tree, and the directory metadata corresponds to a directory node (including a subdirectory node and a root directory node) of the hierarchical directory tree. Then, the scattered metadata can be distributed to at least one metadata server.
[0071] The hierarchical directory tree is used to describe a logical structure of storing data of the file system in an ordered key value manner, in a possible implementation, the file system can currently store metadata in an ordered key value manner, and then the hierarchical directory tree of the file system can be determined based on the current storage manner of the file system.
[0072] The hierarchical directory tree of the file system can be, for example, Figure 2 as shown, including a root directory node, a subdirectory node, and a file node, and specific data information can be stored in the file node. The root directory node and the subdirectory node can have other subdirectory nodes and file nodes.
[0073] Optionally, if the scattered metadata is distributed to multiple metadata servers, the key value (here, the first key value) of each metadata can be determined first, then the metadata server corresponding to each metadata is determined through a consistent hash algorithm and the key value of each metadata, and then the metadata is distributed to the corresponding metadata server.
[0074] In a possible implementation, the scattered metadata includes multiple file metadata and directory metadata, each directory metadata can include directory access metadata and directory content metadata. The first key value of different types of metadata can be different, for example, for file metadata, the first key value can be the identifier of the parent node of the node corresponding to the file metadata; the first key value of the directory access metadata can also be the identifier of the parent node of the node corresponding to the directory access metadata, and the first key value of the directory content metadata can be the identifier of the node corresponding to the directory metadata.
[0075] Optionally, the file metadata can be used to describe basic attributes, data location and access information of the file, and can include file identification information, file attribute information (such as file type, permission information, creation time, modification time, access time, etc.), data location information, etc.; the directory metadata is applied to maintain the organization structure of the files / subdirectories in the directory, and can include directory identification information, directory name, directory attribute information (such as permission information, creation time, modification time, access time, etc.), etc.
[0076] In a possible implementation, the unordered index structure can include a hash table, and when storing the metadata in the file system based on the unordered index structure, the metadata server can first determine the storage location of the metadata in the hash table, and then store the metadata according to the storage location.
[0077] For example, the key value (here, the second key value) corresponding to the metadata can be determined based on the data type of the metadata, and then the key value corresponding to the metadata is subjected to a hash operation to determine the storage location of the metadata in the hash table.
[0078] The keys corresponding to different types of metadata can be different. For example, the key corresponding to the file metadata can be the name of the file metadata and the identifier of the parent node of the corresponding node of the file metadata; the key corresponding to the directory access metadata can be the name of the directory access metadata and the identifier of the parent node of the corresponding node of the directory access metadata; and the key corresponding to the directory content metadata can be the identifier of the corresponding node of the directory content metadata.
[0079] If the node corresponding to a certain directory metadata is a root node, the directory metadata can not be distinguished from the directory access metadata and the directory content metadata, and the key corresponding to the directory metadata can be the identifier of the root node.
[0080] For example, if the hierarchical directory tree is as shown in FIG. 1, the directory metadata corresponding to the root directory can be the root directory metadata, the directory metadata corresponding to the directory A can be the directory A metadata, the directory metadata corresponding to the directory B can be the directory B metadata, the directory metadata corresponding to the directory C can be the directory C metadata, the directory metadata corresponding to the directory D can be the directory D metadata, the directory metadata corresponding to the directory E can be the directory E metadata, the directory metadata corresponding to the directory F can be the directory F metadata, the directory metadata corresponding to the directory G can be the directory G metadata, the directory metadata corresponding to the directory H can be the directory H metadata, the directory metadata corresponding to the directory I can be the directory I metadata, the directory metadata corresponding to the directory J can be the directory J metadata, the directory metadata corresponding to the directory K can be the directory K metadata, and the directory metadata corresponding to the directory L can be the directory L metadata. Figure 3As shown on the left side, it includes a root directory node Dir, for example, its corresponding identifier is 1, and the root directory node has three file nodes, which are F1, F2 and F3, respectively, and after the hierarchical directory tree is scattered, four metadata are included, which are Dir directory metadata, F1 file metadata, F2 file metadata and F3 file metadata, when storing the four metadata to the hash table, for the Dir directory metadata, since it is a root node, there is no parent node, therefore, the identifier "1" of the root node can be subjected to a hash operation (in the figure, h(1) represents a hash operation on 1), and it is determined that its corresponding position in the hash table is "0", therefore, the Dir directory metadata can be stored at the position "0"; correspondingly, for the F1 file metadata, the F2 file metadata and the F3 file metadata, the identifier (that is, the identifier "1" of Dir) of the parent node of the corresponding node and the file name can be subjected to a hash operation, respectively, to obtain the storage position "2" of F3, the storage position "4" of F1 and the storage position "6" of F2.
[0081] Here, it should be noted that when determining the storage position of each metadata in the hash table, the second key value is subjected to a hash operation, and the first key value is also subjected to a hash operation when the metadata is dispersed to each metadata server, but the calculation methods of the two hash operations can be different, and the two hash operations are different hash operations.
[0082] Among them, after the storage of the above metadata, each storage position of the hash table has a corresponding storage position in the persistent memory (PM), and the metadata storage can actually be understood as being stored in the PM storage position corresponding to the position of the hash table.
[0083] In this way, each metadata in the file system can be stored in an unordered index manner, so that the complexity of other metadata operations (such as file state obtaining operation, read operation, etc.) other than directory traversal is O(1), which reduces the time complexity of metadata operation.
[0084] For S102 and S103,
[0085] The directory semantics means what content is included under the directory. When determining the directory semantics, if the ordered key value index manner in the related art can be realized by directory traversal, in the present application, since the metadata are stored in an unordered index manner, the directory semantics query by directory traversal is not available.
[0086] Based on this, when storing metadata, for any metadata, the position information of a first node corresponding to the metadata on the hierarchical directory tree can be determined, a second node associated with the first node in the hierarchical directory tree can be determined, and the storage location information of the associated metadata corresponding to the second node in the metadata server can be linked in the metadata. That is, the metadata is linked with other metadata in the metadata by a chain table, thereby realizing the maintenance of the logical relationship between the metadata.
[0087] Here, the second node associated with the first node includes that the second node is a child node or a sibling node of the first node, and the storage location information of the associated metadata in the metadata server can be a pointer to the storage location information of the associated metadata in the metadata server, where the storage location information of the associated metadata in the metadata server is the storage location information in the PM.
[0088] For example, for file metadata, the second node associated with the first node can be a sibling node, and the file metadata can include a sibling pointer "sibling_ptr" data field, which can point to the associated metadata of the second node. For directory metadata, the second node associated with the first node can include a sibling node and a child node, and the directory metadata can include a sibling pointer "sibling_ptr" and a child node pointer "child_ptr" data field, the sibling pointer can point to the metadata of the sibling node, and the child node pointer can point to the metadata of the child node.
[0089] In implementation, for each directory metadata, a chain table corresponding to the directory metadata can be maintained, the chain table can be a single chain table, the head pointer of the chain table is the child node pointer of the directory metadata, and the head pointer of the chain table points to the storage location information of the metadata of the first created child node of the directory metadata in the metadata server. Wherein, if there are multiple child nodes under the directory metadata, the head pointer of the chain table of the directory metadata points to the storage location information of the metadata of the first created child node of the directory metadata in the metadata server.
[0090] For example, if the child nodes under the directory A include file 1, file 2 and subdirectory 1, and the creation order of file 1, file 2 and subdirectory 1 is file 1> file 2> subdirectory 1, then the chain table of directory A is composed of: the child node pointer of directory A points to the storage location information of the metadata of file 1, the sibling node pointer of file 1 points to the storage location information of the metadata of file 2, and the sibling node pointer of file 2 points to the storage location information of the metadata of subdirectory 1.
[0091] In this way, when determining the directory semantics of directory A, the head pointer of the linked list of directory A can be used to find file 1, and then file 2 can be found through the sibling pointer of file 1, and subdirectory 1 can be found through the sibling pointer of file 2, so as to construct the complete directory semantics of directory A: file 1, file 2, and subdirectory 1.
[0092] By storing the metadata in an unordered index manner and maintaining the directory semantics of the metadata in a linked list manner, the decoupling between the metadata index and the directory semantics can be achieved, and the complexity of metadata operations is reduced.
[0093] In a possible implementation, when storing the metadata in a corresponding storage location, the metadata can be stored in a preset memory layout in the corresponding location. In the preset memory layout, data fields related to the same metadata operation are arranged in the same aligned persistent memory block, and the length of the persistent memory block is a preset length, which is generally 16 bytes, and 16 bytes will be taken as an example in the following description. The data in the same aligned persistent memory block is processed in a same CPU atomic operation.
[0094] Here, the preset memory layout corresponding to different types of metadata can be different. For example, the first memory layout corresponds to file metadata, and the second memory layout corresponds to directory metadata. The first memory layout and the second memory layout are different. In the first memory layout, the file content modification time and the file byte number are arranged in the same aligned persistent memory block. The mode, the user identifier of the file, the file state change time, and the low two bits of the group identifier of the file are arranged in the same aligned persistent memory block. The high two bits of the group identifier of the file, the state information of the file, and the creation time are arranged in the same aligned persistent memory block. In the second memory layout, the file content modification time and the head pointer pointing to the subnode linked list are arranged in the same aligned persistent memory block. The mode, the user identifier of the file, the file state change time, and the low two bits of the group identifier of the file are arranged in the same aligned persistent memory block. The high two bits of the group identifier of the file, the state information of the file, and the creation time are arranged in the same aligned persistent memory block.
[0095] For example, the first memory layout can be as shown in Figure 4 The second memory layout can be as shown in Figure 5 In the two memory layouts, each character represents the meaning shown in the following table:
[0096] ID Globally Unique Identifier, also known as inode number dev Identifies the device on which the file resides mode Identifies the file type and access permissions gid_l Lower two bytes of the file's group ID (group ID) uid Owner of the file (user ID) ctime File status (i.e., file metadata information) change time gid_u Upper two bytes of the file's group ID (group ID) status File status information (normal state, deleted state, etc.) bdtime Creation time sibling_ptr Pointer to the next sibling node metadata parent_ptr Pointer to the parent node metadata mtime File content modification time atime Last access time to the file content
[0097] In the first memory layout corresponding to the file metadata, the metadata fields specific to the first memory layout include the following:
[0098] size File size in bytes blksize File block size (in bytes) nlink Number of hard links to the file rdev If the file is a device file, the rdev field contains the major and minor IDs of the device
[0099] The metadata field unique in the second memory layout corresponding to the directory metadata includes the following:
[0100] child_ptr Head pointer to the child node list ll_len Length of the child node list ll_mid Middle node of the child node list
[0101] In the above table, ctime records the last modification time of the metadata (such as permission, owner, size, etc.), and mtime records the last modification time of the file content. There is a sequential relationship between ctime and mtime, that is, ctime is always monotonically increasing, and since metadata modification (triggering ctime update) is often accompanied or later than content modification (triggering mtime update), ctime must be greater than or equal to mtime. Based on this feature, when the metadata operation needs to update ctime and mtime at the same time, there is no need to explicitly write ctime, only mtime needs to be updated, and any subsequent access can dynamically infer ctime by max(ctime, mtime).
[0102] In a possible implementation, in response to receiving a target metadata operation request of any metadata, in a case where the metadata update field corresponding to the target metadata operation request includes the file content modification time (mtime) and the file state change time (ctime), only the file content modification time of the any metadata is updated; and then after receiving an acquisition request for the file state change time of the any metadata, the maximum value of the updated file content modification time and the file state change time recorded in the any metadata is taken as the file state change time corresponding to the acquisition request.
[0103] Based on the above preset memory layout and the update setting of ctime and mtime, the method provided by the application can guarantee crash consistency.
[0104] Now that the central processing unit (CPU) supports 16-byte aligned atomic memory operations, the above metadata can be stored in the PM after storage. Since the persistence granularity of the PM is performed according to the cache line, the persistence granularity of the metadata after being arranged according to the above preset memory layout is one cache line byte, which also provides atomic update and persistence capability for the aligned 16-byte data.
[0105] Specifically, all single-point metadata operations can complete data update through a single CPU atomic write operation. Exemplarily, the single-point metadata operation can include the following operations:
[0106] Write and truncate:
[0107] Write is divided into two types: overwrite and append. Overwrite operation only needs to update the modification time (mtime) of the file. Append operation needs to update both the modification time (mtime) and the file size (size). Truncate operation needs to update the modification time (mtime), the change time (ctime) and the file size (size), which is similar to the append operation.
[0108] For overwrite operation, only 8-byte atomic write is needed to update mtime. For append operation, both mtime and size need to be updated due to the memory layout according to Figure 4 mtime and size are in an aligned 16-byte persistent memory block, so they can be executed by a single atomic operation. Truncate operation is similar to the append operation, and also needs to update mtime and size.
[0109] Read and readdir:
[0110] Read operation (Read) refers to reading file content, which only needs to update the access time (atime).
[0111] Readdir operation refers to listing files / subdirectories under the directory, which also only needs to update the access time (atime). Updating atime only needs a single atomic operation to complete.
[0112] Chown and chmod:
[0113] Chown operation involves modifying the user ID (uid) and group ID (gid) of the file, and chmod operation involves modifying the permissions (mode) of the file. Both operations need to update the change time (ctime). These fields together occupy 18 bytes (4 bytes of uid, 4 bytes of gid, 2 bytes of mode, 8 bytes of ctime), which exceeds the size of a persistent memory block.
[0114] To support simultaneous update of the above fields by one 16-byte write operation, the gid can be split into a high two-byte field (gid_u) and a low two-byte field (gid_l) of the queue, and the low two-byte field (gid_l) and other required fields (uid, mode, ctime) can be put into one aligned 16-byte persistent memory block, so that when the Chown and Chmod operations are performed, the fields can be updated by one 16-byte atomic write operation when the gid is less than 2 16 .
[0115] In addition, for the multi-point metadata operation, it can also be completed by a few CPU atomic operations. Specifically, the multi-point metadata operation can include the following:
[0116] Create file operation (Create): a new file is created under a specified directory, and the metadata of the file and the metadata of the parent directory need to be updated simultaneously. The process of creating file metadata is as shown in Figure 6 , including the following steps:
[0117] (1) Insert the file metadata into the hash table (atomic write operation 1).
[0118] A new file metadata F2 (including the file name, parent pointer, sibling pointer, etc.) is inserted into the hash table, and the second key value is (1 / F2). At this time, the file status is marked as creating (StatusCreating, SC). In the diagram, "P" represents the parent pointer, and the parent pointer of F2 points to the metadata identified as "1". "S" represents the sibling pointer, and the sibling pointer of F2 points to F1.
[0119] (2) Update the parent directory metadata (atomic write operation 2).
[0120] The parent directory metadata (i.e., the directory metadata identified as "1") needs to update the mtime and child_ptr pointer, and the child_ptr pointer needs to point to the metadata of the new file (i.e., the new file F3 is added to the chain table of the directory). Since the second preset layout is used above, the mtime and child_ptr are in the same aligned persistent memory block, so the update of the parent directory metadata can be completed by one atomic write operation. Figure 5
[0121] (3) Update the file status (atomic write operation 3).
[0122] By one atomic write operation, the file status of F2 can be updated from SC to normal (StatusNormal, SN), indicating that the creation is completed.
[0123] Mkdir operation: create a subdirectory under the specified directory, in addition to updating the directory access metadata and parent directory metadata, it also needs to insert the directory content metadata, specifically, it can include the following steps:
[0124] (1) Insert directory access metadata (same as step (1) of the create file operation).
[0125] The access metadata of the subdirectory (including name, parent directory pointer, etc.) is inserted into the hash table, and the status flag is set to SC.
[0126] (2) Update parent directory metadata (same as step (2) of the create file operation):
[0127] The mtime and child_pt of the parent directory are updated by atomic writing, and the subdirectory is added to the parent directory linked list.
[0128] (3) Insert directory content metadata.
[0129] The content metadata of the subdirectory (used to maintain the linked list of the subdirectory itself) is inserted into the hash table (key is the subdirectory's own ID). Since the content metadata can be generated from the access metadata, even if it crashes, it can be reconstructed without the need for distributed transactions.
[0130] (4) Update directory status (same as step 3 of create):
[0131] The status of the subdirectory access metadata is updated from SC to SN, and the creation is completed.
[0132] Delete file and delete directory operations (Delete and rmdir):
[0133] Delete file (Delete): remove the specified file, mark the file status as "deleting", and update the parent directory metadata, which can include the following steps:
[0134] (1) Mark the file as "deleting" (atomic write operation 1):
[0135] By a 16-byte atomic write, the file status is marked as StatusDeleting (deleting), and the bdtime (create-delete timestamp) is updated to record the deletion initiation time.
[0136] (2) Update parent directory metadata (atomic write operation 2):
[0137] The mtime of the parent directory is updated by 8-byte atomic writing (commit point) - if the system crashes at this point, the file status is Deleting but the parent directory has been updated, and it will be treated as "should be deleted" and completed later.
[0138] (3) Mark the file as "deleted" (atomic write operation 3):
[0139] Update the file status from StatusDeleting to StatusDeleted (deleted). The physical deletion of the file (removing from the hash table and the linked list) is performed later (e.g., updating the linked list pointer when traversing the directory, skipping the deleted node), reducing the real-time operation overhead.
[0140] Delete directory (rmdir): remove the specified subdirectory, first check if the directory is empty, then perform similar steps as file deletion, and additionally delete the directory content metadata, which can include the following steps:
[0141] (1) Check if the directory is empty:
[0142] Access the child_ptr (linked list head pointer) of the subdirectory content metadata to confirm that the linked list is empty (no child node), otherwise the deletion fails.
[0143] (2) Mark the directory access metadata as "deleting" (atomic write operation 1):
[0144] As in step 1 of the delete operation, mark the directory access metadata as StatusDeleting and update the bdtime.
[0145] (3) Update the parent directory metadata (atomic write operation 2):
[0146] As in step 2 of the delete operation, update the mtime of the parent directory.
[0147] (4) Clean up the directory metadata and mark it as "deleted" (atomic write operation 3):
[0148] Specifically, the directory access metadata status can be updated to StatusDeleted, and the directory content metadata can be deleted from the hash table (since the directory is empty, the content metadata is not necessary).
[0149] In the above operations, whether it is a single-point atomic operation or a multi-point atomic operation, it can be completed by a small number of atomic operations using the above memory layout and the update mechanism of mtime.
[0150] In one possible implementation, after receiving a directory traversal request sent by a client, the directory traversal result can be returned to the client based on at least one of the following methods:
[0151] While reading the linked list node, return the read part of the result to the client;
[0152] Splitting the linked list into multiple segments and reading the split multiple segments of the linked list in parallel through multiple threads, returning the merged directory traversal result to the client after merging the read results;
[0153] Determining whether the directory traversal result corresponding to the directory traversal request is cached in the dynamic random access memory (DRAM), if yes, reading the directory traversal result corresponding to the directory traversal request from the dynamic random access memory (DRAM), and returning the read directory traversal result to the client; wherein the DRAM caches the historical directory traversal result within a preset time range before the current time;
[0154] Returning the read directory traversal result to the client after compression.
[0155] Exemplarily, the optimization process of directory traversal is as shown in Figure 7 , which comprises:
[0156] 1. Basic: serial read process without optimization:
[0157] After the client (Client) sends a directory traversal request (readdir) through remote procedure call (RPC), the server serially reads the linked list nodes (accesses the metadata of each file in turn along the sibling_ptr) from the persistent memory (PM), and returns the results to the client at once after reading all.
[0158] However, PM reading and network transmission are serial, and if the directory contains a large number of files (such as millions), the total time consumption = total PM reading time + network transmission time, and the delay is high.
[0159] 2. Pipeline optimization (+Pipeline): overlapping PM reading and network transmission process:
[0160] The server starts returning the read results to the client through the network while reading the first node in the PM; when reading the second node, it continues to transmit the results of the first node, and so on, so that the "PM reading" and "network transmission" are performed in parallel (pipelined overlap), which can hide part of the PM reading delay.
[0161] 3. Parallel reading optimization (+Parallel): split linked list multi-segment reading process:
[0162] For a directory containing a large number of files, the server splits the linked list into multiple segments (such as two segments) according to the length (ll_len) and midpoint pointer (ll_mid), and multiple threads read different segments of nodes in parallel, and after reading is completed, the results are combined and returned to the client. In this way, the problem of "node-by-node serial reading" of a long linked list can be solved, and the traversal time is reduced.
[0163] 4. DRAM cache optimization (+Cache): reuse the traversal result process of the hot directory:
[0164] The server caches the recently traversed (i.e. within a preset time range before the current time) directory linked list nodes to DRAM. When the client traverses the same directory again, the server directly reads the cached results from DRAM, without the need to access the PM again.
[0165] In this way, for frequently accessed directories, the traversal time can be close to the DRAM access speed, avoiding repeated consumption of PM bandwidth.
[0166] 5. Data compression optimization (+Compress): reduce network transmission flow:
[0167] The server compresses the directory entries (such as file names, inode numbers, and other high-repetition information) before returning the traversal results to the client, and then transmits the compressed data, which is then decompressed by the client for use.
[0168] In this way, the amount of data transmitted over the network can be reduced, and network latency can be reduced.
[0169] Corresponding to the above-mentioned embodiments of the metadata storage method, the present application also provides embodiments of a metadata storage device. Figure 8 The schematic diagram of the metadata storage device provided by the present application specifically includes:
[0170] The storage module 801 is configured to store any metadata of the file system in a metadata server, wherein the metadata server is configured to store the metadata in the file system based on an unordered index structure;
[0171] The determination module 802 is configured to determine a second node associated with a first node in a hierarchical directory tree based on position information of the first node in the hierarchical directory tree, wherein the hierarchical directory tree is configured to describe a logical structure of storing data of the file system based on an ordered key value, and the second node is associated with the first node, including that the second node is a child node or a sibling node of the first node.
[0172] The processing module 803 is configured to acquire storage location information of the association metadata corresponding to the second node in the metadata server, and link the storage location information of the association metadata corresponding to the second node in the metadata server in the metadata based on the association relationship between the first node and the second node.
[0173] Optionally, the unordered index structure comprises a hash table.
[0174] The storage module 801 is configured to, when storing the metadata in the metadata server, perform the following steps.
[0175] Determine a key value corresponding to the metadata based on a data type of the metadata.
[0176] Perform hash operation on the key value corresponding to the metadata to determine a storage location of the metadata in the hash table.
[0177] Store the metadata according to the storage location.
[0178] Optionally, the metadata of the file system comprises file metadata and directory metadata, and the directory metadata comprises directory access metadata and directory content metadata.
[0179] The key corresponding to the file metadata is a name of the file metadata and an identifier of a parent node of a corresponding node of the file metadata.
[0180] The key corresponding to the directory access metadata is a name of the directory access metadata and an identifier of a parent node of a corresponding node of the directory access metadata.
[0181] The key corresponding to the directory content metadata is an identifier of a corresponding node of the directory content metadata.
[0182] Optionally, the storage module 801 is configured to, when storing the metadata according to the storage location, perform the following steps.
[0183] Store the metadata according to a preset memory layout in the corresponding storage location.
[0184] In the preset memory layout, data fields related to a same metadata operation are arranged in a same aligned persistent memory block, and data in the same aligned persistent memory block is processed in a same CPU atomic operation.
[0185] Optionally, the data field of the metadata comprises a file content modification time and a file state change time.
[0186] The processing module 803 is further configured to perform the following steps.
[0187] In response to receiving a target metadata operation request of any metadata, if a file content modification time and a file state change time are included in a metadata update field corresponding to the target metadata operation request, only the file content modification time of the any metadata is updated;
[0188] After receiving a request for obtaining a file state change time of the any metadata, a maximum value of the updated file content modification time and a file state change time recorded in the any metadata is taken as a file state change time corresponding to the request for obtaining.
[0189] Optionally, the metadata includes file metadata and directory metadata, the file metadata corresponds to a first memory layout, and the directory metadata corresponds to a second memory layout, the first memory layout and the second memory layout are different.
[0190] In the first memory layout, a file content modification time and a file byte number are arranged in a same alignment persistent memory block; a mode, a user identifier of a file, a file state change time, and a low two-bit field of a group identifier of the file are arranged in a same alignment persistent memory block; a high two-bit field of the group identifier of the file, a state information of the file, and a creation time are arranged in a same alignment persistent memory block.
[0191] In the second memory layout, a file content modification time and a head pointer pointing to a child node list are arranged in a same alignment persistent memory block; a mode, a user identifier of a file, a file state change time, and a low two-bit field of a group identifier of the file are arranged in a same alignment persistent memory block; a high two-bit field of the group identifier of the file, a state information of the file, and a creation time are arranged in a same alignment persistent memory block.
[0192] Optionally, the apparatus further includes a returning module 804, configured to:
[0193] receive a directory traversal request sent by a client, and return a directory traversal result to the client based on at least one of the following manners:
[0194] return a read part of a result to the client while reading a list node;
[0195] split a list into multiple segments, read the split multiple segments of the list in parallel through multiple threads, return a combined directory traversal result to the client after the read result is combined, and
[0196] determining whether the directory traversal result corresponding to the directory traversal request is cached in the dynamic random access memory (DRAM), and if so, reading the directory traversal result corresponding to the directory traversal request from the dynamic random access memory (DRAM) and returning the read directory traversal result to the client; wherein the dynamic random access memory (DRAM) caches historical directory traversal results in a preset time range before the current time;
[0197] returning the read directory traversal result to the client after compression.
[0198] The implementation process of the functions and roles of the units in the above apparatus is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here.
[0199] For the device embodiment, since it basically corresponds to the method embodiment, the relevant part can be referred to the part of the method embodiment. The above described device embodiment is only illustrative, wherein the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place or distributed on multiple network units. According to the actual needs, part or all of the modules can be selected to achieve the purpose of the scheme of the present application. Those skilled in the art can understand and implement it without creative labor.
[0200] The present application also provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the metadata storage method described in the above embodiment.
[0201] The present application also provides a computer device, as shown in Figure 9 The structure schematic diagram of the computer device provided by the present application is shown in the figure, and at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and of course, it can also include other hardware required by the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to realize the metadata storage method described in the above embodiment. Of course, in addition to the software implementation, the present specification does not exclude other implementation modes, such as logic device or software and hardware combined mode, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but also can be hardware or logic device.
[0202] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0203] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit), and the apparatus can be implemented as special purpose logic circuitry.
[0204] Computers suitable for the execution of a computer program include, by way of example, general and / or special purpose microprocessors, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory and / or a random access memory. The essential elements of a computer are a central processing unit for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.
[0205] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0206] While the specification contains many specifics, these should not be construed as limiting the scope of any invention or of what can be claimed, but as merely providing illustrations of some of the embodiments of the inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0207] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0208] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.
[0209] The above description is merely illustrative of the application and does not limit the scope of the application as determined by the appended claims.
Claims
1. A metadata storage method applied to a distributed file system, characterized in that, The method comprises: for any metadata of the file system, storing the metadata in a metadata server, wherein the metadata server is configured to store the metadata in the file system based on an unordered index structure; based on the position information of a first node corresponding to the metadata on a hierarchical directory tree, determining a second node associated with the first node in the hierarchical directory tree, wherein the hierarchical directory tree is configured to describe a logical structure for storing data of the file system in an ordered key value manner, and the second node is associated with the first node, including that the second node is a child node or a sibling node of the first node; obtaining the storage location information of the associated metadata corresponding to the second node in the metadata server; and based on the association relationship between the first node and the second node, linking the storage location information of the associated metadata corresponding to the second node in the metadata server in the metadata.
2. The method of claim 1, wherein, The unordered index structure comprises a hash table; storing the metadata in the metadata server comprises: based on the data type of the metadata, determining the key value corresponding to the metadata; performing hash operation on the key value corresponding to the metadata to determine the storage location of the metadata in the hash table; storing the metadata according to the storage location.
3. The method of claim 2, wherein, The metadata of the file system comprises file metadata and directory metadata, and the directory metadata comprises directory access metadata and directory content metadata; the key corresponding to the file metadata is the name of the file metadata and the identifier of the parent node of the corresponding node of the file metadata; the key corresponding to the directory access metadata is the name of the directory access metadata and the identifier of the parent node of the corresponding node of the directory access metadata; the key corresponding to the directory content metadata is the identifier of the corresponding node of the directory content metadata.
4. The method of claim 2, wherein, storing the metadata according to the storage location comprises: storing the metadata according to the preset memory layout in the corresponding storage location; wherein, in the preset memory layout, the data fields related to the same metadata operation are arranged in the same aligned persistent memory block, and the data in the same aligned persistent memory block is processed in the same CPU atomic operation.
5. The method of claim 4, wherein, The data fields of the metadata comprise file content modification time and file state change time; The method further comprises: in response to receiving a target metadata operation request of any metadata, if the metadata update field corresponding to the target metadata operation request comprises the file content modification time and the file state change time, only updating the file content modification time of the any metadata; after receiving a request for obtaining the file state change time of the any metadata, taking the maximum value of the updated file content modification time and the file state change time recorded in the current any metadata as the file state change time corresponding to the request.
6. The method of claim 4, wherein, The metadata of the file system includes file metadata and directory metadata, the file metadata corresponds to a first memory layout, and the directory metadata corresponds to a second memory layout, the first memory layout and the second memory layout are different; In the first memory layout, the file content modification time and the file byte number are arranged in the same aligned persistent memory block; The mode, the user identifier of the file, the file state change time, and the low two bytes of the group identifier of the file are arranged in the same aligned persistent memory block; The high two bytes of the group identifier of the file, the state information of the file, and the creation time are arranged in the same aligned persistent memory block; In the second memory layout, the file content modification time and the head pointer pointing to the child node list are arranged in the same aligned persistent memory block; The mode, the user identifier of the file, the file state change time, and the low two bytes of the group identifier of the file are arranged in the same aligned persistent memory block; and the high two bytes of the group identifier of the file, the state information of the file, and the creation time are arranged in the same aligned persistent memory block.
7. The method of claim 1, wherein, The method further comprises: receiving a directory traversal request sent by a client, and returning a directory traversal result to the client based on at least one of the following manners: returning the read part of the result to the client while reading the list node; splitting the list into multiple segments and reading the split list segments in parallel through multiple threads, and returning the combined directory traversal result to the client after combining the read results; determining whether the directory traversal result corresponding to the directory traversal request is cached in a dynamic random access memory (DRAM), if yes, reading the directory traversal result corresponding to the directory traversal request from the DRAM, and returning the read directory traversal result to the client; wherein the DRAM caches historical directory traversal results within a preset time range before the current time; returning the read directory traversal result to the client after compression.
8. A metadata storage apparatus applied to a distributed file system, characterized in that, The apparatus comprises: a storage module configured to store any metadata of the file system in a metadata server, wherein the metadata server is configured to store the metadata in the file system based on an unordered index structure; a determination module configured to determine a second node associated with a first node in a hierarchical directory tree based on position information of the first node in the hierarchical directory tree, wherein the hierarchical directory tree is configured to describe a logical structure of storing data of the file system based on an ordered key-value manner, and the second node is associated with the first node, including that the second node is a child node or a sibling node of the first node; a processing module configured to obtain storage position information of associated metadata of the second node in the metadata server, and link the storage position information of the associated metadata of the second node in the metadata server in the metadata based on an association relationship between the first node and the second node.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by a processor to implement the steps of the method of any one of claims 1-7.
10. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the steps of the method of any of claims 1-7.
Citation Information
Patent Citations
Internal memory database system and method and device for implementing internal memory data base
CN101315628A
Index tree building method, Chinese vocabulary searching method and related device
CN103514287A
Decoupling distribution method for metadata of distributed file system
CN106874383A
Declarative Software Application Meta-Model and System for Self-Modification
US20130246996A1
Hash multi-table join implementation method based on grouping vector
WO2020248604A1