Metadata Update Method, Device, Node Device and Computer Readable Storage Medium
By obtaining the file index node identity of the target file in a distributed file system, updating its metadata and metadata of the associated directory, so that it can be stored or unbined on the same node device, the performance impact caused by cross-node transactions when creating hard links is solved, and the hard link function is supported and performance optimization is achieved.
Patent Information
- Application Number
- CN202411983926.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-31
AI Technical Summary
When creating hard links, existing distributed file systems need to modify metadata through cross-node transactions, resulting in performance impacts and cannot support hard link functions.
By obtaining the file index node identification of the target file, update the metadata of the target file and the metadata of the associated directories related to its hard links, so that it is stored on the same node device, or unbinding its metadata to make its updates independent of each other and avoid cross-node transactions.
It realizes the avoidance of cross-node transactions when creating hard links, supports hard link functions, and reduces the impact on the overall performance of distributed file systems.
Smart Images

Figure CN119396794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed file systems, and in particular to a metadata updating method, apparatus, node device and computer-readable storage medium. Background Art
[0002] The metadata of a high-performance distributed file system is usually organized in the form of Key-Value pairs, and the metadata is also stored in the KV database in the form of Key-Value pairs. In order to ensure the reliability of metadata modification, the database transaction model is usually used to ensure that metadata modification meets ACID, namely Atomicity, Consistency, Isolation, and Durability.
[0003] The metadata of a distributed file system is usually stored in the form of Key-Value (KV pairs) in a distributed manner in the KV databases of multiple nodes. When updating files in a distributed file system, it usually involves the simultaneous modification of the corresponding metadata in multiple KV databases. At this time, cross-node distributed transactions are required to ensure that the modification of the corresponding metadata in multiple KV databases meets ACID. However, cross-node distributed transactions will affect the overall performance of the distributed file system.
[0004] In order to avoid the impact of cross-node distributed transactions on overall performance, the existing implementation method is to bind the file and the directory to which it belongs, and determine the node device that stores the file metadata according to the directory to which it belongs, thereby avoiding the modification of metadata across nodes. However, this implementation method requires that a file can only belong to one directory. When creating a hard link for a file, a file may belong to multiple directories. Therefore, this method cannot support hard links. Summary of the invention
[0005] The purpose of the present invention is to provide a metadata updating method, apparatus, node device and computer-readable storage medium, which can not only support hard links, but also avoid modifying metadata through cross-node transactions when creating hard links, ultimately achieving both support for hard links and reducing the impact of cross-node transactions on the overall performance of the distributed file system.
[0006] The embodiments of the present invention can be implemented as follows:
[0007] In a first aspect, the present invention provides a metadata updating method, which is applied to each node device of a plurality of node devices in a distributed file system, and the method comprises:
[0008] Get the file index node identifier of the target file for which a hard link needs to be created;
[0009] Update the metadata of the target file according to the file inode identifier, and update the metadata of the associated directory related to the hard link of the target file;
[0010] In the updated metadata of the target file, the inode identifiers of the associated directory and the target directory do not exist, or the metadata of the target file, the metadata of the associated directory related to the target file, and the metadata of the target directory related to the target file are stored in the same node device among the multiple node devices, where the target directory is the directory specified when creating the target file.
[0011] In an optional implementation manner, the step of updating the metadata of the target file according to the file inode identifier includes:
[0012] Query the metadata of the target file according to the file inode identifier, where the metadata of the target file includes the directory inode identifier of the target directory;
[0013] Update the directory inode identifier to an invalid value;
[0014] Migrate the metadata of the target file to a preset node device among the multiple node devices.
[0015] In an optional implementation manner, the node device caches the mapping relationship between the file inode identifier and the directory inode identifier, and the method further includes:
[0016] Based on a file query operation, obtain the directory inode identifier according to the mapping relationship, where the file query operation is used to query the target file attribute information according to the file inode identifier;
[0017] Determine a target node from the multiple node devices according to the directory inode identifier;
[0018] Search for the Value value corresponding to the Key value in the target node with the file inode identifier as the Key value;
[0019] If the Value value is an invalid value, obtain the target file attribute information from the preset node device, otherwise obtain the target file attribute information from the target node according to the file inode identifier.
[0020] In an optional implementation manner, the method further includes:
[0021] Based on a directory query operation for querying all directory entries of a specified directory, determine the node to be queried for each directory entry from the multiple node devices;
[0022] Query the metadata of the corresponding directory entry from the nodes to be queried in each directory entry, and respond to the directory query operation according to the metadata of all the queried directory entries.
[0023] In an alternative embodiment, the step of determining the nodes to be queried for each directory entry from the multiple node devices includes:
[0024] Obtain the inode identifier of the specified directory, and determine the first node to be queried from the multiple node devices according to the inode identifier of the specified directory;
[0025] Query the type and inode identifier of each directory entry in the specified directory from the first node to be queried;
[0026] For any directory entry, if the type of the directory entry is a directory type, then determine the second node to be queried from the multiple node devices according to the inode identifier of the directory entry, and use the first node to be queried and the second node to be queried as the two nodes to be queried for the directory entry, where the first node to be queried stores the first metadata of the directory entry, the second node to be queried stores the second metadata of the directory entry, and the modification frequency of the first metadata is less than that of the second metadata;
[0027] If the type of the directory entry is a file type and the inode identifier of the directory entry is not an invalid value, then use the first node to be queried as the node to be queried for the directory entry;
[0028] If the type of the directory entry is a file type and the inode identifier of the directory entry is an invalid value, then use the preset node device as the node to be queried for the directory entry.
[0029] In an alternative embodiment, the method further includes:
[0030] When there is no hard link for the target file, obtain the inode identifier of the directory;
[0031] Update the inode identifier of the directory to the metadata of the target file;
[0032] Migrate the metadata of the target file from the preset node device to the node device determined by the inode identifier of the directory.
[0033] In an alternative embodiment, before the step of obtaining the file inode identifier of the target file for which a hard link needs to be created, it includes:
[0034] Based on the creation operation of creating the target file in the target directory, obtain the file inode identifier of the target file;
[0035] Determine a node to be updated from the multiple node devices according to the file inode identifier;
[0036] Create a first transaction on the node to be updated, and update the metadata of the target file and the metadata of the target directory in the first transaction.
[0037] In an optional implementation manner, the step of updating the metadata of the target file according to the file inode identifier and updating the metadata of the associated directory related to the hard link of the target file includes:
[0038] Create a second transaction on the node to be updated, and update the inode identifier of the associated directory to the metadata of the target file and update the metadata of the associated directory in the second transaction.
[0039] In an optional implementation manner, the method further includes:
[0040] Based on a file query operation for querying the attribute information of the target file, obtain the file inode identifier;
[0041] Determine a target node from the multiple node devices according to the file inode identifier;
[0042] Obtain the target file attribute information from the target node.
[0043] In a second aspect, the present invention provides a metadata update device, which is applied to each node device in multiple node devices in a distributed file system. The device includes:
[0044] An acquisition module, configured to acquire the file inode identifier of a target file for which a hard link needs to be created;
[0045] An update module, configured to update the metadata of the target file according to the file inode identifier and update the metadata of the associated directory related to the hard link of the target file;
[0046] In the updated metadata of the target file, the inode identifiers of the associated directory and the target directory do not exist, or the metadata of the target file, the metadata of the associated directory related to the target file, and the metadata of the target directory related to the target file are stored in the same node device among the multiple node devices. The target directory is the directory specified when creating the target file.
[0047] In a third aspect, the present invention provides a node device, including a processor and a memory. The memory is used to store a program, and the processor is used to implement the metadata update method according to any one of the foregoing implementation manners when executing the program.
[0048] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the metadata updating method described in any one of the aforementioned embodiments.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] When creating a hard link to a target file, the present invention uses a file index node identifier to update the metadata of the target file and the metadata of an associated directory related to the hard link to the target file, either by not having the index node identifiers of the associated directory and the target directory in the metadata of the updated target file, so that the metadata of the target file and the directory to which it belongs are unbound, so that updates of the two are independent of each other and do not need to be restricted to the same transaction, thereby avoiding modification of metadata through cross-node transactions; or by storing the metadata of the target file, the metadata related to the target file in the associated directory, and the metadata related to the target file in the target directory in the same node device, so that updates of the three can be completed in the same node device, thereby avoiding modification of metadata through cross-node transactions. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0052] Figure 1 This is an example diagram of an application scenario provided by this embodiment.
[0053] Figure 2 This is a block diagram of an example node device provided in this embodiment.
[0054] Figure 3 This is an example diagram of a file hard link provided in this embodiment.
[0055] Figure 4 This is an example flowchart of the metadata updating method provided in this embodiment.
[0056] Figure 5 The embodiment provided Figure 3 Example of metadata after creating a hard link in Figure 1 .
[0057] Figure 6 The embodiment provided Figure 3 Example of metadata after creating a hard link in Figure 2 .
[0058] Figure 7 It is a block diagram example of the metadata update device provided in this embodiment.
[0059] Icons: 10 - Node device; 11 - Processor; 12 - Memory; 13 - Bus; 20 - Client; 100 - Metadata update device; 110 - Acquisition module; 120 - Update module; 130 - Query module; 140 - Creation module. Detailed implementation manners
[0060] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0061] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] It should be noted that: like reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0063] In the description of the present invention, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0064] In addition, if terms such as "first", "second", etc. are only used for distinguishing descriptions, they cannot be understood as indicating or implying relative importance.
[0065] It should be noted that the features in the embodiments of the present invention can be combined with each other without conflict.
[0066] Please refer to Figure 1 , Figure 1 which is an example diagram of the application scenario provided in this embodiment, Figure 1Among them, the distributed file system includes multiple node devices 10. The distributed file system is communicatively connected to the client 20. The client 20 creates a file in the distributed file system and writes data to the created file. The data written by the client 20 is stored in a distributed form in each node device 10. In addition to the data written by the client 20, the distributed file system also needs to manage the files or directories created by the client 20. The data for managing files or directories is called metadata, and the metadata can also be stored in a distributed form in each node device 10. In specific implementation, the storage of data and metadata can be independent of each other and stored in different multiple node devices respectively. The data and metadata can also be distributed in each node device simultaneously.
[0067] A single-node database, such as RocksDB, can be deployed on a single node device 10, or a distributed cluster database, such as TiKV, can be deployed on multiple node devices 10. The metadata is stored in a distributed form in the database deployed on the node device 10 in the form of Key-Value.
[0068] The node device 10 can be a physical computer device or a virtual device with the same function as a physical computer device. The node device 10 can be a server, a storage array, etc.
[0069] Based on Figure 1 , this embodiment also provides a block diagram of the node device 10. The node device 10 implements the metadata update method of the foregoing embodiment. Please refer to Figure 2 , Figure 2 This is the block diagram of the node device 10 provided in this embodiment. The node device 10 includes a processor 11, a memory 12, and a bus 13. The processor 11 and the memory 12 are connected through the bus 13.
[0070] The processor 11 may be an integrated circuit chip with signal processing capability. In the implementation process, each step of the metadata update method of the above embodiment may be completed by an integrated logic circuit of hardware in the processor 11 or by instructions in the form of software. The above processor 11 may be a general-purpose processor, including a CPU (Central Processing Unit), a NP (Network Processor), etc.; it may also be a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Logic Gate Array) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0071] The memory 12 is used to store a program for implementing the metadata updating method. The program may be a software function module stored in the memory 12 in the form of software or firmware or solidified in the OS (Operating System) of the node device 10 .
[0072] After receiving the execution instruction, the processor 11 executes the program to implement the metadata updating method of the above embodiment.
[0073] Before introducing the metadata updating method provided by this embodiment, several concepts involved in this embodiment are first explained.
[0074] Distributed file system: A system that allows files to be stored and managed across multiple devices (such as servers and storage devices), allowing multiple users or clients to access these files simultaneously, and providing high availability, reliability, and scalability. For example, Windows' NTFS (New Technology File System), Linux's ext4, xfs, etc.
[0075] Directory: Used to organize and store other files and subdirectories.
[0076] Directory Entry: An entry in a directory that contains information such as the name of a file or subdirectory.
[0077] File: A collection of data stored on a storage device, which can contain different types of data such as text, images, audio, video, etc.
[0078] Hard Link: A file link method that allows multiple directory entries to point to the same file in a distributed file system. Hard links do not occupy additional storage space. Please refer to Figure 3 , Figure 3 which is an example diagram of the file hard link provided by this embodiment. Figure 3 In it, the square boxes represent directories, such as D1, D2, and D3, and the round boxes represent files, such as F1, F2, F3, and F4. Before creating a hard link for F3, D3 is an empty directory, and F3 only belongs to directory D2, that is, F3 has only one parent directory D2. After creating a hard link for F3, F3 belongs to both directory D2 and D3, that is, F3 has two parent directories D2 and D3.
[0079] inode: A data structure used to store metadata of a file or directory (such as file size, permissions, creation time, modification time, etc.), and pointers to file data blocks. A directory or a file is an inode.
[0080] inode_id: A number that uniquely identifies an inode. Each file or directory has a unique inode_id, which is used to find the corresponding inode in a distributed file system. inode_id is generally represented by a 64-bit unsigned number.
[0081] Key-Value: A basic model for data storage and retrieval. In this model, each data item consists of a unique key and a corresponding value. Key: The key is a unique identifier used to identify and retrieve a specific data item. The key is usually a string or other immutable type of data to ensure its uniqueness. Value: The value can be data of any type and represents the actual stored content. The value can be a simple data type (such as an integer, a string) or a complex data structure (such as a list, a dictionary). In this embodiment, the metadata is organized according to Key-Value.
[0082] Since some operations trigger updates to multiple metadata, and different metadata may be distributed on different node devices, in order to ensure that the updates of metadata meet atomicity, consistency, isolation, and durability, a transaction mechanism is usually used. When the metadata to be updated is distributed on multiple node devices, a cross-node device transaction mechanism needs to be created.
[0083] Taking the file creation operation as an example, at least three operations are involved: (1) modifying the timestamp of the parent directory of the file; (2) adding a directory entry to the parent directory; (3) adding metadata attribute information of the newly created file. These three operations need to be completed in the same database transaction. If it is a distributed transaction of the entire database cluster, it may involve coordination among multiple database nodes, resulting in high performance overhead, thus affecting the overall performance of the distributed file system.
[0084] To reduce the impact of the cross-node device transaction mechanism on performance, the distributed file system InfiniFS handles it as follows: The attributes of the directory are divided into two parts: (1) the immutable part, such as the access rights and owner of the directory; (2) the mutable part: including the timestamp and directory entries; the mutable part and the attributes of its sub-items are both hashed (using hash) to belong to the inode_id of the directory to which they belong, ensuring that their belonging is to the same node device, making the entire transaction a local transaction rather than a cross-node distributed transaction, which theoretically greatly optimizes the file creation performance.
[0085] The distributed file system SingularFS handles it as follows: The file attributes are stored in the KV pairs of the directory entry. The addressing method of the key is the parent directory inode_id + file name, and the file inode_id and file attribute information are stored in the Value. By binding the file to its parent directory, it is ensured that their belonging is to the same node device, making the entire transaction a local transaction rather than a cross-node distributed transaction.
[0086] Although the above two methods both reduce the impact of the cross-node device transaction mechanism on performance, neither of these two methods can support hard links. After in-depth analysis of their implementation principles, the inventor found that the reason why these two methods do not support hard links is as follows: First, for the addressing of the file Key, there is a definite parent directory inode_id, and the file can only support a single parent directory; Second, the storage rule of the file KV pairs is bound to the parent directory inode_id, which also limits the file to only support a single parent directory. While a hard link means that the same file belongs to two different parent directories at the same time, thus restricting the support for hard links.
[0087] In view of this, this embodiment provides a metadata update method, device, node device, and computer-readable storage medium, which can both support hard links and avoid modifying metadata through cross-node transactions when creating hard links, ultimately achieving both support for hard links and reduction of the impact of cross-node transactions on the overall performance of the distributed file system. The following will describe it in detail.
[0088] Please refer to Figure 4 , Figure 4A flowchart example of the metadata update method provided in this embodiment. This method can be applied to Figure 1 and Figure 2 the node devices therein. The method includes the following steps:
[0089] Step S101: Obtain the file index node identifier of the target file for which a hard link needs to be created.
[0090] In this embodiment, the target file is the file for which a hard link needs to be created, and the file index node identifier is the inode_id of the target file. The distributed file system will assign an inode_id to the target file when it is created.
[0091] Step S102: Update the metadata of the target file according to the file index node identifier, and update the metadata of the associated directory related to the hard link of the target file;
[0092] In the updated metadata of the target file, the index node identifiers of the associated directory and the target directory do not exist, or the metadata of the target file, the metadata of the associated directory related to the target file, and the metadata of the target directory related to the target file are stored in the same node device among multiple node devices. The target directory is the directory specified when creating the target file.
[0093] In this embodiment, creating a hard link for the target file will trigger changes in the metadata of the target file and the directory to which the target file belongs. The associated directory is the directory related to the hard link of the target file, and the target directory is the directory specified when creating the target file. For example, when creating the target file F and specifying to create it under the directory D1, then D1 is the target directory of F. When creating a hard link for the target file and specifying to create a hard link from D2 / F to D1 / F, then D2 is the associated directory.
[0094] In this embodiment, the metadata of the associated directory related to the target file, that is, there is a directory entry, which represents the association relationship between the associated directory and the target file. For example, the Key value of this directory entry is: the inode_id of the associated directory and the directory entry name, and the Value value of this directory entry is: the type of the target file and the inode_id of the target file. The metadata of the target directory related to the target file is similar.
[0095] In this embodiment, the metadata of the target file can be unbound from the target directory. As a result, the updates of the two will be independent of each other and do not need to be restricted by the same transaction. It is also possible to store the metadata of the target file, the metadata related to the target file in the associated directory, and the metadata related to the target file in the target directory on the same node device, avoiding modifying the metadata through a cross-node transaction. Both methods can avoid modifying the metadata through a cross-node transaction, achieving the technical effect of supporting hard links while reducing the impact of cross-node transactions on the overall performance of the distributed file system.
[0096] For these two methods, there are at least two corresponding implementation manners for step S102. Each implementation manner and at least one processing procedure for the metadata information based on each implementation manner will be introduced below in sequence.
[0097] First, an implementation manner of unbinding the metadata of the target file from the target directory (referred to as Method 1 in this article) is introduced. In this implementation manner, the attributes of the directory are divided into two parts: one part is the immutable part, that is, the part with a low modification frequency, such as the access permission and owner of the directory, and the other part is the mutable part, that is, the part with a high modification frequency, such as the timestamp and directory entry. Correspondingly, the metadata of the directory also includes two categories: one category is the metadata of the directory related to the immutable part, and the other category is the metadata of the directory related to the mutable part. In addition, there is also metadata related to the attributes of the directory entry and metadata related to the attributes of the file. The KV of these several types of metadata is organized as follows:
[0098] (1) Immutable part of directory attributes (determined by inode-Stable for naming)
[0099] Key: self_inode_id + stable flag
[0100] Value: parent_inode_id, information related to owner permissions, etc.
[0101] Basis for distribution: parent_inode_id
[0102] Among them, the stable flag is used to identify the immutable part of the directory attributes. self_inode_id is the inode_id of the directory itself, and parent_inode_id is the inode_id of the parent directory of this directory. The basis for distribution means the basis for determining the node device. For example, hash calculation can be performed using the inode_id, and the result obtained from the hash calculation is used to determine the node device. For the immutable part of the directory attributes, the metadata of this directory is stored on the node device determined by the parent_inode_id.
[0103] (2)Volatile part of directory attributes (name determined based on inode-Volatile)
[0104] Key: self_inode_id + volatile flag
[0105] Value: Information such as each timestamp
[0106] Basis for distribution: self_inode_id
[0107] Among them, the volatile flag is used to identify the volatile part of directory attributes. self_inode_id is the inode_id of the directory itself. The basis for distribution means that the metadata of this directory is stored on the node device determined by self_inode_id. Each timestamp includes, but is not limited to, the creation time, the last modification time, etc.
[0108] (3)Directory entry (name determined based on DEntry)
[0109] Key: self_inode_id + name (directory entry name)
[0110] Value: type (child item type) + child_inode_id (child item inode_id)
[0111] Basis for distribution: self_inode_id
[0112] Among them, self_inode_id is the inode_id of each directory entry itself, name is the directory entry name. If the directory entry is a file, then self_inode_id is the inode_id of this file, name is the file name of this file. If the directory entry is a directory, then self_inode_id is the inode_id of this directory, name is the directory name of this directory. type is the child item type, including two types: directory type and file type. child_inode_id is the inode_id of the child item of the directory entry. For example, for a directory entry of directory type, its child items are subdirectories or files belonging to this directory. The basis for distribution means that the metadata of this directory entry is stored on the node device determined by self_inode_id.
[0113] (4)File attributes (name determined based on inode)
[0114] Key: self_inode_id (corresponding to child_inode_id in the Value of the directory entry)
[0115] Value: parent_inode_id, all other attributes of the file (such as permission configuration, time, length, etc.)
[0116] Distribution basis: parent_inode_id
[0117] Among them, self_inode_id is the inode_id of the file itself, which corresponds to the child_inode_id in the directory entry of the directory to which the file belongs, and parent_inode_id is the inode_id of the directory to which the file belongs. The distribution basis means that the metadata of the file is stored in the node device determined by the parent_inode_id.
[0118] Based on the above organization method, when a file is created, the file belongs to only one directory. The node device that stores the file's metadata is determined by the inode_id of the directory to which it belongs (that is, the parent directory). The mutable part of the parent directory attribute of the file is also determined according to the inode_id of the parent directory itself. The directory entry of the parent directory is also determined according to the inode_id of the parent directory itself. All three are determined according to the inode_id of the parent directory, that is, the metadata involved in the operation of creating a file are stored in the same node device, avoiding the creation of cross-node transactions.
[0119] When no hard links are created for a file, the file belongs to only one directory. By managing metadata in the above manner, cross-node transactions can be avoided as much as possible, reducing the impact on the overall performance of the distributed file system. When a hard link is created for a file, the file belongs to two directories at the same time. In order to continue to support the normal management of each metadata, this embodiment provides the following method for updating metadata:
[0120] (1) querying metadata of the target file according to the file index node identifier, where the metadata of the target file includes the directory index node identifier of the target directory;
[0121] (2) Update the directory index node identifier to an invalid value;
[0122] (3) Migrate the metadata of the target file to a preset node device among the multiple node devices.
[0123] In this embodiment, before creating a hard link, the file index node identifier is the inode_id of the file, the directory index node identifier is the inode_id of the target directory, the Key value of the metadata of the target file is the inode_id of the target file, and the Value value is the inode_id of the target directory. Since creating a hard link will cause the target file to have two parent directories at the same time, at this time, in order to continue to support the normal management of metadata, the node device storing the metadata of the target file is unbound from the target directory. One implementation method is to update the directory index node identifier to an invalid value, and the invalid value can be configured according to actual needs. For example, 0 is used to represent the invalid value.
[0124] In this embodiment, since the node device storing the metadata of the target file is unbound from the target directory, therefore, the node device storing the metadata of the target file cannot be determined according to the inode_id of the target directory. In order to be able to manage the metadata of the target file normally, the metadata of the target file is migrated to a preset node device. The preset node device can be selected from multiple node devices in the distributed file system according to a preset principle. The preset principle can be a random selection principle, that is, randomly select, or the principle of the least load pressure, that is, select the node device with the least load pressure.
[0125] To more intuitively display the metadata of the hard link created using Method 1, this embodiment combines Figure 3 the structures of each file and directory after creating the hard link in Figure 5 , Figure 5 to show the metadata of each file and directory. Please refer to Figure 3 which is an example diagram of the metadata after creating the hard link provided in this embodiment. Figure 5 In Figure 3 , for the Key column in Figure 5 is the Key value of each metadata, the Value column is the Value value corresponding to the Key, and the Partition column is the distribution basis of the corresponding metadata. For example, for the D1 directory, Figure 5DEntry_D1_inode_id + "D2" in it represents the directory entry of sub-directory D2 under directory D1. The type of D2 is directory, and the corresponding Value value is D2_inode_id. The distribution basis is D1_inode_id. For file F3 with hard links, the Key value of its metadata is FILE_F3_inode_id, that is, the _inode_id of file F3. The corresponding Value value is all file attributes of F3, and its distribution basis is fixed to the 0th database shard, that is, the database in the 0th node device. The rest of the metadata is similar and will not be elaborated here.
[0126] It should be noted that Figure 5 This is just an example. In fact, those skilled in the art can expand and vary the specific representation forms of the Key value and Value value based on actual needs to meet different requirements of actual scenarios. For example, more or less information can be included in the Value.
[0127] In this embodiment, for the query scenario of querying the attributes of a target file through the inode_id of the target file, since the target file is no longer bound to its parent directory and it is not necessarily determined which node device stores its metadata based on the inode_id of its parent directory, one implementation method is: query the inode_id of the parent directory of the target file on all node devices. Such a query method has poor performance. To improve the query performance of this method, this embodiment provides an improved method, that is, during the process of searching and traversing from the root directory to the target file, cache the parent_inode_id of each inode_id. For example, cache it in the service process of the client of the distributed file system. When receiving a request to modify the parent directory of the target file, such as unlink, rename, etc., synchronously modify the mapping relationship cache of inode_id -> parent_inode_id. Thus, avoid the performance loss caused by querying the parent directory of the target file on all node devices. When there is a mapping relationship cache of file inode identifiers and directory inode identifiers, the method of querying the target file is as follows:
[0128] First, based on the file query operation, obtain the directory inode identifier according to the mapping relationship. The file query operation is used to query the target file attribute information according to the file inode identifier;
[0129] Second, determine the target node from multiple node devices according to the directory inode identifier;
[0130] Third, use the file inode identifier as the Key value to find the Value value corresponding to the Key value in the target node;
[0131] Fourth, if the Value value is an invalid value, obtain the target file attribute information from the preset node device; otherwise, obtain the target file attribute information from the target node according to the file index node identifier.
[0132] In this embodiment, when all hard links of the target file are deleted and the attribute information of the target file is queried in the above manner, the file index node identifier is still used as the Key value to find the corresponding Value value in the target node. If the Value value is found to be invalid, the target file attribute information is obtained from the preset node device. To avoid unnecessary queries and judgments, in this embodiment, when there are no hard links to the target file, the node device of the target file is migrated from the preset node device to the node device determined by the directory index node identifier. The processing method is as follows:
[0133] First, when there are no hard links to the target file, obtain the directory index node identifier.
[0134] Second, update the directory index node identifier to the metadata of the target file.
[0135] Finally, migrate the metadata of the target file from the preset node device to the node device determined by the directory index node identifier.
[0136] Since all hard links of the target file are deleted and the target file returns to the situation of having only one parent directory, its metadata also returns to the situation of having only one parent directory, thus avoiding the need to judge invalid values each time and further avoiding the performance loss caused by obtaining the target file attribute information from the preset node device.
[0137] In this embodiment, in addition to the query operation for files, there are also operations for querying directories and directory entries under the directory, such as ReadDir, ReadDirPlus, etc. For ReadDir, all directory entries under the specified directory are obtained. For ReadDirPlus, all directory entries under the specified directory and the attribute information of each directory entry are obtained. When querying all directory entries of the specified directory, that is, obtaining all directory entries of the specified directory and the attribute information corresponding to each directory entry. The information of all directory entries is in the node device determined by the inode_id of the specified directory (referred to as node device 1). If the directory entry type is a file, it is also in node device 1. If the directory entry type is a directory, the immutable part of its attributes is in node device 1, and the mutable part needs to calculate the hash distribution based on the inode_id of the directory entry itself and may be in other node devices. At this time, cross-node query requests will be involved.
[0138] For a file whose directory entry has a hard link created, the value of the value will be found to be invalid when querying the node device 1 for the first time, and then the attribute information of the file will be obtained from the preset node device. Due to common application scenarios, there are only a small number of directories and most files do not have hard links. Therefore, this query operation has good performance. This embodiment gives a specific implementation method for querying all directory entries of a specified directory:
[0139] First, based on the directory query operation for querying all directory entries of a specified directory, determine the node to be queried for each directory entry from multiple node devices;
[0140] In this embodiment, the node to be queried is different in different cases. For whether the directory entry is a file or a directory, the corresponding node to be queried is also different. For a file, there can be only one node to be queried. For a directory, there can be two nodes to be queried. This embodiment gives an implementation method for determining the node to be queried:
[0141] (1) Obtain the inode identifier of the specified directory, and determine the first node to be queried from multiple node devices according to the inode identifier of the specified directory;
[0142] (2) Query the type and inode identifier of each directory entry in the specified directory from the first node to be queried;
[0143] (3) For any directory entry, if the type of the directory entry is a directory type, then determine the second node to be queried from multiple node devices according to the inode identifier of the directory entry, and use the first node to be queried and the second node to be queried as the two nodes to be queried for the directory entry. Among them, the first node to be queried stores the first metadata of the directory entry, the second node to be queried stores the second metadata of the directory entry, and the modification frequency of the first metadata is less than that of the second metadata;
[0144] In this embodiment, the first metadata corresponds to the immutable part of the directory attributes, and the second metadata corresponds to the mutable part of the directory attributes. The metadata of the directory includes the first metadata and the second metadata.
[0145] (4) If the type of the directory entry is a file type and the inode identifier of the directory entry is not an invalid value, then use the first node to be queried as the node to be queried for the directory entry;
[0146] (5) If the type of the directory entry is a file type and the inode identifier of the directory entry is an invalid value, then use the preset node device as the node to be queried for the directory entry.
[0147] Secondly, query the metadata of the corresponding directory entry from the node to be queried for each directory entry, and respond to the directory query operation according to the metadata of all the directory entries queried.
[0148] The above uses the method of unbinding the target file and the target directory to support hard links of the target file. This embodiment also provides a method of storing the metadata of the target file, the metadata related to the target file in the associated directory, and the metadata related to the target file in the target directory in the same node device among multiple node devices to support hard links (referred to as Method 2 in this article). This method is still based on the method of dividing the attributes of the directory into an immutable part and a mutable part in Method 1, which will not be elaborated here. In addition, the metadata related to the target file in the associated directory may belong to the mutable part of the associated directory, and the metadata related to the target file in the target directory may belong to the mutable part of the target directory. This method still also includes metadata related to the attributes of directory entries and metadata related to the attributes of files. Before introducing this method, the KV organization methods of these several types of metadata will still be introduced first.
[0149] (1) Immutable part of directory attributes (determine the name based on inode-Stable)
[0150] Key: self_inode_id + stable flag
[0151] Value: parent_inode_id, information related to owner permissions, etc.
[0152] Distribution basis: self_inode_id
[0153] (2) Mutable part of directory attributes (determine the name based on inode-Volatile)
[0154] Key: self_inode_id + volatile flag
[0155] Value: various timestamp information
[0156] Distribution basis: child_inode_id
[0157] (3) Directory entry DEntry (determine the name based on DEntry)
[0158] Key: self_inode_id + name (directory entry name)
[0159] Value: type (sub-item type) + child_inode_id (sub-item inode_id)
[0160] Distribution basis: child_inode_id
[0161] (4)File attribute inode (determine the name based on inode)
[0162] Key: self_inode_id (corresponding to child_inode_id in the Value of the directory entry)
[0163] Value: parent_inode_id, all other attributes of the file (such as permission configuration, time, length, etc.)
[0164] Basis for distribution: self_inode_id
[0165] The Key, Value, and distribution basis in the above (1) - (4) have been introduced in the aforementioned Method 1, and will not be elaborated here.
[0166] Since in Method 2, the node device storing the metadata of the target file is determined by the inode_id of the file, therefore, before introducing the processing of updating the metadata when creating a hard link to the target file in this embodiment, first introduce the creation processing method of the target file:
[0167] First, based on the creation operation of creating the target file in the target directory, obtain the file index node identifier of the target file;
[0168] In this embodiment, the file index node identifier is assigned by the distributed file system when creating the target file, and the file index node identifier can be the inode_id of the target file.
[0169] Secondly, determine the node to be updated from multiple node devices according to the file index node identifier;
[0170] In this embodiment, use the file index node identifier for hash calculation, and the node to be updated can be determined according to the hash calculation result. The node to be updated is the node storing the metadata of the target file.
[0171] Finally, create a first transaction on the node to be updated, and update the metadata of the target file and the metadata of the target directory in the first transaction.
[0172] In this embodiment, since the node storing the metadata related to the target directory and the target file is also determined according to the file index node identifier, therefore, both are stored on the node to be updated.
[0173] Based on the above method of creating the target file, this embodiment provides a method for updating metadata when creating a hard link to the target file:
[0174] Create a second transaction at the node to be updated. In the second transaction, update the inode identifier of the associated directory to the metadata of the target file and update the metadata of the associated directory.
[0175] In this embodiment, the metadata related to the target file in the associated directory is also stored at the node to be updated. At this time, the target file has two parent directories: the target directory and the associated directory. Therefore, the associated directory also needs to be updated to the metadata of the target file.
[0176] To more intuitively show the metadata of the hard link created using Method 2, this embodiment combines Figure 3 the structures of each file and directory after creating the hard link in Figure 6 , Figure 6 to show the metadata of each file and directory. Please refer to Figure 3 which is an example diagram of the metadata after creating the hard link provided in this embodiment. Figure 6 The meanings of Key, Value, and Partition in Figure 5 can be referred to
[0177] Here, it will not be elaborated. Figure 6 It should be noted that
[0178] This is just an example. In fact, those skilled in the art can expand and vary the specific representation forms of the Key value and Value value based on actual needs to meet the different requirements of actual scenarios. For example, more or less information can be included in Value.
[0179] First, based on the file query operation for querying the attribute information of the target file, obtain the file inode identifier.
[0180] Secondly, determine the target node from multiple node devices according to the file inode identifier.
[0181] Finally, obtain the target file attribute information from the target node.
[0182] In addition to the query operation on file attribute information, similar to Method 1, Method 2 also has an operation to query directories. If only all directory entries of a specified directory need to be obtained, such as the ReadDir command, since the metadata of the mutable part and the immutable part of the directory are on different node devices, the request needs to be sent to all node devices. If, in addition to obtaining all directory entries of the specified directory, the attribute information of each directory entry also needs to be obtained, for example, the ReadDirPlus command, since the metadata of the mutable part of the directory entry and its attribute information are on the same node device, compared with the ReadDir command, no additional cross-node device requests will be added.
[0183] In addition, for commands that need to be sent to all node devices, such as the lookup command and the ReadDir command, to alleviate the additional performance overhead caused by the operation of sending to all node devices, it can be alleviated through the directory entry caching mechanism. The directory entry caching mechanism stores the recently accessed directory entries in memory to avoid the need for disk access every time a directory is accessed. Through directory entry caching, the system can locate the disk location of the required file faster, thereby accelerating the file access speed.
[0184] To execute the corresponding steps in the above embodiments and each possible implementation manner, an implementation manner of a metadata update device 100 is given below. Please refer to Figure 7 , Figure 7 which is a block diagram of the metadata update device provided in this embodiment. It should be noted that for the metadata update device 100 provided by the present invention, its basic principle and the generated technical effects are the same as those of the corresponding above embodiments. For the sake of brief description, they are not mentioned in this embodiment.
[0185] The metadata update device 100 includes an acquisition module 110 and an update module 120.
[0186] The acquisition module 110 is used to acquire the file inode identifier of the target file for which a hard link needs to be created;
[0187] The update module 120 is used to update the metadata of the target file according to the file inode identifier, and update the metadata of the associated directory related to the hard link of the target file;
[0188] In the updated metadata of the target file, the inode identifiers of the associated directory and the target directory do not exist, or the metadata of the target file, the metadata of the associated directory related to the target file, and the metadata of the target directory related to the target file are stored on the same node device among the multiple node devices, and the target directory is the directory specified when creating the target file.
[0189] In an alternative embodiment, the update module 120 is specifically configured to:
[0190] Query the metadata of the target file according to the file inode identifier, where the metadata of the target file includes the directory inode identifier of the target directory; update the directory inode identifier to an invalid value; and migrate the metadata of the target file to a preset node device among multiple node devices.
[0191] In an alternative embodiment, the metadata update device 100 further includes a query module 130. The mapping relationship between the file inode identifier and the directory inode identifier is cached in the node device. The query module 130 is configured to:
[0192] Based on a file query operation, obtain the directory inode identifier according to the mapping relationship. The file query operation is used to query the target file attribute information according to the file inode identifier; determine the target node from multiple node devices according to the directory inode identifier; use the file inode identifier as the Key value to find the Value value corresponding to the Key value in the target node; if the Value value is an invalid value, obtain the target file attribute information from the preset node device, otherwise obtain the target file attribute information from the target node according to the file inode identifier.
[0193] In an alternative embodiment, the query module 130 is further configured to:
[0194] Based on a directory query operation for querying all directory entries of a specified directory, determine the node to be queried for each directory entry from multiple node devices; query the metadata of the corresponding directory entry from the node to be queried for each directory entry, and respond to the directory query operation according to the metadata of all the directory entries queried.
[0195] In an alternative embodiment, when the query module 130 is specifically configured to determine the node to be queried for each directory entry from multiple node devices, it is further specifically configured to:
[0196] Obtain the inode identifier of the specified directory, and determine the first node to be queried from multiple node devices according to the inode identifier of the specified directory; query the type and inode identifier of each directory entry in the specified directory from the first node to be queried; for any directory entry, if the type of the directory entry is a directory type, determine the second node to be queried from multiple node devices according to the inode identifier of the directory entry, and use the first node to be queried and the second node to be queried as the two nodes to be queried for the directory entry, where the first node to be queried stores the first metadata of the directory entry, the second node to be queried stores the second metadata of the directory entry, and the modification frequency of the first metadata is less than that of the second metadata; if the type of the directory entry is a file type and the inode identifier of the directory entry is not an invalid value, use the first node to be queried as the node to be queried for the directory entry; if the type of the directory entry is a file type and the inode identifier of the directory entry is an invalid value, use the preset node device as the node to be queried for the directory entry.
[0197] In an optional implementation manner, the update module 120 is further configured to:
[0198] When the target file does not have a hard link, obtain the directory inode identifier; update the directory inode identifier to the metadata of the target file; migrate the metadata of the target file from the preset node device to the node device determined by the directory inode identifier.
[0199] In an optional implementation manner, the metadata update device 100 further includes a creation module 140, and the creation module 140 is configured to:
[0200] Based on the creation operation of creating a target file in the target directory, obtain the file inode identifier of the target file; determine the node to be updated from multiple node devices according to the file inode identifier; create a first transaction in the node to be updated, and update the metadata of the target file and the metadata of the target directory in the first transaction.
[0201] In an optional implementation manner, the creation module 140 is specifically configured to:
[0202] Create a second transaction in the node to be updated, and update the inode identifier of the associated directory to the metadata of the target file and update the metadata of the associated directory in the second transaction.
[0203] In an optional implementation manner, the query module 130 is further configured to:
[0204] Based on the file query operation of querying the attribute information of the target file, obtain the file inode identifier; determine the target node from multiple node devices according to the file inode identifier; obtain the target file attribute information from the target node.
[0205] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the metadata updating method in the aforementioned implementation is implemented.
[0206] In summary, an embodiment of the present invention provides a metadata updating method, apparatus, node device and computer-readable storage medium, the method comprising: obtaining a file index node identifier of a target file for which a hard link needs to be created; updating the metadata of the target file according to the file index node identifier, and updating the metadata of an associated directory related to the hard link of the target file; the index node identifiers of the associated directory and the target directory do not exist in the metadata of the updated target file, or the metadata of the target file, the metadata related to the target file in the associated directory and the metadata related to the target file in the target directory are stored in the same node device among multiple node devices, and the target directory is the directory specified when the target file is created. Compared with the prior art, this embodiment has at least the following advantages: (1) by eliminating the index node identifiers of the associated directory and the target directory in the metadata of the updated target file, the metadata of the target file and the directory to which it belongs are unbound, so that the updates of the two are independent of each other and do not need to be restricted by the same transaction, thereby avoiding the modification of metadata through cross-node transactions; (2) by storing the metadata of the target file, the metadata related to the target file in the associated directory, and the metadata related to the target file in the target directory in the same node device, the updates of the three can be completed in the same node device, thereby avoiding the modification of metadata through cross-node transactions; (3) the mapping relationship between the file index node identifier and the directory index node identifier is pre-cached, thereby improving the performance of file query operations; (4) when there is no hard link to the file, the metadata of the target file is migrated from the preset node device to the node device determined by the directory index node identifier, thereby avoiding unnecessary judgment and search, and optimizing the search efficiency.
[0207] The above are only various embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A metadata updating method, characterized in that: Applied to each node device of a plurality of node devices in a distributed file system, the method comprises: Get the file index node identifier of the target file for which a hard link needs to be created; updating the metadata of the target file according to the file index node identifier, and updating the metadata of the associated directory related to the hard link of the target file; The index node identifiers of the associated directory and the target directory do not exist in the updated metadata of the target file, or the metadata of the target file, the metadata related to the target file in the associated directory, and the metadata related to the target file in the target directory are stored in the same node device among the multiple node devices, and the target directory is the directory specified when the target file is created; The step of updating the metadata of the target file according to the file index node identifier comprises: querying metadata of the target file according to the file index node identifier, the metadata of the target file includes a directory index node identifier of the target directory, the file index node identifier is a Key value of the metadata of the target file, the directory index node identifier of the target directory is a Value value corresponding to the Key value, and a node device storing the metadata of the target file is determined according to the index node identifier of the target directory; Updating the directory index node identifier to an invalid value to release the binding relationship between the node device storing the metadata of the target file and the target directory; The metadata of the target file is migrated to a preset node device among the multiple node devices.
2. The metadata updating method according to claim 1, characterized in that: The node device cache has a mapping relationship between the file index node identifier and the directory index node identifier, and the method further includes: Based on a file query operation, obtaining the directory index node identifier according to the mapping relationship, the file query operation is used to query target file attribute information according to the file index node identifier; Determine a target node from the plurality of node devices according to the directory index node identifier; Using the file index node identifier as the Key value, searching the target node for a Value value corresponding to the Key value; If the Value value is an invalid value, the target file attribute information is obtained from the preset node device; otherwise, the target file attribute information is obtained from the target node according to the file index node identifier.
3. The metadata updating method according to claim 1, characterized in that: The method further comprises: Based on a directory query operation for querying all directory items of a specified directory, determining a node to be queried for each directory item from the plurality of node devices; The metadata of the corresponding directory item is queried from the to-be-queried node of each directory item, and the directory query operation is responded to according to the metadata of all the queried directory items.
4. The metadata updating method according to claim 3, characterized in that: The step of determining the node to be queried for each directory entry from the plurality of node devices comprises: Obtaining an index node identifier of the specified directory, and determining a first node to be queried from the multiple node devices according to the index node identifier of the specified directory; Query the type and directory entry index node identifier of each directory entry in the specified directory from the first node to be queried; For any directory entry, if the type of the directory entry is a directory type, a second node to be queried is determined from the multiple node devices according to the directory entry index node identifier, and the first node to be queried and the second node to be queried are used as two nodes to be queried for the directory entry, wherein the first node to be queried stores first metadata of the directory entry, the second node to be queried stores second metadata of the directory entry, and the modification frequency of the first metadata is less than that of the second metadata; If the type of the directory entry is a file type, and the directory entry index node identifier is not an invalid value, taking the first node to be queried as the node to be queried of the directory entry; If the type of the directory entry is a file type and the index node identifier of the directory entry is an invalid value, the preset node device is used as the node to be queried for the directory entry.
5. The metadata updating method according to claim 1, characterized in that: The method further comprises: When the target file does not have a hard link, obtaining the directory index node identifier; Update the directory index node identifier to the metadata of the target file; The metadata of the target file is migrated from the preset node device to the node device determined by the directory index node identifier.
6. The metadata updating method according to claim 1, characterized in that: The step of obtaining the file index node identifier of the target file for which a hard link needs to be created includes: Based on the creation operation of creating the target file in the target directory, obtaining a file index node identifier of the target file; Determine a node to be updated from the plurality of node devices according to the file index node identifier; A first transaction is created at the node to be updated, and metadata of the target file and metadata of the target directory are updated in the first transaction.
7. The metadata updating method according to claim 6, characterized in that: The step of updating the metadata of the target file according to the file index node identifier, and updating the metadata of the associated directory related to the hard link of the target file includes: A second transaction is created at the node to be updated, and in the second transaction, the index node identifier of the associated directory is updated into the metadata of the target file and the metadata of the associated directory is updated.
8. The metadata updating method according to claim 6, characterized in that: The method further comprises: Based on the file query operation of querying the attribute information of the target file, obtaining the file index node identifier; Determine a target node from the plurality of node devices according to the file index node identifier; The target file attribute information is obtained from the target node.
9. A metadata updating device, characterized in that: Applied to each node device of a plurality of node devices in a distributed file system, the device comprises: An acquisition module is used to obtain a file index node identifier of a target file for which a hard link needs to be created; An update module, used to update the metadata of the target file according to the file index node identifier, and update the metadata of the associated directory related to the hard link of the target file; The index node identifiers of the associated directory and the target directory do not exist in the updated metadata of the target file, or the metadata of the target file, the metadata related to the target file in the associated directory, and the metadata related to the target file in the target directory are stored in the same node device among the multiple node devices, and the target directory is the directory specified when the target file is created; The update module is specifically used for: querying the metadata of the target file according to the file index node identifier, the metadata of the target file includes the directory index node identifier of the target directory, the file index node identifier is the Key value of the metadata of the target file, the directory index node identifier of the target directory is the Value value corresponding to the Key value, and the node device storing the metadata of the target file is determined according to the index node identifier of the target directory; updating the directory index node identifier to an invalid value to release the binding relationship between the node device storing the metadata of the target file and the target directory; migrating the metadata of the target file to a preset node device among the multiple node devices.
10. A node device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a program, and the processor is used to implement the metadata updating method according to any one of claims 1 to 8 when executing the program.
11. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the metadata updating method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Metadata processing method and device, equipment and medium
CN115481096A
Distributed file system metadata management method, device and equipment
CN118035200A