Metadata management method and device for distributed file system
By storing the mapping relationship between directory keywords and inode numbers in a distributed file system, and storing metadata using cloud databases, the problem of metadata storage capacity limitation in traditional distributed file systems is solved, and high-scalable file storage is achieved.
Patent Information
- Application Number
- CN202210307777.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-25
AI Technical Summary
In traditional distributed file systems, metadata is stored in the central node and is limited by disk capacity and cannot adapt to the application scenarios of massive files.
By storing the mapping relationship between the keywords of directories at each level and the inode number of the directory metadata in the distributed file system, and combining the cloud database to jointly realize the storage of metadata, so as to obtain the metadata in the cloud database after locally finding the inode number.
It solves the performance bottleneck of stand-alone metadata services, improves the scalability of the system, and can provide file storage of more than one billion.
Smart Images

Figure CN114840487B_ABST
Abstract
Description
Technical Field
[0001] The present specification relates to the field of storage technology, and in particular to a metadata management method and device for a distributed file system. Background Art
[0002] In traditional distributed file systems, metadata is often stored in central nodes. Due to limitations such as the central node's disk capacity, this metadata management method is no longer suitable for application scenarios with massive files. Summary of the invention
[0003] In view of this, the present specification provides a metadata management method and device for a distributed file system.
[0004] Specifically, this specification is implemented through the following technical solutions:
[0005] A metadata management method for a distributed file system is applied to a distributed file system, wherein the mapping relationship between keywords of directories at various levels and index node numbers of directory metadata is stored in the distributed file system, wherein the keywords of the directory at this level are generated based on the index node numbers of the metadata of the upper level directory and the names of the directories at this level, and the method comprises:
[0006] Extract the directory names of all levels of directories from the file path;
[0007] In order of the directories from the upper level to the lower level, for each extracted directory name, a keyword of the lower level directory is generated based on the directory name and the index node number of the metadata of the upper level directory;
[0008] Searching the index node number of the metadata of the directory at this level in the mapping relationship based on the keyword;
[0009] The metadata of the file path is obtained from a cloud database based on the inode number of the file path.
[0010] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0011] When the current level directory is a primary directory, a keyword of the current level directory is generated based on the directory name and preset characters.
[0012] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0013] The directory name and the index node number of the upper-level directory metadata are concatenated based on a preset order to generate a keyword of the current-level directory.
[0014] Optionally, acquiring metadata of the file path from a cloud database based on the inode number of the file path includes:
[0015] Searching the cloud database for the index node numbers of the directory metadata corresponding to the directory keywords at all levels of the file path;
[0016] Determine whether the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database;
[0017] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database, the metadata of the file path is obtained from the cloud database based on the index node number of the file path.
[0018] Optionally, also include:
[0019] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are different from the index node numbers found in the cloud database, the mapping relationship is updated based on the index node numbers found in the cloud database.
[0020] Optionally, also include:
[0021] The metadata of the file path is obtained from the cloud database based on the inode number of the file path found in the cloud database.
[0022] Optionally, also include:
[0023] If the index node number of the metadata of the current level directory cannot be found in the mapping relationship based on the keyword, the index node number of the metadata of the current level directory is found from the cloud database based on the keyword, and the mapping relationship is updated based on the found index node number.
[0024] A data access method for a distributed file system, applied to a distributed file system, comprising:
[0025] In response to a data access request sent by a client, query corresponding metadata according to the path of the data to be accessed;
[0026] Returning the metadata to the client so that the client can access data based on the metadata;
[0027] The aforementioned metadata management method is used to query metadata based on the path.
[0028] A metadata management device for a distributed file system is applied to a distributed file system, wherein the mapping relationship between keywords of directories at various levels and index node numbers of directory metadata is stored in the distributed file system, wherein the keywords of the directory at this level are generated based on the index node numbers of the metadata of the upper level directory and the names of the directories at this level, and the device comprises:
[0029] A name acquisition unit extracts directory names of directories at all levels from the file path;
[0030] A keyword generating unit generates a keyword of the directory at the current level based on the directory name and the index node number of the metadata of the upper directory for each extracted directory name in the order of the directories from the upper level to the lower level;
[0031] A number search unit, which searches for the index node number of the metadata of the directory at this level in the mapping relationship based on the keyword;
[0032] The metadata acquisition unit acquires the metadata of the file path from a cloud database based on the index node number of the file path.
[0033] A data access device for a distributed file system, applied to a distributed file system, comprising:
[0034] A metadata query unit, in response to a data access request sent by a client, queries corresponding metadata according to a path of the data to be accessed;
[0035] The data access unit returns the metadata to the client so that the client can access the data based on the metadata;
[0036] The aforementioned metadata management method is used to query metadata based on the path.
[0037] A metadata management device for a distributed file system, comprising:
[0038] processor;
[0039] memory for storing machine-executable instructions;
[0040] The distributed file system stores a mapping relationship between keywords of directories at all levels and index node numbers of directory metadata, the keyword of the directory at this level is generated based on the index node number of the metadata of the upper level directory and the name of the directory at this level, and the processor is prompted to:
[0041] Extract the directory names of all levels of directories from the file path;
[0042] In order of the directories from the upper level to the lower level, for each extracted directory name, a keyword of the lower level directory is generated based on the directory name and the index node number of the metadata of the upper level directory;
[0043] Searching the index node number of the metadata of the directory at this level in the mapping relationship based on the keyword;
[0044] The metadata of the file path is obtained from a cloud database based on the inode number of the file path.
[0045] A computer-readable storage medium stores a computer program, wherein the computer program is used to enable a processor to execute the above metadata management method.
[0046] By adopting the above-mentioned distributed file system metadata management solution provided in this specification, the mapping relationship between the directory keywords at all levels and the index node numbers of the directory metadata is stored locally in the distributed file system. When obtaining metadata, the index node numbers of the directory metadata at all levels can be found locally first, and then the metadata can be found in the cloud database based on the index node numbers. The method of jointly storing metadata in the distributed file system locally and in the cloud database solves the performance bottleneck of the stand-alone metadata service, improves the scalability of the system, and can provide file storage of more than one billion scales. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the architecture of a distributed file system in related technologies.
[0048] Figure 2 It is a flowchart of a metadata management method for a distributed file system shown in an exemplary embodiment of this specification.
[0049] Figure 3 It is a flowchart of another metadata management method of a distributed file system shown in an exemplary embodiment of this specification.
[0050] Figure 4 It is a schematic diagram of the architecture of a distributed file system shown in an exemplary embodiment of this specification.
[0051] Figure 5 It is a flowchart of a data access method of a distributed file system shown in an exemplary embodiment of this specification.
[0052] Figure 6 It is a hardware structure diagram of an electronic device where a metadata management device of a distributed file system is located, shown as an exemplary embodiment of this specification.
[0053] Figure 7It is a block diagram of a metadata management device for a distributed file system shown in an exemplary embodiment of this specification.
[0054] Figure 8 It is a block diagram of a data access device of a distributed file system shown in an exemplary embodiment of this specification. DETAILED DESCRIPTION
[0055] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this specification. Instead, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0056] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a", "the" and "the" used in this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0057] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0058] Figure 1 It is a schematic diagram of the architecture of a distributed file system shown in an exemplary embodiment of this specification.
[0059] Please refer to Figure 1 , a distributed file system may include name nodes and data nodes.
[0060] Among them, the name node is responsible for managing the namespace of the distributed file system, maintaining the file system tree, and the metadata of each file and folder in the file system.
[0061] Datanodes are used to store data in the form of blocks.
[0062] When a distributed file system client accesses data, it can send an access request to the name node. Taking a data read request as an example, the name node will search for the corresponding metadata and return the metadata to the client. The client can then obtain the data block where the data is located based on the metadata, and then read the corresponding data from the data node based on the data block. Taking a data write request as an example, the name node will also search for the corresponding metadata. If the metadata is found, the metadata can be returned to the client. The client can then obtain the data block where the data is located based on the metadata, and then write the corresponding data to the data node based on the data block; if the metadata is not found, a new index node number can be created, and metadata such as the data block location can be written, and then these metadata are returned to the client. The client can then write data in the corresponding data block.
[0063] In traditional distributed file systems, metadata is often stored in name nodes. Due to limitations such as the name node's disk capacity, this metadata management method is no longer suitable for application scenarios with massive files.
[0064] This specification provides a metadata management solution for a distributed file system. The distributed file system can be combined with a cloud database to store metadata, thereby solving the limitation of disk capacity on metadata storage.
[0065] Among them, metadata is data about data, which can be used to describe data attributes and to support functions such as indicating storage location, historical data, resource search, and file records.
[0066] In a distributed file system, paths, directories, files, links, etc. can all have metadata, which includes descriptive information necessary for reading and writing, such as the real path, size, creation time, permissions, etc.
[0067] A file path usually points to a specific file, and the file can be accessed through the file path. A path usually includes multiple directories, and each directory can correspond to a directory name.
[0068] Table of contents Directory Level Directory Name / user First level directory user / user / hive Secondary directory hive / user / hive / warehouse Level 3 Directory warehouse / user / hive / warehouse / file Level 4 Directory file
[0069] Table 1
[0070] For example, assuming that a file path is / user / hive / warehouse / file, please refer to the example in Table 1. The file path includes four directories. The names of the four directories are folder names, namely user, hive, warehouse, and file.
[0071] In this specification, the mapping relationship between the keywords of directories at all levels and the index node numbers of the directory metadata can be stored in the distributed file system, without storing the full metadata.
[0072] The mapping relationship may be stored in a key-value format, for example, in a name node.
[0073] The index node number is an Inode (index node) number, that is, an Inode ID. An index node is a data structure, and metadata can be found based on the index node number.
[0074] The keyword may be generated based on the index node number of the metadata of the upper-level directory and the name of the current-level directory.
[0075] For example, the keyword of the directory at this level is obtained by concatenating the name of the directory at this level and the index node number of the metadata of the upper directory based on a preset order.
[0076] For example, assuming that the name of the current directory is hive and the index node number of the metadata of the parent directory is 100, the keyword 100hive can be generated.
[0077] For another example, the keyword of the current level directory is obtained by calculating the name of the current level directory and the index node number of the metadata of the upper level directory based on a preset algorithm.
[0078] Of course, other methods may be used to generate keywords for directories at all levels, and this specification does not impose any special restrictions on this.
[0079] For a first-level directory, there is no parent directory. When generating a keyword for the first-level directory, the keyword for the first-level directory can be generated based on the first-level directory name and preset characters.
[0080] Taking the file path / user / hive / warehouse / file as an example, its first-level directory is / user, and the keyword 0user can be generated based on the preset character 0 and the directory name user.
[0081] In this specification, the cloud database can store the full metadata, and can also store the mapping relationship between the keywords of directories at all levels and the index node numbers of the directory metadata, so that the distributed file system can update the stored mapping relationship.
[0082] In this manual, the distributed file system and the cloud database jointly implement metadata storage. There is no need to store the full metadata in the distributed file system. This distributed metadata storage method can effectively solve the storage limitation of the distributed file system disk capacity on metadata and is suitable for application scenarios with massive files such as data lakes.
[0083] Figure 2 It is a flowchart of a metadata management method for a distributed file system shown in an exemplary embodiment of this specification.
[0084] Please refer to Figure 2 The metadata management method of the distributed file system can be applied to a distributed file system, for example, to a name node in a distributed file system, and includes the following steps:
[0085] Step 202: extract directory names of directories at all levels from the file path.
[0086] In this manual, when reading and writing files, the user-side client can send a read and write request to the distributed file system. The distributed file system usually needs to find the metadata of the file and the metadata of the file path. Sometimes it is also necessary to find the metadata of the directory where the file is located, the metadata of the file's parent directory, etc., and then based on these metadata, it can obtain information such as file type, file size, creation time, modification time, user, executable permissions, etc.
[0087] In this specification, when performing metadata search, the distributed file system may first extract directory names of directories at all levels from the file path.
[0088] Taking the aforementioned file path / user / hive / warehouse / file as an example, the directory names user, hive, warehouse, and file at all levels can be extracted.
[0089] Step 204 , in the order of directories from upper level to lower level, for each extracted directory name, generate a keyword of the current directory based on the directory name and the index node number of the metadata of the upper directory.
[0090] Step 206: Search the index node number of the metadata of the directory at this level in the mapping relationship based on the keyword.
[0091] In this specification, the distributed file system can find the index node numbers of the directories at all levels on the file path based on the mapping relationship between the keywords of the directories at all levels stored locally and the index node numbers of the directory metadata.
[0092] Before performing a search, the distributed file system may generate the key required to search for the index node number.
[0093] Since the keywords in this manual are generated based on the index node number of the parent directory metadata, when querying the index node number, the keywords of each level of directory can be generated in sequence from the upper level to the lower level to query the index node number of each level of directory.
[0094] For example, the keyword of the first-level directory can be generated first, and then the index node number of the metadata of the first-level directory can be found in the above mapping relationship stored locally in the distributed file system based on the keyword of the first-level directory. Then, the index node number of the metadata of the second-level directory can be found in the above mapping relationship based on the directory name of the second-level directory and the index node number of the metadata of the first-level directory. Next, the index node number of the metadata of the third-level directory can be found in the above mapping relationship based on the directory name of the third-level directory and the index node number of the metadata of the second-level directory. And so on, the index node numbers of the metadata of directories at all levels on the file path can be found.
[0095] Taking the aforementioned file path / user / hive / warehouse / file as an example, assuming that the index node numbers of the first-level directory metadata to the fourth-level directory metadata are 100-103, the distributed file system can store the mapping relationship shown in Table 2 in the form of key-value.
[0096] Table of contents Key Value / user 0user id:100 / user / hive 100hive id:101 / user / hive / warehouse 101warehouse id:102 / user / hive / warehouse / file 102file id:103
[0097] Table 2
[0098] It is worth noting that Table 2 is only an example. In actual applications, there is no need to store the left directory column. In addition, in addition to storing the index node number, the value field can also store some metadata of the directory, such as the directory name, directory size, etc.
[0099] In this embodiment, when performing an index node number query, the keyword 0user of the first-level directory can be generated based on the first-level directory name user and the preset character 0, and the mapping relationship shown in Table 2 can be queried based on the keyword 0user to find the index node number 100 of the first-level directory metadata.
[0100] Then, the secondary directory keyword 100hive can be generated based on the secondary directory name hive and the index node number 100 of the primary directory metadata, and the mapping relationship shown in Table 2 can be queried based on the keyword 100hive to find the index node number 101 of the secondary directory metadata.
[0101] Next, the third-level directory keyword 101warehouse can be generated based on the third-level directory name warehouse and the index node number 101 of the second-level directory metadata, and the mapping relationship shown in Table 2 is queried based on the keyword 101warehouse to find the index node number 102 of the third-level directory metadata.
[0102] Finally, the fourth-level directory keyword 102file can be generated based on the fourth-level directory name file and the index node number 102 of the third-level directory metadata, and the mapping relationship shown in Table 2 is queried based on the keyword 102file to find the index node number 103 of the fourth-level directory metadata.
[0103] It should be noted that, in this embodiment, step 202 can be executed before step 204, that is, before generating the keyword, the directory names of the directories at all levels are extracted from the file path. Step 202 can also be executed in conjunction with the loop process of steps 204-206, that is, in step 202, the first-level directory name is first extracted from the file path, and then steps 204-206 are executed to generate the first-level directory keyword and find the index node number of the first-level directory metadata; then the process can return to step 202 to extract the second-level directory name from the file path, and then steps 204-206 are executed to generate the second-level directory keyword and find the index node number of the second-level directory metadata, and so on, and steps 202-206 are executed in a loop, and this specification does not impose any special restrictions on this.
[0104] Step 208: Obtain metadata of the file path from a cloud database based on the inode number of the file path.
[0105] Based on the above steps, after finding the index node numbers of the metadata of the directories at all levels on the file path, the distributed file system can obtain the full metadata pointed to by the index node numbers from the cloud database.
[0106] In this embodiment, metadata may be obtained in the cloud database based on access requirements.
[0107] Taking the aforementioned file path / user / hive / warehouse / file as an example, if there is no need to obtain the metadata of the parent directory of the file, the metadata of the file path / user / hive / warehouse / file can be obtained based on index number 103. If the metadata of the parent directory of the file needs to be obtained, the metadata of the third-level directory / user / hive / warehouse can be obtained based on index node number 102. On this basis, the metadata of the second-level directory / user / hive can be obtained based on index node number 101, and so on.
[0108] By adopting the above-mentioned distributed file system metadata management solution provided in this specification, the mapping relationship between the directory keywords at all levels and the index node numbers of the directory metadata is stored locally in the distributed file system. When obtaining metadata, the index node numbers of the directory metadata at all levels can be found locally first, and then the metadata can be found in the cloud database based on the index node numbers. The method of jointly storing metadata in the distributed file system locally and in the cloud database solves the performance bottleneck of the stand-alone metadata service, improves the scalability of the system, and can provide file storage of more than one billion scales.
[0109] Moreover, after finding the index node numbers of the directories at all levels based on the mapping relationship stored locally, batch processing can be used to merge multiple index node numbers that need to be queried, and then the metadata pointed to by these index node numbers can be obtained from the cloud database at one time. Compared with the traditional technology of storing metadata in the cloud database, which requires multiple recursive queries for metadata of directories at all levels from the cloud database, the use of batch processing to obtain metadata from the cloud database can greatly save the query process overhead, improve the efficiency of metadata acquisition, and thus improve the efficiency of subsequent file access.
[0110] In this specification, if the index node number of the directory metadata cannot be found in the mapping relationship based on the generated keyword in the aforementioned step 206, it may mean that the distributed file system has not yet stored the mapping relationship between the directory keywords at all levels stored in the cloud database and the index node numbers of the directory metadata locally; or, the original directory name has been modified, and the distributed file system cannot find the corresponding index node number using the keyword generated by the new directory name.
[0111] When the distributed file system cannot find the index node number of the directory metadata in the local mapping relationship based on the generated keyword, it can find the index node number from the cloud database based on the generated keyword to update the local mapping relationship, and obtain the directory metadata based on the index node number found in the cloud database.
[0112] Still taking the aforementioned file path / user / hive / warehouse / file as an example, suppose the secondary directory name hive is changed to hive001, and the mapping relationship of the distributed file system local storage is not updated and remains as Table 2.
[0113] The cloud database stores the latest mapping relationship. In the example of changing hive to hive001, the storage method of the mapping relationship between keywords and index node numbers provided in this manual is adopted. It is only necessary to modify the keyword (key value) of the secondary directory in the cloud database mapping relationship, that is, to change 100hive to 100hive001. Compared with the mapping relationship storage method of using directories at all levels as keywords, there is no need to recursively modify the keywords of directories at all levels, which greatly reduces the keyword modification overhead caused by renaming.
[0114] In this embodiment, the latest mapping relationship stored in the cloud database is shown in Table 3 below.
[0115] Key Value 0user id:100 100hive001 id:101 101warehouse id:102 102file id:103
[0116] Table 3
[0117] In this embodiment, when the distributed file system obtains the index node numbers of the metadata of directories at all levels on the new file path / user / hive001 / warehouse / file, it first generates the keyword 0user of the first-level directory based on the first-level directory name user and the preset character 0, and queries the mapping relationship shown in Table 2 stored locally based on the keyword 0user, and then finds the index node number 100 of the first-level directory metadata.
[0118] Then, based on the secondary directory name hive001 and the index node number 100 of the primary directory metadata, the secondary directory keyword 100hive001 is generated. Based on the keyword 100hive001, the corresponding index node number cannot be queried in Table 2 stored locally. The distributed file system can then query the index node number in the cloud database. That is, the index node number 101 corresponding to the keyword 100hive001 is queried in the mapping relationship shown in Table 3 stored in the cloud database.
[0119] The distributed file system can also update the mapping relationship stored locally based on the corresponding relationship between the keyword 100hive001 and the index node number 101 queried in the cloud database, that is, update the mapping relationship shown in Table 2 stored locally to the mapping relationship shown in Table 3. For the distributed file system, when the directory name is modified, only the keyword of the corresponding directory needs to be modified.
[0120] It should be noted that to ensure the accuracy of the query results, the distributed file system can also query the index node numbers of the metadata of each subordinate directory of the second-level directory in the cloud database, that is, further query the index node numbers of the metadata of the third-level directory and the fourth-level directory in the cloud database, and update the mapping relationship stored locally based on the query results to avoid problems such as local query failure or inaccurate query caused by the modification of the subordinate directory name.
[0121] Optionally, in other examples, the distributed file system may also periodically obtain the latest mapping relationship from the cloud database and update the latest mapping relationship locally, and this specification does not impose any special restrictions on this.
[0122] By adopting the metadata management solution for the distributed file system provided in this specification, when the metadata changes, it can also ensure that the distributed file system obtains accurate metadata, avoiding the problem of obtaining erroneous metadata due to failure to update the mapping relationship of local storage in time.
[0123] Figure 3 It is a flowchart of another metadata management method of a distributed file system shown in an exemplary embodiment of this specification.
[0124] Please refer to Figure 3 The metadata management method of the distributed file system can be applied to the distributed file system, and includes the following steps:
[0125] Step 302: extract directory names of directories at all levels from the file path.
[0126] Step 304 , in the order of directories from upper level to lower level, for each extracted directory name, generate a keyword of the current directory based on the directory name and the index node number of the metadata of the upper directory.
[0127] Step 306: Search the local mapping relationship for the index node number of the metadata of the directory at this level based on the keyword.
[0128] In this embodiment, the implementation of steps 302-306 can refer to the aforementioned Figure 2 The implementation method of steps 202-206 in the illustrated embodiment will not be described in detail in this specification.
[0129] Step 308: Search the cloud database for the index node numbers of the directory metadata corresponding to the directory keywords at each level of the file path.
[0130] In this embodiment, after the distributed file system finds the index node numbers of the metadata of each level of directories on the file path based on the mapping relationship stored locally, the index node numbers of the directories at each level can also be queried in the cloud database based on the keywords generated for the directories at each level.
[0131] For example, batch processing is used to merge multiple index node numbers that need to be queried, and then the cloud database is queried.
[0132] Still taking the file path / user / hive / warehouse / file as an example, after the distributed file system finds the index node numbers 100-103 of the metadata of the directories at all levels in the mapping relationship stored locally, it can query the index node numbers of the metadata of the directories at all levels in the cloud database based on the keywords 0user, 100hive, 101warehouse, and 102file of the directories at all levels. That is, the index node numbers are queried based on the mapping relationship between the keywords of the directories at all levels and the index node numbers stored in the cloud database.
[0133] Step 310: determine whether the index node numbers of the directory metadata at each level found based on the mapping relationship are the same as the index node numbers found in the cloud database.
[0134] Based on the aforementioned step 308, after the distributed file system finds the index node numbers of the directory metadata at each level in the cloud database, it determines whether the index node numbers found in the local mapping relationship are the same as the index node numbers found in the cloud database.
[0135] If they are the same, step 312 may be executed.
[0136] If they are not the same, step 314 may be executed.
[0137] Step 312: When the index node numbers of the metadata of each level of directories found based on the mapping relationship are the same as the index node numbers found in the cloud database, the metadata of the file path is obtained from the cloud database based on the index node number of the file path.
[0138] Based on the judgment result of the aforementioned step 310, if the index node number found in the local mapping relationship is the same as the index node number found in the cloud database, it can be explained that the mapping relationship stored locally is the latest mapping relationship, and the metadata stored in the cloud database has not changed. The metadata can be obtained from the cloud database based on the index node number.
[0139] Step 314, when the index node numbers of the directory metadata at all levels found based on the mapping relationship are different from the index node numbers found in the cloud database, update the mapping relationship based on the index node numbers found in the cloud database, and obtain metadata based on the index node numbers queried in the cloud database.
[0140] Based on the judgment result of the aforementioned step 310, if the index node number found in the local mapping relationship is different from the index node number found in the cloud database, that is, the metadata index node numbers corresponding to the same directory keyword are different, it means that the directory name in the cloud database may be updated, and the local mapping relationship has not been updated in time. The index node number found in the local mapping relationship based on the new directory name is not the index node number of the updated directory metadata you want to find, and may be the index node number of the historical directory metadata in the original cloud database.
[0141] Taking the aforementioned directory name hive being changed to hive001 as an example, if the index node number corresponding to the keyword 100hive001 can be found in the locally stored mapping relationship, for example, the index node number found in the locally stored mapping relationship is 200, which is different from the index node number 101 of 100hive001 stored in the cloud database, it can be explained that the local mapping relationship has not been updated in time. 200 may be the index node number of the historical directory / user / hive001 in the cloud database, which may no longer exist or has been modified.
[0142] In this case, the distributed file system can update the mapping relationship stored locally based on the index node number queried in the cloud database, and can also obtain metadata based on the index node number queried in the cloud database.
[0143] Key Value (local mapping relationship) Value (Cloud Database) 0user id:100 id:100 100hive id:101 id:101 101warehouse id:102 id:105 102file id:103 id:106
[0144] Table 4
[0145] For example, please refer to the example in Table 4. The distributed file system finds that the index node numbers of the directory metadata at each level in the mapping relationship stored locally are 100-103, while the index node numbers found in the cloud database are 100, 101, 105 and 106, that is, the index node numbers of the third-level directory metadata and the fourth-level directory metadata are different from those stored locally. Based on the query results of the cloud database, the distributed file system can modify the index node number 102 of the third-level directory metadata stored in the local mapping relationship to 105, and modify the index node number 103 of the fourth-level directory metadata stored in the local mapping relationship to 106.
[0146] Of course, if the value field also stores other metadata, if other metadata changes, they also need to be updated synchronously.
[0147] The distributed file system can also obtain the third-level directory metadata and the fourth-level directory metadata based on the index node numbers 105 and 106 .
[0148] Using the metadata management solution of the distributed file system provided in this specification, before obtaining metadata from the cloud database, the distributed file system determines whether the index node number queried in the cloud database is the same as the index node number queried based on the local mapping relationship, and obtains metadata when the index node numbers are the same. When the metadata in the cloud database changes, accurate metadata can still be obtained, which can effectively avoid problems such as metadata acquisition errors caused by the failure to update the local mapping relationship of the distributed file system in a timely manner in high-concurrency scenarios.
[0149] Based on the above-mentioned distributed file system metadata management method, this specification also provides a distributed file system data access method, which can be applied to the name node in the distributed file system. Please refer to Figure 4 and Figure 5 , including the following steps:
[0150] Step 502: In response to a data access request sent by the client, query corresponding metadata according to the path of the data to be accessed.
[0151] In this embodiment, the data access request may be a data read request or a data write request. Taking a data read request as an example, the name node queries the corresponding metadata according to the path of the data to be read. The query of the metadata may be based on the aforementioned Figure 2 or Figure 3 The metadata query solution described in the embodiment is implemented. For example, the name node first queries the index node number of the path in the mapping relationship between the directory keywords at all levels and the metadata index node numbers stored locally, and then obtains the corresponding metadata from the cloud database.
[0152] Step 504: Return the metadata to the client so that the client can access data based on the metadata.
[0153] Based on the aforementioned step 502, after obtaining the metadata from the cloud database, the metadata can be returned to the client. Taking the data read request as an example, the client can then obtain the data block where the data is located based on the metadata, and then read the corresponding data from the data node based on the data block.
[0154] In this embodiment, for data write requests, the name node can also be based on the aforementioned Figure 2 or Figure 3 The metadata query solution described in the embodiment implements metadata query, and other data writing processes can refer to related technologies, which will not be described in detail in this specification.
[0155] Corresponding to the above-mentioned embodiment of the metadata management method of the distributed file system, this specification also provides an embodiment of the metadata management device of the distributed file system.
[0156] The embodiments of the metadata management device of the distributed file system of this specification can be applied in electronic devices. The device embodiments can be implemented by software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the electronic device in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 6 The figure is a hardware structure diagram of an electronic device where the metadata management device of the distributed file system of this specification is located, except Figure 6 In addition to the processor, memory, network interface, and non-volatile memory shown, the electronic device in which the device in the embodiment is located may also include other hardware according to the actual function of the electronic device, which will not be described in detail.
[0157] Figure 7 It is a block diagram of a metadata management device for a distributed file system shown in an exemplary embodiment of this specification.
[0158] Please refer to Figure 7 The metadata management device 700 of the distributed file system can be applied to the aforementioned Figure 3 On the electronic device shown, the electronic device can be a name node of a distributed file system. The distributed file system stores a mapping relationship between keywords of directories at all levels and index node numbers of directory metadata, and the keywords of the directory at this level are generated based on the index node numbers of the upper directory metadata and the names of the directories at this level. The device 700 includes:
[0159] The name acquisition unit 701 extracts the directory names of the directories at all levels from the file path;
[0160] The keyword generating unit 702 generates keywords of the directory at the current level based on the directory name and the index node number of the metadata of the upper directory for each extracted directory name in the order of the directories from the upper level to the lower level;
[0161] A number search unit 703 searches for an index node number of metadata of the current level directory in the mapping relationship based on the keyword;
[0162] The metadata acquisition unit 704 acquires metadata of the file path from a cloud database based on the index node number of the file path.
[0163] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0164] When the current level directory is a primary directory, a keyword of the current level directory is generated based on the directory name and preset characters.
[0165] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0166] The directory name and the index node number of the upper-level directory metadata are concatenated based on a preset order to generate a keyword of the current-level directory.
[0167] Optionally, acquiring metadata of the file path from a cloud database based on the inode number of the file path includes:
[0168] Searching the cloud database for the index node numbers of the directory metadata corresponding to the directory keywords at all levels of the file path;
[0169] Determine whether the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database;
[0170] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database, the metadata of the file path is obtained from the cloud database based on the index node number of the file path.
[0171] Optionally, also include:
[0172] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are different from the index node numbers found in the cloud database, the mapping relationship is updated based on the index node numbers found in the cloud database.
[0173] Optionally, also include:
[0174] The metadata of the file path is obtained from the cloud database based on the inode number of the file path found in the cloud database.
[0175] Optionally, also include:
[0176] If the index node number of the metadata of the current level directory cannot be found in the mapping relationship based on the keyword, the index node number of the metadata of the current level directory is found from the cloud database based on the keyword, and the mapping relationship is updated based on the found index node number.
[0177] Corresponding to the embodiment of the data access method of the distributed file system described above, this specification also provides an embodiment of a data access device of the distributed file system.
[0178] The embodiments of the data access device for the distributed file system of this specification can be applied in electronic devices. The device embodiments can be implemented by software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the electronic device in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, the hardware structure of the electronic device in which the data access device for the distributed file system of this specification is located can be similar to Figure 6 The electronic devices shown are similar, and this specification does not impose any particular limitation on this.
[0179] Figure 8 It is a block diagram of a data access device of a distributed file system shown in an exemplary embodiment of this specification.
[0180] Please refer to Figure 8 The metadata management device 800 of the distributed file system can be applied in the name node of the distributed file system, and includes:
[0181] The metadata query unit 801 queries corresponding metadata according to the path of the data to be accessed in response to the data access request sent by the client.
[0182] The query of the metadata may be implemented by using the metadata management method provided in this specification.
[0183] The data access unit 802 returns the metadata to the client, so that the client can access the data based on the metadata.
[0184] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0185] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial description of the method embodiments. The device embodiments described above are only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. Ordinary technicians in this field can understand and implement it without paying creative work.
[0186] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device or a combination of any of these devices.
[0187] Corresponding to the embodiment of the metadata management method of the aforementioned distributed file system, this specification also provides a metadata management device of a distributed file system, the device comprising: a processor and a memory for storing machine executable instructions. The processor and the memory are usually connected to each other via an internal bus. In other possible implementations, the device may also include an external interface to enable communication with other devices or components.
[0188] In this embodiment, the mapping relationship between the keywords of directories at all levels and the index node numbers of the directory metadata is stored in the distributed file system, wherein the keywords of the directory at this level are generated based on the index node numbers of the metadata of the upper level directory and the names of the directories at this level. By reading and executing the machine executable instructions corresponding to the metadata management logic of the distributed file system stored in the memory, the processor is prompted to:
[0189] Extract the directory names of all levels of directories from the file path;
[0190] In order of the directories from the upper level to the lower level, for each extracted directory name, a keyword of the lower level directory is generated based on the directory name and the index node number of the metadata of the upper level directory;
[0191] Searching the index node number of the metadata of the directory at this level in the mapping relationship based on the keyword;
[0192] The metadata of the file path is obtained from a cloud database based on the inode number of the file path.
[0193] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0194] When the current level directory is a primary directory, a keyword of the current level directory is generated based on the directory name and preset characters.
[0195] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0196] The directory name and the index node number of the upper-level directory metadata are concatenated based on a preset order to generate a keyword of the current-level directory.
[0197] Optionally, acquiring metadata of the file path from a cloud database based on the inode number of the file path includes:
[0198] Searching the cloud database for the index node numbers of the directory metadata corresponding to the directory keywords at all levels of the file path;
[0199] Determine whether the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database;
[0200] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database, the metadata of the file path is obtained from the cloud database based on the index node number of the file path.
[0201] Optionally, also include:
[0202] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are different from the index node numbers found in the cloud database, the mapping relationship is updated based on the index node numbers found in the cloud database.
[0203] Optionally, also include:
[0204] The metadata of the file path is obtained from the cloud database based on the inode number of the file path found in the cloud database.
[0205] Optionally, also include:
[0206] If the index node number of the metadata of the current level directory cannot be found in the mapping relationship based on the keyword, the index node number of the metadata of the current level directory is found from the cloud database based on the keyword, and the mapping relationship is updated based on the found index node number.
[0207] Corresponding to the embodiment of the data access method of the aforementioned distributed file system, this specification also provides a data access device of a distributed file system, the device comprising: a processor and a memory for storing machine executable instructions. The processor and the memory are usually connected to each other via an internal bus. In other possible implementations, the device may also include an external interface to enable communication with other devices or components.
[0208] In this embodiment, by reading and executing the machine executable instructions corresponding to the data access logic of the distributed file system stored in the memory, the processor is prompted to:
[0209] In response to a data access request sent by a client, query corresponding metadata according to the path of the data to be accessed;
[0210] Returning the metadata to the client so that the client can access data based on the metadata;
[0211] The query of the metadata may be implemented by using the metadata management method provided in this specification.
[0212] Corresponding to the embodiment of the metadata management method of the aforementioned distributed file system, the mapping relationship between the keywords of the directories at each level and the index node numbers of the directory metadata is stored in the distributed file system, wherein the keywords of the directory at this level are generated based on the index node numbers of the metadata of the upper level directory and the names of the directories at this level. This specification also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented:
[0213] Extract the directory names of all levels of directories from the file path;
[0214] In order of the directories from the upper level to the lower level, for each extracted directory name, a keyword of the lower level directory is generated based on the directory name and the index node number of the metadata of the upper level directory;
[0215] Searching the index node number of the metadata of the directory at this level in the mapping relationship based on the keyword;
[0216] The metadata of the file path is obtained from a cloud database based on the inode number of the file path.
[0217] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0218] When the current level directory is a primary directory, a keyword of the current level directory is generated based on the directory name and preset characters.
[0219] Optionally, generating a keyword of the primary directory based on the directory name and the index node number of the metadata of the upper directory includes:
[0220] The directory name and the index node number of the upper-level directory metadata are concatenated based on a preset order to generate a keyword of the current-level directory.
[0221] Optionally, acquiring metadata of the file path from a cloud database based on the inode number of the file path includes:
[0222] Searching the cloud database for the index node numbers of the directory metadata corresponding to the directory keywords at all levels of the file path;
[0223] Determine whether the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database;
[0224] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are the same as the index node numbers found in the cloud database, the metadata of the file path is obtained from the cloud database based on the index node number of the file path.
[0225] Optionally, also include:
[0226] When the index node numbers of the metadata of the directories at all levels found based on the mapping relationship are different from the index node numbers found in the cloud database, the mapping relationship is updated based on the index node numbers found in the cloud database.
[0227] Optionally, also include:
[0228] The metadata of the file path is obtained from the cloud database based on the inode number of the file path found in the cloud database.
[0229] Optionally, also include:
[0230] If the index node number of the metadata of the current level directory cannot be found in the mapping relationship based on the keyword, the index node number of the metadata of the current level directory is found from the cloud database based on the keyword, and the mapping relationship is updated based on the found index node number.
[0231] Corresponding to the embodiment of the data access method of the aforementioned distributed file system, this specification also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented:
[0232] In response to a data access request sent by a client, query corresponding metadata according to the path of the data to be accessed;
[0233] Returning the metadata to the client so that the client can access data based on the metadata;
[0234] The query of the metadata may be implemented by using the metadata management method provided in this specification.
[0235] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0236] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A metadata management method for a distributed file system, applied to a distributed file system, wherein the distributed file system stores a first mapping relationship between keywords of directories at various levels and index node numbers of directory metadata, and a cloud database stores a second mapping relationship between keywords of directories at various levels and index node numbers of directory metadata, in, The cloud database stores full metadata, which is used to describe data attributes; The keyword of the current directory is generated based on the index node number of the metadata of the upper directory and the name of the current directory. The method includes: Extract the directory names of all levels of directories from the file path; In order of the directories from the upper level to the lower level, for each extracted directory name, a keyword of the lower level directory is generated based on the directory name and the index node number of the metadata of the upper level directory; Based on the keyword, searching the first mapping relationship of the distributed file system for the index node number of the directory metadata at the current level; searching the second mapping relationship of the cloud database for the index node number of the directory metadata corresponding to the directory keywords at each level of the file path; Determine whether the index node numbers of the metadata of the directories at all levels found based on the first mapping relationship are the same as the index node numbers found in the second mapping relationship of the cloud database; When the index node numbers of the metadata of each level of directories found based on the first mapping relationship are the same as the index node numbers found in the second mapping relationship of the cloud database, the metadata of the file path is obtained from the cloud database based on the index node number of the file path.
2. The method according to claim 1, wherein the keyword of the primary directory is generated based on the directory name and the index node number of the metadata of the upper directory. include: When the current level directory is a primary directory, a keyword of the current level directory is generated based on the directory name and preset characters.
3. The method according to claim 1, wherein the keyword of the primary directory is generated based on the directory name and the index node number of the metadata of the upper directory. include: The directory name and the index node number of the upper-level directory metadata are concatenated based on a preset order to generate a keyword of the current-level directory.
4. The method according to claim 1, further comprising: include: When the index node numbers of the directory metadata at all levels found based on the first mapping relationship are different from the index node numbers found in the second mapping relationship of the cloud database, the mapping relationship is updated based on the index node numbers found in the second mapping relationship of the cloud database.
5. The method according to claim 4, further comprising: include: The metadata of the file path is obtained from the cloud database based on the index node number of the file path found in the second mapping relationship of the cloud database.
6. The method according to claim 1, further comprising: include: If the index node number of the metadata of the current level directory cannot be found in the first mapping relationship based on the keyword, the index node number of the metadata of the current level directory is found in the cloud database based on the keyword, and the first mapping relationship is updated based on the found index node number.
7. A data access method for a distributed file system, applied to a distributed file system, include: In response to a data access request sent by a client, query corresponding metadata according to the path of the data to be accessed; Returning the metadata to the client so that the client can access data based on the metadata; Wherein, the metadata is queried based on the path using the method described in any one of claims 1 to 6.
8. A metadata management device for a distributed file system, applied to a distributed file system, wherein the distributed file system stores a first mapping relationship between keywords of directories at various levels and index node numbers of directory metadata, and a cloud database stores a second mapping relationship between keywords of directories at various levels and index node numbers of directory metadata, in, The cloud database stores full metadata, which is used to describe data attributes; The keyword of the current directory is generated based on the index node number of the metadata of the upper directory and the name of the current directory, and the device includes: A name acquisition unit extracts directory names of directories at all levels from the file path; A keyword generating unit generates a keyword of the directory at the current level based on the directory name and the index node number of the metadata of the upper directory for each extracted directory name in the order of the directories from the upper level to the lower level; A number search unit searches for the index node number of the directory metadata of the current level in the first mapping relationship of the distributed file system based on the keyword; searches for the index node number of the directory metadata corresponding to the directory keywords of each level of the file path in the second mapping relationship of the cloud database; The metadata acquisition unit determines whether the index node numbers of the metadata of the directories at all levels found based on the first mapping relationship are the same as the index node numbers found in the second mapping relationship of the cloud database; if the index node numbers of the metadata of the directories at all levels found based on the first mapping relationship are the same as the index node numbers found in the second mapping relationship of the cloud database, the metadata of the file path is acquired from the cloud database based on the index node number of the file path.
9. A data access device for a distributed file system, applied to a distributed file system, include: A metadata query unit, in response to a data access request sent by a client, queries corresponding metadata according to a path of the data to be accessed; The data access unit returns the metadata to the client so that the client can access the data based on the metadata; Wherein, the metadata is queried based on the path using the method described in any one of claims 1 to 6.
10. A metadata management device for a distributed file system, include: processor; memory for storing machine-executable instructions; The distributed file system stores a first mapping relationship between keywords of directories at all levels and index node numbers of directory metadata, and the cloud database stores a second mapping relationship between keywords of directories at all levels and index node numbers of directory metadata. The cloud database stores full metadata, and metadata is used to describe data attributes. The keyword of the directory at this level is generated based on the index node number of the metadata of the upper directory and the name of the directory at this level. By reading and executing machine executable instructions corresponding to the metadata management logic of the distributed file system stored in the memory, the processor is prompted to: Extract the directory names of all levels of directories from the file path; In order of the directories from the upper level to the lower level, for each extracted directory name, a keyword of the lower level directory is generated based on the directory name and the index node number of the metadata of the upper level directory; Based on the keyword, searching the first mapping relationship of the distributed file system for the index node number of the directory metadata at the current level; searching the second mapping relationship of the cloud database for the index node number of the directory metadata corresponding to the directory keywords at each level of the file path; Determine whether the index node numbers of the metadata of the directories at all levels found based on the first mapping relationship are the same as the index node numbers found in the second mapping relationship of the cloud database; When the index node numbers of the metadata of each level of directories found based on the first mapping relationship are the same as the index node numbers found in the second mapping relationship of the cloud database, the metadata of the file path is obtained from the cloud database based on the index node number of the file path.
11. A computer-readable storage medium storing a computer program, wherein the computer program is used to enable a processor to execute the metadata management method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Metadata searching method, device and equipment and computer readable storage medium
CN113010476A
Metadata query method and device based on distributed file system and storage medium
CN114116613A