Metadata access method, electronic equipment, storage medium and program product

By using snapshot sequence numbers in a distributed file system to determine the latest metadata version of the target file, the metadata loading delay problem caused by multiple snapshot versions is solved, and more efficient metadata access performance is achieved.

CN119938599APending Publication Date: 2025-05-06SANGFOR TECH INC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411849005.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In a distributed file system, after a directory creates a snapshot, multiple snapshot versions will be generated for the original metadata, resulting in an increase in delay in loading metadata from disk.

Method used

Determine the target metadata corresponding to the target file on disk by snapshot sequence number, and only the latest version of the metadata is loaded to reduce unnecessary metadata loading.

Benefits of technology

Reduces latency for loading metadata from disk and improves metadata access performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938599A_ABST
    Figure CN119938599A_ABST
Patent Text Reader

Abstract

The embodiment of the invention is suitable for the technical field of computers, and provides a metadata access method, electronic equipment, a storage medium and a program product.The metadata access method comprises the steps that target metadata corresponding to a target file is determined in a disk through a snapshot sequence number, and the target metadata is determined as response data of a reading request; the target metadata is metadata of the latest version of the target file, each metadata in the disk corresponds to a snapshot sequence number, and the snapshot sequence number represents the version of the metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a metadata access method, electronic equipment, storage medium and program product. Background Art

[0002] In the current distributed file system, after creating a snapshot of a directory, a snapshot version is generated for the original metadata (such as inode (index node) information and dentry (directory entry) information). When accessing directories and files, the file system needs to load the head version and all snapshot versions of metadata from the disk, which increases the latency of loading metadata from the disk. Summary of the invention

[0003] In order to solve the above problems, embodiments of the present invention provide a metadata access method, an electronic device, a storage medium and a program product.

[0004] The technical solution of the present invention is achieved in this way:

[0005] In one aspect, an embodiment of the present invention provides a metadata access method, the method comprising:

[0006] receiving a read request for metadata of a target file;

[0007] The target metadata corresponding to the target file is determined in the disk through the snapshot serial number, and the target metadata is determined as the response data of the read request; the target metadata is the metadata of the latest version of the target file, each metadata in the disk corresponds to a snapshot serial number, and the snapshot serial number represents the version of the metadata.

[0008] In the above solution, determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number includes:

[0009] The metadata whose snapshot sequence number is a set value is determined as the target metadata, and the set value indicates that the corresponding metadata is the metadata of the latest version of the target file.

[0010] In the above solution, determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number includes:

[0011] Based on the snapshot serial number, the index node identifier of the parent directory of the target file, the identifier of the directory shard where the target file is located, and the name of the target file, the target metadata corresponding to the target file is determined from the disk; the metadata in the disk is stored in the form of key-value pairs.

[0012] In the above solution, the method further includes:

[0013] Receiving a snapshot creation request for the target file;

[0014] Generate a new snapshot sequence number based on the snapshot creation instruction;

[0015] Generate metadata corresponding to the new snapshot sequence number in the in-memory database.

[0016] In the above solution, the method further includes:

[0017] Receive an operation request to modify target metadata;

[0018] Determine whether the snapshot sequence number carried in the operation request is greater than the snapshot sequence number of the target metadata;

[0019] If it is greater, the target metadata is copied on-write in the memory database to obtain first metadata and second metadata split from the target metadata, where the first metadata is metadata changed based on the operation request, and the second metadata is snapshot metadata of the target metadata.

[0020] In the above scheme, the metadata includes a first variable and a second variable; the value of the first variable of the first metadata is the newly generated snapshot serial number, and the value of the second variable of the first metadata is the set value; the value of the first variable of the second metadata is the snapshot serial number of the target metadata, and the value of the second variable of the second metadata is the newly generated snapshot serial number.

[0021] In the above solution, before determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number, the method further includes:

[0022] Determine whether the metadata in the memory database matches the read request;

[0023] If the metadata in the memory database does not match the read request, the target metadata corresponding to the target file is determined in the disk.

[0024] In the above solution, the method further includes:

[0025] Determine whether the cache space occupied by the metadata stored in the memory database is greater than a threshold;

[0026] When the cache space occupied by the metadata stored in the memory database is greater than the threshold, the metadata of the snapshot version that has not been used in the memory database is eliminated.

[0027] In the above solution, after eliminating the unused snapshot metadata in the memory database, the method further includes:

[0028] If the cache space occupied by the metadata stored in the memory database is still greater than the threshold, metadata in the memory database that has not been used and whose snapshot sequence number is the set value is eliminated.

[0029] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the metadata access method in the above scheme.

[0030] On the other hand, an embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the processor and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the steps of the metadata access method provided in the first aspect of the embodiment of the present invention.

[0031] In another aspect, an embodiment of the present invention provides a computer-readable storage medium, including: the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the metadata access method provided in the first aspect of the embodiment of the present invention are implemented.

[0032] The embodiment of the present application receives a read request for metadata of a target file, and determines the target metadata corresponding to the target file in the disk through a snapshot sequence number. Each metadata in the disk corresponds to a snapshot sequence number, and the snapshot sequence number represents the version of the metadata. The target metadata is determined as the response data of the read request, and the target metadata is the metadata of the latest version of the target file. The embodiment of the present application loads metadata according to the snapshot sequence number, and only loads the required target metadata, and the remaining metadata is not loaded at the same time, which can reduce the amount of metadata data loaded from the disk and reduce the delay in accessing the metadata. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of an implementation flow of a metadata access method provided by an embodiment of the present invention;

[0034] Figure 2 is a schematic diagram of a snapshot creation process provided by an embodiment of the present invention;

[0035] Figure 3 is a schematic diagram of a copy-on-write process provided by an embodiment of the present invention;

[0036] Figure 4 is a schematic diagram of a metadata access process provided by an embodiment of the present invention;

[0037] Figure 5 is a schematic diagram of a metadata elimination process provided by an embodiment of the present invention;

[0038] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] In the file system, metadata is data that describes data and is used to support functions such as indicating storage location, historical data, resource search, and file records. For example, metadata includes inode and dentry. Inode is the file system index node, which contains metadata information such as file size, permissions, user information, and modification time. Each file has a corresponding inode. Dentry is a directory entry, and dentry contains information such as the file name and the file's inode number.

[0041] In the file system, when a snapshot is created for a directory, a snapshot version is generated for the original inode and dentry, causing the number of inodes and dentry in the system to increase as the number of snapshots increases. When too many snapshots are created, when accessing directories and files, the file system needs to load the head version (latest version) and the inode and dentry information of all snapshots from the disk, which increases the delay in loading metadata from the disk.

[0042] As metadata is continuously written to the memory, for the same size of memory, the number of inodes and dentry stored in the head version decreases due to the presence of a large number of snapshot versions of inode and dentry information in the cache. When accessing metadata, there is a greater probability that the head version of metadata needs to be loaded from the disk rather than hitting the memory cache, resulting in a longer latency in accessing metadata.

[0043] The head version refers to the metadata of the latest version currently in the file system.

[0044] In view of the shortcomings of the above-mentioned related technologies, an embodiment of the present invention provides a metadata access method, which can reduce metadata access latency and improve access performance. In order to illustrate the technical solution of the present invention, a specific embodiment is used for description below.

[0045] Figure 1 is a schematic diagram of an implementation flow of a metadata access method provided by an embodiment of the present invention, refer to Figure 1 , metadata access methods include:

[0046] S101, receiving a request to read metadata of a target file.

[0047] Here, the read request may include an identifier of the target file, which may be a file name or other information that can represent the target file, so as to find the storage location of the metadata of the target file in the server. The read request may also include version information of the metadata to be read.

[0048] In this embodiment, metadata are all stored in the disk, and the metadata in the disk can be identified by the snapshot sequence number.

[0049] S102, determining the target metadata corresponding to the target file in the disk through the snapshot serial number, and determining the target metadata as the response data of the read request; the target metadata is the metadata of the latest version of the target file, each metadata in the disk corresponds to a snapshot serial number, and the snapshot serial number represents the version of the metadata.

[0050] The target file may include multiple metadata, for example, metadata of a head version and multiple snapshot metadata.

[0051] Each metadata has a snapshot sequence number, which may be assigned to the metadata when it is generated.

[0052] The snapshot sequence number may be a version number of the metadata, and each metadata snapshot sequence number may be unique.

[0053] For example, the snapshot sequence number of metadata can be incremented according to the update of metadata, so the metadata with the largest snapshot sequence number is the metadata of the current latest version. When reading the target metadata, the metadata with the largest snapshot sequence number can be used as the target metadata.

[0054] The embodiment of the present application receives a read request for metadata of a target file, and determines the target metadata corresponding to the target file in the disk through a snapshot sequence number. Each metadata in the disk corresponds to a snapshot sequence number, and the snapshot sequence number represents the version of the metadata. The target metadata is determined as the response data of the read request, and the target metadata is the metadata of the latest version of the target file. The embodiment of the present application loads metadata according to the snapshot sequence number, and only loads the required target metadata, and the remaining metadata is not loaded at the same time, which can reduce the amount of metadata data loaded from the disk and reduce the delay in accessing the metadata.

[0055] In one embodiment, determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number includes:

[0056] The metadata whose snapshot sequence number is a set value is determined as the target metadata, and the set value indicates that the corresponding metadata is the metadata of the latest version of the target file.

[0057] In this embodiment, the snapshot number of the latest version of metadata is always fixed to a set value. When a new version of metadata is generated, its snapshot number is automatically set to the set value, and the snapshot number of the previous version of metadata is changed from the set value to another value. In this way, the server can automatically read the metadata with the snapshot number set as the target metadata.

[0058] In one embodiment, determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number includes:

[0059] Based on the snapshot serial number, the index node identifier of the parent directory of the target file, the identifier of the directory shard where the target file is located, and the name of the target file, the target metadata corresponding to the target file is determined from the disk; the metadata in the disk is stored in the form of key-value pairs.

[0060] The index node is inode, and the inode identifier is inode number. Each file and directory has a unique inode number, which is used to identify the entity of the file or directory. When there are too many files in a directory, it will be split into multiple virtual dir (virtual directories) for storage. The identifier of the directory fragment is the identifier of the virtual dir.

[0061] In this embodiment, the metadata of the file is stored in a distributed key-value (KV) database, where the key of the KV pair is designed to be a quadruple consisting of the snapshot number, the index node identifier of the parent directory of the target file, the identifier of the directory shard where the target file is located, and the name of the target file, and the value of the KV pair is the metadata information.

[0062] For example, in the database, the key value is composed of a four-tuple of Parentid+Fragid+Snapid+Dname, where Parentid is the inode number of the file's parent directory, fragid is the id of the directory fragment where the file is located, Snapid is the snapshot sequence number, and Dname is the name corresponding to the file.

[0063] The metadata corresponding to the key value can be read from the database according to the key value composed of Parentid+Fragid+Snapid+Dname. In this embodiment, specific metadata items can be loaded on demand according to the key value composed of the quadruple, without loading all metadata items.

[0064] In another embodiment, the metadata is stored in a disk database in a scattered manner. The metadata can be directly indexed by an identifier. The identifier is also configured with a snapshot number. The identifier can be a hash value (the distributed system stores metadata randomly on the disk through hash calculation). The embodiment of the present application does not limit the specific type of the identifier.

[0065] In one embodiment, the method further comprises:

[0066] Receiving a snapshot creation request for the target file;

[0067] Generate a new snapshot sequence number based on the snapshot creation instruction;

[0068] Generate metadata corresponding to the new snapshot sequence number in the in-memory database.

[0069] refer to Figure 2 , Figure 2 Schematic diagram of a snapshot creation process provided by an embodiment of the present invention, including the following steps:

[0070] S201, the client sends a command to create a snapshot of a directory to the metadata server.

[0071] Here, the target file is a directory and the command corresponds to a snapshot creation request.

[0072] S202: The metadata server generates a snapshot sequence number for the snapshot.

[0073] The metadata server receives the message to create a snapshot and obtains a snapshot inside the metadata server to ensure that the snapshot is unique in the cluster. The snapshot is the snapshot serial number. After receiving the request to create a snapshot, the metadata server assigns a new snapshot serial number to the snapshot metadata to be created.

[0074] S203, generating a temporary inode version, and modifying the snapshot information of the directory snapshot based on the temporary inode version.

[0075] Generate a temporary version of the inode in memory, change the newly created snapshot information to the temporary version of the inode, and the structure that saves the snapshot information is snaprealm.

[0076] This embodiment creates snapshot metadata in memory, first generates a temporary version of the inode in memory, and then changes the newly created snapshot information (new snapshot sequence number and metadata information) to the temporary version of the inode. Snapshots are organized into a tree structure through SnapRealm. Each inode node with snapshot information will have a corresponding SnapRealm. Inodes without snapshot information use the nearest SnapRealm on the parent node path. The root node has a SnapRealm by default.

[0077] S204: Generate a journal log for creating the snapshot and write it to the disk.

[0078] The journal log is a system log used to record system behavior. The journal log for creating a snapshot is written to the disk.

[0079] S205, after the journal log is written successfully, the temporary version of the inode in the memory becomes the official version, and the snapshot creation success is returned to the client.

[0080] In the callback function, a successful packet creation is sent back to the client, and the modified temporary version in the memory is changed to the official version.

[0081] After the journal log is successfully written to the disk, the callback function replies to the client that the snapshot has been successfully created, and the temporary version of the inode in the memory is changed to the official version.

[0082] S206, trim log, update the metadata of the snapshot directory to the dkv database.

[0083] In the subsequent tick function, when trimming the created journal log, the metadata of the directory is written into the dkv database.

[0084] Among them, dkv refers to a distributed database, and the tick function refers to a data elimination function that is executed regularly by a timer within a specified time interval. The journal log is deleted according to the tick function. When the journal log is deleted, the metadata corresponding to the journal log is written into the dkv database to achieve permanent storage.

[0085] In one embodiment, the method further comprises:

[0086] Receive an operation request to modify target metadata;

[0087] Determine whether the snapshot sequence number carried in the operation request is greater than the snapshot sequence number of the target metadata;

[0088] If it is greater, the target metadata is copied on-write in the memory database to obtain first metadata and second metadata split from the target metadata, where the first metadata is metadata changed based on the operation request, and the second metadata is snapshot metadata of the target metadata.

[0089] Users can modify metadata, and the operation request carries the snapshot number of the target metadata to be modified. The snapshot number carried in the operation request is greater than the snapshot number of the target metadata, indicating that the previous version of metadata has not been generated before the snapshot, so a copy on write (Copy On Write, COW) is performed once, and the snapshot number will be added after COW.

[0090] Copy On Write (COW) is called copy on write or copy before write. After creating a snapshot, if the data of the source volume changes, the snapshot system will first copy the original data to the corresponding data block on the snapshot volume, and then rewrite the source volume.

[0091] In one embodiment, the metadata includes a first variable and a second variable; the value of the first variable of the first metadata is a newly generated snapshot serial number, and the value of the second variable of the first metadata is the set value; the value of the first variable of the second metadata is the snapshot serial number of the target metadata, and the value of the second variable of the second metadata is the newly generated snapshot serial number.

[0092] The embodiment of the present invention sets two member variables (a first variable and a second variable) in the inode structure, the first variable represents the snapshot number when the metadata is created, and the second variable is set to a set value when created, and the set value represents the latest version of the metadata.

[0093] For example, when COW is performed on inode and dentry, the original inode and dentry will be split into two, for example:

[0094] Inode(first, head)->snap_inode(first, new_snapid), head_inode(new_snapid, head)

[0095] dentry(first, head)->snap_dentry(first, new_snapid), head_dentry(new_snapid, head).

[0096] Among them, new_snapid is the latest snapshot serial number, first is the first variable, and last is the second variable. After COW, the target metadata Inode (first, head) is split into the second metadata snap_inode (first, new_snapid) and the first metadata head_inode (new_snapid, head). Among them, head is a set value, indicating that the corresponding metadata is the latest version of metadata, that is, head_inode (new_snapid, head) is the latest version of metadata, for example, head can be set to -2.

[0097] In combination with the above four-tuple embodiment, the snapid in the four-tuple stored in the DKV database is the second variable in this embodiment.

[0098] refer to Figure 3 , Figure 3 Schematic diagram of a copy-on-write process provided by an embodiment of the present invention, comprising the following steps:

[0099] S301: The metadata server receives an operation request for modifying metadata.

[0100] The operation request carries the snapshot number (snapid) of the target metadata to be changed.

[0101] S302, determining whether the snapshot carried in the operation request is greater than the first stored in the inode.

[0102] Compare the snapshot carried in the operation request with the first stored in the inode. If it is not greater than the first, no COW is required. Otherwise, a COW operation is performed.

[0103] In this embodiment, two member variables, first and last, are set in the metadata. First indicates the snapshot when the inode is created, and Last is set to the set value head when it is created. If the snapshot version is greater than fisrt, it means that the previous version has not been generated before the snapshot, so COW is performed once, and first will be added after COW. Each time a snapshot is created, deleted, or renamed, the snapshot will increase by 1.

[0104] S303, executing the COW operation, the original metadata will be split into two.

[0105] For example, Inode(first, head)->snap_inode(first, new_snapid), head_inode(new_snapid, head).

[0106] Among them, head is a set value, indicating that the corresponding metadata is the metadata of the latest version, head_inode (new_snapid, head) is the metadata of the current latest version, and new_snapid is the latest snapshot sequence number.

[0107] S304, write the journal log to the disk.

[0108] The journal log is a system log that records system behavior. The journal log indicates that the metadata change is successful.

[0109] S305, the journal is written successfully and a success message is returned to the client.

[0110] After the Journal log is written to disk, a message indicating that the metadata has been successfully changed is returned to the client.

[0111] S306, trim log, update the modified metadata to the DKV database.

[0112] In the subsequent tick function, when trimming the COW journal log, the COW inode information is written to the dkv database.

[0113] Among them, dkv refers to a distributed database, and the tick function refers to a data elimination function that is executed regularly by a timer within a specified time interval. The journal log is deleted according to the tick function. When the journal log is deleted, the metadata corresponding to the journal log is written into the dkv database to achieve permanent storage.

[0114] According to the above embodiment, it can be seen that the generated metadata are all in the memory. When reading the metadata, directly reading in the memory can improve the metadata access speed. Therefore, in one embodiment, before determining the target metadata corresponding to the target file in the disk by the snapshot sequence number, the method also includes:

[0115] Determine whether the metadata in the memory database matches the read request;

[0116] If the metadata in the memory database does not match the read request, the target metadata corresponding to the target file is determined in the disk.

[0117] This embodiment first obtains the target metadata from the memory database. If the target metadata is not in the memory database, the target metadata is obtained from the disk.

[0118] refer to Figure 4 , Figure 4 Schematic diagram of a metadata access process provided by an embodiment of the present invention, including the following steps:

[0119] S401: The metadata server receives a metadata request of readdir.

[0120] Among them, readdir is used to read the metadata of the directory.

[0121] S402, determining whether the metadata request is hit in the memory.

[0122] If yes, the metadata hit is read directly from the memory cache as a response, otherwise the corresponding metadata is obtained from the disk.

[0123] Traverse the memory cache to see if the memory cache hits. If it hits, read the metadata directly from the memory cache. If it does not hit, read the metadata from the DKV database on the disk, suspend the request, wait for the callback and try again.

[0124] S403, use parentid+fragid+snapid to get the metadata of the head version from dkv.

[0125] Among them, Parentid is the inode number of the parent directory of the file; fragid is the id of the directory fragment where the file is located; Snapid is the snapshot sequence number.

[0126] The metadata of the head version can be obtained from the dkv database based on the key value composed of Parentid+Fragid+Snapid.

[0127] S404: Reading is successful, and the readdir request is called again in the callback.

[0128] The metadata of the head version is read successfully. In the callback function, the previously suspended readdir request is called again.

[0129] If the metadata of the head version is obtained from the DKV database, the metadata of the head version obtained from the DKV database is placed in the memory, and the readdir request is called again through the callback function.

[0130] S405, the required metadata information exists in the memory.

[0131] Because the metadata of the head version obtained from the DKV database is placed in the memory, the required metadata information exists in the memory.

[0132] S406, reading metadata from the memory and returning it to the client.

[0133] The readdir request reads the head version of metadata from the memory and returns it to the client, completing the readdir request call.

[0134] The embodiment of the present application preferentially reads metadata from the memory cache. If the cache does not hit, the metadata is pulled from the disk into the memory and then returned to the client.

[0135] In one embodiment, the method further comprises:

[0136] Determine whether the cache space occupied by the metadata stored in the memory database is greater than a threshold;

[0137] When the cache space occupied by the metadata stored in the memory database is greater than the threshold, the metadata of the snapshot version that has not been used in the memory database is eliminated.

[0138] A metadata server (MDS) is a server used to store and manage metadata. MDS servers are usually used in distributed file systems, such as Hadoop Distributed File System (HDFS). In such a system, data is stored on multiple nodes, while metadata is stored on MDS servers. MDS servers are responsible for managing file names, locations, copies, access permissions, and other information, and are also responsible for processing file metadata operations, such as creation, deletion, movement, and renaming.

[0139] In the metadata service MDS, in order to optimize the access speed of metadata, a layer of mdcache is deployed in the memory to save the recently accessed metadata, and the cache is eliminated and updated through the least recently used (Least Recently Used, LRU) algorithm. To avoid data loss, metadata is written to the memory cache and disk at the same time.

[0140] Since a large amount of snapshot metadata is stored in the memory database, when accessing metadata, it is more likely that the metadata needs to be loaded from the disk instead of being hit in the memory, resulting in a longer delay in accessing metadata. Therefore, in the case where the memory database capacity is insufficient, this embodiment eliminates the metadata of the snapshots therein, so that the memory database stores the metadata of the head version, thereby increasing the probability of metadata being hit in the cache.

[0141] Because the metadata of the head version is the metadata required for the current access, the more metadata of the head version in the memory database, the higher the access hit rate. When the cache space occupied by the metadata stored in the memory database is greater than the first set value, by eliminating the metadata of the snapshot version used in the memory database, the metadata of the redundant snapshots in the memory database will be quickly eliminated in the memory and will not occupy additional cache space, and the access hit rate of the memory will be greatly increased.

[0142] In the embodiment of the present application, when a new snapshot is created, or when COW is used, the newly generated metadata is first cached in the memory database and then persistently stored on the disk. The embodiment of the present application stores metadata in both the memory and the disk. When new metadata is generated in the memory, the metadata is written to the disk persistent cache. Since the metadata of the eliminated snapshot has been written to the disk, even if it is deleted, it will not affect normal access and can be obtained again from the disk.

[0143] In one embodiment, after eliminating unused snapshot metadata from the in-memory database, the method further includes:

[0144] If the cache space occupied by the metadata stored in the memory database is still greater than the threshold, metadata in the memory database that has not been used and whose snapshot sequence number is the set value is eliminated.

[0145] If, after unused snapshot metadata in the memory database is eliminated, the cache space occupied by the metadata stored in the memory database is still greater than the threshold, the unused metadata in the memory database whose snapshot sequence number is the set value is eliminated until the cache space occupied by the metadata in the memory database is no greater than the first set value.

[0146] refer to Figure 5 , Figure 5 Schematic diagram of a metadata elimination process provided by an embodiment of the present invention, including the following steps:

[0147] S501, the metadata server starts the tick function to eliminate redundant metadata.

[0148] Here, what is eliminated is the metadata in the memory. The tick function refers to the data deletion function that is executed periodically by the timer at a specified time interval. For example, the remaining storage space of the memory database is checked every 1 minute. If it is less than the threshold, the metadata elimination is triggered.

[0149] S502: Determine whether the memory or metadata entries occupied by metadata access exceeds a set threshold.

[0150] If it does not exceed, return directly; if it exceeds, execute the subsequent steps.

[0151] If the memory or metadata entries occupied by metadata access exceed the set threshold, the elimination mechanism is executed.

[0152] S503, traverse the lru queue and eliminate unused snapshot metadata in the lru queue.

[0153] Based on the least recently used principle, unused snapshot metadata is eliminated first.

[0154] S504: Determine whether the memory or metadata entries occupied by metadata access exceeds a set threshold.

[0155] After the snap metadata of the snapshot is eliminated, it is determined again whether the memory or metadata entries occupied by the metadata access exceeds the set threshold. If the memory or metadata entries occupied by the metadata access still exceeds the set threshold, S505 is executed.

[0156] S505, eliminating unused head metadata in the lru queue until the metadata access occupies memory or the metadata entries do not exceed the set threshold, then stopping the elimination.

[0157] This embodiment uses an elimination mechanism to eliminate the metadata of the snapshot version in the cache, so that the unused snapshot metadata stored in the database is quickly eliminated without occupying additional cache space, which can improve the hit probability of the metadata in the memory.

[0158] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0159] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0160] It should be noted that the technical solutions described in the embodiments of the present invention can be arbitrarily combined without conflict.

[0161] In addition, in the embodiments of the present invention, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0162] An embodiment of the present invention provides a metadata access device, the device comprising:

[0163] A receiving module, used for receiving a read request for metadata of a target file;

[0164] A determination module is used to determine the target metadata corresponding to the target file in the disk through the snapshot serial number, and determine the target metadata as the response data of the read request; the target metadata is the metadata of the latest version of the target file, each metadata in the disk corresponds to a snapshot serial number, and the snapshot serial number represents the version of the metadata.

[0165] In one embodiment, the determining module is specifically used to:

[0166] The metadata whose snapshot sequence number is a set value is determined as the target metadata, and the set value indicates that the corresponding metadata is the metadata of the latest version of the target file.

[0167] In one embodiment, the determining module is specifically used to:

[0168] Based on the snapshot serial number, the index node identifier of the parent directory of the target file, the identifier of the directory shard where the target file is located, and the name of the target file, the target metadata corresponding to the target file is determined from the disk; the metadata in the disk is stored in the form of key-value pairs.

[0169] In one embodiment, the device further includes: a creation module, configured to: receive a snapshot creation request for the target file; generate a new snapshot serial number based on the snapshot creation instruction; and generate metadata corresponding to the new snapshot serial number in a memory database.

[0170] In one embodiment, the device also includes: a modification module, which is used to: receive an operation request to modify target metadata; determine whether the snapshot sequence number carried by the operation request is greater than the snapshot sequence number of the target metadata; if greater, perform a write-time copy on the target metadata in a memory database to obtain first metadata and second metadata split from the target metadata, wherein the first metadata is metadata changed based on the operation request, and the second metadata is snapshot metadata of the target metadata.

[0171] In one embodiment, the metadata includes a first variable and a second variable; the value of the first variable of the first metadata is a newly generated snapshot serial number, and the value of the second variable of the first metadata is the set value; the value of the first variable of the second metadata is the snapshot serial number of the target metadata, and the value of the second variable of the second metadata is the newly generated snapshot serial number.

[0172] In one embodiment, the determination module is further used to: determine whether the metadata in the memory database hits the read request; if the metadata in the memory database does not hit the read request, determine the target metadata corresponding to the target file in the disk.

[0173] In one embodiment, the device also includes: an elimination module, which is used to: determine whether the cache space occupied by the metadata stored in the memory database is greater than a threshold; if the cache space occupied by the metadata stored in the memory database is greater than the threshold, eliminate the metadata of the unused snapshot version in the memory database.

[0174] In one embodiment, the elimination module is further used to eliminate metadata in the memory database that has not been used and whose snapshot sequence number is the set value if the cache space occupied by the metadata stored in the memory database is still greater than the threshold.

[0175] In actual application, the determination module and the receiving module can be implemented by a processor in an electronic device, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU) or a programmable gate array (FPGA).

[0176] It should be noted that: when the device provided in the above embodiment performs metadata access, only the division of the above modules is used as an example. In actual applications, the above processing can be assigned to different modules as needed, that is, the internal structure of the device is divided into different modules to complete all or part of the above-described processing. In addition, the metadata access device provided in the above embodiment and the metadata access method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0177] The metadata access device can be in the form of an image file. After the image file is executed, it can be run in the form of a container or a virtual machine to implement the metadata access method described in this application. Of course, it is not limited to the image file form. As long as some software forms that can implement the metadata access method described in this application are within the protection scope of this application.

[0178] Based on the hardware implementation of the above program modules and in order to implement the method of the embodiment of the present application, the embodiment of the present application further provides an electronic device, and the above method is implemented by a processor of the electronic device. Figure 6 Schematic diagram of the hardware structure of the electronic device of the present application embodiment. Figure 6 As shown, the electronic equipment includes:

[0179] Communication interface, which can exchange information with other devices such as network equipment;

[0180] The processor is connected to the communication interface to realize information exchange with other devices, and is used to execute the method provided by one or more technical solutions of the electronic device side when running the computer program. The computer program is stored in the memory.

[0181] Of course, in actual applications, the various components in the electronic device are coupled together through the bus system. It is understandable that the bus system is used to achieve connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 6 In the specification, various buses are labeled as bus systems.

[0182] In addition, the electronic device of the present application may be in the form of a cluster, such as a cloud computing platform composed of clusters. The so-called cloud computing platform is a business form that uses virtualization technology to pool the resources of multiple terminals and then provides the required virtual resources and services.

[0183] The memory in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program used to operate on the electronic device.

[0184] It can be understood that the memory can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), direct memory bus random access memory (DRRAM). The memory described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memory.

[0185] The method disclosed in the above embodiment of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the program in the memory and completes the steps of the above method in combination with its hardware.

[0186] Optionally, when the processor executes the program, the corresponding processes implemented by the electronic device in each method of the embodiments of the present application are implemented, which will not be described in detail here for the sake of brevity.

[0187] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, and the computer program can be executed by a processor of an electronic device to complete the steps described in the metadata access method in the embodiment of the present application.

[0188] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a first memory storing a computer program, and the computer program can be executed by a processor of an electronic device to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.

[0189] In the several embodiments provided in the present application, it should be understood that the disclosed devices, electronic devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0190] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0191] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0192] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.

[0193] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0194] It should be noted that the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.

[0195] In addition, in the examples of the present application, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0196] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A metadata access method, characterized in that: The method comprises: receiving a read request for metadata of a target file; The target metadata corresponding to the target file is determined in the disk through the snapshot serial number, and the target metadata is determined as the response data of the read request; the target metadata is the metadata of the latest version of the target file, each metadata in the disk corresponds to a snapshot serial number, and the snapshot serial number represents the version of the metadata.

2. The method according to claim 1, characterized in that Determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number includes: The metadata whose snapshot sequence number is a set value is determined as the target metadata, and the set value indicates that the corresponding metadata is the metadata of the latest version of the target file.

3. The method according to claim 1, characterized in that Determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number includes: Based on the snapshot serial number, the index node identifier of the parent directory of the target file, the identifier of the directory shard where the target file is located, and the name of the target file, the target metadata corresponding to the target file is determined from the disk; the metadata in the disk is stored in the form of key-value pairs.

4. The method according to claim 1, characterized in that: The method further comprises: Receiving a snapshot creation request for the target file; Generate a new snapshot sequence number based on the snapshot creation instruction; Generate metadata corresponding to the new snapshot sequence number in the in-memory database.

5. The method according to claim 1, characterized in that The method further comprises: Receive an operation request to modify target metadata; Determine whether the snapshot sequence number carried in the operation request is greater than the snapshot sequence number of the target metadata; If it is greater, the target metadata is copied on-write in the memory database to obtain first metadata and second metadata split from the target metadata, where the first metadata is metadata changed based on the operation request, and the second metadata is snapshot metadata of the target metadata.

6. The method according to claim 5, characterized in that The metadata includes a first variable and a second variable; the value of the first variable of the first metadata is a newly generated snapshot serial number, and the value of the second variable of the first metadata is the set value; the value of the first variable of the second metadata is the snapshot serial number of the target metadata, and the value of the second variable of the second metadata is the newly generated snapshot serial number.

7. The method according to claim 4 or 5, characterized in that: Before determining the target metadata corresponding to the target file in the disk by using the snapshot sequence number, the method further includes: Determine whether the metadata in the memory database matches the read request; If the metadata in the memory database does not match the read request, the target metadata corresponding to the target file is determined in the disk.

8. The method according to claim 1, characterized in that The method further comprises: Determine whether the cache space occupied by the metadata stored in the memory database is greater than a threshold; When the cache space occupied by the metadata stored in the memory database is greater than the threshold, the metadata of the snapshot version that has not been used in the memory database is eliminated.

9. The method according to claim 8, characterized in that After eliminating the unused snapshot metadata in the memory database, the method further includes: If the cache space occupied by the metadata stored in the memory database is still greater than the threshold, metadata in the memory database that has not been used and whose snapshot sequence number is the set value is eliminated.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the metadata access method according to any one of claims 1 to 9 are implemented.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the metadata access method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the metadata access method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Data access control method and device, storage medium and electronic equipment

    CN120891991A

  • Data access control method and device, storage medium and electronic equipment

    CN120891991B