Metadata management method and device, equipment, medium and product

By determining the priority of flash writes based on the cache space utilization and hit rate of the metadata node and aggregating the flash write nodes, the problem of low disk write efficiency of metadata nodes is solved, and more efficient disk writes and reduce write amplification is achieved.

CN120144064AActive Publication Date: 2025-06-13INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510615239.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-13
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the prior art, the disk writing efficiency of metadata nodes is low, mainly due to the complex process of metadata flushing, which increases the amount of data and write amplification.

Method used

By obtaining the priority of brushing the memory by the cache space utilization and node hit rate of each metadata node, determining the metadata node to be brushed, and aggregating it, forming a metadata block and then writing it to disk.

Benefits of technology

This method effectively avoids excessive cache space, reduces the number of I/O operations, improves disk write efficiency, and reduces write amplification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144064A_ABST
    Figure CN120144064A_ABST
Patent Text Reader

Abstract

The invention discloses a metadata management method and device, equipment, a medium and a product, the flashing priority of metadata nodes is determined through the cache space utilization rate and the node hit rate, the metadata nodes which occupy large cache space and are low in utilization rate are flashed, data which are not frequently accessed are written into a disk in time, and the efficiency of the metadata management is improved. Data which are more frequently accessed are reserved in the cache, and the reasonable flashing strategy avoids excessive occupation of the cache space. In the embodiment of the invention, the metadata nodes to be flashed are subjected to aggregation processing and then written into the disk instead of being written into single nodes respectively, so that multiple small I / O operations can be merged into less large I / O operations, the I / O frequency is reduced, the burden of the disk can be effectively reduced, the writing efficiency of the disk is improved, and the writing amplification of the disk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and particularly to a metadata management method, apparatus, device, medium and product. Background Art

[0002] All-flash storage systems use solid-state drives as the main storage medium. In an all-flash storage system, metadata access is a key area for optimization because the effective management of metadata directly affects the performance of data access and the stability of the storage system. Metadata includes the directory structure of the file system, file attributes, index information, etc., and it is crucial for quickly locating and accessing data.

[0003] In the related art, metadata is usually maintained by a B+Tree. Each metadata node will orderly maintain multiple mapping relationships from logical addresses to physical addresses. Since the amount of metadata far exceeds the amount of random access memory, the metadata nodes in the metadata need to periodically organize metadata to be flushed to disk. For example, when modifying a piece of metadata during data flushing to disk and inserting it into the corresponding leaf node for flushing to disk, the mapping relationships from the leaf node to the root node will all change. Therefore, all nodes (root node, intermediate nodes, and leaf nodes) need to be flushed to disk. However, this way of processing metadata flushing to disk requires reorganizing and correlating the merged leaf node metadata to determine their hierarchical relationships and storage orders, etc., which increases the process of metadata flushing, generates additional data, exacerbates the write amplification phenomenon, and results in a low disk write efficiency of the metadata nodes.

[0004] Therefore, how to improve the disk write efficiency of metadata nodes is an urgent problem to be solved currently. Summary of the Invention

[0005] This application provides a metadata management method, apparatus, device, medium and product to at least solve the problem of low disk write efficiency of metadata nodes in the related art.

[0006] This application provides a metadata management method, and the method includes: When the metadata stored in the cache meets the condition of being written to disk, obtain the flushing priorities of each metadata node according to the cache space utilization rate and node hit rate of each metadata node; Determine the metadata nodes to be flushed according to the flushing priorities of each metadata node; Perform aggregation processing on the metadata nodes to be flushed to obtain a metadata block; Write the metadata block to disk.

[0007] This application also provides a metadata management apparatus, and the apparatus includes: An acquisition module, configured to, when the metadata stored in the cache meets the condition for writing to the disk, obtain the flushing priorities of each metadata node according to the cache space utilization rate and node hit rate of each metadata node; A determination module, configured to determine the metadata nodes to be flushed according to the flushing priorities of the metadata nodes; An aggregation module, configured to perform aggregation processing on the metadata nodes to be flushed to obtain metadata blocks; A writing module, configured to write the metadata blocks to the disk.

[0008] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above metadata management methods when executing the computer program.

[0009] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above metadata management methods are implemented.

[0010] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above metadata management methods are implemented.

[0011] Through this application, when the metadata stored in the cache meets the condition for writing to the disk, the flushing priorities of each metadata node are obtained according to the cache space utilization rate and node hit rate of each metadata node. According to the flushing priorities of each metadata node, the metadata nodes to be flushed are determined, the metadata nodes to be flushed are aggregated to obtain metadata blocks, and the metadata blocks are written to the disk. By determining the flushing priorities of metadata nodes through the cache space utilization rate and node hit rate, the metadata nodes that occupy a large amount of cache space and have a low utilization rate are flushed, the data that is not frequently accessed is written to the disk in a timely manner, and the data that is more frequently accessed is retained in the cache. This reasonable flushing strategy avoids excessive occupation of the cache space. After aggregating the metadata nodes to be flushed and then writing to the disk, instead of writing each node separately, multiple small I / O operations can be combined into fewer large I / O operations. Reducing the number of I / O operations can effectively reduce the burden on the disk, improve the disk writing efficiency, and reduce disk write amplification. Description of the Drawings

[0012] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0013] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0014] Figure 1 It is a schematic flowchart of a metadata management method provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a metadata management device provided by an embodiment of the present application. Detailed implementation manners

[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0016] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0017] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0018] Term explanation: LBA: Logical Block Addressing, logical block address, which is an address location method for hard disk data storage. LBA is a linear address allocation method. It divides the storage space on the hard disk into logical blocks of the same size (usually each logical block is 512 bytes or 4096 bytes), and assigns a unique number (address) to each logical block. These numbers start from 0 and increase sequentially. For example, if a hard disk has 1000 logical blocks, then the LBA address range of these logical blocks is from 0 to 999.

[0019] PBA: Physical Block Addressing, the physical block address, refers to the actual storage location of data in storage devices (such as hard disks, solid-state drives, etc.). It corresponds to LBA (Logical Block Addressing) and is the address used within the storage device to identify physical storage units.

[0020] Metadata: It is data that describes other data and is used to provide information such as the structure, attributes, and location of data. For example, in a file system, metadata may include information such as the size of a file, creation time, modification time, permissions, file type, and storage location. In a database, metadata includes table structures, index information, data types, etc. Through metadata, the file system can quickly locate the storage location of a file on the disk, and the database system can quickly query table structures and index information.

[0021] Metadata node: In a all-flash storage system or related distributed system environment, a metadata node usually refers to a node responsible for storing and managing metadata, with functions such as storing, processing, and providing metadata. The metadata node provides metadata access services for other components or clients in the entire storage system. When an application needs to access the metadata of a file, it sends a request to the metadata node, and the metadata node will extract relevant information from the stored metadata according to the request content and return it to the application, enabling the application to understand the relevant attributes and data storage location of the file and thus perform subsequent operations.

[0022] Flushing to disk: It means forcing the data or metadata stored in volatile storage media such as memory to be written into non-volatile storage media such as flash memory to ensure that the data will not be lost when the system powers off or fails. An all-flash storage system usually sets up a cache in memory to temporarily store frequently accessed data and metadata to improve the read and write performance of the system. When the data in the cache reaches a certain threshold (for example, the cache is full or the residence time of some data in the cache exceeds the set period), or the system receives a specific instruction (such as a regular disk flushing instruction, the system is about to shut down, etc.), the flushing to disk operation will be triggered.

[0023] The cache hit rate refers to the ratio of the number of requests for which the required data exists in the cache to the total number of requests. When the cache hit rate for metadata read and write is low, it means that most read and write requests for metadata cannot find the required data in the cache and still need to be retrieved from the original storage location (disk).

[0024] A low effective utilization rate of the cache indicates that the cache is not fully utilized to store valuable metadata. It is possible that there is a lot of metadata stored in the cache that is not frequently accessed or will hardly be accessed again, while the frequently accessed metadata is not cached or not updated to the cache in a timely manner, resulting in waste of cache space and failure to achieve the best performance improvement effect.

[0025] A multi-level queue is a data structure composed of multiple independent queues. Each queue usually has a different priority, and data items can be assigned to different queues according to their priorities, access frequencies, or other business rules. This structure allows the system to classify and prioritize data items, thereby achieving more efficient management and scheduling.

[0026] In a cache system, a multi-level queue can be used to manage the cache policy of metadata nodes. By retaining frequently accessed data items in a high-priority queue, the cache hit rate can be improved. When a cache hit occurs, the corresponding metadata node may be moved to a higher-priority queue to reflect its popularity.

[0027] Cache eviction: When the cache space is insufficient, the system may start evicting metadata nodes from the lowest-priority queue.

[0028] Prefetching and warm-up: The system can preload the data that may be needed into the cache according to the priorities and access patterns of metadata nodes.

[0029] Meta Cache: Metadata cache is a cache mechanism used to store and manage metadata. By caching frequently accessed metadata, the number of accesses to the underlying storage device is reduced, thereby improving the read performance and response speed of the system.

[0030] All-flash storage system: It is a high-performance storage system completely based on flash memory technology. It abandons the traditional hard disk drive (HDD) and only uses solid-state drives (SSDs) as the storage medium, thereby significantly improving the read and write performance, response speed, and reliability of the storage system.

[0031] Based on the above problems, the embodiments of the present application provide a metadata management method, which will be described in detail in combination with the execution process of the metadata management method.

[0032] Refer to Figure 1 As shown in the figure, the metadata management method provided by the embodiments of the present invention includes the following steps: S11. When the metadata stored in the cache meets the condition for writing to the disk, obtain the write priorities of each metadata node according to the cache space utilization rate and node hit rate of each metadata node.

[0033] Specifically, when the metadata stored in the cache meets the conditions for writing to the disk, the flushing priority of each metadata node is obtained according to the cache space utilization and node hit rate of each metadata node.

[0034] Optionally, the above step S11 can be implemented in the following manner: (1) Based on the preset data structure, each metadata node is classified into levels and the initial flushing priority of each metadata node is determined.

[0035] Among them, the initial flushing priority of each metadata node is from high to low: leaf node, intermediate node, root node.

[0036] Specifically, the default initial flushing priority order is leaf node > intermediate node > root node. Since leaf nodes directly store actual metadata information, flushing them to disk first can make the latest data more persistent faster, ensure data security and consistency, and reduce the risk of data loss. Intermediate nodes and root nodes are mainly used to index and organize leaf nodes, and their updates are relatively less urgent, so their flushing priority is lower.

[0037] (2) Calculate the flushing priority weight of each metadata node based on the cache space utilization and node hit rate of each metadata node.

[0038] The node hit rate refers to the ratio of the number of visits to a node to the total number of visits within a certain period of time. A high hit rate means that the node is frequently visited and may be hot data. If the space utilization rate of hot data is also low, it means that it occupies more cache space but the amount of data is relatively small. At this time, reduce its disk flushing priority and let it continue to stay in the cache so as to quickly respond to subsequent access requests, reduce the number of reads from the disk, and improve system performance.

[0039] Node space utilization refers to the ratio of the cache space occupied by a node to the total space allocated to it. A high space utilization indicates that the node fully utilizes the cache space to store data. When a node has a low hit rate but high space utilization, it indicates that the node stores a lot of infrequently accessed data, occupying valuable cache resources. In this case, its flushing priority should be increased and its data should be flushed to disk to free up cache space for nodes that need it more.

[0040] Optionally, the above step (2) can be implemented as follows: Obtain the cache space utilization and node hit rate of each metadata node; Determine the cache space utilization weight and node hit rate weight of each metadata node; Calculate the write priority weights of each metadata node according to the cache space utilization rate, the cache space utilization rate weight, the node hit rate, and the node hit rate weight of each metadata node.

[0041] Specifically, the write priority weight of each node is calculated through the formula: write priority weight = node hit rate × node hit rate weight + space utilization rate × space utilization rate weight. The node hit rate weight and the space utilization rate weight are parameters preset according to the actual requirements and characteristics of the system, and are used to adjust the influence degree of the node hit rate and the space utilization rate on the weight. For example, if the system has higher requirements for the access performance of hot data, the node hit rate weight can be appropriately increased; if the cache resources of the system are relatively tight and more attention needs to be paid to the space utilization rate, the space utilization rate weight can be increased.

[0042] (3) Based on the changes in the write priority weights of the respective metadata nodes, adjust the initial write priorities of the respective metadata nodes to determine the write priorities of the respective metadata nodes.

[0043] Specifically, the higher the node hit rate and the lower the space utilization rate, the lower the write priority weight of the node at this time, and the corresponding disk write priority of the node is reduced; the lower the node hit rate and the higher the space utilization rate, the higher the write priority weight of the node at this time, and the corresponding disk write priority of the node is increased.

[0044] A possible situation is that when the cache resources are insufficient, increasing the weight of the space utilization rate is to be more inclined to write the data of the nodes with high space utilization rate but low hit rate to the disk, so as to release the cache space in time, reduce the waiting time of metadata in the cache, and avoid data loss or system performance degradation caused by cache overflow. In this way, the disk write priority can be dynamically adjusted according to the usage of cache resources, so that the system can maintain better performance and stability under different load conditions.

[0045] Exemplarily, assume that there are three metadata nodes A, B, and C in the system. The hit rate of node A is 80% and the space utilization rate is 30%; the hit rate of node B is 30% and the space utilization rate is 70%; the hit rate of node C is 50% and the space utilization rate is 50%. Let the weight of the node hit rate be 0.4 and the weight of the space utilization rate be 0.6. Then the weight of node A = 0.8×0.4 + 0.3×0.6 = 0.5; the weight of node B = 0.3×0.4 + 0.7×0.6 = 0.54; the weight of node C = 0.5×0.4 + 0.5×0.6 = 0.5. At this time, the weight of node B is the highest, and the disk writing priority is relatively high. If the cache resources are insufficient, increase the weight of the space utilization rate to 0.8 and decrease the weight of the node hit rate to 0.2. Then the weight of node A = 0.8×0.2 + 0.3×0.8 = 0.4; the weight of node B = 0.3×0.2 + 0.7×0.8 = 0.62; the weight of node C = 0.5×0.2 + 0.5×0.8 = 0.5. It can be seen that the weight of node B is further increased, and its disk writing priority is also further improved, and it is more likely to be preferentially written to the disk to release the cache space.

[0046] In some embodiments, before determining that the metadata stored in the cache meets the condition for writing to the disk, the following steps are further performed: Determine whether the metadata stored in the cache is consistent with the metadata read from the disk; If the metadata stored in the cache is inconsistent with the metadata read from the disk, mark the metadata stored in the cache as dirty metadata.

[0047] Specifically, when data is written, it not only involves the storage of actual data, but also generates operations for adding, deleting, and modifying metadata. Since the operations for adding, deleting, and modifying metadata are only written into the cache space first and not immediately persisted to storage devices such as disks, the metadata in the cache is inconsistent with the data state on the disk at this time. This metadata with inconsistent state is called dirty metadata (dirty-meta). That is, dirty metadata refers to the metadata that has been modified in the cache but has not been synchronized to the persistent storage (such as a disk).

[0048] Exemplarily, the metadata of a file has its file size information modified in the cache, but this modification has not been written to the disk. Then the metadata of this file is in the state of dirty metadata in the cache.

[0049] Optionally, the condition that the metadata stored in the cache meets the requirement for writing to the disk includes: the proportion of the dirty metadata occupied in the cache is greater than a preset proportion.

[0050] Among them, the preset ratio can be set according to the actual situation. For example, when the dirty metadata stored in the cache reaches 80% of the cache space, or other reasonable values, there is no specific limitation here.

[0051] Specifically, the cache space has a certain capacity limit. When the dirty metadata stored in it reaches a certain water level (that is, a certain proportion or quantity threshold of the cache capacity), a cache write operation will be triggered. Cache writing is to write the accumulated dirty metadata in the cache to a persistent storage device (such as a disk) so that the data in the storage device (such as a disk) is consistent with the data state in the cache.

[0052] Exemplarily, when the dirty metadata occupancy in the cache space reaches 80%, the system will automatically start the cache writing process, batch write these dirty metadata to the disk to release the cache space and ensure data persistence.

[0053] In some embodiments, before determining that the metadata stored in the cache meets the condition for writing to the disk, the following steps are also performed: In response to a data write instruction, aggregate the data to be written into target data; Allocate continuous space for the target data and sequentially write the target data into the storage units of the data space by means of append writing; When the target metadata is generated when writing the target data, query whether there is a metadata node corresponding to the target metadata in the preset data structure; If there is a metadata node corresponding to the target metadata in the preset data structure, update the metadata node; If there is no metadata node corresponding to the target metadata in the preset data structure, apply for a metadata node corresponding to the target metadata.

[0054] Among them, the preset data structure can be a B+ tree. A B+ tree is a balanced tree data structure, often used in storage systems to maintain indexes. It supports fast data search, insertion, and deletion operations.

[0055] Specifically, when writing data, in response to a data write instruction, aggregate the data to be written into target data; allocate continuous space for the target data and write the target data into the storage units of the data space by means of append writing. When the target metadata is generated when writing the target data, query whether there is a metadata node corresponding to the target metadata in the preset data structure; if there is a metadata node corresponding to the target metadata in the preset data structure, update the metadata node; if there is no metadata node corresponding to the target metadata in the preset data structure, apply for a metadata node corresponding to the target metadata.

[0056] When applying for a metadata node, a cache space of a preset size will be allocated for it, and the size of the cache space is 4K or 8K. Setting the size of the metadata node to 4KB or 8KB (each node can index more child nodes), compared with the cache space of 512 bytes for each metadata node in the prior art, the cache space of the metadata node in this solution can more easily perform node splitting and merging operations, and is also convenient for caching and prefetching in memory; larger nodes can reduce the depth of the tree, thereby reducing the number of cache misses required when searching for data; larger nodes can store more key values and pointers, thereby reducing disk I / O operations when searching for data. B+ trees are usually optimized for sequential I / O operations, and larger nodes can make better use of this feature to improve I / O efficiency.

[0057] In a log-structured storage system, all write operations (including data and metadata) are recorded in a log area, which is usually located in the storage device and is written sequentially. This sequential writing method can make full use of the writing performance of the storage device. Log aggregation refers to combining multiple small write operations into one large write operation. The advantage of doing this is that it can reduce the number of write operations, thereby reducing the wear on the storage device and improving the write efficiency. In a log-structured storage system, these small write operations are first recorded in the log, and then a background process (such as a garbage collector) periodically aggregates these logs and writes them into the data space.

[0058] When writing the aggregated data into the data space, the system will allocate continuous space for the aggregated data. This can reduce the addressing time of the storage device when writing data and improve the writing performance.

[0059] Through the above strategies of log aggregation and allocating continuous space, write amplification can be effectively reduced. Specifically, by combining multiple small write operations into one large write operation, the actual amount of data written can be reduced, thereby reducing write amplification.

[0060] In addition, when data is written into an all-flash storage system, new metadata is usually generated. Taking the insertion of the logical address to physical address (LP) mapping as an example, the processing of new metadata through an indexed B+ tree is described. Among them, the LP mapping is the mapping from the logical address to the physical address, which indicates the actual location of the data on the storage device and involves updating the node to include the logical address and the corresponding physical address.

[0061] When a user or application initiates a data write request, the storage system writes the data to the specified location; associated with the data write operation, the system generates new metadata. For example, in a file system, this may include the creation time, owner, size, etc. of a file; in a database, this may include index information of a table, identifiers of rows, etc. The system uses a B+ tree to index and query metadata. When new metadata needs to be inserted, the system first looks up the corresponding metadata node (Node) in the B+ tree. The system queries in the B+ tree whether there is a node related to the new metadata (i.e., the target metadata), for example, it can be achieved by looking up the file name, logical address of a data block, etc. in the B+ tree. If the corresponding node is found, the system updates the information in that node to reflect the new data write operation. If the corresponding node is not found, the system needs to create a target metadata node and insert it into the B+ tree. After inserting a new LP mapping, the system needs to maintain the balance of the B+ tree, which may involve splitting or merging of nodes. To maintain the performance of the B+ tree, the system needs to ensure that the height of the tree is as low as possible.

[0062] Querying and inserting metadata LP mappings through an indexed B+ tree is an effective method for managing metadata in a storage system. This method not only supports fast data lookup and update, but also helps to maintain data consistency and integrity. By optimizing the structure of the B+ tree and managing the cache, the performance and reliability of the storage system can be further improved.

[0063] In the embodiments of the present disclosure, the allocation of the metadata space is uniformly responsible for by the data space management system, which can achieve more efficient resource management and scheduling.

[0064] Among them, the metadata memory space refers to the area allocated in the memory for temporarily storing metadata for fast access and management of metadata. The metadata memory space is used to store metadata, and the metadata is used to describe the attributes and locations of data in the data space, including the name, size, creation time, modification time, permissions, storage location, etc. of a file, which can help the system quickly locate and access the data in the data space and maintain data consistency and integrity. The data in the metadata memory space is dynamically changing and is updated as data is added, deleted, modified, or queried. Since the memory access speed is much faster than that of the disk, the metadata access speed can be improved. For frequently updated metadata, caching it in the memory can reduce the number of writes to the disk because multiple write operations can be aggregated in the memory first and then written to the disk uniformly.

[0065] The data space refers to the area in the storage system for storing actual user data. The data space is used to store user data such as file contents, database records, pictures, videos, etc.

[0066] In addition, when data is added, deleted, or modified, the system does not immediately apply these operations directly to the metadata in the storage device. Instead, these operations are first merged and then written to the cache space. The cache space is a region in memory. Since the read and write speeds of memory are much faster than those of storage devices (such as solid-state drives), this can reduce the number of direct writes to the storage device and improve the write performance. For example, the add, delete, and modify operations of the metadata of multiple files may be merged into a batch and written to the cache, rather than each operation being written to the storage device individually.

[0067] In some embodiments, the data of the metadata node is synchronously copied to the mirror node.

[0068] To ensure the security and reliability of the data, a synchronous mirroring method is adopted. That is, the data of the metadata node is synchronously copied to the mirror node for backup. When the primary node (the node storing the original metadata) fails, such as due to hardware damage, software errors, etc., resulting in data loss or inaccessibility, a complete and consistent metadata copy can be obtained from the mirror node, thus ensuring that the system can continue to operate normally and reducing the risk of data loss. This cache mirroring function is similar to a data redundancy strategy, which improves the fault tolerance of the system and the availability of the data by saving the same data on different nodes.

[0069] S12. Determine the metadata node to be flushed according to the flush priorities of the respective metadata nodes.

[0070] In some embodiments, after performing the above step S12, the following steps may further be performed: Within a first preset duration, determine whether the number of writes of the first metadata node in the preset address space of the cache is greater than or equal to a preset number; If the number of writes of the first metadata node is greater than or equal to the preset number, reduce the flush priority of the first metadata node; If the number of writes of the first metadata node is less than the preset number, increase the flush priority of the first metadata node.

[0071] Specifically, before writing metadata, it is necessary to determine whether to write a certain metadata node based on the hit situation within the local address space and time range. The preset address space refers to the address area adjacent to or related to the metadata node to be judged currently, which includes the metadata node itself and some surrounding nodes. The first preset duration is a set time period used to count the hit situation of nodes during this period. By comprehensively considering these two factors, the system can more accurately determine which nodes need to be written and which can be temporarily not processed. Within the first preset duration, by judging whether the write count of the first metadata node in the preset address space of the cache is greater than or equal to the preset count, if the write count of the first metadata node is greater than or equal to the preset count, the write priority of the first metadata node is reduced; otherwise, the write priority of the first metadata node is increased.

[0072] When it is found that the first metadata node and its related metadata nodes before and after are written multiple times in a short period, it can be judged that there is hot write data in the preset address space. In some cases, the preset duration can be defined as a time interval of several minutes or even shorter.

[0073] For the related metadata nodes with hot write data, the system will reduce their disk write priorities. This is because if frequently updated data is immediately written to disk (writing data in the cache to a persistent storage device such as a disk), it may cause a large number of disk I / O operations, affecting system performance. After reducing the disk write priority, the data of these nodes will stay in the cache for a longer time, waiting for an appropriate time to be written to disk to reduce unnecessary I / O overhead.

[0074] Contrary to the situation of hot write data, if the metadata node within the preset address space range is not modified within the first preset duration, the system will increase the write priority of the first metadata node. The first preset duration is also a configurable time parameter, for example, it may be several hours or even several days. For such nodes that have not been modified for a long time, writing them to a storage device such as a disk in a timely manner can ensure data persistence and release cache space to cache other required node data. After increasing the write priority, these nodes will be preferentially processed in the cache write operation to complete the disk write operation as soon as possible.

[0075] A strategy for dynamically adjusting the metadata node flushing priority based on the hit situation within the local address space and time range, the main purpose of which is to optimize the data transfer process between the cache and persistent storage (such as disks), and improve the overall performance and resource utilization rate of the system. By reasonably arranging which nodes are flushed first and which nodes are flushed later, it can not only ensure that hot data is processed more effectively in the cache, but also persistently store stable data in a timely manner, avoiding waste of cache space, so that the system can maintain efficient operation when processing different types of data access and updates.

[0076] As the metadata is continuously updated and deleted, there may be many unused metadata spaces. If these garbage spaces are not recycled in time, it will reduce the utilization rate of the storage space. Based on this, the embodiments of the present disclosure provide the following metadata space garbage collection mechanism to ensure the continuous effectiveness of the metadata space.

[0077] Optionally, after performing data aggregation processing on multiple metadata nodes in the cache to form multiple metadata blocks, the following method can also be executed: Periodically obtain the space utilization rate of the multiple metadata blocks; Recycle the space of the metadata blocks with a space utilization rate lower than the preset space utilization rate among the multiple metadata blocks.

[0078] Among them, the preset space utilization rate can be set according to the actual situation. For example, it can be set to 50%, 30%, 20%, etc., and no specific limitation is made here.

[0079] Specifically, periodically obtain the space utilization rate of multiple metadata blocks, and recycle the space of the metadata blocks with a space utilization rate lower than the preset space utilization rate among the multiple metadata blocks. That is, when the usage efficiency of the metadata space drops to a certain level (for example, there is too much free space or fragmentation is serious), the system should start metadata garbage collection. Among them, metadata garbage collection includes: identifying unused metadata spaces, marking them as recyclable, and then reorganizing the storage space to release these spaces for subsequent use.

[0080] Optionally, after performing data aggregation processing on multiple metadata nodes in the cache to form multiple metadata blocks, the following method can also be executed: Real-time obtain the overall inefficiency of the metadata space, the usage rate of the metadata space, and the dynamic change speed of the metadata space.

[0081] Determine the preset recycling threshold of the metadata space according to the overall inefficiency of the metadata space, the usage rate of the metadata space, and the dynamic change speed of the metadata space.

[0082] Among them, the overall inefficiency of the metadata space is used to represent the proportion of the unutilized part existing in the metadata space. For example, some metadata may become irrelevant due to changes in the data structure but still occupy storage space, which will reduce the overall efficiency of the metadata space. By evaluating the overall inefficiency, it is possible to understand how much potential space can be recycled in the metadata space.

[0083] The usage rate of the metadata space is used to represent the proportion of the occupied part in the current metadata space to the total space. If the usage rate of the metadata space is too high, it indicates that the metadata space may face the pressure of insufficient space and some space needs to be recycled in a timely manner to avoid system failures or performance degradation. On the contrary, if the usage rate is too low, it may mean that there is an over-allocation situation in the metadata space, and the allocation strategy needs to be adjusted to improve the space utilization rate. By monitoring the usage rate, the usage status of the metadata space can be understood, providing an important basis for adjusting the recycling threshold.

[0084] The dynamic change speed of the metadata space is used to represent the speed at which the usage of the metadata space changes over time. The dynamic change speed includes the frequencies of operations such as adding, deleting, and modifying metadata. If the dynamic change speed of the metadata space is relatively fast, it indicates that the data activities in the system are relatively frequent, and the recycling threshold may need to be adjusted more flexibly to adapt to the changing space requirements. For example, in some systems with high real-time requirements, metadata may be updated and deleted frequently. At this time, the recycling strategy needs to be adjusted in a timely manner according to the dynamic change speed to ensure that the metadata space always maintains a high utilization state.

[0085] Specifically, by comprehensively considering the overall inefficiency of the metadata space, the usage rate of the metadata space, and the dynamic change speed of the metadata space, a preset recycling threshold is determined through a certain algorithm. For example, corresponding weights can be assigned according to the importance of different factors and then weighted calculations are performed. For example, for a system with high requirements for space utilization, the weight of the overall inefficiency may be set to a relatively high value; while for a system with high requirements for data real-time performance, the dynamic change speed may be more emphasized.

[0086] Through this comprehensive judgment method, a relatively reasonable recycling threshold can be obtained, enabling the metadata space to maintain a high recycling utilization rate under different workloads and system states, while ensuring the stability and performance of the system.

[0087] When it is detected that the usage rate of the metadata space is greater than or equal to the preset recycling threshold, the space recycling mechanism is triggered.

[0088] The preset recycling threshold of the metadata space adopts a dynamic adjustment strategy, which can be flexibly adjusted according to the actual usage of the metadata space, and more efficient metadata space management can be achieved.

[0089] S13. Aggregate the metadata nodes to be rewritten to obtain metadata blocks.

[0090] Specifically, aggregate the metadata nodes to be rewritten to obtain metadata blocks. First, classify these metadata nodes to be rewritten and group them according to the type, relevance or other attributes of the metadata. During the classification process, check whether there is redundant or duplicate information in the metadata nodes. If duplicate metadata records are found, merge or remove operations will be performed, and only one valid piece of data will be retained. This can reduce the data volume, improve the storage efficiency, and also avoid repeatedly writing the same data during the rewrite process.

[0091] The aggregation process also includes constructing the logical relationships between metadata nodes. For example, in a tree - structured metadata organization, determine the parent - child relationships, hierarchical relationships, etc. between each node. By clarifying these logical relationships, after rewriting the metadata to persistent storage, the integrity and correctness of the data structure can be maintained, facilitating subsequent data access and management.

[0092] S14. Write the metadata blocks to the disk.

[0093] Specifically, when a metadata rewrite operation is triggered, first add the metadata nodes to be rewritten to the write - to list. This write - to list can be understood as a task queue, which records the metadata node information that is about to be rewritten to the storage device (such as a disk) so that the system can process it in order.

[0094] After adding the metadata nodes to the above - mentioned write - to list, the system will apply for block space for these metadata. Metadata blocks are the basic units for storing metadata and are used to store metadata information. The process of applying for block space is actually to reserve a continuous storage space on the storage device for subsequent writing of metadata.

[0095] To improve storage efficiency and performance, the system will organize the metadata into stripes for rewriting. Striping is a technique that divides data into multiple small pieces and distributes them at different storage locations. When constructing a full - stripe rewrite of data, the system will fill the metadata nodes in the destage list into the stripe according to certain rules to make it reach the full - stripe state. This can make full use of the read - write bandwidth of the storage device, reduce the number of I / O operations, and improve the rewrite efficiency.

[0096] After constructing the full stripe and writing data, the system needs to allocate physical addresses for each stripe. The physical address is the actual storage location on the storage device. By mapping the stripe to the physical address, the system can accurately know where to write data to the storage device and where to read data from. The process of allocating physical addresses is usually completed by the address mapping module of the storage system, which will allocate appropriate physical addresses for the stripes according to the layout of the storage device and the available space situation.

[0097] Write the constructed stripe data to the allocated physical address. This process is the metadata write operation. Once the write is completed, the metadata changes from the original dirty state to the clean state. Dirty metadata refers to data that has been modified in memory but has not yet been written to the storage device, while clean metadata refers to data that has been successfully written to the storage device and is consistent with the data on the storage device.

[0098] In some embodiments, after performing the above step S14 (writing the metadata block to the disk), the following steps are further performed: 1). Obtain the access popularity of the metadata written to the disk according to the access frequency and access time of the metadata; 2). Dynamically adjust the positions of each metadata in the multi-level queue according to the access popularity of the metadata written to the disk; 3). Periodically evaluate the access popularity of the metadata stored in the multi-level queue to identify hot metadata; 4). Prefetch the hot data in the multi-level queue and the associated data of the hot data into the cache.

[0099] Among them, the associated data of the hot data includes: the metadata adjacent to the hot data and the metadata logically related to the hot data.

[0100] Exemplarily, if the currently accessed metadata is a file or directory in the file system, prefetch its adjacent metadata. For example, if a file is accessed, the metadata of other files in the same directory can be prefetched. It is also possible to prefetch according to the logical relationship between metadata. For example, if the metadata of a directory is accessed, the metadata of all subdirectories and files under the directory can be prefetched. Another example is that if a user frequently accesses the files in a certain directory, the metadata of other files in the directory can be prefetched.

[0101] Optionally, the multi-level queue includes a first multi-level queue and a second multi-level queue. Among them, the first multi-level queue is used to count the metadata heat of intermediate nodes, and the second multi-level queue is used to count the metadata heat of leaf nodes; the access priority of the first multi-level queue is higher than that of the second multi-level queue; the first multi-level queue and the second multi-level queue each include at least two priority queues.

[0102] Specifically, the multi-level queue may include multiple priority queues, and each priority queue is used to store metadata with different heat levels. In the embodiments of the present disclosure, the multi-level queue includes a first multi-level queue and a second multi-level queue. The first multi-level queue is used to count the metadata heat of intermediate nodes, and the second multi-level queue is used to count the metadata heat of leaf nodes. The access priority of the first multi-level queue is higher than that of the second multi-level queue; the first multi-level queue and the second multi-level queue each include at least two priority queues. For example, both the first multi-level queue and the second multi-level queue may include a low-priority queue, a medium-priority queue, and a high-priority queue. Among them, the low-priority queue is used to store metadata with a low access frequency, the medium-priority queue is used to store metadata with a medium access frequency, and the high-priority queue is used to store metadata with a high access frequency.

[0103] Among them, the access frequency refers to the number of times the metadata is accessed, and the access time refers to the time when the metadata was last accessed.

[0104] Specifically, based on the access count and access time, the access heat of the metadata can be calculated. When the metadata is first accessed, it can be added to the lowest-priority queue, and at the same time, the access count and access time of the metadata are recorded. Each time the metadata is accessed, its access count and access time are updated. For example, a counter can be used to record the access count, and the current timestamp can be recorded as the access time. According to the calculated access heat, the position of the metadata in the multi-level queue is dynamically adjusted. If the access heat of the metadata exceeds the access heat threshold, it is promoted from the current queue to a higher-priority queue. For example, if the access heat of the metadata exceeds the threshold of the low-priority queue, it is promoted to the medium-priority queue; if the access heat of the metadata exceeds the threshold of the medium-priority queue, it is promoted to the high-priority queue. Periodically evaluate the heat of the metadata in the multi-level queue, identify the metadata that may become hot, and prefetch these hot metadata and their associated metadata into the cache. For example, when the metadata in the high-priority queue is accessed, the system believes that these metadata and their associated metadata may be hot data and need to be prefetched. Additionally, if the cache space is insufficient, the metadata in the low-priority queue is evicted.

[0105] In the embodiments of the present disclosure, by combining the access frequency and time, the popularity of metadata can be more accurately counted. The structure of the multi-level queue makes cache eviction and prefetching more efficient, improving the cache hit rate. By dynamically adjusting the priority of metadata, it is ensured that hot data always remains in the high-priority queue. By prefetching hot data and its associated data, it is ensured that this data is in the cache, thereby improving the cache hit rate. The prefetch operation can load the data into the cache in advance, reducing the latency when the user accesses the data. The clean metadata after the write operation is maintained by the read cache. It can be understood that after the write operation is completed, the original write cache becomes a read cache for use. Putting the clean metadata into the cache can enable subsequent read operations on this metadata to directly obtain it from the cache without having to read it from the storage device again, thus greatly improving the performance of the system.

[0106] In some embodiments, after the leaf node is flushed to disk, update the mapping relationships of the intermediate node, the root node, and the leaf node.

[0107] Specifically, after the metadata write operation is completed, it is necessary to update the mapping relationships of the intermediate node and the root node with the flushed leaf node. In the tree structure of the metadata, the leaf node stores the actual metadata information, while the intermediate node and the root node are used to organize and index these leaf nodes. After the leaf node is flushed to disk, its physical location on the storage device may change, so it is necessary to update the mapping information in the intermediate node and the root node to ensure that the flushed leaf node can be correctly accessed.

[0108] Since the intermediate node and the root node are frequently changed due to the flushing of the leaf node, they belong to hot data. Hot data refers to data that is frequently accessed or modified. For this hot data, the system sets their flushing priority to the lowest. This is because frequent flushing of these nodes will generate a large number of I / O operations, affecting the system performance. If the cache space permits, these intermediate nodes and root nodes will remain in memory permanently to quickly respond to frequent access requests, reducing the number of reads from the storage device and improving the overall performance of the system.

[0109] The metadata management method provided by the embodiments of the present disclosure determines the flushing priority of metadata nodes based on the cache space utilization rate and the node hit rate. For metadata nodes that occupy a large amount of cache space and have a low utilization rate, they are flushed, writing the data that is not frequently accessed to the disk in a timely manner and retaining the data that is more frequently accessed in the cache. This reasonable flushing strategy avoids over-occupying the cache space. Aggregate the metadata nodes to be flushed and then write them to the disk instead of writing them separately for each single node. In this way, multiple small I / O operations can be combined into fewer large I / O operations. Reducing the number of I / O operations can effectively reduce the burden on the disk, improve the disk write efficiency, and reduce the disk write amplification.

[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0111] An embodiment of the present application further provides a metadata management device 200. Figure 2 FIG. 262 is a schematic structural diagram of a metadata management device 200 provided by the present disclosure, including: An acquisition module 210, configured to obtain the write priorities of each metadata node according to the cache space utilization rate and node hit rate of each metadata node when the metadata stored in the cache meets the condition for writing to the disk. A determination module 220, configured to determine the metadata nodes to be flushed according to the write priorities of the respective metadata nodes. An aggregation module 230, configured to perform aggregation processing on the metadata nodes to be flushed to obtain metadata blocks. A write module 240, configured to write the metadata blocks to the disk.

[0112] As an optional implementation manner of an embodiment of the present disclosure, the device further includes a judgment module, and the judgment module is configured to: Judge whether the metadata stored in the cache is consistent with the original metadata read from the disk; If the metadata stored in the cache is inconsistent with the original metadata read from the disk, mark the metadata stored in the cache as dirty metadata.

[0113] As an optional implementation manner of an embodiment of the present disclosure, the disk write condition includes: the condition that the metadata stored in the cache meets the condition for writing to the disk includes: the proportion of dirty metadata occupied in the cache is greater than a preset proportion.

[0114] As an optional implementation manner of an embodiment of the present disclosure, the device further includes a prefetch module, and the prefetch module includes: An acquisition unit, configured to obtain the access popularity of the metadata written to the disk according to the access times and access time of the metadata. A statistics unit, configured to dynamically adjust the positions of each metadata in the multi-level queue according to the access popularity of the metadata written to the disk. An identification unit, configured to periodically evaluate the access popularity of the metadata stored in the multi-level queue and identify hot metadata. A prefetching unit for prefetching the hot data in the multi-level queue and the associated data of the hot data into a cache; the associated data of the hot data includes: metadata adjacent to the hot data and metadata logically related to the hot data.

[0115] As an optional implementation manner of an embodiment of the present disclosure, the multi-level queue includes a first multi-level queue and a second multi-level queue; the first multi-level queue is used to count the metadata heat of intermediate nodes, and the second multi-level queue is used to count the metadata heat of leaf nodes; the access priority of the first multi-level queue is higher than that of the second multi-level queue; the first multi-level queue and the second multi-level queue each include at least two priority queues.

[0116] As an optional implementation manner of an embodiment of the present disclosure, the apparatus further includes: a write response module, and the write response module is specifically configured to: In response to a data write instruction, aggregate the data to be written into target data; Allocate continuous space for the target data, and sequentially write the target data into the storage unit of the data space by means of append writing; When target metadata is generated when writing the target data, query whether there is a metadata node corresponding to the target metadata in a preset data structure; If there is a metadata node corresponding to the target metadata in the preset data structure, update the metadata node; If there is no metadata node corresponding to the target metadata in the preset data structure, apply for a metadata node corresponding to the target metadata.

[0117] As an optional implementation manner of an embodiment of the present disclosure, the obtaining module includes: A determination unit for hierarchically dividing each metadata node based on a preset data structure to determine the initial write priority of each metadata node; the initial write priorities of the metadata nodes from high to low are: leaf nodes, intermediate nodes, and root nodes; A calculation unit for calculating the write priority weights of each metadata node according to the cache space utilization rate and node hit rate of each metadata node; An adjustment unit for adjusting the initial write priority of each metadata node based on the change in the write priority weights of each metadata node to determine the write priority of each metadata node.

[0118] As an optional implementation manner of an embodiment of the present disclosure, the calculation unit is specifically configured to: Obtain the cache space utilization rate and node hit rate of each metadata node; Determine the cache space utilization weight and the node hit rate weight of each metadata node; Calculate the write priority weight of each metadata node according to the cache space utilization, the cache space utilization weight, the node hit rate, and the node hit rate weight of each metadata node.

[0119] As an optional implementation manner of an embodiment of the present disclosure, the apparatus further includes an update module, configured to: When the leaf node is flushed to the disk, update the mapping relationship between the intermediate node, the root node, and the leaf node.

[0120] As an optional implementation manner of an embodiment of the present disclosure, the apparatus further includes a priority adjustment module, and the priority adjustment module is configured to: Within a first preset duration, determine whether the write times of a first metadata node in a preset address space of the cache are greater than or equal to a preset number of times; If the write times of the first metadata node are greater than or equal to the preset number of times, reduce the write priority of the first metadata node; If the write times of the first metadata node are less than the preset number of times, increase the write priority of the first metadata node.

[0121] As an optional implementation manner of an embodiment of the present disclosure, the apparatus further includes a space recovery module, and the space recovery module is configured to: Periodically obtain the space utilization rate of the metadata block; When the space utilization rate of the metadata block is lower than a preset space utilization rate, perform metadata block space recovery.

[0122] For the description of the features in the corresponding embodiment of the metadata management apparatus 200, reference may be made to the relevant description in the corresponding embodiment of the metadata management method, which will not be elaborated here one by one.

[0123] The metadata management apparatus provided by the embodiments of the present disclosure determines the write priority of metadata nodes through the cache space utilization rate and the node hit rate, flushes the metadata nodes that occupy a large cache space and have a low utilization rate, writes the data that is not frequently accessed to the disk in time, and retains the data that is more frequently accessed in the cache. This reasonable write strategy avoids excessive occupation of the cache space. After aggregating the metadata nodes to be written and then writing them to the disk instead of writing them separately for each node, multiple small I / O operations can be combined into fewer large I / O operations. Reducing the number of I / O operations can effectively reduce the burden on the disk, improve the disk write efficiency, and reduce the disk write amplification.

[0124] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the metadata management method.

[0125] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the metadata management method when running.

[0126] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs for short), random access memories (RAMs for short), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0127] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the metadata management method are implemented.

[0128] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the metadata management method are implemented.

[0129] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0130] The above has introduced in detail a metadata management method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A metadata management method, characterized in that: The method comprises: When the metadata stored in the cache meets the conditions for writing to the disk, the flushing priority of each metadata node is obtained according to the cache space utilization and node hit rate of each metadata node; Determine the metadata nodes to be flushed according to the flushing priorities of the metadata nodes; Aggregating the metadata nodes to be written to obtain metadata blocks; Write the metadata block to disk.

2. The metadata management method according to claim 1, characterized in that: The method further comprises: Determining whether the metadata stored in the cache is consistent with the metadata read from the disk; If the metadata stored in the cache is inconsistent with the metadata read from the disk, the metadata stored in the cache is marked as dirty metadata.

3. The metadata management method according to claim 2, characterized in that: The condition that the metadata stored in the cache meets the writing condition to the disk includes: the proportion of dirty metadata in the cache is greater than a preset proportion.

4. The metadata management method according to claim 1, characterized in that: After writing the metadata block to the disk, the method further includes: Obtain the access popularity of metadata written to disk based on the number of metadata accesses and access time; Dynamically adjust the position of each metadata in the multi-level queue according to the access popularity of the metadata written to the disk; Periodically evaluating the access heat of the metadata stored in the multi-level queues to identify hot metadata; The hot data in the multi-level queue and the associated data of the hot data are pre-read into the cache; the associated data of the hot data includes: metadata adjacent to the hot data and metadata logically related to the hot data.

5. The metadata management method according to claim 4, characterized in that: The multi-level queue includes a first multi-level queue and a second multi-level queue; the first multi-level queue is used to count the metadata heat of intermediate nodes, and the second multi-level queue is used to count the metadata heat of leaf nodes; the access priority of the first multi-level queue is higher than that of the second multi-level queue; the first multi-level queue and the second multi-level queue each include at least two priority queues.

6. The metadata management method according to claim 1, characterized in that: The method further comprises: In response to a data writing instruction, aggregating the data to be written into target data; Allocating continuous space for the target data, and sequentially writing the target data into the storage unit of the data space by appending writing; When writing the target data to generate target metadata, querying whether there is a metadata node corresponding to the target metadata in the preset data structure; If a metadata node corresponding to the target metadata exists in the preset data structure, updating the metadata node; If the metadata node corresponding to the target metadata does not exist in the preset data structure, then apply for the metadata node corresponding to the target metadata.

7. The metadata management method according to claim 1, characterized in that: The step of obtaining the flushing priority of each metadata node according to the cache space utilization rate and the node hit rate of each metadata node includes: Based on the preset data structure, each metadata node is graded to determine the initial flushing priority of each metadata node; the initial flushing priority of each metadata node is from high to low: leaf node, intermediate node, root node; Calculate the flushing priority weight of each metadata node based on the cache space utilization and node hit rate of each metadata node; Based on the change of the flushing priority weight of each metadata node, the initial flushing priority of each metadata node is adjusted to determine the flushing priority of each metadata node.

8. The metadata management method according to claim 7, characterized in that: The step of calculating the flushing priority weight of each metadata node according to the cache space utilization and node hit rate of each metadata node includes: Obtain the cache space utilization and node hit rate of each metadata node; Determine the cache space utilization weight and node hit rate weight of each metadata node; The flushing priority weight of each metadata node is calculated based on the cache space utilization, the cache space utilization weight, the node hit rate and the node hit rate weight of each metadata node.

9. The metadata management method according to claim 7, characterized in that: After writing the metadata block to the disk, the method further includes: When the leaf node is refreshed, the mapping relationship among the intermediate node, the root node and the leaf node is updated.

10. The metadata management method according to claim 1, characterized in that: After determining the metadata nodes to be flushed according to the flushing priorities of the metadata nodes, the method further includes: Within a first preset time period, determining whether the number of writes to the first metadata node in the preset address space of the cache is greater than or equal to a preset number; If the number of writes to the first metadata node is greater than or equal to a preset number, lowering the flushing priority of the first metadata node; If the number of write times of the first metadata node is less than the preset number, the flushing priority of the first metadata node is increased.

11. The metadata management method according to claim 1, characterized in that: After the metadata nodes to be flushed are aggregated to obtain metadata blocks, the method further includes: Periodically obtaining the space utilization of the metadata block; When the space utilization rate of the metadata block is lower than the preset space utilization rate, metadata block space recovery is performed.

12. A metadata management device, characterized in that: The device comprises: The acquisition module is used to obtain the flushing priority of each metadata node according to the cache space utilization rate and node hit rate of each metadata node when the metadata stored in the cache meets the conditions for writing to the disk; A determination module, used to determine the metadata nodes to be flushed according to the flushing priorities of the metadata nodes; An aggregation module, used for performing aggregation processing on the metadata nodes to be flushed to obtain metadata blocks; The writing module is used to write the metadata block to the disk.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the metadata management method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the metadata management method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the metadata management method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Remote file data access performance optimization method based on efficient caching of client

    CN110188080A

  • Data processing method and device, terminal and medium

    CN110555001A

  • All-flash storage system metadata write cache disk refreshing method and related components

    CN110795042A

  • Data updating method, device and equipment and readable storage medium

    CN115168244A

  • Metadata flashing method, electronic equipment and computer program product

    CN115840663A