File metadata multi-version storage management method and device, equipment and storage medium

By generating a snapshot version set and storing it in an encoded data structure, decoupling snapshot information from the current version, and optimizing the memory query strategy, the memory and storage burden issues of the file system in high-frequency snapshot scenarios are resolved, thereby improving system performance and resource utilization.

CN120804030APending Publication Date: 2025-10-17JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510936906.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In existing file systems, multi-version metadata storage increases the burden on memory cache and persistent storage in high-frequency or large-scale snapshot scenarios, affecting performance and efficiency. Query operations are also complex and time-consuming.

Method used

By generating a snapshot version set, extracting key snapshot versions and storing them in encoded data structures, decoupling snapshot information from the current version, prioritizing in-memory queries, and dynamically managing memory resources in combination with the LRU strategy.

Benefits of technology

It achieves efficient management of multiple versions of file metadata, improves system performance and resource utilization, reduces memory and storage burdens, and is suitable for high-frequency snapshots and high-concurrency access scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804030A_ABST
    Figure CN120804030A_ABST
Patent Text Reader

Abstract

The invention discloses a file metadata multi-version storage management method and device, computer equipment and a storage medium. The method comprises the steps of obtaining file metadata; in response to the received generation instruction, generating a snapshot version for the file metadata to obtain a snapshot version set; extracting a first snapshot version and a second snapshot version from the snapshot version set, updating the snapshot version set, and storing the snapshot version set in a memory area; packaging the file metadata, the first snapshot version and the second snapshot version to obtain a coded data structure, and storing the coded data structure in a memory area; storing the updated snapshot version set and the coded data structure in a storage device; querying the memory area according to the query instruction, querying the storage device if the query fails, and returning a query result; by means of the method, decoupling storage of multiple versions of file metadata can be achieved, and the system memory utilization rate and the snapshot version management efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and in particular to a file metadata multi-version storage management method and device, equipment and a storage medium. BACKGROUND

[0002] In modern file systems, multi-version storage management of metadata is a basic capability to implement key functions such as snapshot creation, data recovery, and historical version rollback. As the requirements of business systems for data integrity, auditability, and operation traceability continue to increase, file systems must be able to support high-frequency, massive version snapshot operations for a long time. In particular, in actual production environments, metadata snapshots at different time points need to be recorded and managed completely, which puts higher requirements on the architecture design and resource scheduling of file systems.

[0003] Current file systems generally use the copy-on-write (COW) mechanism to implement the snapshot function of metadata. Whenever data changes, in order to preserve the original state, the system will copy the current inode structure to generate a snapshot version. The inode structure usually contains fields such as file permission information, user and group identification, time attributes, and extended attributes. In existing designs, both the current metadata version and the historical snapshot version generated through COW are uniformly encapsulated in the same inode structure body for management. The system loads these versions in memory as a whole and writes them together in the storage layer. For systems that use key-value databases for metadata persistence, the implementation is usually to serialize the complete inode structure, including all snapshot content, into a single record and write it into the value field, thereby completing unified storage.

[0004] This approach can meet the basic needs when the number of snapshots is small, but it is not suitable for high-frequency snapshot or large-scale snapshot scenarios. Centralized storage of multi-version data significantly increases the data volume of a single inode, causing additional burden on memory caching and persistent storage; the system needs to load all version data when accessing the inode, even if only the current version is actually accessed, which will cause unnecessary memory overhead and reduce performance; any metadata modification needs to process the entire structure, making persistent operations complex and time-consuming; merging multiple version data into a single KV data record will cause the record volume to expand, affecting the read-write efficiency and maintainability of the key-value database. Therefore, the existing design is difficult to efficiently support system operation in scenarios with many snapshot versions and frequent access. SUMMARY

[0005] Therefore, it is necessary to provide a file metadata multi-version storage management method and device, equipment and a storage medium that can not only achieve decoupled storage of file metadata multi-versions, but also improve system memory utilization and snapshot version management efficiency.

[0006] In a first aspect, a management method of file metadata multi-version storage is provided, comprising:

[0007] obtaining file metadata;

[0008] in response to receiving a generation instruction, generating a snapshot version of the file metadata to obtain a snapshot version set;

[0009] extracting a first snapshot version and a second snapshot version from the snapshot version set, updating the snapshot version set, and storing the updated snapshot version set to a memory area;

[0010] packing the file metadata, the first snapshot version, and the second snapshot version to obtain an encoded data structure, and storing the encoded data structure to the memory area;

[0011] storing the updated snapshot version set and the encoded data structure to a storage device, respectively;

[0012] in response to receiving a query instruction, querying the memory area according to the query instruction, querying the storage device if the query fails, and returning a query result.

[0013] In a second aspect, a model inference performance optimization apparatus is provided, which is applied to the management method of file metadata multi-version storage in the first aspect, comprising:

[0014] a data obtaining module for obtaining file metadata;

[0015] a snapshot generation module for, in response to receiving a generation instruction, generating a snapshot version of the file metadata to obtain a snapshot version set;

[0016] a snapshot update module for extracting a first snapshot version and a second snapshot version from the snapshot version set, updating the snapshot version set, and storing the updated snapshot version set to a memory area;

[0017] a first storage module for packing the file metadata, the first snapshot version, and the second snapshot version to obtain an encoded data structure, and storing the encoded data structure to the memory area;

[0018] a second storage module for storing the updated snapshot version set and the encoded data structure to a storage device, respectively;

[0019] a query module for, in response to receiving a query instruction, querying the memory area according to the query instruction, querying the storage device if the query fails, and returning a query result.

[0020] On the third aspect, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the management method for multi-version storage of file metadata recorded in the first aspect is implemented.

[0021] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the management method for multi-version storage of file metadata described in the first aspect is implemented.

[0022] By implementing the aforementioned method, apparatus, device, and storage medium for managing multi-version file metadata, the method obtains file metadata information and, upon receiving a snapshot generation instruction, promptly generates corresponding snapshot versions and constructs a snapshot version set. This snapshot version set is further updated by extracting the first and last snapshot versions and storing them uniformly in memory, ensuring complete tracking and management of the snapshot history. Simultaneously, the file's current metadata and the extracted key snapshot versions are structured and encapsulated to generate an encoded data structure, which is then stored in memory. This enables efficient management of core metadata without requiring the full loading of all snapshot versions. Furthermore, the snapshot version set and the encoded data structure are stored separately on disk, effectively decoupling the data structures and reducing the persistence burden. During queries, the system prioritizes fast responses in the memory area, re-accessing the underlying storage device if a query misses, ensuring query efficiency while maintaining data integrity. Furthermore, the system dynamically manages the memory area based on access frequency and a least recently used (LRU) strategy, ensuring that frequently accessed metadata remains in memory while less frequently accessed data can be replaced or offloaded, effectively controlling memory usage. Overall, this method not only improves the flexibility and response performance of file metadata multi-version management, but also optimizes the resource utilization efficiency of the storage system. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0024] Figure 1 A flowchart of a method for managing multi-version storage of file metadata provided in an embodiment of the present application;

[0025] Figure 2 A structural block diagram of a file metadata multi-version storage management device provided in an embodiment of the present application;

[0026] Figure 3A time sequence diagram for managing a difference field of a file metadata multi-version storage is provided in an embodiment of the present application.

[0027] Figure 4 An internal structure diagram of a computer device in an embodiment of the present application is provided. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0029] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or system. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0030] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0031] In one embodiment, as shown in Figure 1 a file metadata multi-version storage management method is provided, comprising:

[0032] S100: obtaining file metadata;

[0033] S200: in response to receiving a generation instruction, generating a snapshot version of the file metadata to obtain a snapshot version set;

[0034] S300: extracting a first snapshot version and a second snapshot version from the snapshot version set, updating the snapshot version set, and storing the updated snapshot version set to a memory area;

[0035] S400: packaging the file metadata, the first snapshot version and the second snapshot version to obtain an encoded data structure, and storing the encoded data structure to the memory area;

[0036] S500: storing the updated snapshot version set and the encoded data structure to a storage device, respectively;

[0037] S600: In response to receiving the query instruction, querying the memory region according to the query instruction, querying the storage device if the query fails, and returning the query result.

[0038] Wherein, the file metadata is a collection of information describing the file attributes, usually including the file's permissions (such as read-write permissions), owner (UID), user group (GID), timestamp (creation time, access time, modification time), extended attributes (xattr), file size, inode number, etc.; the generation instruction refers to the command or event triggered by the system or user for creating a file snapshot version; the snapshot version refers to a complete copy of the file metadata at a certain point in time, used to record the historical state of the file. The collection of multiple snapshot versions is called a snapshot version set, which is used to preserve the state of the file at different historical moments; the first snapshot version refers to the earliest created snapshot in the snapshot version set, and the second snapshot version refers to the latest generated snapshot; the encoded data structure refers to a structured storage form formed by packaging and combining file metadata and key snapshot versions (such as the current version, the first snapshot, and the last snapshot) according to a certain format; the memory region is a memory space specially used to cache file metadata and snapshot data during the running of the file system; the storage device refers to a non-volatile medium used to persistently save file metadata and its snapshot versions, usually including hard disk (HDD), solid state disk (SSD), etc.; the query instruction is an information access instruction for requesting a certain file or its snapshot version metadata, which comes from user operation or system service; the access frequency represents the number of times the system accesses a certain file or its metadata within a period of time; the least recently used algorithm (LRU) is a commonly used cache eviction strategy. This algorithm considers that the data that has not been used recently is unlikely to be accessed in the future, so it is preferentially removed from memory to make room for new data.

[0039] Specifically, the system first acquires the metadata information of the target file as the basis for subsequent version management. When receiving a snapshot generation instruction, the system performs a snapshot operation on the current file metadata, generates a corresponding snapshot version, and builds a snapshot version set to comprehensively record the historical state information of the metadata. Subsequently, the earliest generated snapshot version (i.e., the first snapshot version) and the latest generated snapshot version (i.e., the second snapshot version) are extracted from the snapshot version set, and the snapshot set is simplified and updated, so that the system only needs to maintain key version nodes to realize ordered tracking of the history, thereby reducing the storage and computing burden. The updated snapshot version set is written to the memory area to support efficient access, and the current file metadata and the second snapshot version are packaged to form an encoded data structure, which is also cached in the memory for fast response to query requests. To ensure the safety and consistency of the data, the system separately stores the updated snapshot version set and the encoded data structure to the persistent device, so that the snapshot information and the data of the current version are logically independent and physically separated, reducing the coupling degree and improving the storage efficiency. When processing a query instruction, the system first searches for the required metadata information in the memory area to achieve high-speed access; if the query is not hit, the system further accesses the underlying storage device to obtain the required data, ensuring the integrity and correctness of the query. In addition, considering the limited memory resources, the system also introduces a dynamic management mechanism based on access frequency and least recently used (LRU) strategy, which periodically adjusts the retention and elimination strategy of the data in the memory area, so that high-frequency access data is retained in the memory, while low-frequency data is replaced or stored, thereby improving the query response speed while effectively controlling the memory occupancy.

[0040] In summary, the method optimizes the management of snapshot version structures and reasonably designs the memory caching mechanism, realizes efficient storage and fast access of multiple versions of file metadata, not only improves the running performance and response efficiency of the system, but also effectively reduces the consumption of memory and storage resources, and is particularly suitable for file system environments with frequent snapshot version generation and high concurrency access.

[0041] In one embodiment, in response to receiving a generation instruction, a snapshot version is generated for the file metadata, and a snapshot version set is obtained, including:

[0042] A file metadata structure body is copied to generate a snapshot version corresponding to the file metadata;

[0043] The first snapshot version generated for the file metadata structure body contains a full file metadata structure body, and a preset mark is set for the snapshot version;

[0044] The other snapshot versions generated after the first snapshot version only record the difference fields of the previous snapshot version;

[0045] According to the generated first snapshot version and other snapshot versions, a snapshot version set is obtained.

[0046] The structure refers to a set of ordered fields of file metadata in memory or a persistent layer for describing file attributes, including information such as permission information, a timestamp, an extended attribute, and the like, and constitutes a complete description of a file state.

[0047] Specifically, the system completely retains all field contents of the current file metadata in the generated first snapshot version, and marks the snapshot with a preset mark to identify it as the starting version of the snapshot sequence. The mark can be used as an anchor point for reconstructing the snapshot chain or restoring the metadata. In other snapshot versions generated after the first version, the system no longer redundantly saves complete structure data, but only records the fields that have changed since the last snapshot version. Such differential fields refer to data items that have changed between two adjacent snapshot versions, such as a modification time and access permissions, thereby avoiding unnecessary data duplication. By organizing and collecting these snapshots in chronological order, a snapshot version set can be formed to completely record the change history of file metadata and support subsequent query, rollback, or audit operations. Through the differential storage strategy, the required storage space for each snapshot version is greatly compressed, and especially in the scenario of frequent snapshot generation, the total data volume can be significantly reduced. By using the structure copy combined with differential field extraction, the traceability of the complete snapshot is retained, and the efficiency of snapshot generation and persistence is improved, thereby reducing the delay and write pressure of the system when performing snapshot operations.

[0048] In one embodiment, the method further comprises:

[0049] In response to receiving the generation instruction and reaching the preset interval number, a snapshot version containing the full file metadata structure is generated, and a preset mark is generated for the snapshot version.

[0050] The preset interval number refers to a numerical threshold value set in advance by the system according to performance, storage load, or business needs, for example, generating 1 full snapshot for every 10 differential snapshots. The preset mark is used to identify the version as a full snapshot.

[0051] Specifically, the system maintains a snapshot counter while continuously performing the snapshot version difference record; whenever the snapshot counter reaches the preset interval value, the system stops the difference extraction operation and instead re-replicates the current complete file metadata structure to generate a new full snapshot version. This process ensures that in a long snapshot chain, there will be a snapshot node that can be used independently and contains complete information at a certain distance, effectively preventing the problem of exponential growth of reconstruction overhead due to a long snapshot chain. When restoring a certain historical state, the system can preferentially locate the nearest full snapshot version as the starting point and only need to play back the subsequent difference snapshots, significantly reducing the calculation and I / O cost of reconstruction.

[0052] In one embodiment, extracting a first snapshot version and a second snapshot version from a snapshot version set comprises:

[0053] generating a corresponding timestamp for the snapshot version while generating the snapshot version for the file metadata;

[0054] traversing each timestamp in the snapshot version set and comparing the timestamps to determine a maximum timestamp and a minimum timestamp;

[0055] taking the snapshot version corresponding to the minimum timestamp as the first snapshot version;

[0056] taking the snapshot version corresponding to the maximum timestamp as the second snapshot version.

[0057] In the present scheme, the timestamp refers to the time label attached when each snapshot version is generated, usually expressed in system time accurate to seconds or milliseconds, used to identify the time point of snapshot generation.

[0058] Specifically, the system attaches a unique timestamp to each file metadata snapshot version when generating the snapshot, which identifies the time when the snapshot is generated. The timestamp is automatically generated by the system time and has the characteristics of global monotonic increase, ensuring the time sequence consistency of the snapshot versions. The system traverses the entire snapshot version set, sequentially extracts and compares the timestamp bound to each snapshot, and quickly determines the earliest snapshot version (i.e., the version with the smallest timestamp) and the latest snapshot version (i.e., the version with the largest timestamp) through maximum and minimum value calculation logic. The snapshot version corresponding to the minimum timestamp is regarded as the first snapshot version, representing the starting point of the snapshot history, while the snapshot version corresponding to the maximum timestamp is confirmed as the second snapshot version, representing the latest state record. By extracting these two boundary snapshot versions and managing them uniformly, the number of snapshot versions involved in coding can be greatly reduced, thereby reducing memory pressure. The snapshot boundary identification implemented based on the timestamp can also provide a good sorting basis and version evolution path for the system, which can be used in snapshot aging, cleaning and archiving management scenarios, enhancing the controllability and maintainability of the entire snapshot system.

[0059] In one embodiment, in response to receiving the query instruction, the memory region is queried according to the query instruction, the storage device is queried if the query fails, and the query result is returned. The query instruction includes at least one of the following: current file metadata query and snapshot version query, including:

[0060] In response to the instruction being a current file metadata query, it is determined whether the target current file metadata identifier exists in the encoding data structure in the memory region according to the target current file metadata identifier;

[0061] In response to the encoding data structure in the memory region existing, the current file metadata is returned;

[0062] In response to the encoding data structure in the memory region not existing, it is determined whether the target current file metadata identifier exists in the data packet in the storage device;

[0063] In response to the encoding data structure in the storage device existing, the current file metadata is returned;

[0064] In response to the encoding data structure in the storage device not existing, an error alarm is returned;

[0065] In response to the instruction being a snapshot version query, the target search range is determined according to the timestamp of the first snapshot version and the timestamp of the second snapshot version in the encoding data structure, and it is determined whether the target snapshot version identifier exists in the target search range;

[0066] In response to the target snapshot version identifier existing in the updated snapshot set in the memory region, the current snapshot version is returned;

[0067] In response to the absence of the updated snapshot set in the memory region, the updated snapshot set in the storage device is searched for the target snapshot version identifier;

[0068] In response to the presence of the updated snapshot set in the storage device, the current snapshot version is returned;

[0069] In response to the absence of the updated snapshot set in the storage device, an error alarm is returned.

[0070] The query instruction is an access request initiated by a user or a system for obtaining a certain metadata or snapshot version, mainly including two types of current file metadata query and snapshot version query. The current file metadata identifier is identification information for uniquely identifying a certain file metadata structure, usually including a file path, a file ID, or a file handle, etc., and is used for accurately locating the data content of a target file. The target snapshot version identifier is data identification for uniquely identifying a specific snapshot version, which can include a version number, a timestamp, or a version generation number, etc., and is used for locating a specific historical state.

[0071] Specifically, by judging the type of the query instruction, the corresponding data structure and range are selected for processing according to different query types. When the query instruction is a current file metadata query, the system will use the target file metadata identifier carried in the query to locate in the encoded data structure in the memory region. The encoded data structure usually contains the core field index of the current metadata and the first and second snapshot versions, so fast matching can be achieved. If the metadata corresponding to the identifier is found in the memory, the system directly returns the current file metadata, with fast response speed and low delay. If the metadata does not exist in the memory, the system automatically switches to the encoded data structure that has been persisted in the storage device to further match the target identifier. If the matching is successful in the device, the result is returned; if it is still not hit, an error alarm is triggered to prompt that the data does not exist or has been cleaned up.

[0072] When the query instruction is a snapshot version query, the system first determines the effective time window of the current version chain according to the first snapshot version timestamp and the second snapshot version timestamp recorded in the current encoded data structure, which is used as the target query range. Within this range, the system first searches the updated snapshot version set in the memory region for the target snapshot version identifier, and if it is hit, the target snapshot version is directly returned; if it is not hit in the memory, the corresponding snapshot version set in the storage device is searched again, and the result is returned when the target version is successfully located. If the target version does not exist in the snapshot version set, an error alarm is returned.

[0073] By taking the memory region as the primary target of the query, the response speed is effectively improved, meeting the low latency requirement under high-frequency access; at the same time, with the help of the structured coding data structure and the time stamp boundary limited query, the query range is greatly reduced, avoiding full traversal and improving the overall query efficiency. Secondly, in the case of query failure, the system has the ability to fall back to the storage device, thereby guaranteeing the data integrity and accessibility. Even if part of the data is temporarily unavailable due to memory eviction, it can be recovered from the lower layer storage. Through the introduction of the error alarm mechanism, the system can provide explicit feedback for the invalid identification behavior, which is beneficial for problem positioning and data integrity verification.

[0074] By constructing a double-layer storage query path, combining structured coding and snapshot time stamp constraints, efficient, accurate and rollbackable access control of file metadata and snapshot versions is realized, which not only improves the system query performance, but also enhances the stability and reliability.

[0075] In one embodiment, as shown in Figure 3 in response to the existence of the updated snapshot in the memory region, the current snapshot version is returned, including:

[0076] determining whether the current snapshot version exists a preset mark;

[0077] in response to the current snapshot version existing the preset mark, the current snapshot version is returned;

[0078] in response to the current snapshot version not existing the preset mark, the current snapshot version is recursively queried forward according to the current snapshot version, until a snapshot version with the preset mark is queried, the structure body of the queried snapshot version is sequentially covered to the structure body of the current snapshot version, the full-amount structure body of the current snapshot version is obtained, the current snapshot version is updated, and the updated current snapshot version is returned.

[0079] wherein, the preset mark refers to a special mark in the snapshot version for identifying "full-amount snapshot", which is used to distinguish the snapshot containing complete metadata structure body from the incremental snapshot only recording the difference field; in the present application, the recursive query refers to that when the system cannot directly restore the complete metadata from the current snapshot, the last snapshot is sequentially found forward according to the snapshot time sequence, and the difference is gradually applied until a full-amount snapshot with the preset mark is found.

[0080] Specifically, by judging whether the snapshot version is marked with a preset mark, i.e., whether it is a full snapshot. If the mark exists, it means that the current snapshot already contains a complete structure, and no further restoration processing is needed. The system can directly return the version as a query response, which is fast in response speed and simple in logic. If the current snapshot version does not contain the preset mark, the system enters a recursive restoration mechanism. At this time, the system takes the current snapshot version as the starting point and gradually searches for a snapshot version with an earlier timestamp, until the latest snapshot version with the preset mark is found. Each time a superior snapshot is found, the system combines the fields contained in the structure of the snapshot with the difference fields of the current snapshot version, and constantly complements the missing fields of the current snapshot version by means of successive covering or filling. Finally, the complemented version structure will have the same field completeness as the full snapshot, and constitute the "full structure" of the current snapshot version. This structure is then marked as an updated state and returned to the inquirer as a complete snapshot version.

[0081] The incremental structure of the snapshot is maintained at the storage level, greatly saving the overall storage space of the system, while the completeness of the returned data is ensured through the on-demand construction mechanism during the query phase, achieving a balance between performance and resource utilization. Secondly, the recursive restoration mechanism has high fault tolerance. Even if part of the difference snapshot is missing, as long as the latest complete snapshot exists on the chain, most of the data can be recovered, enhancing the stability of the system.

[0082] In one embodiment, the memory area is dynamically managed based on the access frequency of the memory area and the least recently used algorithm, including:

[0083] In response to the occupied space of the memory area reaching a preset capacity threshold, an eviction mechanism is triggered;

[0084] Based on the access frequency and the latest access time of each data item, an eviction weight is calculated for each data item, which is used to comprehensively reflect the lag degree of the access time and the sparsity degree of the access frequency;

[0085] The data items in the memory area are sorted in descending order of the eviction weight, and the data item with the highest weight is selected as the evicted object, wherein the data items include file metadata and each snapshot version in the snapshot set;

[0086] The version data corresponding to the data item is deleted to release the storage resources of the memory area.

[0087] The eviction weight is a comprehensive index calculated for each data item, reflecting its priority for eviction.

[0088] Specifically, when the actual occupied space of the memory region reaches the system preset capacity threshold, the mechanism automatically triggers the eviction process. At this time, the system traverses the data items currently residing in the memory region, including the current file metadata and multiple snapshot versions in the snapshot set, and calculates a corresponding eviction weight for each data item.

[0089] The calculation model of the eviction weight combines two core indicators: one is the access frequency of the data item in a unit of time, reflecting its hotness; the other is the last access timestamp, used to evaluate the degree of access lag. The lower the access frequency and the more distant the last access time, the lower the importance and activity of the data item, and therefore the higher the corresponding eviction weight. The system sorts all data items according to the weight values from high to low according to the calculation results, and selects a batch of data items with the highest weight as the eviction objects.

[0090] The eviction operation is carried out in a structured manner to ensure that data integrity is not damaged. Once a data item (such as a snapshot version or current file metadata) is selected, the system removes the item from the memory region and retains its corresponding persistent copy in the storage device, so that it can be reloaded from the disk when needed, ensuring that query integrity is not affected. With the completion of the eviction, the released memory space can be reused by newly loaded high-priority data, forming an efficient memory system that is resource adaptive and dynamically updated.

[0091] Through the dual-factor judgment of access hotness and time dimension, a more accurate and intelligent cache eviction than traditional LRU is realized, which maximizes the retention of high-frequency hot data and guarantees the response efficiency of the system. The eviction behavior occurs before the resource criticality, effectively avoiding the performance fluctuations caused by memory overflow or frequent GC, and improving the stability of the system. Since the eviction objects include not only temporary file metadata but also historical snapshot versions, the system can intelligently optimize the residence strategy of different data categories according to the business use scenarios, further improving memory utilization and processing throughput.

[0092] In one embodiment, the updated snapshot version set and the encoded data structure are stored in the storage device respectively, including:

[0093] Serializing the encoded data structure to obtain a binary data stream;

[0094] Writing the binary data stream into the storage device through an interface.

[0095] Wherein, serialization refers to the process of converting the encoded data structure (such as structure, object, dictionary, etc.) in memory into a format that can be stored or transmitted, which is a byte stream or binary data stream;

[0096] Specifically, the structure is serialized, i.e., its structured layout in memory is converted into a binary data stream in a standard format that can completely preserve the field content, data type, and hierarchical relationship of the original structure. The serialization process can be performed according to a predefined protocol (such as ProtocolBuffers, FlatBuffers, custom byte format, etc.) to ensure a balance between data consistency and transmission efficiency. After serialization is completed, the system writes the binary data stream into the designated storage device through the underlying interface call to form a persistent data packet that can be quickly located and decoded. The storage process usually combines metadata identification, timestamps, version numbers, and other information to provide index support for subsequent queries. Serialization compresses complex data structures into compact binary streams, significantly reducing storage space occupation, while improving write speed and bandwidth utilization efficiency. Through a unified serialization standard, the system can implement cross-module transmission or persistent restoration of metadata between different modules and different nodes, enhancing the portability and module decoupling ability of the system.

[0097] In one embodiment, in response to the absence of the encoded data structure in the memory region, the system searches for the target current file metadata identifier in the encoded data structure in the storage device, including:

[0098] In response to detecting the encoded data structure matching the target current file metadata identifier in the data packet in the storage device, the system loads the encoded data structure;

[0099] The encoded data structure is written into the memory region and a decoding operation is performed thereon;

[0100] The decoding operation extracts the current file metadata, the first snapshot version, and the tail snapshot version from the encoded data structure and returns the current file metadata.

[0101] Specifically, when the target current file metadata identifier is not hit in memory, the system queries the data packet set in the storage device. By comparing the file identifier field carried by the encoded data structure in the storage structure, the system locates whether there is a structure matching the target identifier. Once a match is detected, the system immediately loads the encoded data structure and writes it into the memory region for subsequent use. Subsequently, the system performs a decoding operation on the structure to disassemble its contents into three parts of the current file metadata, the first snapshot version, and the second snapshot version. Through the decoding operation, the system not only restores the use state of the current metadata, but also synchronously loads the boundary data at both ends of the snapshot chain to provide support for subsequent snapshot range queries. The system returns the current file metadata as the query result to the upper layer business. Even if the memory data is invalidated or cleaned up, the system can quickly recover the target structure from the storage device, ensuring business continuity and query integrity.

[0102] In one embodiment, as shown in Figure 2 A model inference performance optimization apparatus is provided, comprising: a data acquisition module 710, a snapshot generation module 720, a snapshot update module 730, a first storage module 740, a second storage module 750, and a query module 760, configured to:

[0103] The data acquisition module 710 is configured to acquire file metadata.

[0104] The snapshot generation module 720 is configured to, in response to receiving a generation instruction, generate snapshot versions for the file metadata to obtain a snapshot version set.

[0105] The snapshot update module 730 is configured to extract a first snapshot version and a second snapshot version from the snapshot version set, update the snapshot version set, and store the updated snapshot version set in a memory area.

[0106] The first storage module 740 is configured to package the file metadata, the first snapshot version, and the second snapshot version to obtain an encoded data structure, and store the encoded data structure in the memory area.

[0107] The second storage module 750 is configured to store the updated snapshot version set and the encoded data structure in a storage device, respectively.

[0108] The query module 760 is configured to, in response to receiving a query instruction, query the memory area according to the query instruction, query the storage device if the query fails, and return a query result.

[0109] In one embodiment, the snapshot generation module 720 is configured to:

[0110] Copy a structure of the file metadata to generate a snapshot version corresponding to the file metadata.

[0111] Include full file metadata structures in a first snapshot version generated for the file metadata and generate a preset marker for the snapshot version.

[0112] The snapshot versions generated after the first snapshot version only record difference fields from a previous snapshot version.

[0113] Obtain the snapshot version set according to the generated snapshot versions.

[0114] In one embodiment, the query module 760 is configured to

[0115] In response to an instruction being a current file metadata query, search for whether a target current file metadata identifier exists in the encoded data structure in the memory area according to the target current file metadata identifier.

[0116] in response to the existence in the encoding data structure in the memory region, returning the current file metadata;

[0117] in response to the non-existence in the encoding data structure in the memory region, searching the encoding data structure in the storage device for whether the target current file metadata identifier exists;

[0118] in response to the existence in the encoding data structure in the storage device, returning the current file metadata;

[0119] in response to the non-existence in the encoding data structure in the storage device, returning an error alarm;

[0120] in response to the instruction being a snapshot version query, determining a target search range according to a timestamp of a first snapshot version and a timestamp of a second snapshot version in the encoding data structure, and searching the target search range for whether the target snapshot version identifier exists;

[0121] in response to the existence in the updated snapshot set in the memory region, returning the current snapshot version;

[0122] in response to the non-existence in the updated snapshot set in the memory region, searching the updated snapshot set in the storage device for whether the target snapshot version identifier exists;

[0123] in response to the existence in the updated snapshot set in the storage device, returning the current snapshot version;

[0124] in response to the non-existence in the updated snapshot set in the storage device, returning an error alarm.

[0125] It should be understood that, although Figure 2 the steps in the device structure block diagram are displayed in sequence according to the arrows, these steps are not necessarily executed in sequence according to the arrows. Unless explicitly stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, Figure 2 at least a part of the steps in the device structure block diagram can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or at least a part of sub-steps or stages of other steps.

[0126] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above-mentioned file metadata multi-version storage management method embodiments when running.

[0127] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0128] Embodiments of the present application also provide a computer program product including a computer program, which, when executed by a processor, implements the steps in any of the embodiments of the management method for multi-version storage of file metadata.

[0129] Those skilled in the art will further appreciate that the functions of the examples described herein can be implemented using electronic hardware, computer software, or any combination thereof. Whether such functions are implemented in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application. Figure 4 As shown, in order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0130] The above describes in detail the management method for multi-version storage of file metadata provided by the present application. The principles and implementation methods of the present application are described by applying specific examples. The above description of the examples is only to help understand the method of the present application and its core idea. It should be noted that for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for managing multi-version storage of file metadata, characterized in that: include: Get file metadata; In response to receiving the generation instruction, generating a snapshot version for the file metadata to obtain a snapshot version set; Extracting a first snapshot version and a second snapshot version from the snapshot version set, updating the snapshot version set, and storing the updated snapshot version set in a memory area; Packing the file metadata, the first snapshot version, and the second snapshot version to obtain a coded data structure, and storing the coded data structure in the memory area; storing the updated snapshot version set and the encoding data structure in a storage device respectively; In response to receiving the query instruction, the memory area is queried according to the query instruction. If the query fails, the storage device is queried and the query result is returned.

2. A method for managing multi-version storage of file metadata according to claim 1, characterized in that: In response to receiving the generation instruction, generating a snapshot version for the file metadata to obtain a snapshot version set includes: Copying the file metadata structure to generate a snapshot version corresponding to the file metadata; The first snapshot version generated for the file metadata structure includes a full file metadata structure and a preset mark is set for the snapshot version; Other snapshot versions generated after the first snapshot version only record the difference fields with the previous snapshot version; The snapshot version set is obtained according to the generated first snapshot version and the other snapshot versions.

3. A method for managing multi-version storage of file metadata according to claim 2, characterized in that: The method further comprises: In response to receiving the generation instruction and reaching the preset interval number, a snapshot version including the full file metadata structure is generated and the preset mark is generated for the snapshot version.

4. A method for managing multi-version storage of file metadata according to claim 2, characterized in that: The extracting the first snapshot version and the second snapshot version from the snapshot version set includes: While generating a snapshot version of the file metadata, generating a corresponding timestamp for the snapshot version; Traversing each of the timestamps in the snapshot version set, and comparing the timestamps to determine a maximum timestamp and a minimum timestamp; Using the snapshot version corresponding to the minimum timestamp as the first snapshot version; The snapshot version corresponding to the maximum timestamp is used as the second snapshot version.

5. A method for managing multi-version storage of file metadata according to claim 4, characterized in that: In response to receiving a query instruction, querying the memory area according to the query instruction, querying the storage device if the query fails, and returning a query result, wherein the query instruction includes at least one of the following: current file metadata query and snapshot version query, including: In response to the instruction being a current file metadata query, searching the encoded data structure in the memory area for a target current file metadata identifier based on the target current file metadata identifier; In response to the existence of the encoded data structure in the memory area, returning the current file metadata; In response to the target current file metadata identifier not existing in the encoded data structure in the memory area, searching the encoded data structure in the storage device for the target current file metadata identifier; In response to the existence of the encoded data structure in the storage device, returning the current file metadata; In response to the encoding data structure not existing in the storage device, returning an error warning; In response to the instruction being a snapshot version query, determining a target search range according to the timestamp of the first snapshot version and the timestamp of the second snapshot version in the encoded data structure, and searching within the target search range for the target snapshot version identifier; In response to the updated snapshot set existing in the memory area, returning the current snapshot version; In response to the updated snapshot set not existing in the memory area, searching the updated snapshot set in the storage device for a target snapshot version identifier; In response to the updated snapshot set existing in the storage device, returning the current snapshot version; In response to the updated snapshot set not existing in the storage device, an error alarm is returned.

6. A method for managing multi-version storage of file metadata according to claim 5, characterized in that: In response to the updated snapshot set existing in the memory area, returning the current snapshot version includes: Determine whether the current snapshot version has the preset mark; In response to the presence of the preset mark in the current snapshot version, returning the current snapshot version; In response to the fact that the preset mark does not exist in the current snapshot version, a recursive query is performed forward based on the current snapshot version until the snapshot version with the preset mark is queried, the structure of the queried snapshot version is overwritten with the structure of the current snapshot version in sequence to obtain the full structure of the current snapshot version, the current snapshot version is updated, and the updated current snapshot version is returned.

7. A method for managing multi-version storage of file metadata according to claim 1, characterized in that: The method further includes dynamically managing the memory area based on the access frequency of the memory area and a least recently used algorithm, including: In response to the occupied space of the memory area reaching a preset capacity threshold, triggering an elimination mechanism; Based on the access frequency and the most recent access time of each data item, an elimination weight is calculated for each data item, wherein the elimination weight is used to comprehensively reflect the lag degree of access time and the sparseness of access frequency; Sort the data items in the memory area according to the elimination weights from high to low, and select the data item with the highest weight as the elimination object, wherein the data items include: the file metadata and each of the snapshot versions in the snapshot set; The version data corresponding to the data item is deleted to release storage resources of the memory area.

8. A management device for multi-version storage of file metadata, characterized in that: The device comprises: Data acquisition module, used to obtain file metadata; A snapshot generation module, configured to generate a snapshot version of the file metadata in response to receiving a generation instruction, to obtain a snapshot version set; a snapshot update module, configured to extract the first snapshot version and the second snapshot version from a snapshot version set, update the snapshot version set, and store the updated snapshot version set in a memory area; a first storage module, configured to package the file metadata, the first snapshot version, and the second snapshot version to obtain a coded data structure, and store the coded data structure in a memory area; A second storage module, configured to store the updated snapshot version set and the encoding data structure in a storage device respectively; The query module is used to query the memory area according to the query instruction in response to receiving the query instruction, query the storage device if the query fails, and return the query result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Metadata version control method

    CN121979865A

  • A metadata versioning method

    CN121979865B