Distributed parallel file system metadata management method and system capable of being perceived by NUMA (Non Uniform Memory Access)
By using a hierarchical metadata management structure and CXL-SSD's memory semantic operations, the problems of high latency in cross-node access and low persistence efficiency of storage devices in distributed parallel file systems are solved, achieving high-performance, low-latency metadata management that adapts to dynamic load changes.
Patent Information
- Application Number
- CN202510894866.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-07
AI Technical Summary
Distributed parallel file systems suffer from high latency in cross-node access and low persistence efficiency of traditional storage devices, which limits system performance, scalability, and dynamic load adaptability.
A hierarchical metadata management structure is adopted, including a local metadata view and a global metadata view. Memory semantic operations are implemented through CXL-SSD. Combined with metadata partitioning strategy and consistency guarantee mechanism, metadata is dynamically allocated to NUMA nodes. The high write throughput of CXL-SSD and MESI cache consistency protocol are used to optimize cross-node access.
It significantly reduces the overhead of remote access across NUMA nodes, improves the performance and scalability of metadata management, achieves low-latency persistence and efficient consistency, and adapts to dynamic load changes.
Smart Images

Figure CN120910012A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of distributed parallel file system data management, and in particular to a NUMA-aware distributed parallel file system metadata management method and system. BACKGROUND
[0002] With the rapid development of distributed systems and big data applications, metadata management has become one of the core bottlenecks of system performance. In a distributed parallel file system, based on NUMA (Non-Uniform Memory Access), the storage and computing resources are deployed locally according to the NUMA node to achieve high bandwidth and low latency access. However, the memory access delay varies significantly due to the node location difference, and the traditional metadata management scheme usually adopts a centralized or static distribution strategy, which is difficult to meet the performance requirements in high concurrency and dynamic load scenarios. Specifically, the current metadata management in the distributed parallel file system has the following problems: 1. High latency of cross-node access: In the traditional NUMA architecture, remote access of metadata needs to pass through the cross-node interconnection channel (such as Intel UPI or AMD Infinity Fabric), and its delay is usually 2-5 times that of local memory access. When metadata requests involve multiple nodes (such as cross-node directory traversal or file movement), frequent remote access leads to a sharp increase in operation delay, becoming the main bottleneck of the overall throughput of the system.
[0003] 2. Inefficient persistence of traditional storage devices: Existing solutions rely on traditional SSDs to persist metadata through PCIe / NVMe protocol stacks, which requires multiple kernel / user mode switches, DMA transmission, and block device driver processing, resulting in additional protocol stack overhead (usually increasing the delay by 5-10us). In addition, traditional SSDs lack memory semantic access capabilities and are difficult to combine with the locality optimization depth of the NUMA architecture.
[0004] In summary, the current metadata management in the distributed parallel file system has high cross-node access delay and low persistence efficiency of storage devices, which limits the system performance, scalability, and dynamic load adaptability. Therefore, there is an urgent need to provide a new metadata management scheme and system to reduce the remote access overhead across NUMA nodes and improve performance, scalability, and dynamic adaptability. SUMMARY
[0005] The technical problem to be solved by the present application: In view of the above problems of the prior art, a NUMA-aware distributed parallel file system metadata management method and system are provided to reduce the remote access overhead across NUMA nodes and improve performance, scalability, and dynamic adaptability, providing efficient metadata management support for distributed file systems.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A NUMA-aware distributed parallel file system metadata management method, comprising the following steps: A hierarchical metadata management structure is constructed, comprising a local metadata view and a global metadata view. The local metadata view, serving as the local view layer for NUMA local nodes, dynamically allocates metadata to local NUMA nodes according to access modes to partition the metadata by NUMA node, and maintains an independent metadata copy on each local NUMA node. When the processor performs I / O operations on peripherals, it intercepts metadata requests and converts the peripheral access semantics from block device operations to memory semantic operations, thereby directly mapping the CXL-SSD storage space to the NUMA local address space through user-space memory mapping. The global metadata view, serving as the global view layer, stores metadata in a hierarchical manner by metadata key through a distributed global LSM tree based on CXL-SSD, providing consistency and persistence support for cross-NUMA node operations on metadata. A metadata partitioning strategy is adopted. The local metadata view dynamically distributes metadata to each NUMA node according to the access mode. For frequently accessed directory subtrees, the creation of directories and files of these directory subtrees is completed on the local NUMA node according to the local priority rule. For other metadata, the consistent hashing algorithm is used to dynamically distribute the metadata under the same directory so that the metadata belongs to the same NUMA node.
[0007] Furthermore, for metadata distributed across multiple local NUMA nodes with primary replicas, when metadata operations cause changes to the local metadata view, a consistency guarantee operation is triggered. The specific steps are as follows: Local metadata updates ensure transactionality through atomic multi-block write metadata operations; A global version number is maintained for metadata items accessed across nodes. When one node modifies its local metadata, the version number is incremented and notified to other replica nodes via CXL's cache consistency broadcast or global metadata view update. When other nodes read remote metadata, they verify that the version number matches the global metadata view. Figure One If there is any inconsistency, a metadata synchronization process will be triggered to ensure consistency.
[0008] Furthermore, the method also includes a dynamic migration strategy for hot metadata. The dynamic migration strategy for hot metadata includes: during the operation of each metadata, determining local hot metadata based on the access frequency of the metadata, and selecting metadata accessed by more than a preset proportion of NUMA nodes from all local hot metadata as global hot metadata, triggering a global hot migration operation, copying the global hot metadata to the NUMA node with the highest access frequency, or migrating the main copy of metadata that needs to be frequently modified to the NUMA node that frequently accesses the metadata, and updating the global metadata view and local metadata view accordingly after the global hot migration operation is completed.
[0009] Furthermore, local hotspot metadata is determined based on the access frequency of metadata, specifically including: Each NUMA node counts the access frequency of metadata in real time and stores the access frequency in the local metadata view as the metadata access counter Access Count. The revenue metrics of metadata are calculated based on the access frequency and the cumulative average access frequency. Based on the revenue metrics of metadata, a metadata hotspot score is calculated using a multi-armed slot machine exploration strategy. Metadata with scores higher than a preset value is designated as local hotspot metadata.
[0010] Furthermore, the expression for calculating the revenue metrics of metadata is as follows:
[0011] In the above formula, This is metadata, where t is the time period. and These represent the access frequency and cumulative average access frequency of metadata within the time period t, respectively. The expression for calculating the metadata hotspot score is:
[0012] In the above formula, Metadata at time t The revenue, metadata at time t The upper limit, where γ is an adjustable parameter. For metadata The number of times it was selected.
[0013] Further, if the metadata is marked as global hot metadata and the global hot migration is triggered, and the form of hot migration is to move the global hot metadata to the NUMA node with the highest access frequency, then multiple volumes are divided in the CXL-SSD, multiple copies are maintained, and after the hot metadata is copied to the primary volume, the data is backed up to the copy volume through a background asynchronous operation; if the form of hot migration is to copy the global hot metadata to the NUMA node with the highest access frequency, then a locking operation is performed during the writing process, the original node needs to synchronize the uncommitted modification log to the new primary copy, and the ownership mapping table is updated through the global view, and then the specific metadata is migrated, and the primary copy of the metadata is migrated to the NUMA node with the highest access frequency.
[0014] Further, the method further comprises optimizing cross-node operations through batch transactions and lock-free reading, combining batch metadata operations into a single transaction submission, utilizing the multi-stream writing of the CXL-SSD to process multiple operation streams in parallel, and allowing lock-free access to the read operation through the MESI cache consistency protocol of CXL, so that the node directly reads the cache copy of the remote CXL-SSD, and only accesses the persistent storage layer when the cache is invalid. All metadata update operations are written in the form of logs to a continuous area in memory, and when batch metadata operations are combined, similar operations are combined through batch transaction combination, and a submission is made when the current operation frequency is greater than the historical frequency or the cumulative number of operations is greater than the preset number, wherein a fast comparison strategy based on hash is used to compare the similarity of operations with a similarity greater than a preset percentage.
[0015] Further, the write operation of the global metadata view is appended to the log layer of the LSM tree, and the log data is periodically merged to the higher level; the merge operation writes the entries in order to the first level of the LSM tree; the query operation quickly determines whether a certain metadata exists in a specific level by maintaining a global Bloom filter in DRAM.
[0016] Further, the consistency guarantee operation further comprises: if multiple NUMA nodes concurrently modify the same metadata, then the global metadata view uses the last write wins principle to take the operation corresponding to the latest timestamp that modifies the metadata as the final operation.
[0017] A NUMA-aware distributed parallel file system metadata management system, the system is based on a CXL-SSD architecture, and the system comprises: A local metadata view, which is used to dynamically allocate metadata to local NUMA nodes according to access patterns to partition the metadata according to NUMA nodes. a global metadata view for providing consistency and persistence support for cross- NUMA node operations of metadata; The entry in the local metadata view comprises a key Key, a value Value, a latest version number Global Version in the global metadata view, a local modification version number Local Version, an access counter Access Count, a last access timestamp Last Access Time, state flag bits State Flags, a pointer to the corresponding physical address Pointer to Global in the global metadata view, a pointer to a log area Log Pointer, and a checksum Checksum. The entry in the global metadata view comprises a key Key, a value Value, a global version number Version, a NUMA node where a primary copy of metadata is located Location, a metadata replica distribution bitmask Replica Mask, and state flag bits State Flags.
[0018] Compared with the prior art, the application has the following advantages: By adopting the hierarchical metadata management structure, the local metadata view dynamically allocates metadata to each NUMA node according to an access mode and maintains an independent metadata copy in each NUMA node, and by adopting a metadata partitioning strategy, the metadata can be partitioned according to NUMA nodes, so that each metadata I / O operation can be performed at the performance of local memory access, reducing the overhead of remote access across NUMA nodes; by intercepting metadata requests and converting the access semantics of peripherals from block device operations to memory semantic operations, the implicit overhead of cross-node access can be eliminated by using the memory semantic access of CXL-SSD; by the global metadata view based on the distributed global LSM tree of CXL-SSD for hierarchical storage according to metadata keys, consistency and persistence support can be provided for cross- NUMA node operations of metadata, ensuring the atomicity of cross-node metadata updates, while the high write throughput characteristics of CXL-SSD and the MESI cache coherence protocol are used to optimize the cross-node access efficiency and realize low-latency persistence. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 The figure is a schematic diagram of the architecture of the NUMA-aware distributed parallel file system metadata management system in the embodiment of the application.
[0020] Figure 2 The figure is the entry format of the local metadata view in the specific application embodiment.
[0021] Figure 3 The figure is the entry format of the global metadata view in the specific application embodiment. Detailed Implementation
[0022] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0023] The following is a brief explanation of some of the terminology used in this invention and the relevant technical background.
[0024] NUMA is a memory design architecture for multiprocessor systems designed to address memory access bottlenecks in traditional symmetric multiprocessing (SMP) systems. In a NUMA architecture, each processor node (CPU group) has local memory and can access remote memory on other nodes. Local memory access latency is significantly lower than remote memory latency (typically 2-4 times lower), and remote access requires inter-node interconnects (such as Intel's QPI and AMD's Infinity Fabric), making bandwidth and latency key limiting factors for system scalability.
[0025] Metadata includes information such as file attributes, directory structure, permissions, and data distribution. Its access efficiency directly affects the scalability and throughput of the file system.
[0026] In existing technologies, NUMA-based solutions mainly include the following categories: ① Database optimization based on NUMA architecture (GaussDB) Huawei GaussDB uses NUMA-Aware technology to achieve data partitioning and load balancing, assigning global variables to specific CPUs for processing and reducing cross-core memory access conflicts. Its GTM-Lite technology decouples transaction management from global timestamps, synchronizing global timestamps only during cross-shard access, reducing distributed transaction overhead.
[0027] Limitations: Primarily designed for database transaction processing, it does not deeply optimize dynamic hotspot migration for metadata management; it relies on traditional persistent storage and does not fully utilize the memory semantics of CXL-SSD.
[0028] ②NUMA-aware connection table management This scheme pre-allocates memory resources for NUMA nodes and maps the join table to local nodes using 5-tuple hashes, reducing remote access. It manages join table consistency using version numbers (age field) and status flags, and reduces the overhead of frequent memory allocations through asynchronous synchronization mechanisms.
[0029] Limitations: The design targets network device connection management and does not involve metadata persistence and global consistency guarantee; lacks adaptation to CXL-SSD characteristics, unable to optimize hierarchical storage of metadata.
[0030] ③Kubernetes Topology Manager Kubernetes' Topology Manager coordinates CPU, device manager, and other resource allocation by collecting topology information (TopologyHint) of NUMA nodes, achieving NUMA alignment of container resources. Its strategies include "best match" and "strict match", prioritizing resource allocation to the same NUMA node to optimize latency-sensitive application performance.
[0031] Limitations: Focuses on resource scheduling and lacks dedicated mechanisms for metadata management; lacks support for persistent storage (such as CXL-SSD), unable to solve metadata persistence efficiency problems.
[0032] ④Technical metadata solutions in metadata management In data warehouse scenarios, technical metadata (such as table structure, field definition) is stored through Hive metadata database (such as MySQL), and combined with automated scripts to implement data consistency verification. For example, by comparing field values using MD5 hashing, quickly verifying the consistency of new and old table data.
[0033] Limitations: Relies on centralized metadata database, high cross-node access latency; static partitioning strategy cannot adapt to dynamic load changes, hot data can easily cause performance bottlenecks.
[0034] In addition, although NUMA architecture can effectively solve memory access bottlenecks, its architecture characteristics will introduce potential performance overhead at every step of I / O operations (such as data read and write of storage devices), as follows: When an I / O operation needs to transfer data from a storage device (such as an SSD) to memory, the operating system will allocate a DMA buffer. If the buffer is allocated in the memory of a remote node, the data needs to be transmitted through the inter-node interconnection channel, adding additional latency. For example, assume that node A's CPU initiates a read request for an NVMe SSD, but the DMA buffer is located in node B's memory. At this time, data needs to be transmitted from the SSD → node A's PCIe bus → node B's memory → node A's CPU. The cross-node transmission delay can be as high as hundreds of nanoseconds, much higher than local memory access (about 100 ns).
[0035] Traditional I / O interrupts can be handled by any CPU core. If the interrupt handling core is inconsistent with the NUMA node where the data resides, it can lead to cache coherency issues and overhead from remote memory access. For example, after an NVMe SSD completes an I / O operation, it may trigger an interrupt on node B, but the data processing task is running on node A. In this case, node A needs to remotely access the memory of node B to retrieve the data, resulting in performance degradation.
[0036] NUMA-based cache coherency protocols (such as MESI) need to maintain cross-node cache consistency. Frequent I / O operations can cause cache lines to migrate frequently between nodes, increasing protocol overhead. For example, if multiple nodes concurrently modify the same I / O buffer, cache lines will be constantly passed between nodes, causing "cache thrashing" and significantly reducing effective bandwidth.
[0037] If the operating system or application is unaware of the NUMA topology, it may centrally allocate I / O queues and device driver resources (such as NVMe queue pairs) to a single node, leading to resource contention. For example, if the I / O queues of all NVMe devices are bound to node 0, while the computing tasks are distributed across nodes 1-3, node 0 becomes the I / O bottleneck, while other nodes remain idle.
[0038] Furthermore, in distributed file systems, the I / O operations for metadata management (such as file path resolution and permission verification) have the following characteristics, which further amplify the impact of NUMA architecture overhead: High concurrency and fine-grained operations: Metadata operations (such as file creation and directory traversal) usually involve a large number of fine-grained requests, and frequent cross-node memory accesses (such as metadata caching and lock management) can accumulate significant latency.
[0039] Centralized access to hot data: Frequently accessed metadata (such as shared directories) may be concentrated on a few NUMA nodes, leading to congestion of local memory bandwidth and interconnect channels, thus creating a performance bottleneck.
[0040] Persistence and consistency overhead: Metadata persistence (such as log writing) must be done through storage devices. If the persistence path involves cross-node memory access (such as a log buffer on a remote node), it will further increase I / O latency.
[0041] In summary, existing distributed parallel file system metadata management technologies have the following problems: Low efficiency of cross-node access: Traditional solutions do not fully combine NUMA locality with the low latency characteristics of CXL-SSD; Poor adaptability to dynamic load: Static partitioning and centralized management are unable to cope with sudden hotspot access. Inadequate balance between persistence and consistency: traditional SSD protocol stack has large overhead and lacks efficient global consistency mechanism.
[0042] In view of the problems of high access delay across NUMA nodes and low efficiency of traditional storage persistence in a distributed parallel file system, the layered metadata management structure is constructed, the metadata of a plurality of local metadata views is dynamically distributed to each NUMA node according to an access mode, and an independent metadata copy is maintained in each NUMA node, the memory semantic access of the CXL-SSD is used to take over the traditional I / O stack, the metadata operation is converted into a local memory instruction, and the operation delay is significantly reduced (for example, from the microsecond level to the sub-microsecond level); the global metadata view is used to provide consistency and persistence support for the cross- NUMA node operation of the metadata, and the cache consistency of the CXL-SSD is used to realize the lock-free read operation and atomic persistence.
[0043] The NUMA-aware distributed parallel file system metadata management method of the embodiment of the application comprises the following steps: The layered metadata management structure comprises a local metadata view and a global metadata view, the local metadata view is a local view layer of a local NUMA node, is used for dynamically distributing metadata to the local NUMA node according to an access mode to partition the metadata according to the NUMA node, and maintains an independent metadata copy in each local NUMA node; when a processor performs an I / O operation on an external device, the metadata request is intercepted, and the access semantic of the external device is converted from a block device operation into a memory semantic operation, so as to directly map the CXL-SSD storage space to the local address space of the NUMA through the user-mode memory mapping; the global metadata view is a global view layer, is used for storing metadata keys in a layered manner through a distributed global LSM tree based on the CXL-SSD, and provides consistency and persistence support for the cross- NUMA node operation of the metadata; The metadata partition strategy is adopted, the local metadata view dynamically distributes metadata to each NUMA node according to an access mode through the metadata partition strategy, for a high-frequency access directory subtree, the creation of the directories and files of the directory subtree is completed in the local NUMA node through a local priority rule; and for other metadata, the same metadata under the same directory is attributed to the same NUMA node through a consistent hashing algorithm.
[0044] It can be understood that the embodiment adopts a hierarchical metadata management structure, dynamically allocates metadata to each NUMA node through a local metadata view according to an access mode, and maintains an independent metadata copy at each NUMA node, and adopts a metadata partitioning strategy, which can partition metadata according to NUMA nodes, so that each metadata I / O operation can be performed at the performance of local memory access, reducing the overhead of remote access across NUMA nodes; by intercepting metadata requests and converting the access semantics of peripherals from block device operations to memory semantic operations, the implicit overhead of cross-node access can be eliminated by using the memory semantic access of CXL-SSD; through the global metadata view of the distributed global LSM tree based on CXL-SSD, the metadata key is layered stored, which can provide consistency and persistence support for cross- NUMA node operations of metadata, ensure the atomicity of cross-node metadata update, and at the same time, use the high write throughput characteristics of CXL-SSD and the MESI cache coherence protocol to optimize the cross-node access efficiency and realize low-latency persistence. The local metadata view and the global metadata view can realize high-performance and low-latency metadata storage, and the performance, reliability and dynamic load adaptability are better, which can provide efficient support for high-performance computing, AI training and other scenarios.
[0045] In a specific application embodiment, as shown in Figure 1 FIG. 1 is a schematic diagram of a NUMA-aware distributed parallel file system metadata management system, which is specifically implemented based on CXL-SSD (Compute Express Link SSD, a high-performance solid state drive using CXL protocol), and the following describes a local metadata view thereof.
[0046] In a NUMA architecture, the delay of cross-node memory access is usually several times that of local access, and the metadata operations (such as file creation and directory traversal) of a distributed file system usually present the characteristics of high concurrency and fine granularity. In order to reduce the overhead of cross-node access, the core of the design of the local metadata view is to partition metadata according to the locality of the NUMA node, and maintain an independent metadata copy at each node, so that each metadata I / O operation can be performed at the performance of local memory access. In order to reduce the overhead of remote access across NUMA nodes, the core is to dynamically allocate metadata to the local node according to the access mode, which mainly includes the following contents: Take over traditional I / O stack: when the processor performs I / O operation on the peripheral, the local metadata view will take over the traditional I / O stack, and unify the potential cross-node memory access into local access. In the traditional distributed file system, metadata operation needs to interact with the peripheral through the kernel mode I / O stack, which leads to frequent kernel-user state switching and potential cross- NUMA node memory access. The embodiment takes over the core path of the traditional I / O stack through the local metadata view layer, converts the access semantics of the peripheral from block device operation to memory semantic operation, and thus eliminates the implicit overhead of cross-node access. For example, in the file creation operation, the metadata update (such as allocating inode, updating directory entry) can be directly submitted to the mapped region through the atomic memory write instruction, instead of delivering the I / O request through the cross-node NVMe queue. To achieve this optimization, the system introduces an address redirection layer, dynamically intercepts the metadata request of the traditional I / O stack, and according to the NUMA node distribution rule, remaps the logical address to the local physical memory region, and finally ensures that all potential peripheral accesses converge to the local operation of the node. It can be understood that by mapping the CXL-SSD storage space to the NUMA local address space through the user mode memory mapping (the transparent mapping of logical address to local storage is realized through the hardware address redirection layer ARL), the metadata operation is completed through the atomic memory instruction, which can enable the local metadata view layer of the NUMA local node to directly access the CXL-SSD through the memory mapping, bypass the traditional block device protocol stack, and eliminate the overhead of the traditional block device protocol stack.
[0047] In addition, for different metadata processing methods, a hybrid algorithm is adopted by using a metadata partitioning strategy, combining consistent hashing and directory tree level division. For high-frequency access directory subtrees, through local priority rules (i.e., for these hotspots, directory and file creation prefer to be completed on the local server or local NUMA node), they are allocated to a specific NUMA node, ensuring that metadata operations (such as file creation, attribute update) under the directory are completed locally. For other metadata, they are dynamically distributed through a consistent hashing algorithm, and the hash key is generated based on the file path, ensuring that the metadata under the same directory belongs to the same node as much as possible. The address space is divided into fixed-size slots, each of which is used to store metadata related to directories or files, and all slots are organized through a simple and efficient Hash Table data structure, where the Key is the ID of the directory or file, and the Value is the address where the metadata is stored. For example, the hash value of the path / user / alice / docs is mapped to Node2, and all file metadata under it is stored in the local metadata view of Node 2. It can be understood that the hybrid partitioning strategy combining consistent hashing and directory tree level division fixes the high-frequency directory subtrees to the specified node, and the other metadata is dynamically distributed according to the hash value. The above dynamic allocation of directories and files to NUMA nodes according to the directory tree level and hash value can maximize the proportion of local operations (such as reducing the remote access overhead of cross- NUMA nodes by about 80% or more), thereby improving the throughput.
[0048] Local Metadata View Entry Format: The entry of the local metadata view needs to support low-latency access and dynamic optimization, as shown in the following table: Figure 2 Key-Value: Key-value pair, aligned with the global metadata view, accelerating local queries.
[0049] Global / Local Version: The latest version number in the global metadata view / local modification version number, used to implement the "Read Validation" mechanism, returning data only when the Global Version is consistent with the version in the global metadata view, otherwise triggering synchronization.
[0050] Access Count & Last Access Time: Access counter and last access timestamp, used to drive hotspot migration strategies (such as marking as a hotspot when the access count exceeds the threshold).
[0051] Pointer to Global: Pointer to the corresponding physical address in the global metadata view, supporting fast positioning of the global entry and reducing the overhead of consistency checking.
[0052] Log Pointer: A pointer to a log region used to associate log entries with the global metadata view LSM tree, ensuring the persistence of atomic operations.
[0053] Checksum: A checksum used to prevent hardware errors or silent data corruption.
[0054] In this embodiment, for metadata distributed across multiple local NUMA nodes with primary replicas, when a metadata operation causes a change in the local metadata view, a consistency guarantee operation is triggered (implemented through atomic operations and a version number mechanism). The specific steps are as follows: Local metadata updates (such as file renaming) ensure transactionality through atomic multi-block write metadata operations; for example, a renaming operation needs to atomically execute the deletion of the old path and the creation of the new path to avoid inconsistencies caused by partial writes. A global version number is maintained for metadata items accessed across nodes. When one node modifies its local metadata, the version number is incremented and notified to other replica nodes via CXL's cache consistency broadcast or global metadata view update. When other nodes read remote metadata, they verify that the version number matches the global metadata view. Figure One If inconsistencies are found, a metadata synchronization process is triggered to ensure consistency. To reduce synchronization frequency, the system introduces a lazy failure mechanism, allowing nodes to cache remote metadata copies and setting a timeout threshold (e.g., 1 second). After the timeout, a forced refresh is performed to ensure consistency.
[0055] In this embodiment, the write operation of the global metadata view is appended to the log layer of the LSM tree, and the log data is periodically merged to the higher level; the merge operation writes the entries into the first level of the LSM tree after sorting them by key; the query operation quickly determines whether a certain metadata exists in a specific level by maintaining a global Bloom filter in DRAM.
[0056] In specific application embodiments, continue with Figure 1 The architecture diagram of the NUMA-aware distributed parallel file system metadata management system shown illustrates the global metadata view.
[0057] The core objective of the global metadata view is to provide consistency and persistence support for cross-node operations such as file movement across nodes and distributed lock management. Leveraging the large capacity and persistent access characteristics of CXL-SSD's memory mode, a distributed LSM (Log-Structured Merge-Tree) is constructed, utilizing its high write throughput and combining it with CXL's cache consistency protocol to optimize cross-node access efficiency.
[0058] In terms of storage structure design, the global LSM tree is stored in layers according to metadata keys (such as file path hashes). Write operations are first appended to the log layer of the CXL-SSD, and background threads periodically merge log data to higher levels to reduce random write overhead. For example, a file creation operation generates a log entry containing the path, inode number, and version number, which is appended to the log layer; the merge operation sorts these entries by key and writes them to the Level 1 layer of the LSM tree. To speed up queries, the system maintains a global Bloom filter in DRAM to quickly determine whether a certain metadata exists in a specific level, avoiding unnecessary CXL-SSD scans. It can be understood that metadata updates are written in the form of log entries to the CXL-SSD, and are merged to the LSM tree levels in the background, which can better support atomic persistence, such as reducing the persistence delay to 1 / 5 of traditional SSDs while ensuring high reliability.
[0059] Global metadata view entry format: The entry of the global metadata view needs to support cross-node consistency and other functions, as shown in Figure 3 Key-Value: Key-value pair, used to support efficient hash lookup and range query (such as directory traversal).
[0060] Version: Global version number, used to implement "multi-version concurrency control" to solve cross-node write conflicts.
[0061] Location&Replica Mask: NUMA node where the metadata primary copy is located and replica distribution bit mask, used to record the location and distribution of the metadata primary copy, support hot migration and fault recovery.
[0062] In distributed file systems, there are also problems of uneven load and congestion caused by hot metadata. High-frequency access metadata (such as shared directories and active file attributes) often concentrate in a few nodes, causing hot spot effect. Traditional load balancing strategies (such as static hash partitioning) cannot dynamically adapt to changes in access patterns, causing some nodes to be overloaded, cross-node traffic to surge, and even triggering a cascade of performance degradation. In addition, there is a lack of adaptive ability in dynamic load scenarios. Existing metadata management systems lack real-time hot spot detection and migration mechanisms, and cannot dynamically adjust data distribution according to load changes. In the case of sudden high concurrency (such as AI training tasks accessing the same data set), the system may experience performance jitter due to hot spot migration lag, affecting the quality of service (QoS).
[0063] Therefore, this embodiment also introduces a dynamic migration strategy for hot metadata. During the operation of various metadata, local hot metadata is further determined based on the access frequency of the metadata. From all local hot metadata, metadata accessed by more than a preset proportion of NUMA nodes is selected as global hot metadata, triggering a global hot metadata migration operation. This global hot metadata is copied to the NUMA node with the highest access frequency, or the primary copy of metadata that needs frequent modification is migrated to the NUMA node that frequently accesses that metadata. After the global hot metadata migration operation is completed, the global metadata view and the local metadata view are updated accordingly. It can be understood that the dynamic migration of hot metadata aims to move frequently accessed metadata closer to the requesting node through real-time detection and intelligent migration strategies, reducing the proportion of remote access.
[0064] In this embodiment, local hotspot metadata is determined based on the access frequency of metadata, specifically including: Hotspot detection employs a heuristic mechanism. Each NUMA node continuously monitors the access frequency (v) of metadata and stores this frequency in its local metadata view as an access counter (Access Count). Metadata access is calculated based on the access frequency and the cumulative average access frequency. Profitability metrics:
[0065] In the above formula, This is metadata, where t is the cumulative time period. and These represent the access frequency and cumulative average access frequency of metadata within the time period t, respectively. After calculating the ratio of the access frequency of a certain metadata to the cumulative average access frequency within time t, the metadata hotspot score is further calculated based on the metadata revenue metrics and a multi-armed slot machine exploration strategy.
[0066] In the above formula, Metadata at time t The revenue, metadata at time t The upper limit, where γ is an adjustable parameter. For metadata The number of times it was selected.
[0067] Finally, based on the score ranking, hot metadata can be selected, with metadata scoring above a preset value designated as local hot metadata. The central manager aggregates the hotspot information reported by each node. If a piece of metadata is accessed frequently by more than a set proportion of nodes, it is marked as a global hotspot. For example, if the shared directory / shared / dataset is accessed frequently by four NUMA nodes simultaneously, a global hotspot migration process is triggered.
[0068] Specifically, the migration strategy is divided into two categories: copy replication and ownership transfer: Copy replication: Global hot metadata is copied to the node with the highest access frequency, and multiple volumes are divided in CXL-SSD to maintain multiple copies. The specific process is that after the hot metadata is copied to the main volume, the data is backed up to the copy volume through the background asynchronous operation. For example, the metadata of the directory / shared / dataset is copied to the CXL-SSD of Node 1 and Node 3, and subsequent access can be preferentially routed to the nearest copy.
[0069] Ownership transfer: For write-intensive hotspots (such as frequently modified log files), the "main copy" of the metadata is migrated to the access node. During the migration process, the write process is first locked to prevent data inconsistency. Second, the specific metadata, including directories and files, is migrated through a blocking thread. The original node needs to synchronize the uncommitted modification log to the new main copy and update the ownership mapping table through the global metadata view. For example, if Node 2 frequently modifies the file / logs / system.log, the system migrates its main copy to the CXL-SSD of Node 2, and subsequent write operations are performed locally. It can be understood that the "main copy" of the write-intensive hotspot is migrated to the access node through the atomic version number, and the read-intensive hotspot maintains multiple copies in multiple CXL-SSDs, which are synchronized and updated through the cache consistency protocol, thereby optimizing the access path of hot metadata, balancing the load of the dynamic migration mechanism, adapting to sudden high-concurrency scenarios, avoiding system performance degradation caused by node congestion, and improving the quality of service (QoS).
[0070] In this embodiment, cross-node operation optimization is also included through batch transactions and lock-free reading. Batch metadata operations are combined into a single transaction submission, multiple operations are processed in parallel using the multi-stream write of CXL-SSD, read operations are accessed lock-free through the MESI cache consistency protocol of CXL, and the node directly reads the cache copy of the remote CXL-SSD, and only when the cache is invalid, access the persistent storage layer; All metadata update operations are written in the form of logs to the continuous area of memory. When batch metadata operations are combined, similar operations are combined through batch transactions, and a submission is made when the current operation frequency is greater than the historical frequency or the cumulative operation number is greater than the preset number. Among them, a fast comparison strategy based on hash is used to compare the similarity of operations with operation type and parameter as Key and output Value as Value, and operations with similarity greater than a preset percentage are combined.
[0071] In a specific application embodiment, in view of the crash consistency problem that may exist in the local metadata view, including persistence and fault recovery, it is implemented through log-structured storage and periodic snapshots. All metadata update operations are appended to a log in a memory continuous area, and the log entry contains the operation type, version number and timestamp. The system generates a snapshot of the global LSM tree every fixed period, and clears the processed log entries. During fault recovery, the latest snapshot is loaded first, and then the logs after the snapshot time point are played back to ensure the complete reconstruction of the metadata state. For example, if the system fails at 10:05, and the latest snapshot is generated at 10:00, the recovery process needs to play back the log entries from 10:00 to 10:05 to reconstruct the final state. By playing back only the logs within the fault time window, combined with the high bandwidth of CXL-SSD, the fault recovery time is compressed from hours to minutes.
[0072] To improve the performance of the log, the compression process of the log storage is optimized. First, the logs of metadata operations are classified according to the time window, and further classified according to extended attributes such as operation type, importance, etc. in the time window. High-value logs use low-compression-ratio algorithms to ensure query speed, and archived logs use high-compression algorithms. Second, the association between the compression algorithm and the memory pressure is optimized, and a background thread is used to continuously update the memory usage, which is used as an indicator to further optimize the compression process. As shown in the following formula:
[0073] It can be understood that the embodiment utilizes the cache coherence protocol of CXL-SSD and log-structured storage to achieve low-latency persistence and fast fault recovery.
[0074] Cross-node operation optimization is achieved through batch transactions and lock-free reading. For batch metadata operations (such as batch deleting 1000 files), the system combines them into a single transaction submission, and uses the multi-stream write feature of CXL-SSD to handle multiple operation streams in parallel. The read operation is implemented through the MESI cache coherence protocol of CXL to achieve lock-free access: the node can directly read the cache copy of the remote CXL-SSD, and access the persistent storage layer only when the cache is invalid. For example, when node A needs to read the metadata maintained by node B, if there is a valid copy in the local cache of node A (guaranteed by the MESI protocol), the result is returned directly, otherwise the data is loaded from the global LSM tree. It can be understood that the lock-free read of metadata operations of the global metadata view is achieved by using the MESI cache coherence protocol of CXL, which allows the node to directly read the cache copy of the remote CXL-SSD, and only accesses the persistent layer when the cache is invalid, which can effectively reduce the cross-node read delay (e.g. less than 100 ns).
[0075] For the conditions of batch transaction merging, optimization is performed through a corresponding heuristic algorithm. First, determine which operations need to be merged, specifically through a fast comparison strategy based on hashing, with operation type and parameters as Key and output Value for similarity comparison. Operations with a similarity greater than 90% can be merged. Second, for the timing of merging and committing, a comprehensive judgment is made in combination with the current operation frequency and threshold, as follows:
[0076] In this embodiment, if the metadata is marked as global hot metadata and triggers global hot migration, and the hot migration form is to move the global hot metadata to the NUMA node with the highest access frequency, then multiple volumes are divided in the CXL-SSD, multiple copies are maintained, and after the hot metadata is copied to the primary volume, the data is backed up to the replica volume through a background asynchronous operation; if the hot migration form is to copy the global hot metadata to the NUMA node with the highest access frequency, then a locking operation is performed during the writing process, the original node needs to synchronize the uncommitted modification log to the new primary copy, and the ownership mapping table is updated through the global view, and then the specific metadata is migrated, and the primary copy of the metadata is migrated to the NUMA node with the highest access frequency.
[0077] In one exemplary specific application, the process of modifying metadata across NUMA nodes is illustrated by taking file renaming as an example.
[0078] Step 1, local metadata view query: find the local metadata view entry according to the Key, if it hits and the global metadata view is valid, directly read the Value; if it does not hit or the version is expired, access the global metadata view entry through the Pointer to Global. Figure One
[0079] Step 2, local metadata view modification: lock the State Flags of the local metadata view entry (set the Locked bit), prevent concurrent write conflict; append a new operation (old path deletion + new path creation) in the log, update the LogPointer; increment the Version, update the Location (if hot ownership migration is triggered), and clear the Dirty bit.
[0080] Step 3, global metadata view synchronization: write the new Value and Global Version back to the global metadata view and mark it as Valid; if the current node is the location of the primary copy, update the Replica Mask and broadcast it to other replica nodes.
[0081] In the embodiment, the consistency guarantee operation further includes: if multiple NUMA nodes concurrently modify the same metadata, the global metadata view regards the operation corresponding to the timestamp of the latest modification of the metadata as the final operation according to the last-write-wins principle.
[0082] Specifically, the consistency synchronization is guaranteed by the version number and write propagation mechanism. The master replica node increments the version number after modifying the metadata and propagates the update to other replica nodes through the cache coherence broadcast of the CXL-SSD. If multiple nodes concurrently modify the same metadata (such as concurrent renaming conflicts), the global metadata view determines the final value according to the version number timestamp (the last-write-wins principle). For example, node A and node B concurrently modify file / conf / config.ini, and the global manager compares the submission time stamps of the two to select the later version as the final value.
[0083] The embodiment further provides a NUMA-aware distributed parallel file system metadata management system, which is based on the CXL-SSD architecture and includes: a local metadata view configured to dynamically allocate metadata to local NUMA nodes according to access patterns to partition the metadata according to NUMA nodes; a global metadata view configured to provide consistency and persistence support for cross- NUMA node operations of the metadata; The entry in the local metadata view includes a key Key, a value Value, a latest version number Global Version in the global metadata view, a local modification version number Local Version, an access counter Access Count, a last access timestamp Last Access Time, state flag bits State Flags, a pointer to the corresponding physical address Pointer to Global in the global metadata view, a pointer to a log area Log Pointer, and a checksum Checksum. The entry in the global metadata view includes a key Key, a value Value, a global version number Version, a NUMA node Location where the metadata master replica is located, a metadata replica distribution bitmask Replica Mask, and state flag bits State Flags.
[0084] The system of the embodiment corresponds to the method described above and also has the advantages of the method described above.
[0085] The above merely describes the preferred embodiments of the present application, and the protection scope of the present application is not limited to the above-described embodiments. Any technical solution falling within the concept of the present application shall fall within the protection scope of the present application. It should be noted that, for ordinary skilled persons in the art, some improvements and refinements without departing from the principles of the present application shall also be considered as falling within the protection scope of the present application.
Claims
1. A NUMA-aware distributed parallel file system metadata management method, characterized in that, The method comprises the steps of: building a hierarchical metadata management structure, the hierarchical metadata management structure comprising a local metadata view and a global metadata view, the local metadata view serving as a local view layer of a NUMA local node, and being used for dynamically allocating metadata to the local NUMA node according to an access mode to partition the metadata according to the NUMA node, and maintaining an independent metadata copy at each local NUMA node, when a processor performs an I / O operation on an external device, intercepting a metadata request and converting an access semantic of the external device from a block device operation to a memory semantic operation, so as to directly map a CXL-SSD storage space to a NUMA local address space through a user-mode memory mapping; the global metadata view serving as a global view layer, and being used for hierarchical storage according to a metadata key through a distributed global LSM tree based on the CXL-SSD, so as to provide consistency and persistence support for cross- NUMA node operations of the metadata; adopting a metadata partitioning strategy, the local metadata view dynamically allocating metadata to each NUMA node according to the access mode through the metadata partitioning strategy, and creating directories and files of high-frequency-accessed directory subtrees in the local NUMA node through a local priority rule; and distributing other metadata through a consistent hashing algorithm to make metadata under the same directory belong to the same NUMA node.
2. The NUMA-aware distributed parallel file system metadata management method of claim 1, wherein, for metadata having a primary copy distributed in multiple local NUMA nodes, when a metadata operation causes a change in the local metadata view, triggering a consistency guarantee operation, and the specific steps are as follows: the local metadata update is ensured to be transactional through an atomic multi-block write metadata operation; maintaining a global version number for a metadata item accessed across nodes, when one of the nodes modifies the local metadata, incrementing the version number and broadcasting or updating a notification of the other copy nodes through a cache consistency of the CXL or a global metadata view; when the other nodes read remote metadata, verifying whether the version number is consistent with the global metadata view, and if not, triggering a metadata synchronization process to ensure consistency.
3. The NUMA-aware distributed parallel file system metadata management method of claim 1, wherein, the method further comprises a hot metadata dynamic migration strategy, the hot metadata dynamic migration strategy comprising: further determining local hot metadata according to an access frequency of the metadata in a process of operating the metadata, screening out metadata accessed by more than a preset proportion of NUMA nodes from all the local hot metadata as global hot metadata, triggering a global hot migration operation, copying the global hot metadata to a NUMA node with the highest access frequency, or migrating a primary copy of metadata that needs to be frequently modified to a NUMA node frequently accessing the metadata, and updating the global metadata view and the local metadata view after the global hot migration operation is completed.
4. The NUMA-aware distributed parallel file system metadata management method of claim 3, wherein, determining the local hot metadata according to the access frequency of the metadata, and specifically comprising: each NUMA node statistically determines the access frequency of the metadata in real time, stores the access frequency in an access counter Access Count of the metadata in the local metadata view, and calculates a benefit index of the metadata according to the access frequency and a cumulative average access frequency; The benefit index of the metadata is calculated based on a multi-armed bandit exploration strategy to obtain a metadata hotspot score, and metadata with a score higher than a preset value is regarded as local hotspot metadata.
5. The NUMA-aware distributed parallel file system metadata management method of claim 4, wherein, The expression for calculating the benefit index of the metadata is: In the above formula, is metadata, t is a time period, and are the access frequency and the cumulative average access frequency of the metadata in the time period t, respectively. The expression for calculating the metadata hotspot score is: In the above formula, the metadata at time t the benefit of the metadata at time t the upper bound of the metadata at time t, γ is a tunable parameter, the metadata at time t the number of times the metadata is selected.
6. The NUMA-aware distributed parallel file system metadata management method of claim 3, wherein, If the metadata is marked as global hotspot metadata and global hotspot migration is triggered, and the hotspot migration form is to move the global hotspot metadata to the NUMA node with the highest access frequency, then multiple volumes are divided in the CXL-SSD, multiple copies are maintained, and after the hotspot metadata is copied to the primary volume, the data is backed up to the copy volume through a background asynchronous operation; If the hotspot migration form is to copy the global hotspot metadata to the NUMA node with the highest access frequency, then a locking operation is performed during the write process, the original node needs to synchronize the uncommitted modification log to the new primary copy, and the ownership mapping table is updated through the global view, and then the specific metadata is migrated, and the primary copy of the metadata is migrated to the NUMA node with the highest access frequency.
7. The NUMA-aware distributed parallel file system metadata management method of any of claims 1-6, wherein, It also includes cross-node operation optimization through batch transactions and lock-free reading, combining batch metadata operations into a single transaction submission, using the multi-stream write of CXL-SSD to process multiple operations in parallel, and using the MESI cache consistency protocol of CXL to perform lock-free access, so that the node directly reads the cache copy of the remote CXL-SSD, and only when the cache is invalid, access the persistent storage layer; All metadata update operations are written in the form of logs to a continuous area in memory, and when batch metadata operations are combined, similar operations are combined through batch transactions, and a submission is made when the current operation frequency is greater than the historical frequency or the cumulative number of operations is greater than a preset number, wherein a fast comparison strategy based on hash is used to compare the similarity of operations with a similarity greater than a preset percentage.
8. The NUMA-aware distributed parallel file system metadata management method of any of claims 1-6, wherein, The write operation of the global metadata view is appended to the log layer of the LSM tree, and the log data is merged to the higher level regularly; the merge operation writes the entries in order to the first level of the LSM tree; the query operation quickly judges whether a certain metadata exists in a specific level by maintaining a global Bloom filter in DRAM.
9. The NUMA-aware distributed parallel file system metadata management method of any of claims 2-6, wherein, The consistency guarantee operation also includes: if multiple NUMA nodes concurrently modify the same metadata, the global metadata view uses the last write wins principle to regard the operation corresponding to the latest timestamp of modifying the metadata as the final operation.
10. A system for NUMA-aware distributed parallel file system metadata management method according to any one of claims 1-9, characterized in that, The system is based on a CXL-SSD architecture, and the system comprises: a local metadata view, which is used to dynamically allocate metadata to local NUMA nodes according to access patterns to partition the metadata according to NUMA nodes; a global metadata view, which is used to provide consistency and persistence support for cross- NUMA node operations of the metadata; The entry in the local metadata view includes a key Key, a value Value, a latest version number Global Version in the global metadata view, a local modification version number Local Version, an access counter Access Count, a last access timestamp Last Access Time, state flag bits State Flags, a pointer to the corresponding physical address Pointer to Global in the global metadata view, a pointer to the log area Log Pointer, and a checksum Checksum. The entry in the global metadata view includes the key Key, the value Value, a global version number Version, a NUMA node where the metadata primary copy is located Location, a metadata replica distribution bit mask Replica Mask, and state flag bits State Flags.
Citation Information
Cited By
Concurrent counting system and method based on hierarchical conflict perception
CN121412075A