A file storage management system of a cloud storage system

CN122593707BActive Publication Date: 2026-09-18SICHUAN LEWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611050422.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-18
Estimated Expiration
2046-07-15

AI Technical Summary

Technical Problem

[0005]针对现有技术的不足,本发明提供了一种云存储系统的文件存储管理系统,解决了现有分布式存储架构中因海量数据块映射表膨胀导致的元数据寻址瓶颈、多租户离散小文件写入引发的磁盘随机I/O性能下降,以及依赖分布式锁处理并发写入所带来的系统高延迟的问题

Benefits of technology

1、本发明依据文件标识符和数据块索引号,结合集群拓扑视图计算出数据块的目标容器及物理位置,无需针对每个数据块查询中心化的元数据记录。从而消除了大规模集群中因元数据频繁访问及记录膨胀导致的性能瓶颈,降低了元数据服务器的内存与计算压力,使得存储系统能够支持海量文件的快速寻址以及集群规模的线性扩展。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122593707B_ABST
    Figure CN122593707B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of file storage management, and discloses a file storage management system of a cloud storage system, which comprises an access adaptation module, a metadata module, a slice version module, a deterministic mapping module, a container management module and a persistence control module. The deterministic mapping module performs stateless calculation based on a file identifier and a data block index number to generate a target container identifier and a virtual slot number. The container management module constructs a buffer container in the memory, ignores the discreteness of the virtual slot, and writes data blocks into a data storage area in a compact appending mode. The persistence control module drives effective data to be written into a disk according to a comparison result of a global version identifier and a historical version. The application eliminates the metadata bottleneck through deterministic routing, converts random I / O into sequential I / O through memory compact reorganization, and solves concurrent conflicts based on a lock-free mechanism, thereby significantly improving the throughput and expansibility of the storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of file storage management technology, specifically a file storage management system for a cloud storage system. Background Technology

[0002] With the rapid development of mobile internet and IoT technologies, the scale of unstructured data is growing exponentially. Business scenarios such as high-definition video surveillance, autonomous driving data collection, and massive log archiving generate hundreds of millions of files every day. Cloud storage systems provide highly reliable data access services by integrating large-scale general-purpose server resources.

[0003] In mainstream distributed file system architectures, data addressing largely relies on centralized metadata services. The system not only records the file directory structure but also maintains a large block mapping table to record the specific location of logical data blocks on the physical disk. For data persistence, the storage engine typically maps logical blocks directly to physical disk sectors. However, when dealing with consistency issues arising from concurrent writes by multiple clients, existing technologies generally employ a distributed lock-based control strategy. Clients must request permission over the network before writing and release the lock resource after the operation is complete.

[0004] Existing technical architectures have shortcomings when dealing with massive small file scenarios. As storage clusters expand, the number of data blocks surges, causing block mapping tables to rapidly expand and consume significant memory. If metadata cannot be fully loaded, disk swapping operations lead to a sharp increase in addressing latency. This table-lookup mechanism limits the linear scalability of the cluster. The random read / write performance of mechanical hard drives is far lower than sequential read / write. In multi-tenant scenarios, write requests are highly discrete in time and space, and direct mapping causes frequent and significant disk head seeks, resulting in severe mechanical overhead and extremely low effective system bandwidth. Therefore, this invention provides a file storage management system for cloud storage systems to address the shortcomings of existing technologies. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a file storage management system for cloud storage systems, which solves the problems of metadata addressing bottlenecks caused by the expansion of massive data block mapping tables, the performance degradation of disk random I / O caused by multi-tenant discrete small file writing, and the high system latency caused by relying on distributed locks to handle concurrent writes in existing distributed storage architectures.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a file storage management system for a cloud storage system, comprising: The access adaptation module is used to reorganize external write requests into a standardized data stream object that contains file paths and binary data carriers within the system. The metadata module is used to parse the file path in the data stream object to obtain a globally unique file identifier, providing a logical identifier for data addressing; The slice version module is used to discretize the data stream object into fixed-length data blocks and generate a global version identifier for each data block to identify the timing state. The deterministic mapping module is used to perform stateless computation based on the file identifier and data block index number to generate the target container identifier and virtual slot number; The container management module is used to build buffer containers in memory, compactly append data blocks to the data storage area according to the target container identifier, and establish a mapping relationship between virtual slot numbers and physical offsets in memory. The persistence control module is used to drive the physical disk write operation of valid data based on the comparison result between the global version identifier and the historical version in the physical storage medium.

[0007] Preferably, the access adaptation module includes a buffer aggregation unit; The buffer aggregation unit is used to build a buffer layer in memory. By monitoring the total amount of data accumulated in the current cache area or the data residence time, it triggers a data flushing operation when a preset threshold is reached. The buffer aggregation unit, based on the triggered data flushing operation, concatenates multiple discrete small block write requests into a single continuous large block of data or encapsulates it into a vector I / O list to construct batch write requests.

[0008] Preferably, the metadata module includes a namespace management unit; The namespace management unit is used to employ a key-value pair-based distributed storage architecture. It calculates the logical index number of the target metadata shard node by calculating the bitwise XOR result of the globally unique identifier of the parent directory and the hash value of the file name string, and then applies a consistent hash function and a modulo operation on the bitwise XOR result to the total number of available shard nodes. This determines the physical storage location of the directory entry.

[0009] Preferably, the slice version module includes a version management unit; The version management unit is used to construct a global version identifier by combining the high-order physical timestamp and the low-order logical sequence number; The version management unit ensures, based on the incrementing logic of the atomic counter, that the global version identifier generated by the later write request for the same block index of the same file is numerically greater than the identifier of the earlier write request, so as to establish the global temporal order of the data blocks.

[0010] Preferably, the deterministic mapping module includes an object key calculation unit; The object key calculation unit is used to perform a left shift operation on the data block index number, perform a bitwise XOR operation on the shift result and the file identifier, and apply a hybrid hash function to the operation result to generate a logical object key that is independent of physical location. The logical object key serves as the logical digital fingerprint of the data block and is used to drive the topology-aware routing unit mapping calculation of the deterministic mapping module.

[0011] Preferably, the deterministic mapping module further includes a topology-aware routing unit and a virtual slot allocation unit; The topology-aware routing unit is used to calculate the target container identifier and the list of physical nodes carrying the target container based on the logical object key, combined with the cluster topology diagram with version number and preset data placement rules. The virtual slot allocation unit is used to perform a modulo operation on the data block index number for the maximum logical slot capacity preset for a single container, and determine the remainder obtained by the operation as the virtual slot number of the data block inside the container.

[0012] Preferably, the container management module includes a buffer allocation unit; The buffer allocation unit is used to pre-allocate a memory page pool using the Slab allocation mechanism and allocate independent memory buffer objects according to the target container identifier. The memory buffer object is constructed as two physically isolated areas: a data storage area for sequentially storing the file content byte stream and a metadata index area for recording the block mapping relationship inside the container.

[0013] Preferably, the container management module further includes a compact reorganization unit and a dynamic micro-mapping unit; The compact reorganization unit is used to perform the transformation from logically sparse to physically compact, ignoring the discreteness of the virtual slot number corresponding to the data block, and appending the data block to the physical location pointed to by the current write pointer of the data storage area. The dynamic micro-mapping unit is used to update the index entries in the mapping table with virtual slot number as the subscript according to the physical location after writing, and to construct a direct mapping from virtual slot number to memory physical offset, data effective length and global version identifier.

[0014] Preferably, the persistence control module includes a version verification unit; The version verification unit is used to traverse the index entries of the buffer container before writing to disk, extract the global version identifier of the data block to be written as the new version, and read the historical version of the same virtual slot on the storage medium. The version verification unit generates operation instructions based on the comparison logic: when the new version value is greater than the historical version, an execution write instruction is generated; when the new version value is less than or equal to the historical version, a discard instruction is generated and the data block is removed, so as to resolve concurrent write conflicts in a lock-free environment.

[0015] Preferably, the persistence control module further includes a serialization encapsulation unit and an atomic brush unit; The serialization encapsulation unit is used to add a container header and a cyclic redundancy check code based on polynomial calculation to the beginning and end of the verified memory buffer object, and encapsulate it into a linear binary format suitable for disk storage. The atomic brush unit is used to utilize the append write feature of the underlying file system to bypass the operating system page cache and directly write the encapsulated container data to the pre-allocated area of ​​the physical disk.

[0016] This invention provides a file storage management system for a cloud storage system. It has the following beneficial effects: 1. This invention calculates the target container and physical location of a data block based on the file identifier and data block index number, combined with a cluster topology view, eliminating the need to query centralized metadata records for each data block. This eliminates performance bottlenecks caused by frequent metadata access and record bloat in large-scale clusters, reduces the memory and computational pressure on the metadata server, and enables the storage system to support fast addressing of massive files and linear scaling of the cluster.

[0017] 2. This invention utilizes a container management module to decouple logical addresses from physical addresses and achieve compact storage in memory. Through compact reassembly units, regardless of how discrete or sparse the virtual slot numbers corresponding to upper-layer write requests are, data blocks are sequentially appended to the data storage area of ​​the memory buffer container. This mechanism transforms a large number of random, scattered small-block write requests into a sequential, large-block write stream on the physical medium, effectively reducing the seek overhead of the underlying disk head and improving the data write throughput of storage media such as hard disk drives.

[0018] 3. This invention provides an efficient concurrency conflict resolution mechanism based on the persistence control module. By assigning a monotonically increasing global version identifier to each data block and comparing the version information of the old and new data during the final disk write phase, optimistic concurrency control is achieved. This avoids introducing high-latency distributed lock protocols during distributed writes, and while ensuring strong data consistency and atomicity, it reduces the system waiting time during concurrent writes by multiple clients, thereby improving the overall system response efficiency. Attached Figure Description

[0019] Figure 1 This is a system architecture diagram of the present invention; Figure 2This is a flowchart of the method of the present invention; Figure 3 This is a schematic diagram of the deterministic mapping module of the present invention; Figure 4 This is a schematic diagram comparing the write throughput under different concurrency scales in Experiment 1 of this invention; Figure 5 This is a schematic diagram comparing the cumulative write latency distribution in Experiment 2 of this invention. Detailed Implementation

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] See attached document Figure 1 , Figure 1 This is a system architecture diagram according to an embodiment of the present invention. The present invention provides a file storage management system for a cloud storage system, the system including an access adaptation module, a metadata module, a slice version module, a deterministic mapping module, a container management module, and a persistence control module. The above modules are interconnected and work together to complete the writing, reading, and storage space management of file data.

[0022] The access adaptation module is configured as the system's communication entry point, used to receive file operation requests initiated by external clients or applications. File operation requests include file creation, writing, reading, and deletion operations. The access adaptation module converts the received requests conforming to standard file system protocols into internal system data stream instructions and performs preliminary buffering processing on the input data.

[0023] The metadata module connects to the access adaptation module and is used to manage the directory tree structure and file attribute information of the file system. The metadata module stores the mapping relationship between filenames, directory levels, and globally unique file identifiers. When a file operation request is received, the metadata module parses the corresponding file identifier based on the file path in the request. The file identifier is used for subsequent data addressing.

[0024] The slice versioning module connects to the metadata module and is used to perform logical discretization processing on the file data stream. Based on a preset slice size, the slice versioning module divides the file data stream to be written into several fixed-length data blocks and calculates the index number corresponding to each data block. Simultaneously, the slice versioning module assigns a monotonically increasing global version identifier to each data block to be written; this global version identifier is used to identify the temporal state of the data block.

[0025] The deterministic mapping module connects to the slice version module and is used to perform stateless data addressing computation. The deterministic mapping module incorporates a topology-aware deterministic hash algorithm, using the file identifier and data block index as input parameters to calculate the target container identifier and the virtual slot number within the container corresponding to the data block in the distributed cluster. The deterministic mapping module does not rely on a centralized block mapping table to record the physical location of data blocks.

[0026] The container management module resides on each storage node and is used to build buffer containers in memory that correspond to the physical storage layout. The container management module receives data blocks routed by the deterministic mapping module and does not perform sparse storage according to virtual slot numbers; instead, it appends the data blocks to a compact data area in the memory buffer. The container management module maintains the mapping relationship between virtual slot numbers and physical offsets in memory.

[0027] The persistence control module connects to the container management module and is used to write data blocks from memory to physical storage media. When performing a write operation, the persistence control module reads the metadata mapping table of the corresponding container on the physical media and compares the global version identifier of the data block to be written with the version information of the existing data on the physical media. The persistence control module only performs a physical overwrite or append write operation if the global version identifier of the data block to be written is greater than the version information of the existing data on the physical media.

[0028] See attached document Figure 2 , Figure 2 This is a flowchart of a method according to an embodiment of the present invention. The present invention provides a file storage management method for a cloud storage system, comprising the following steps: S10, Receive write request and identifier parsing. The access adaptation module receives a write request containing the file path and data content, and the metadata module parses the file path to obtain a unique file identifier; S20, Data Slicing and Version Assignment. The slice version module divides the data content into fixed-length data blocks and assigns a global version identifier for the current moment, generating a data packet containing a file identifier, data block index number, global version identifier, and data body; S30, Deterministic Route Calculation. The deterministic mapping module receives the metadata of the data packet, calculates the physical node and target container identifier where the data block should fall, as well as the logical location of the data block within the container, i.e., the virtual slot number; S40, Memory Buffering and Compact Reassembly. After the data packet is transmitted to the target physical node, the container management module stores the data block into a buffer container in memory and records the virtual slot number pointing to the current memory address in the memory mapping table; S50, version conflict detection and persistence. When the preset disk flushing conditions are met, the persistence control module locks the target physical file, performs conflict detection based on the global version identifier, and writes valid data that passes the detection to the physical disk and updates the slot mapping table on the disk to complete the persistent storage of the data.

[0029] The access adaptation module is configured on the edge side or gateway node of the cloud storage system to shield the differences in underlying heterogeneous storage protocols and transform external unstructured file operation commands into standardized data stream objects within the system. Logically, the access adaptation module includes a protocol interaction unit, a request parsing unit, and a buffer aggregation unit.

[0030] The protocol interaction unit is responsible for establishing a communication connection with the client. This unit supports various standard file transfer protocols, including but not limited to Network File System Protocol (NFSP), Server Message Block Protocol (SLP), Object Storage Protocol (OSP), and local mount protocols compatible with portable operating system interfaces. For each established connection, the NFSP maintains a session context, which records the client's authentication information, connection status, and the root directory handle of the mount point. The specific network communication handshake, authentication, and encryption processes are well-known technologies in the field and will not be elaborated upon here.

[0031] The request parsing unit connects to the protocol interaction unit and is used to unpack and extract semantics from received protocol data packets. When a client initiates a write operation, the request parsing unit extracts key four-tuple information from the protocol message: target file path, write start offset, data length, and binary data carrier. During this process, the request parsing unit verifies the validity of input parameters, eliminates illegal requests that go out of bounds or have incorrect formats, and ensures that the data entering the storage core path conforms to preset specifications.

[0032] The buffer aggregation unit is configured to optimize the write performance of small data blocks. In cloud storage scenarios, frequent small I / O (Input / Output) operations can lead to significant random addressing overhead on the backend disk. The buffer aggregation unit works by leveraging the high-speed read / write capabilities of memory to build a buffer layer. By introducing controllable latency, it trades increased throughput for higher throughput, transforming high-frequency, random small data block write requests into low-frequency, sequential, or batch write requests. This reduces the context switching frequency and disk seek count of the backend storage engine.

[0033] The buffer aggregation unit allocates a volatile or non-volatile buffer in memory to execute an input / output merging strategy. Specifically, the buffer aggregation unit is set with a preset aggregation threshold. and maximum delay time threshold .in, The value is usually set to an integer multiple of the physical sector size of the back-end storage medium or consistent with the slice size of the subsequent logical slicing module (e.g., 4MB to 64MB) to ensure data alignment when writing to disk; The value is determined based on the write latency service level agreement (SLA) tolerable by the business, and the value range is usually between 200 milliseconds and 5 seconds.

[0034] When the length of the received write request is less than the aggregation threshold, the system does not immediately trigger the backend storage process. Instead, it writes the data to the cache and updates the current accumulated data volume in the cache. The system uses the following logic to determine whether to trigger a data flush operation: ; In the formula, This is a boolean signal that triggers the flush operation; flushing is performed when its value is True. This indicates the total amount of data to be written that has accumulated in the current cache, in bytes. This represents the preset spatial aggregation threshold; This indicates the longest time that currently cached data can reside in the buffer, which is the difference between the current time and the time when the oldest piece of data in the buffer was written. This indicates the maximum allowable delay time threshold of the system.

[0035] Through the above mechanism, the buffer aggregation unit can merge multiple discrete small block write requests. For small block requests with contiguous addresses, the buffer aggregation unit concatenates them into a single contiguous large block of data in memory; for small block requests with non-contiguous addresses, the buffer aggregation unit encapsulates them into a list of vector I / Os containing multiple address segments, and passes them down as a single batch request, thereby improving the overall throughput of the system.

[0036] After parsing and aggregation, the data is finally encapsulated by the access adaptation module into a system-wide generic I / O object. The data structure of this I / O object includes a globally unique session identifier (SessI / OnID), a normalized file logical path, the aggregated data body, and corresponding data length and offset descriptors. This I / O object is then passed to the downstream metadata module for further processing. This design decouples the access layer from the storage logic layer, allowing the backend storage engine to be unaware of the specific access protocol type from the frontend.

[0037] The metadata module is configured to maintain the mapping between the logical and physical views of the file system and to implement lightweight metadata management. Logically, the metadata module is divided into a namespace management unit, an identifier generation unit, and an attribute storage unit.

[0038] The namespace management unit is configured to maintain the hierarchical directory tree structure of the file system. Unlike traditional file systems that store directory entries and inodes in a mixed manner, this embodiment uses a key-value pair-based distributed storage architecture to manage the directory hierarchy in a flattened way. The underlying storage engine uses a log-structured merged tree data structure to support high-throughput metadata write operations. The namespace management unit maps each directory entry in the file system to an independent key-value record. Specifically, the system uses a combination of the parent directory identifier and the filename as the search key, and the identifier and object type of the child object as the search value.

[0039] To achieve load balancing in the distributed metadata cluster, the namespace management unit employs a metadata sharding strategy. Its core principle lies in using multiple hash calculations to distribute file metadata with directory locality across different physical nodes, thereby avoiding single-node performance bottlenecks caused by numerous file operations in a single hotspot directory. The system determines the metadata sharding node to which a directory entry belongs based on a preset hash algorithm. This sharding logic satisfies the following relationship: ; In the formula, The logical index number representing the target metadata shard node; A globally unique identifier representing the parent directory; A string representing the filename or directory name to be operated on; This represents a bitwise XOR operation, used to combine the parent directory ID with the characteristic value of the filename to increase the dispersion of the hash result; This indicates the total number of available shard nodes in the current metadata cluster. This number is usually configured as a power of 2 (such as 1024 or 4096) to facilitate dynamic scaling up and down of nodes through a consistent hash ring. Represents modulo operation; This indicates the system's default consistent hash function, preferably the MurmurHash3 or CityHash64 algorithm, which is fast to compute and has good collision resistance.

[0040] Through the above calculations, the namespace management unit can evenly distribute massive amounts of file directory entries across different nodes in the cluster. When a file path access request is received, the namespace management unit resolves the character-encoded file path into a unique identifier used internally by the system through recursive lookup or path caching mechanisms.

[0041] The identifier generation unit is connected to the namespace management unit and is used to generate globally unique file identifiers. This file identifier is a fixed-length integer that serves as a unique identity credential throughout the file's entire lifecycle. The identifier generation unit does not depend on the file's physical storage location; that is, the file identifier is completely decoupled from the physical address of the data block.

[0042] In its implementation, the identifier generation unit employs a bit-segment-based ID generation strategy. The generated 64-bit integer, from most to least significant bit, includes: a 41-bit timestamp difference (accurate to milliseconds), a 10-bit machine identifier (Machine ID), and a 12-bit sequence number. The machine identifier, ranging from 0 to 1023, distinguishes different metadata service nodes from which the IDs are generated. The sequence number resolves concurrent conflicts within the same millisecond, supporting a single node to generate 4096 unique IDs per millisecond. This structure ensures that IDs generated under high-concurrency write scenarios do not conflict and maintain an overall monotonically increasing trend, thereby optimizing the primary key index performance of the backend database and reducing page splits.

[0043] The attribute storage unit is used to manage the metadata attribute information of files. This attribute storage unit maintains an attribute object for each file, which records the file's access permissions, owner information, file size, creation time, modification time, and access time.

[0044] In this embodiment, the attribute storage unit improves the inode structure. The attribute objects managed by the attribute storage unit eliminate the physical address index table of data blocks. Since the system subsequently employs a holographic deterministic routing mechanism, the attribute storage unit does not need to record the specific location of each data block, but only the logical total size of the file. This structural design ensures that the metadata size of a single file is constant and extremely small, typically fixed between 128 and 512 bytes, and does not expand with the growth of file data. This significantly reduces the memory footprint of the metadata server, supporting the system to manage billions of files.

[0045] When the access adaptation module initiates a file creation request, the identifier generation unit generates a new file identifier, the attribute storage unit initializes the corresponding attribute object, and the namespace management unit writes a new directory entry record in the parent directory. The three work together to complete the logical file creation process.

[0046] The slice versioning module is used to convert continuous, variable-length byte streams into discrete, fixed-length standard data blocks, and add a timing credential for distributed consistency verification to each individual data block. This slice versioning module includes a slice calculation unit and a version management unit.

[0047] The slice computation unit is configured to perform discretization processing of the data stream, establishing a mapping mechanism from the linear byte address space in the user view to the discrete block address space in the system view. Through this mapping, the system can decompose operations on arbitrarily large files into atomic operations on standard fixed-size storage units, thereby eliminating external fragmentation caused by varying file sizes in the underlying storage.

[0048] The system has a pre-defined standard block size. The standard block size is typically set to a fixed value between 4 megabytes (MB) and 64 megabytes (MB), for example, 4MB. This value is set to balance the amount of metadata with transmission latency; blocks that are too small will cause metadata bloat, while blocks that are too large will increase the write amplification effect of small file reads and writes.

[0049] When a file identifier and write start offset are received and data length When a write request is received, the slice calculation unit maps the write range to a logical block index space starting with 0, based on the standard block size. For variable-length data processing, the slice calculation unit first calculates the starting block index to be covered by this write operation. and end block index The calculation logic is shown in the following formula: ; ; In the formula, This indicates the floor function; Indicates the starting offset of the data to be written within the file in bytes; Indicates the total length of the data written in bytes; This indicates the system's default standard block size.

[0050] Based on the calculated index range, the slice calculation unit executes a boundary splitting strategy to decompose the original write request into... Individual slice task For each sub-slice task (in from Traversal to The system calculates the initial offset within each block. and write length : Head slice (when) hour): ; ; In the formula, The function is used to handle situations where "a single write operation does not fill the remaining space of the current block" or "the total amount of data written is extremely small and only occupies a portion of the current block".

[0051] Intermediate slice (when) hour): ; ; Among them, the middle slice must be a complete standard block.

[0052] Tail slice (when) and hour): ; ; This formula, by subtracting 1, taking the modulo, and then adding 1, rigorously covers the boundary case where "the data just fills the entire block," avoiding the ambiguity of a zero value that might result from direct modulo. Through this processing, write requests of arbitrary length and position are transformed into standard slice operations targeting a specific block index.

[0053] The version management unit works in conjunction with the slice calculation unit to generate a global version identifier with globally monotonically increasing characteristics. This identifier is used to perform total ordering of concurrent write operations in a distributed environment.

[0054] The version management unit connects to the system's high-availability timestamp generation service. For each slice task output by the slice calculation unit, the version management unit assigns a global version identifier for the current moment. To ensure the uniqueness and monotonicity of the identifier in environments with high concurrency and potential deviations in distributed clocks, the identifier adopts a 64-bit binary structure, specifically defined as follows: The high 48 bits: physical timestamp, taken from the system's current UTC time, with a bit width sufficient to cover a time range of thousands of years.

[0055] The lower 16 bits: a logical sequence number used to distinguish different write requests within the same millisecond. The system maintains an atomic counter that resets at the millisecond level, incrementing by one for each request processed. If the number of requests per millisecond on a single node exceeds 65535 (i.e., 2^35), the system will not be able to detect this. 16 If -1 is encountered, the system will be forced to wait until the next millisecond or borrow from the high bit of the logic clock.

[0056] Specifically, the version control unit ensures that the following timing constraints are met: for two different write requests targeting the same index block of the same file... and ,like Later in physical time Initiate, or and If a causal dependency exists, then the version identifier assigned by the system must satisfy [the following condition]. .

[0057] After slice calculation and version assignment, the slice versioning module reassembles the original request into a set of standardized slice vectors. Each element in the vector contains a four-tuple of information: file identifier, data block index, global version identifier, and data fragment. This four-tuple contains not only the data content but also the logical location coordinates of the data and the time dimension coordinates of the data. This self-contained data structure allows the subsequent routing and storage modules to independently complete the data routing and disk persistence decisions without having to backtrack to query the metadata server.

[0058] See attached document Figure 3 The deterministic mapping module leverages high-density CPU computing power in exchange for low-latency memory and input / output resources. It replaces traditional metadata server lookup operations with deterministic algorithm computation, thereby eliminating metadata access bottlenecks in large-scale clusters. The deterministic mapping module includes an object key calculation unit, a topology-aware routing unit, and a virtual slot allocation unit.

[0059] The object key calculation unit performs the first stage of the holographic mapping calculation, which generates logical object keys independent of specific physical locations. This unit receives the file identifier and data block index number from the metadata quadruple output by the slice version module as input seeds.

[0060] To ensure uniform data distribution across the entire storage cluster, the object key calculation unit employs a highly discrete hash algorithm to perform mixed operations on the input seed. The calculation logic aims to eliminate the influence of file naming or write order on the distribution results. The formula for calculating the object key is shown below: ; In the formula, The calculated logical object key is a 64-bit unsigned integer. A globally unique file identifier; This indicates the logical index number of the data block currently being processed; Indicates left shift operation; This indicates the preset shift step size. The value of this step size is usually set to half the effective bit width of the file identifier (e.g., 32). The technical purpose is to move the monotonically increasing or small-changing data block index features to the high bit region of the binary representation, so that when it is XORed with the file identifier, it can sufficiently interfere with the entropy value of the hash seed and avoid low-bit collisions. This indicates a bitwise XOR operation; This refers to a hybrid hash function exhibiting an avalanche effect, preferably using the WyHash or XXHash algorithm to ensure computation is completed within microseconds and a collision rate of less than 10. 12 .

[0061] Through the above calculations, the system generates a unique and deterministic digital fingerprint for each logical data block. This digital fingerprint uniquely identifies the logical address space of the data block without relying on any external database records.

[0062] The topology-aware routing unit is connected to the object key calculation unit and is used to map logical object keys to physical storage topology. The topology-aware routing unit maintains a cluster topology map with a version number. This topology map records the active storage nodes, rack structure, and fault domain division in the current cluster in the form of a hierarchical tree.

[0063] The topology-aware routing unit, based on a preset replication strategy, uses a consistent hashing algorithm or a pseudo-random placement algorithm to calculate the target container identifier and its corresponding list of physical nodes where a data block should fall. This process follows the "fault domain isolation" principle, meaning that the calculated multiple replica target locations must be located in different racks or power domains. Specifically, the selection of the target container satisfies the following mapping relationship: ; In the formula, An identifier representing the target logical container; This represents the list of physical node IP addresses that host the primary and secondary replicas of the container; Represents the topology mapping function; This represents the key of the logical object calculated in the previous step; This represents a globally consistent view of the cluster topology at the current moment. This indicates the preset data placement rules, including the number of replicas (e.g., 3 replicas) and the minimum fault isolation level (e.g., Host level or Rack level).

[0064] The mapping calculation process is stateless. As long as the nodes participating in the calculation hold the same version of the cluster topology graph (same Map Epoch), the calculation results of any client or storage node on the same data block will necessarily be the same, thereby eliminating the query dependency on the centralized metadata server.

[0065] The virtual slot allocation unit is configured to determine the logical location of a data block within the target container, i.e., the virtual slot number. The virtual slot allocation unit calculates based on the relative position of the data block index number within the container's capacity. This calculation ensures that consecutive data blocks of the same file are distributed across different containers or different slots within the same container according to a preset striping pattern. The virtual slot essentially acts as an index for the hash buckets within the container, and its calculation follows modular arithmetic logic. ; In the formula, This indicates the virtual slot number of the data block within the target container; The logical index number representing the data block; This parameter represents the maximum logical slot capacity preset for a single container. The value is determined based on the design of the container's metadata index structure, and is usually set to a power of 2 (e.g., 4096 or 8192) to facilitate fast location via bitwise operations. Larger capacity values ​​reduce hash collisions but increase the overhead of memory indexing; smaller capacity values ​​have the opposite effect.

[0066] Through the collaborative computation of the three units mentioned above, the deterministic mapping module outputs a routing instruction containing the target physical node IP, container identifier, and virtual slot number. This routing instruction specifies the logical location of the data block storage, rather than the actual physical offset of the data block on the disk. The actual physical disk location will be dynamically determined by the subsequent container management module based on the current disk write pointer, while the virtual slot number will serve as a constant anchor point connecting the deterministic route and the dynamic physical storage.

[0067] The container management module is deployed on physical storage nodes, serving as the execution engine between deterministic routing and the underlying physical media. It is used to build a hardware-software hybrid memory container architecture, shaping random, discrete write requests from the upper layer into a disk-friendly, sequential, and compact data stream by performing logical-physical address decoupling at the memory stage. The container management module includes a buffer allocation unit, a compact reassembly unit, and a dynamic micro-mapping unit.

[0068] The buffer allocation unit manages the memory resources of physical nodes, employing a Slab allocation mechanism to pre-allocate and maintain a fixed-size pool of memory pages. This mechanism avoids the context switching overhead and memory fragmentation issues caused by frequent calls to the operating system kernel (such as malloc / free) by pre-allocating large, contiguous blocks of physical memory. Furthermore, the starting address of the memory allocated by the buffer allocation unit strictly adheres to a 4KB or 8KB page boundary alignment principle to meet the hardware requirements for memory address alignment in subsequent Direct I / O operations, thereby bypassing the operating system page cache and directly writing the memory to disk.

[0069] When the system starts up or a write request arrives, the buffer allocation unit allocates a separate memory buffer object for each active container identifier. To balance memory utilization with I / O efficiency during disk writes, the physical capacity of the memory buffer object is set to a fixed value. This capacity value is typically set between 16 megabytes and 128 megabytes, for example, 64MB. The selection of this value is based on the bandwidth characteristics of the underlying disk and the stripe width of the erasure coding, ensuring that subsequent disk write operations can form a full-strip write, reducing the performance penalty caused by read-modify-write.

[0070] Internally, the memory buffer object is strictly divided into two physically isolated areas: a data storage area and a metadata index area. The data storage area stores the actual file content byte stream and occupies the majority of the buffer object's space; the metadata index area stores the block mapping relationships within the container, and its size is fixed and much smaller than the data area. For example, for a container with a capacity of 64MB and a maximum number of slots of 4096, the metadata index area only requires approximately 64KB to 128KB of space.

[0071] The compact reorganization unit is used to execute a "logically sparse - physically compact" reorganization strategy. In the logic of the deterministic mapping module, data blocks are discretely distributed in virtual slots based on hash values, resulting in a highly sparse logical data distribution. If data is stored directly according to the offset corresponding to the virtual slot number, it will lead to a large number of gaps in memory, greatly wasting storage resources.

[0072] To this end, the compact reassembly unit employs a log-structured append-only write mode. When a write command containing a virtual slot number and a data fragment is received, regardless of the specific value of the virtual slot number, the compact reassembly unit appends the data fragment to the physical location pointed to by the current write pointer in the data storage area. After the write is complete, the write pointer moves forward by the corresponding number of bytes.

[0073] The dynamic micro-mapping unit and the compact reassembly unit work together to maintain a mapping table from virtual slots to physical offsets in memory in real time. This mapping table is physically implemented as a linear array with virtual slot numbers as indices. If a hash collision occurs (i.e., different data block indices map to the same virtual slot number), this embodiment maintains a linked list or red-black tree structure in the corresponding slot of the array to store the conflicting index entries, or uses open addressing to redirect the conflicting entries to subsequent free slots. Since the virtual slot numbers are consecutive integers (0 to N-1), the system can locate the metadata of any slot in O(1) time complexity through array indexing operations, without performing complex hash lookups or tree traversals.

[0074] For each successfully written data block, the dynamic micro-mapping unit updates the index entry at the corresponding index position in the mapping table array. This index entry is a fixed-length structure (e.g., 16 bytes or 32 bytes) containing three key fields: memory physical offset, valid data length, and global version identifier. The calculation logic for the memory physical offset follows a recursive relationship: ; In the formula, Indicates the first The starting physical byte offset of the data to be appended in the data storage area; Indicates the base address of the data storage area; Indicates the first The length of data in each write operation; This indicates the sequence number of the write request currently received by the container.

[0075] Through the above mechanism, the container management module achieves compact data storage. Regardless of how discrete the virtual slot numbers of upper-layer write requests are, their physical storage in memory remains contiguous and without gaps. The technical advantages of this design are: Maximize space efficiency: For sparse files or random write scenarios, only the memory space of the actual data size is occupied, and empty slots that are not written do not occupy any data area space.

[0076] Sequential I / O optimization: Randomly arriving small I / O operations are converted into sequential arrangements in memory, laying the foundation for subsequently flushing the entire container to disk sequentially at once.

[0077] Furthermore, to handle overwrite or update operations on the same virtual slot, the compact reorganization unit does not perform in-situ memory updates. When a new write request for a virtual slot containing existing data is detected, the system still uses an append-only approach, writing the new version of the data to the end of the data storage area and updating the offset pointer of the corresponding slot in the dynamic micro-mapping unit to point to the new location. The old version of the data still exists in the data storage area, but becomes invalid because it has lost its reference in the mapping table.

[0078] Finally, when the data storage area of ​​the memory-buffered object is full, or a preset time threshold is reached, the container management module will freeze the current container state and mark it as pending persistence. This time threshold is set according to the system's recovery point target, typically ranging from 500 milliseconds to 30 seconds, to ensure that the maximum data loss in a power outage scenario is within a controllable range. At this time, the mapping table in the metadata index area accurately records the physical distribution of all valid data within the container at that moment. Subsequently, the system packages the data area and index area and passes them to the next-level persistence control module.

[0079] The persistence control module, as the interaction interface component between volatile memory data and the underlying non-volatile physical medium, is mainly configured to ensure the atomicity, consistency and durability of data writing. The persistence control module includes a version verification unit, a serialization encapsulation unit and an atomic flushing unit.

[0080] The version verification unit is used to handle data version conflicts that may occur due to concurrent writes from multiple clients. Its core working principle is based on an optimistic concurrency control mechanism. In a distributed storage environment, to avoid introducing high-latency distributed locks, the system assumes that conflicts are low-probability events, allowing concurrent writes, but performing strict version checks before final disk writes.

[0081] The version verification unit maintains a persistent global version index. Update operations to this global version index and disk persistence operations for container data are ensured atomicity through Write-Ahead Log (WAL) or atomic batch processing. This index is physically stored using a log-structured merged tree or B+ tree structure, recording the latest version number of all persisted virtual slots on the current physical node, and maintaining a cache of hot data in memory. When a container to be persisted is received from the container management module, the version verification unit traverses each index entry in the container's metadata index area and extracts the global version identifier. And compare it with the historical version identifier of that slot recorded in the current storage engine. A comparison is performed. For each data block to be written, the system executes the following write decision logic: ; In the formula, This indicates the operation instructions for the current data block; This indicates the global version identifier carried by the data block to be written in the current memory container; This indicates the version identifier for the same virtual slot that already exists on the storage medium. If there is no historical data for this slot in the system (i.e., the first write), then the default is used. It is 0.

[0082] This indicates that if the judgment result is "execute", the data block is marked as valid and will be... Update to the global version index; If the judgment result is to discard, it means that the data in the current memory is outdated "old data" (usually caused by network out-of-order delivery or retry mechanism). The system will directly remove the data block from the container, thus solving the distributed write conflict problem in a lock-free environment.

[0083] The serialization and encapsulation unit is connected to the version verification unit and is used to convert the verified memory container into a linear binary format suitable for disk storage. To ensure the atomicity of the disk write operation and prevent errors caused by system crashes or power outages, the serialization and encapsulation unit adopts a checksum-based packetization mechanism.

[0084] Specifically, the serialization encapsulation unit adds a container header to the beginning of the memory buffer object and a cyclic redundancy check (CRC) code to the end. The container header records the container's globally unique ID, creation timestamp, number of valid data blocks, and data layout descriptor. The CRC code at the end is generated using a hardware-accelerated CRC32C algorithm, which is a 32-bit checksum obtained by polynomial calculation of the entire container content (including the header, data area, and index area).

[0085] The atomic flush unit is used to perform the final disk I / O operation. This atomic flush unit adopts the Direct I / O mode, which bypasses the operating system's page cache by setting the file open flag and writes the serialized container data directly to the pre-allocated area of ​​the physical disk.

[0086] The atomic flush unit leverages the append-only write capabilities of the underlying file system to ensure that an I / O operation either succeeds completely or fails completely. After the write operation is complete, the atomic flush unit executes synchronization instructions (such as fsync or fdatasync), forcing the storage device's controller to flush the data in the volatile cache to the physical magnetic media or flash memory chips.

[0087] To verify data integrity, the atomic write unit performs lightweight verification after writing. If the system supports atomic write functionality, it directly uses the return code for verification; otherwise, it reads the checksum at the end of the just-written data and verifies whether it matches the one calculated in memory. If they do not match, it determines that a power outage has occurred, and the system discards the write operation and triggers an error recovery process.

[0088] To further optimize disk seek performance, the atomic brush unit employs a multi-stream merging sequential write strategy during physical disk write. Based on the physical partition characteristics of the disk, the system merges data streams from different containers into several large, consecutive write streams. In the specific implementation, the system sets a minimum write unit. and maximum waiting time .

[0089] It is usually set to align with the size of the underlying RAID stripe, such as 256KB or 512KB; The default value is typically set between 1 and 50 milliseconds. If the current container size has not reached the minimum disk write unit, the atomic flush unit will temporarily hold the container in the aggregation buffer, waiting for subsequent containers to be merged, until the data volume meets the minimum requirement. Or the waiting time exceeds This triggers a batch disk write.

[0090] Through the above mechanism, the persistence control module completely transforms logically random small I / O into physical sequential large I / O. At the same time, by using CRC checksum and version comparison mechanism, it achieves strong consistency data persistence without introducing distributed locks.

[0091] To more clearly illustrate the practical application process of this invention, the following description uses a specific scenario of "city-level high-definition video surveillance aggregation and storage" as an example.

[0092] Scenario: A city's security system comprises 10,000 high-definition cameras, each continuously generating H.264 encoded video streams at a bitrate of 4 Mbps. The system requires all video data to be written to a cloud storage cluster in real time and must support playback for at least 30 days.

[0093] Implementation process: The access node receives a video stream from camera ID Cam_001. When the accumulated video stream reaches the system's preset slice size... At that time, the slice calculation unit truncates it into the first slice. Data blocks. Assume the current absolute time is... The version management unit assigns a global version identifier to the data block. (high position) (The lower digit is the serial number).

[0094] Instead of querying the metadata server, the system directly extracts the file ID and current block number of Cam_001. Calculate object key Based on the current cluster topology, it is calculated that the data block should be stored in container_55 of physical node_19 and belong to virtual slot_88.

[0095] When a data block is transferred to Node_19, Container_55 may simultaneously receive data blocks from Cam_002 (mapped to Slot_12) and Cam_999 (mapped to Slot_4096). Although Slot_12, Slot_88, and Slot_4096 are logically very scattered (sparse), the container management module sequentially and compactly appends these three data blocks to a 64MB memory Slab Buffer and records them in an array in the metadata index area. Index

[12] ->Offset=0MB; Index

[88] ->Offset=4MB; Index

[4096] ->Offset=8MB.

[0096] When Container_55 is full or reaches the 30-second threshold, disk write is triggered. The system detected that Slot_88 has a previous version. (For example, old data caused by network retransmission). After comparison, the current... The system performs the write operation and calculates the CRC32C checksum of the entire 64MB container, appending it to the end. The entire container is then flushed to the physical sectors of the disk in one go via Direct I / O.

[0097] Experimental verification: Experimental description: This experiment verifies the performance of DHT storage or Ceph-type storage in this embodiment and the comparative scheme.

[0098] Experimental environment setup: 5 storage servers, each configured with an Intel Xeon Gold CPU (32 cores), 128GB DDR4 memory, and 4 10TB 7200RPM SATA enterprise-grade hard drives (HDDs). The network environment is 10 Gigabit Ethernet (10GbE).

[0099] Software variables: This embodiment enables logical slicing, deterministic mapping, and 64MB memory container aggregation. The contrasting solution disables container aggregation, writes each 4MB slice directly to disk as an independent object, and requires accessing a separate metadata service to obtain routing information before each write operation.

[0100] Test dataset: Simulates 100 to 10,000 concurrent clients continuously writing variable-length files ranging from 4MB to 128MB in size.

[0101] Analysis of experimental results: Experiment 1: Refer to Appendix Figure 4 , Figure 4 The curves show how the system's aggregate write throughput changes as the number of concurrent clients increases.

[0102] Results: In the low-concurrency phase (<500 clients), the performance difference between this solution and the comparison solution is not significant, both limited by network bandwidth. As concurrency increases (>1000 clients), the throughput of the comparison solution exhibits a clear "saturation oscillation" phenomenon, maintaining at approximately 1.2 GB / s. This is due to the frequent disk head seeks caused by massive small I / O operations and the increased response latency of the metadata server.

[0103] This embodiment exhibits a near-linear growth trend, reaching a peak of 3.8 GB / s at 5000 concurrent connections, close to the theoretical bandwidth limit of the physical disk. The container management module in this embodiment successfully reorganizes a large number of random I / O requests into a sequential stream in memory, enabling the disk to operate in a continuous write mode, and the deterministic mapping eliminates metadata interaction overhead.

[0104] Experiment 2: Refer to Appendix Figure 5 , Figure 5 It shows the probability density distribution of write request response time under full load conditions.

[0105] Results: The latency curve of this embodiment is extremely steep, with 99% of requests (P99) having a latency of less than 50ms. This indicates that the system's performance has extremely high determinism. The comparison scheme exhibits a significant "long tail effect," with approximately 15% of requests having a latency exceeding 500ms, and some reaching 2 seconds. Long tail latency is typically caused by metadata lock contention and disk I / O queue congestion. The lock-free design (resolving conflicts through version numbers) and batch disk write mechanism of this embodiment effectively eliminate these jitter sources.

Claims

1. A file storage management system for a cloud storage system, characterized in that, include: The access adaptation module is used to reorganize external write requests into a standardized data stream object that contains file paths and binary data carriers within the system. The metadata module is used to parse the file path in the data stream object to obtain a globally unique file identifier, providing a logical identifier for data addressing; The slice version module is used to discretize the data stream object into fixed-length data blocks and generate a global version identifier for each data block to identify the timing state. The deterministic mapping module is used to perform stateless computation based on the file identifier and data block index number to generate the target container identifier and virtual slot number; The container management module is used to build buffer containers in memory, compactly append data blocks to the data storage area according to the target container identifier, and establish a mapping relationship between virtual slot numbers and physical offsets in memory. The persistence control module is used to drive the physical disk write operation of valid data based on the comparison result between the global version identifier and the historical version in the physical storage medium. The deterministic mapping module includes an object key calculation unit; The object key calculation unit is used to perform a left shift operation on the data block index number, perform a bitwise XOR operation on the shift result and the file identifier, and apply a hybrid hash function to the operation result to generate a logical object key that is independent of physical location. The logical object key serves as the logical digital fingerprint of the data block, and is used to drive the topology-aware routing unit mapping calculation of the deterministic mapping module. The deterministic mapping module also includes a topology-aware routing unit and a virtual slot allocation unit; The topology-aware routing unit is used to calculate the target container identifier and the list of physical nodes carrying the target container based on the logical object key, combined with the cluster topology diagram with version number and preset data placement rules. The virtual slot allocation unit is used to perform a modulo operation on the data block index number for the maximum logical slot capacity preset for a single container, and determine the remainder obtained by the operation as the virtual slot number of the data block inside the container.

2. The file storage management system for a cloud storage system according to claim 1, characterized in that, The access adaptation module includes a buffer aggregation unit; The buffer aggregation unit is used to build a buffer layer in memory. By monitoring the total amount of data accumulated in the current cache area or the data residence time, it triggers a data flushing operation when a preset threshold is reached. The buffer aggregation unit, based on the triggered data flushing operation, concatenates multiple discrete small block write requests into a single continuous large block of data or encapsulates it into a vector I / O list to construct batch write requests.

3. The file storage management system for a cloud storage system according to claim 1, characterized in that, The metadata module includes a namespace management unit; The namespace management unit is used to employ a key-value pair-based distributed storage architecture. It calculates the logical index number of the target metadata shard node by calculating the bitwise XOR result of the globally unique identifier of the parent directory and the hash value of the file name string, and then applies a consistent hash function and a modulo operation on the bitwise XOR result to the total number of available shard nodes. This determines the physical storage location of the directory entry.

4. The file storage management system for a cloud storage system according to claim 1, characterized in that, The slice version module includes a version management unit; The version management unit is used to construct a global version identifier by combining the high-order physical timestamp and the low-order logical sequence number; The version management unit ensures, based on the incrementing logic of the atomic counter, that the global version identifier generated by the later write request for the same block index of the same file is numerically greater than the identifier of the earlier write request, so as to establish the global temporal order of the data blocks.

5. The file storage management system for a cloud storage system according to claim 1, characterized in that, The container management module includes a buffer allocation unit; The buffer allocation unit is used to pre-allocate a memory page pool using the Slab allocation mechanism and allocate independent memory buffer objects according to the target container identifier. The memory buffer object is constructed as two physically isolated areas: a data storage area for sequentially storing the file content byte stream and a metadata index area for recording the block mapping relationship inside the container.

6. The file storage management system for a cloud storage system according to claim 5, characterized in that, The container management module also includes a compact reorganization unit and a dynamic micro-mapping unit; The compact reorganization unit is used to perform the transformation from logically sparse to physically compact, ignoring the discreteness of the virtual slot number corresponding to the data block, and appending the data block to the physical location pointed to by the current write pointer of the data storage area. The dynamic micro-mapping unit is used to update the index entries in the mapping table with virtual slot number as the subscript according to the physical location after writing, and to construct a direct mapping from virtual slot number to memory physical offset, data effective length and global version identifier.

7. The file storage management system for a cloud storage system according to claim 1, characterized in that, The persistence control module includes a version verification unit; The version verification unit is used to traverse the index entries of the buffer container before writing to disk, extract the global version identifier of the data block to be written as the new version, and read the historical version of the same virtual slot on the storage medium. The version verification unit generates operation instructions based on the comparison logic: when the new version value is greater than the historical version, an execution write instruction is generated; when the new version value is less than or equal to the historical version, a discard instruction is generated and the data block is removed, so as to resolve concurrent write conflicts in a lock-free environment.

8. The file storage management system for a cloud storage system according to claim 7, characterized in that, The persistence control module also includes a serialization encapsulation unit and an atomic brush unit; The serialization encapsulation unit is used to add a container header and a cyclic redundancy check code based on polynomial calculation to the beginning and end of the verified memory buffer object, and encapsulate it into a linear binary format suitable for disk storage. The atomic brush unit is used to utilize the append write feature of the underlying file system to bypass the operating system page cache and directly write the encapsulated container data to the pre-allocated area of ​​the physical disk.

Citation Information

Patent Citations

  • Metadata management method in distributed object storage

    CN111258508A

  • Code-free workflow platform-oriented visual service function packaging method and system

    CN122308817A