Data processing method for block storage, electronic device and program product
By building a multi-level cache architecture that maps index trees, the mapping of logical block addresses to physical block addresses is unified, which solves the problem of redundant management in multi-layer virtualization environments, and improves I/O access efficiency and storage system stability.
Patent Information
- Application Number
- CN202510778418.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In multi-layer virtualization scenarios, the host, virtual machine/container and SSD firmware layers each maintain independent logical block address-to-physical block address mapping mechanisms, resulting in redundant logging, multiple garbage collections and cross-layer scheduling missing, causing write amplification, storage fragmentation and I/O performance fluctuations.
Build a mapping index tree, adopting a multi-level cache architecture that includes a cache layer, segmented memory table and multi-layer ordered table that is set in sequence, and uniformly manages the mapping relationship between logical block addresses and physical block addresses. Writes data to the cache layer or segmented memory table according to the storage priority of the data, and migrates data to the next level of storage space when the merge migration rules are met.
Implement unified scheduling of cross-layer storage mapping, reduce multi-layer redundancy management, improve I/O access efficiency, reduce write amplification effect, optimize storage space utilization and query efficiency, and improve system stability.
Smart Images

Figure CN120295944A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a data processing method, an electronic device, and a program product for block storage. Background Art
[0002] Block storage manages data in units of logical blocks and is widely used in high-performance computing, virtualization, and containerization environments. In a multi-layer virtualization scenario, the host, virtual machine / container, and SSD firmware layer each maintain independent logical block address to physical block address mapping mechanisms, resulting in redundant logging, multiple garbage collections, and lack of cross-layer scheduling, which in turn causes write amplification, storage fragmentation, and I / O performance fluctuations. Summary of the Invention
[0003] The present disclosure provides a data processing method, an electronic device, and a program product for block storage.
[0004] According to one aspect of the present disclosure, there is provided a data processing method for block storage, including: Constructing a mapping index tree for uniformly managing the mapping relationship between logical block addresses and physical block addresses, the mapping index tree being a multi-level cache architecture including a cache layer, a segmented memory table, and a multi-level ordered table arranged in sequence; Determining the storage priority of the data to be written according to the received data write request, the storage priority being used to represent the importance of the data to be written; Writing the data to be written into the cache layer or the segmented memory table for storage according to the storage priority; and When the stored data in the segmented memory table or the multi-level ordered table meets the merge migration rule, merging and migrating the stored data to the next-level storage space.
[0005] According to the technical solution of one aspect, by constructing a mapping index tree to uniformly manage the mapping relationship between logical block addresses and physical block addresses, the mapping index tree being a multi-level cache architecture including a cache layer, a segmented memory table, and a multi-level ordered table arranged in sequence. Then, according to the received data write request, determine the storage priority of the data to be written, the storage priority being used to represent the importance of the data to be written, and write the data to be written into the cache layer or the segmented memory table for storage according to the storage priority. And when the stored data in the segmented memory table or the multi-level ordered table meets the merge migration rule, merge and migrate the stored data to the next-level storage space.
[0006] Thus, based on the mapping index tree, unified scheduling of cross-layer storage mapping can be achieved, reducing multi-layer redundant management, and thereby improving I / O access efficiency.
[0007] Moreover, according to the storage priority of the data to be written, the data to be written is written into the cache layer or the segmented memory table for storage respectively. Non-urgent writes are allowed to be backlogged in the cache layer first to avoid frequent access to the memory table, thereby reducing the write amplification effect.
[0008] In at least one embodiment of the present disclosure, writing the data to be written into the cache layer or the segmented memory table for storage according to the storage priority includes: When the storage priority is the first level, writing the data to be written into the cache layer for storage; When the storage priority is the second level, writing the data to be written into the segmented memory table for storage, and the importance level of the second level is higher than that of the first level.
[0009] According to the technical solution of this embodiment, frequent access to the memory table can be avoided, and the write amplification effect can be reduced.
[0010] In at least one embodiment of the present disclosure, when the stored data in the segmented memory table or the multi-level ordered table meets the merge migration rule and the stored data is merged and migrated to the next-level storage space, the method includes: Determining a first ratio of the stored data in the segmented memory table to its own total storage capacity; Merging and migrating the stored data in the segmented memory table with the first ratio greater than or equal to the first capacity threshold to the multi-level ordered table.
[0011] According to the technical solution of this embodiment, the merge migration operation can be independently triggered when a single segmented memory table is full, and other segmented memory tables can continue to receive new write requests, avoiding blocking.
[0012] In at least one embodiment of the present disclosure, the method further includes: Obtaining system load information; Determining the first capacity threshold according to the system load information.
[0013] According to the technical solution of this embodiment, the first capacity threshold can be dynamically adjusted based on the real-time system load, thereby reducing write amplification, reducing I / O oscillation, and improving system stability.
[0014] In at least one embodiment of the present disclosure, when the stored data in the segmented memory table or the multi-level ordered table meets the merge migration rule and the stored data is merged and migrated to the next-level storage space, the method further includes: For the multi-level ordered table, when the number of ordered tables at the L n level reaches the number threshold, and L nWhen the storage data in the ordered list of each level accounts for a second ratio of its own total storage capacity that is greater than or equal to the second capacity threshold, at least some of the ordered lists are selected from the L n level as the ordered lists to be merged, where n is an integer ≥ 0; Merge the storage data of the ordered lists to be merged at the L n level with the storage data of the ordered lists to be merged at the L n+1 level, and perform merge sorting in the order of logical block addresses to obtain a temporary ordered list. Among them, the ordered lists to be merged at the L n+1 level are the ordered lists whose key value ranges intersect with the key value ranges of the ordered lists to be merged at the L n level; Write the temporary ordered list into the L n+1 level, and discard the ordered lists to be merged at the L n level and the L n+1 level.
[0015] According to the technical solution of this embodiment, by merging the ordered lists whose key value ranges intersect, duplicate data storage can be reduced, thereby improving space utilization. Moreover, the merged ordered lists are continuously distributed in the order of logical block addresses, which can improve query efficiency.
[0016] In at least one embodiment of the present disclosure, the method further includes: Determine the data failure ratio corresponding to the ordered list according to the number of invalid key value pairs and the total number of key value pairs in the ordered list; When the data failure ratio reaches the failure ratio threshold, merge and migrate the storage data in the ordered list to the ordered list of the next lower level.
[0017] According to the technical solution of this embodiment, resources can be effectively recycled and the overhead of reading invalid records can be reduced. Partial merging ensures that the multi-level ordered lists are not occupied by a large amount of invalid data or outdated versions for a long time.
[0018] In at least one embodiment of the present disclosure, the mapping index tree is integrated into the kernel block layer to eliminate duplicate log records and mapping management.
[0019] According to the technical solution of this embodiment, duplicate log records and mapping management between the original layers can be effectively eliminated, avoiding waste of storage resources and write amplification caused by redundant indexes and frequent updates.
[0020] In at least one embodiment of the present disclosure, the method further includes: Determine the logical block address to be queried according to the received data query request; Query in each storage space in the order of the cache layer, the segmented memory table, and the multi-level ordered table according to the logical block address; Feedback the physical block address corresponding to the logical block address according to the query result.
[0021] According to the technical solution of this embodiment, data query based on the mapping index tree can improve the query efficiency.
[0022] According to another aspect of the present disclosure, a data processing device for block storage is provided, including: A construction module for constructing a mapping index tree, which is used to uniformly manage the unified mapping relationship between the logical block address and the physical block address. The mapping index tree is a multi-level cache architecture including a cache layer, a segmented memory table, and a multi-level ordered table arranged in sequence; A determination module for determining the storage priority of the data to be written according to the received data write request, where the storage priority is used to characterize the importance of the data to be written; A write module for writing the data to be written into the cache layer or the segmented memory table for storage according to the storage priority; and, A processing module for merging and migrating the stored data to the next-level storage space when the stored data in the segmented memory table or the multi-level ordered table meets the merge migration rule.
[0023] According to another aspect of the present disclosure, an electronic device is provided, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the data processing method for block storage according to any embodiment of the present disclosure.
[0024] According to still another aspect of the present disclosure, a readable storage medium is provided, in which execution instructions are stored, and when the execution instructions are executed by a processor, they are used to implement the data processing method for block storage according to any embodiment of the present disclosure.
[0025] According to yet another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the data processing method for block storage according to any embodiment of the present disclosure. Description of the Drawings
[0026] The drawings illustrate exemplary embodiments of the present disclosure and are used together with the description to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.
[0027] Figure 1The flowchart shows a data processing method for block storage according to an embodiment of the present disclosure.
[0028] Figure 2 The structural schematic diagram of a mapping index tree according to an embodiment of the present disclosure is shown.
[0029] Figure 3 The structural schematic diagram of a system applicable to the embodiment of the present disclosure is shown.
[0030] Figure 4 The flowchart shows the flowchart of step S130 in the data processing method for block storage according to an embodiment of the present disclosure.
[0031] Figure 5 The flowchart shows the flowchart of step S140 in the data processing method for block storage according to an embodiment of the present disclosure.
[0032] Figure 6 The flowchart shows the flowchart of determining a first capacity threshold further included in the data processing method for block storage according to an embodiment of the present disclosure.
[0033] Figure 7 The flowchart shows the flowchart of triggering the merge and migration operation of a segmented memory table in the data processing method for block storage according to an embodiment of the present disclosure.
[0034] Figure 8 The flowchart shows the flowchart of triggering the merge and migration operation of a multi-layer ordered table further included in step S140 in the data processing method for block storage according to an embodiment of the present disclosure.
[0035] Figure 9 The flowchart shows the flowchart of triggering the partial merge and migration operation of a multi-layer ordered table further included in step S140 in the data processing method for block storage according to an embodiment of the present disclosure.
[0036] Figure 10 The flowchart shows the flowchart of data query further included in the data processing method for block storage according to an embodiment of the present disclosure.
[0037] Figure 11 The technical architecture diagram applicable to the embodiment of the present disclosure is shown.
[0038] Figure 12 The flowchart shows a data processing method for block storage according to another embodiment of the present disclosure.
[0039] Figure 13 The block diagram shows a data processing device for block storage according to an embodiment of the present disclosure.
[0040] Figure 14 It is a schematic structural block diagram of an electronic device configured with a data processing device for block storage according to an embodiment of the present disclosure. Detailed implementation manners
[0041] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and examples. It can be understood that the specific examples described herein are only for explaining the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of convenience of description, only the parts related to the present disclosure are shown in the accompanying drawings.
[0042] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0043] Block storage is a storage technology that manages and reads and writes data in units of fixed-size data blocks (Blocks). Its core mechanism is to achieve direct access to the storage medium through the mapping relationship between the logical block address (Logical Block Address, LBA) and the physical block address (Physical Block Address, PBA). Block storage has been widely used in high-performance computing (HPC), large-scale model training, virtualization platforms and other fields due to its efficient data management capabilities, good compatibility and flexible deployment methods.
[0044] By managing and reading and writing data in units of blocks (Blocks), block storage can provide a unified storage interface between different systems and applications, thereby reducing the adaptation overhead of the file system or the distributed storage layer. In the high-performance computing (HPC) scenario, block storage is usually used as the underlying support for parallel file systems (such as Lustre, GPFS) to provide high-throughput data access capabilities for large-scale scientific computing. In large-scale deep learning model training, the massive parameter updates and high-concurrency I / O requests pose extreme requirements of low latency and high throughput on the storage system. In virtualization and containerization environments, due to the multi-tenant requirements and the elastic expansion of computing resources, the efficient management and isolation of block devices become particularly critical.
[0045] Meanwhile, multi-layer virtualization technologies (such as KVM, Docker) are also widely used in cloud computing and large-scale data centers, providing resource isolation, elastic scaling, and efficient management capabilities. However, multi-layer virtualization introduces multiple data management layers with similar functions but independent of each other, including the host operating system, the block device mapping mechanism inside virtual machines or containers, and the Flash Translation Layer (FTL) of the underlying storage device (such as SSD). Each of these layers maintains an independent mapping from logical block number (LBA) to physical block address (PBA), resulting in increased redundancy in data management and causing problems such as write amplification, storage fragmentation, and decreased I / O performance, thus affecting the overall storage efficiency and lifespan of the system.
[0046] Therefore, the present disclosure proposes the following technical solutions. In this technical solution, based on the mapping index tree, unified scheduling of cross-layer storage mapping can be achieved, reducing multi-layer redundant management, and thus improving I / O access efficiency. Moreover, according to the storage priority of the data to be written, the data to be written is respectively written into the cache layer or the segmented memory table for storage, allowing non-urgent writes to be buffered in the cache layer first, avoiding frequent access to the memory table, and thus reducing the write amplification effect.
[0047] Figure 1 The flowchart of the data processing method for block storage according to an embodiment of the present disclosure is shown. As Figure 1 shown, the method at least includes steps S110 to S140, which are introduced in detail as follows: In step S110, a mapping index tree is constructed. The mapping index tree is used to uniformly manage the unified mapping relationship between the logical block address and the physical block address. The mapping index tree is a multi-level cache architecture including a cache layer, a segmented memory table, and a multi-level ordered table arranged in sequence.
[0048] In this embodiment, the mapping index tree (Mapping Tree, M-Tree) can be a multi-level cache architecture integrated into the operating system, which is used to uniformly manage the mapping relationship between the logical block address and the physical block address. Its core structure includes a cache layer, a segmented memory layer, and a multi-level ordered table arranged in sequence.
[0049] Specifically, as Figure 2 shown, the cache layer can be implemented using a lock-free circular buffer, supporting multi-threaded concurrent writing, and is used to temporarily store newly generated LBA→PBA mapping relationships or write requests to be processed. That is to say, when a write request is received, the LBA→PBA mapping can be temporarily stored in the cache layer, thus avoiding frequent I / O operations caused by directly writing to the persistent structure. Thus, by aggregating small-scale random write requests through the cache layer, the direct impact on the segmented memory table can be reduced, significantly reducing the write latency and improving the throughput in high-concurrency scenarios.
[0050] The segmented memory table maps LBAs to multiple independent segments through hash sharding. Inside each segment, a skip list structure can be used to maintain an ordered LBA→PBA mapping. In one example, the LBA can be hashed and modulo-operated to be assigned to a specific segment. The segments in the segmented memory table do not interfere with each other, improving concurrency. At the same time, hash sharding can achieve load balancing, enabling concurrent writing and querying and avoiding global lock contention.
[0051] The multi-layer ordered table consists of multiple persistent structures (such as SSTable) sorted by LBA. Each layer can be attached with a Bloom Filter (BF) to accelerate queries. Data is flushed from the segmented memory table into the highest layer (i.e., L0 layer) of the multi-layer ordered table to form persistent storage. And when the number of ordered tables in a certain layer or the proportion of invalid data exceeds a threshold (please refer to the following), a Compaction operation can be triggered, that is, integrating and sorting the data in adjacent layers and sinking them to the next layer, and cleaning up invalid records. In this way, through hierarchical storage and dynamic merging, redundant data writing is reduced (write amplification is reduced) and storage space utilization is optimized. Secondly, the Bloom Filter can quickly filter invalid query paths and improve retrieval efficiency; furthermore, some merging strategies reduce I / O resource occupancy and alleviate performance jitter.
[0052] In some embodiments of the present disclosure, the mapping index tree can be integrated into the kernel block layer of the operating system. The kernel block layer is an important intermediate layer in the operating system for processing block device I / O requests, such as Figure 3 As shown, read and write requests submitted by the upper layer (such as file system, user-space application) are uniformly connected to the kernel block layer. After queuing, scheduling, and forwarding, the data access operation is finally completed. By providing a unified abstract interface and adaptation function, the kernel block layer ensures stable and efficient access to storage devices by upper-layer applications.
[0053] Compared with placing the mapping index tree in a higher layer (file system or user space) or a lower layer (such as SSD firmware FTL), since the kernel block layer is located at a key node in the data flow path, it naturally has the advantage of implementing cross-layer unified mapping scheduling and can cover the multi-layer data management requirements of the host, virtual machine / container, and storage medium. Therefore, setting the mapping index tree in the kernel block layer can effectively eliminate duplicate log records and mapping management between the original layers, and avoid storage resource waste and write amplification caused by redundant indexes and frequent updates.
[0054] At the same time, the unified mapping tree can more efficiently identify hot data, stratify cold data, and optimize data local merging, so as to achieve load adaptive adjustment in high-concurrency scenarios and better balance performance stability and device life extension.
[0055] Please continue to refer to Figure 1, in step S120, according to the received data writing request, determine the storage priority of the data to be written, where the storage priority is used to characterize the importance of the data to be written.
[0056] Among them, the storage priority can be used to characterize the importance or urgency of the data to be written. In one example, those skilled in the art can pre-assign corresponding storage priorities to different data according to the importance of the data. For example, large sequential writes, metadata updates, display synchronization requests, crash consistency guarantee operations (such as transaction log submissions), etc. can be determined as high priority levels. Small-scale random writes (such as container temporary logs, edge computing intermediate results, etc.) are determined as low priority levels to allow delayed writes to aggregate batch operations.
[0057] In this embodiment, after receiving the data writing request, the data writing request can be parsed to determine the storage priority corresponding to the data to be written. In one example, it can be determined according to the metadata in the data writing request (such as I / O type, data size, source application, etc.) to determine the corresponding storage priority for the data to be written. For example, according to the received data writing request, if the bi_opf field included in the struct bio in the data writing request is displayed as REQ_META or REQ_SYNC, it indicates that this write belongs to metadata, log writing, or explicit synchronization request, which is a write with relatively high urgency, so it can be marked as high priority.
[0058] In step S130, according to the storage priority, write the data to be written into the cache layer or the segmented memory table for storage.
[0059] In this embodiment, based on the determined storage priority, the data to be written with low priority can be temporarily stored in the cache layer, while the data to be written with high priority is directly written into the segmented memory table. In this way, dynamically selecting the write path according to the storage priority of the data to be written can achieve efficient allocation of I / O resources and system performance optimization.
[0060] In step S140, when the stored data in the segmented memory table or the multi-level ordered table meets the merge and migration rule, merge and migrate the stored data to the next-level storage space.
[0061] Among them, the merge and migration rule can be a determination rule pre-determined by those skilled in the art according to prior experience for triggering data merging. In one example, this merge and migration rule can be based on the system load and the amount of stored data as the determination basis for whether to start data merging.
[0062] In this embodiment, the combined migration of data may occur at two locations. One is the combined migration from the segmented inner-layer table to the top layer (i.e., L0 layer) of the multi-layer ordered table, and the other is the combined migration from the L n layer ordered table in the multi-layer ordered table to the L n+1 layer ordered table.
[0063] In one example, the segmented in-memory table is written sequentially into the L0 layer of the multi-layer ordered table to form an ordered table. There is a quantity limit for the number of ordered tables in each layer. When the limit is exceeded, compaction is triggered. After compaction, a new file is generated and placed in the next layer, and so on until the data is completely cooled or no longer updated frequently. In this way, newer or popular data often stays in the upper layers (such as L0 layer, L1 layer), while cold data sinks to deeper layers.
[0064] It should be noted that for the combination of the segmented in-memory table to the L0 layer of the multi-layer ordered table, the purpose is to quickly release the cache space so that new data can quickly enter persistent storage. Therefore, when the capacity of a certain segmented in-memory table reaches a certain threshold, the combined migration of this segmented in-memory table is triggered, and other segments can continue to accept new requests, thus preventing the combined migration process from blocking the overall write throughput. That is, through segmented flushing, it is possible to avoid blocking all writes when a single large in-memory table reaches its upper limit, and improve the scalability and throughput in a high-concurrency environment.
[0065] For the combination from the L n layer to the L n+1 layer in the multi-layer ordered table, it is mainly to reclaim space and remove invalid data, and not much adjustment space is required. Based on the combined migration rules, hot data with high-frequency updates can be retained in the upper layers (such as L0 layer, L1 layer), and cold data with low-frequency access sinks to the deeper layers (such as L2 layer, L n layer), so as to adapt to the hybrid load scenario.
[0066] In this way, based on Figure 1In the illustrated embodiment, by constructing a mapping index tree to uniformly manage the mapping relationship between the logical block address and the physical block address, the mapping index tree is a multi-level cache architecture including a cache layer, a segmented memory table, and a multi-level ordered table arranged in sequence. Next, according to the received data write request, determine the storage priority of the data to be written, where the storage priority is used to characterize the importance of the data to be written. According to the storage priority, write the data to be written into the cache layer or the segmented memory table for storage. Moreover, when the stored data in the segmented memory table or the multi-level ordered table meets the merge migration rule, migrate the stored data to the next-level storage space for merging. Thus, based on the mapping index tree, unified scheduling of cross-layer storage mapping can be achieved, reducing multi-level redundant management, and thereby improving the I / O access efficiency. Also, according to the storage priority of the data to be written, write the data to be written into the cache layer or the segmented memory table for storage respectively, allowing non-urgent writes to be queued in the cache layer first, avoiding frequent access to the memory table, and thus reducing the write amplification effect.
[0067] Figure 4 FIG. shows a schematic flowchart of step S130 in a data processing method for block storage according to an embodiment of the present disclosure. As Figure 4 shown, step S130 at least includes steps S131 to S132, which are introduced in detail as follows.
[0068] In step S131, when the storage priority is the first level, write the data to be written into the cache layer for storage.
[0069] In this embodiment, the first level may be a low priority level, which is applicable to write requests that allow short delays but require high throughput, such as container temporary logs, edge computing intermediate results, etc. After determining that the storage priority of the data to be written is the first level, the data to be written can be directly allocated to the cache layer for storage. In one example, the cache layer can locally aggregate data according to the logical block address range, combining multiple small random writes into a batch operation, thereby reducing the number of single I / O operations.
[0070] In step S132, when the storage priority is the second level, write the data to be written into the segmented memory table for storage, where the importance of the second level is higher than that of the first level.
[0071] In this embodiment, the importance level of the second level is higher than that of the first level, that is, the second level can be a high priority level. It should be understood that the storage priority is determined for the data to be written at the second level, indicating that it has high importance or strong real-time requirements and needs to be preferentially guaranteed to be quickly written to disk. Therefore, the data to be written with the storage priority of the second level can be directly written into the segmented memory table for storage, so as to ensure the real-time performance and reliability of high-importance data.
[0072] In one example, when writing the data to be written into the segmented memory table, a hash operation can be performed based on the logical block address to determine the corresponding segmented memory table number (such as segmented memory table 3). Then, the LBA→PBA mapping to be written is inserted into the skip list structure in the segmented memory table corresponding to the segmented memory table number in ascending order, so as to ensure the query efficiency.
[0073] Thus, the data of the first level is written into the cache layer, and the data of the second level bypasses the cache layer and is directly written into the segmented memory table, which can allow non-urgent writes to be backlogged in the cache layer first, avoid frequent access to the memory table, and thus reduce the write amplification effect.
[0074] Figure 5 The flowchart of step S140 in the data processing method for block storage according to an embodiment of the present disclosure is shown. As Figure 5 shown, this step S140 at least includes steps S141 to S142, which are introduced in detail as follows.
[0075] In step S141, determine a first ratio of the stored data in the segmented memory table to its own total storable capacity.
[0076] In this embodiment, for each segmented memory table, the amount of data stored in the segmented memory table and the total storable capacity corresponding to the segmented memory table can be obtained. Then, by dividing the amount of data stored in the segmented memory table by its own total storable capacity, the corresponding first ratio can be obtained. For example, if the total storable capacity of the segmented memory table is 1000 mapping records and 800 have been stored currently, the corresponding first ratio is 80%.
[0077] In one example, the total storable capacity of each segmented memory table can be fixedly allocated, or it can be allocated proportionally according to the total system memory (i.e., dynamic capacity), and the present disclosure does not make special limitations on this.
[0078] In step S142, the stored data in the segmented memory table with the first ratio greater than or equal to the first capacity threshold is merged and migrated to the multi-layer ordered table.
[0079] In this embodiment, the first capacity threshold may be a pre-set ratio threshold for triggering the merging and migration of stored data in the segmented memory table. For example, the first capacity threshold may be 70% or 80%, etc.
[0080] After determining the first ratio corresponding to the segmented memory table, the first ratio can be compared with the first capacity threshold. If the first ratio threshold is greater than or equal to the first capacity threshold, the merging and migration operation of the segmented memory table can be triggered. That is, the stored data in the segmented memory table is merged and migrated to the highest layer (i.e., L0 layer) of the multi-layer ordered table.
[0081] Therefore, when the first ratio corresponding to a single segmented memory table reaches the first capacity threshold, the merging and migration operation of the stored data in the segmented memory table can be triggered, and other segmented memory tables can continue to receive new write requests, thus avoiding blocking all writes after a single large memory table reaches its upper limit and improving the scalability and throughput in a high-concurrency environment.
[0082] Based on the foregoing embodiment, Figure 6 shows a schematic flow chart of determining the first capacity threshold further included in the data processing method for block storage according to an embodiment of the present disclosure. As Figure 6 shown, determining the first capacity threshold includes at least step S143 to step S144, which are introduced in detail as follows.
[0083] In step S143, system load information is obtained.
[0084] Among them, the system load information may be parameter information for characterizing the usage status of system resources, which may include but is not limited to at least one of the CPU utilization rate of the system (for example, the percentage of the non-idle state of the CPU within a unit time), the I / O queue depth (the number of I / O requests to be processed by the storage device, measuring the congestion degree of the storage layer), and the memory utilization rate (the ratio of the occupied memory to the total memory capacity, indicating the tightness of memory resources).
[0085] In this embodiment, the above system load information can be periodically collected for subsequent determination of the first capacity threshold.
[0086] In step S144, according to the system load information, the first capacity threshold is determined.
[0087] In this embodiment, the first capacity threshold corresponding to the segmented memory table can be dynamically adjusted according to the collected system load information. It should be noted that the greater the system load pressure, the greater the first capacity threshold, so that the merging and migration process can be delayed to ensure the read and write priorities of the foreground. On the contrary, if the system load pressure is smaller, the first capacity threshold is smaller, so as to quickly clean up the memory and accelerate the speed of new data entering persistent storage.
[0088] In one embodiment, based on the collected system load information, the overall system load level can be determined. For example, weighted calculation can be performed according to the collected system load information to obtain the corresponding load score. Then, the load score is compared with a preset load threshold, and the corresponding first capacity threshold is determined according to the comparison result.
[0089] For example, when the system load score < 30%, the first capacity threshold for triggering the merge migration can be reduced to 60%, so as to quickly clean up the memory and accelerate the speed of new data entering the persistent storage. When the system load score > 80%, the first capacity threshold for triggering the merge migration can be increased to 90% to delay the merge migration process to ensure the foreground I / O priority.
[0090] Based on the foregoing embodiment, Figure 7 FIG. shows a schematic flow chart of triggering the merge migration operation of the segmented memory table in the data processing method for block storage according to an embodiment of the present disclosure. As Figure 7 shown, triggering the merge migration operation of the segmented memory table at least includes the following steps: In step S710, the system load information is periodically collected.
[0091] That is to say, according to the set period, the load information of the system is collected to provide a basis for subsequent threshold adjustment and merge trigger decision-making.
[0092] In step S720, it is determined whether the system load < 30%, that is, according to the collected system load information, the system load level is quantified, and it is determined whether the current system load is lower than 30%.
[0093] In step S730, if the system load is lower than 30%, the first capacity threshold is adjusted to 60%.
[0094] In step S740, if the system load is greater than or equal to 30%, it is further determined whether the system load is greater than 80%.
[0095] In step S750, if the system load is less than or equal to 80%, the first capacity threshold is adjusted to 70%.
[0096] In step S760, if the system load is greater than 80%, the first capacity threshold is adjusted to 90%.
[0097] In step S770, according to the determined first capacity threshold, it is checked whether the first ratio corresponding to a certain segmented memory table reaches the first capacity threshold.
[0098] In step S780, if the first ratio corresponding to the segmented memory table reaches the first capacity threshold (i.e., ≥ the first capacity threshold), a merge migration to layer L0 of the multi-level ordered table is triggered.
[0099] In step S790, if the first ratio corresponding to the segmented memory table does not reach the first capacity threshold, wait for the next cycle, restart the periodic collection of system load information, and then perform subsequent judgments and operations.
[0100] It should be noted that the above numbers are only exemplary examples. Those skilled in the art can set corresponding first capacity thresholds for different load conditions according to specific implementation requirements, and the present disclosure does not make special limitations on this.
[0101] Figure 8 The flowchart shows the process of triggering the merge migration operation of the multi-level ordered table included in step S140 in the data processing method for block storage according to an embodiment of the present disclosure. As Figure 8 shown, triggering the merge migration operation of the multi-level ordered table includes at least steps S145 to S147, which are introduced in detail as follows.
[0102] In step S145, for the multi-level ordered table, when the number of ordered tables at level L n reaches the number threshold, and the second ratio of the stored data in the ordered tables at level L n to their own total storable capacity is greater than or equal to the second capacity threshold, at least some ordered tables are selected from level L n as the ordered tables to be merged, where n is an integer ≥ 0.
[0103] Among them, the number threshold can be a pre-set upper limit of the number of ordered tables at level L n used to determine whether to trigger the merge migration. It should be understood that the number thresholds corresponding to different levels can be different. For example, the number threshold corresponding to the ordered tables at level L0 is 1, and the number threshold corresponding to the ordered tables at level L1 is 2, and so on.
[0104] The second capacity threshold is the lower limit of the second ratio for the ordered table to trigger the merge migration operation.
[0105] In this embodiment, the system can periodically count the number of ordered tables at level L n If it reaches the corresponding number threshold (i.e., greater than or equal to the number threshold), it can enter the capacity ratio verification stage. In this stage, the system can traverse all the ordered tables at level L n and calculate the second ratio corresponding to each ordered table. This second ratio is the ratio of the stored data in the ordered table to its own total storable capacity.
[0106] If L nIf the second ratio corresponding to all the ordered lists in the level is greater than the second capacity threshold (e.g., 70%), the merge and migration operation of the ordered list is triggered. Part of the ordered lists can be selected from the ordered lists in the L n level as the ordered lists to be merged. For example, 1 or 2 ordered lists can be randomly selected from the L n level as the ordered lists to be merged for merge and migration.
[0107] In step S146, the stored data of the ordered lists to be merged in the L n level and the stored data of the ordered lists to be merged in the L n+1 level are merged and sorted in the order of the logical block addresses to obtain a temporary ordered list. Among them, the ordered lists to be merged in the L n+1 level are the ordered lists whose key value ranges intersect with the key value ranges of the ordered lists to be merged in the L n level.
[0108] In this embodiment, ordered lists whose key value ranges overlap with the key value ranges of the ordered lists to be merged in the L n+1 level can be selected in the L n level as the ordered lists to be merged in the L n+1 level. For example, the key value range of the ordered lists to be merged in the L n level covers LBA 0x1000 - 0x2000, while the key value range of the ordered lists to be merged in the L n+1 level covers LBA 0x1800 - 0x2800.
[0109] After determining the ordered lists to be merged in the L n level and the L n+1 level, the stored data in the ordered lists to be merged in the two levels can be merged and sorted in the order of the logical block addresses to obtain a temporary ordered list.
[0110] In step S147, the temporary ordered list is written into the L n+1 level, and the ordered lists to be merged in the L n level and the L n+1 level are discarded.
[0111] In this embodiment, after obtaining the temporary ordered list, the temporary ordered list can be written into the L n+1 level, and several ordered lists (i.e., the ordered lists to be merged determined in the two levels mentioned above) designed in the previous merge and migration operation in the L n level and the L n+1 level are discarded.
[0112] In this way, by merging ordered lists with overlapping key-value ranges, duplicate data storage can be reduced, thereby improving space utilization. Moreover, the merged ordered lists are continuously distributed in the order of logical block addresses, which can improve query efficiency.
[0113] Based on the foregoing embodiments, Figure 9 FIG. shows a schematic flowchart of triggering a partial merge and migration operation of a multi-layer ordered list included in step S140 in a data processing method for block storage according to an embodiment of the present disclosure. As Figure 9 shown, triggering the partial merge and migration operation of the multi-layer ordered list includes at least steps S148 to S149, which are introduced in detail as follows.
[0114] In step S148, according to the number of invalid key-value pairs and the total number of key-value pairs in the ordered list, determine the data invalidation ratio corresponding to the ordered list.
[0115] In this embodiment, each ordered list can maintain a data invalidation ratio (number of invalidations / total number of key-value pairs). The number of invalid key-value pairs is obtained during previous merges. Each newly written (LBA, PBA) from layer L n is the latest data. The key-value pairs corresponding to the original LBA in L n+1 are recorded as invalid. Thus, the number of invalidations of each ordered list is obtained, and combined with the total number of key-value pairs, the data invalidation ratio is calculated.
[0116] In step S149, when the data invalidation ratio reaches the invalidation ratio threshold, merge and migrate the stored data in the ordered list to the ordered list of the next lower level.
[0117] In this embodiment, those skilled in the art can preset an invalidation ratio threshold (for example, 50%) in advance as the upper limit of the data invalidation ratio for triggering the partial merge and migration operation. The system can determine the data invalidation ratio of each ordered list at a certain period and compare the data invalidation ratio with the invalidation ratio threshold. If the data invalidation ratio of a certain ordered list reaches the invalidation ratio threshold (that is, greater than or equal to the invalidation ratio threshold), the stored data of this ordered list can be separately merged downward to the lower level. The steps of this merge operation can refer to the foregoing embodiments and will not be elaborated here.
[0118] In this way, by detecting and promptly sorting out the ordered list with a high failure ratio, resources can be effectively recycled, and the overhead of reading invalid records can be reduced. Partial merging ensures that the multi-layer ordered list is not occupied by a large amount of invalid data or outdated versions for a long time. It should be understood that in the multi-layer ordered list, the data capacity of the upper layer is smaller, the data stored in it is more "fresh", and the corresponding data failure ratio is smaller. And through the above-mentioned method of partial merging and migration, space can be vacated in the middle layer of the multi-layer ordered list, avoiding the multi-layer chain merging caused by the upper layer merging and migrating to the middle and lower layers. Thus, the overall performance and storage efficiency of the system can be maintained in high-concurrency and frequent-update scenarios.
[0119] Specifically, for high-concurrency small random write and mixed load scenarios, the embodiments of the present disclosure propose an adaptive merging mechanism, which can dynamically adjust the merging strategy, and in combination with the partial merging mechanism, can reduce write amplification and reduce I / O oscillation, improving the stability of the storage system.
[0120] Figure 10 The flowchart of data query included in the data processing method for block storage according to an embodiment of the present disclosure is shown. As Figure 10 shown, the data query includes at least steps S150 to S170, which are introduced in detail as follows.
[0121] In step S150, according to the received data query request, determine the logical block address to be queried.
[0122] In this embodiment, the data query request may be information used to request to query the physical block address corresponding to a certain logical block address. After receiving the data query request, the data query request can be parsed to obtain the logical block address to be queried.
[0123] In step S160, according to the logical block address, query in each storage space in the order of the cache layer, the segmented memory table, and the multi-layer ordered list.
[0124] In this embodiment, when querying, query can be performed in each storage space in the order of the cache layer, the segmented memory table, and the multi-layer ordered list. That is, first query the cache layer, if a hit occurs, feedback the result, if not, then query the segmented memory table. Among them, before querying the segmented memory table, the segment corresponding to the physical block address corresponding to the logical block address to be queried can be determined first, so as to determine the segmented memory table corresponding to this segment for query, thereby reducing unnecessary queries.
[0125] If a hit occurs in the segmented memory table, feedback the result, if not, query the multi-layer ordered list layer by layer. And when querying the multi-layer ordered list, a Bloom filter can be used to skip non-target tables, thereby accelerating the query speed.
[0126] In step S170, according to the query result, the physical block address corresponding to the logical block address is fed back.
[0127] In this embodiment, after querying in the mapping index tree, based on the query result, the physical block address corresponding to the queried logical block address can be fed back to the user or the system. In this way, the query efficiency can be improved and the query latency can be reduced.
[0128] Based on the technical solution of the above embodiment, a specific application scenario of the embodiment of the present application is introduced below: The present disclosure proposes an efficient block storage data organization method applicable to a multi-layer virtualization environment, called AdaptiDisk. Its core goal is to achieve efficient mapping index management, dynamic data merging optimization, cache optimization, and I / O resource scheduling at the block layer to reduce storage management redundancy and improve the overall performance of the storage system. As Figure 11 shown, the AdaptiDisk mainly includes three parts: an M-Tree mapping index, an adaptive merge scheduler, and a multi-level cache architecture, which are explained in detail as follows.
[0129] User space / virtualization layer 1110: This layer is the upper-layer application or virtualization management tool in the virtualization environment, responsible for initiating read and write requests, and is the entry of the entire block storage system.
[0130] Kernel block layer 1120: As a key intermediate layer in the operating system for processing block device I / O requests, it receives, queues, schedules, and forwards read and write requests submitted by the upper layer (such as file systems, user-space applications), ensuring stable and efficient access to storage devices by upper-layer applications. And it provides a basis for the mapping index management, dynamic data merging optimization, cache optimization, and I / O resource scheduling of the lower layer.
[0131] M-Tree mapping index 1130: Integrated into the kernel block layer, it is used to manage the mapping of logical block numbers (LBAs) to physical block numbers (PBAs), realizing unified scheduling of cross-layer storage mapping to reduce multi-layer redundant management and improve I / O access efficiency.
[0132] Cache layer (circular buffer) 1131: Adopting a lock-free circular buffer, supporting multi-threaded concurrent writing, it is used to receive the latest LBA→PBA mapping or write requests, temporarily aggregate a large number of small random writes, reduce the number of direct writes to the segmented memory table, relieve the impact pressure of foreground random writes, and write to the segmented memory table when the buffer layer is full.
[0133] Segmented Memory Table (Skip List) 1132: Implemented using a skip list and designed to be multi-segmented. By performing a hash function operation and a modulo operation on the LBA, it determines which segment to place in. The segments of the segmented memory table do not interfere with each other, improving concurrency. At the same time, hash sharding enables load balancing and is beneficial for range queries.
[0134] Multi-level Ordered Table (SSTable) 1133: Organized in a hierarchical structure, persistently stores the mapping relationships flushed from the segmented memory table. The data in each layer is sorted by LBA. Each level of the ordered table contains a corresponding Bloom Filter to quickly locate data.
[0135] Adaptive Merge Scheduler 1140: For high-concurrency small random write and mixed load scenarios, dynamically adjusts the merge strategy to reduce write amplification and reduce I / O oscillation, improving the stability of the storage system.
[0136] Adaptive Trigger Strategy 1141: By collecting the CPU load rate, I / O queue depth, and memory usage, it periodically evaluates the memory usage rate and the real-time load of the system, and dynamically modifies the trigger threshold according to the real-time situation.
[0137] Partial Merge Strategy 1142: In the multi-level ordered table, when merging from level L n to level L n+1 mainly for reclaiming space and removing invalid data. Triggered when a level of the ordered table is full or the ratio of invalid data is too high. That is, for each ordered table in the multi-level ordered table, the invalid ratio of the key-value pairs it stores is saved. When the invalid ratio reaches 50%, the ordered table is merged separately to the lower level.
[0138] Multi-level Cache Structure 1150: Includes this hierarchical structure of buffer, segmented memory table, and multi-level ordered table, combines cache optimization, intelligent write strategy, and crash consistency management to improve the response speed and stability of the storage system.
[0139] Write Strategy 1151: When writing, only need to write to the circular buffer of the M-Tree. The background uses a worker queue thread to periodically detect the cache status, so that when the system is idle or the cache is close to saturation, the aggregated data is transferred to the segmented memory table in batches, reducing the overhead caused by frequent small-scale writes. In special cases, such as high-priority or urgent requests, large sequential writes, they bypass the buffer and directly write to the memory table segment.
[0140] Based on Figure 11 the system architecture shown, Figure 12 shows a flowchart of a data processing method for block storage according to another embodiment of the present disclosure. As Figure 11As shown, the data processing method at least includes steps S1210 to S1290, which are introduced in detail as follows.
[0141] In step S1210, according to the received data write request, it is judged whether the data is of high priority.
[0142] In step S1220, if it is of high priority, a hash operation is performed on the data to determine the serial number of the segmented memory table to which it should be written.
[0143] In step S1230, the data is directly inserted into the segmented memory table corresponding to the determined serial number of the segmented memory table.
[0144] In step S1240, if it is not of high priority, the data is written into the circular buffer.
[0145] In step S1250, it is judged whether the data in the buffer is full.
[0146] In step S1260, if the data in the buffer is full, the data in the buffer is transferred to the segmented memory table in batches.
[0147] In step S1270, it is judged whether the data volume in the segmented memory table has reached a preset threshold.
[0148] In step S1280, if the data volume in the segmented memory table has reached the preset threshold, a flush operation is triggered to persistently store the data on the disk or the next-level storage structure.
[0149] In step S1290, if the data volume in the segmented memory table has not reached the preset threshold, the relevant lock resources are released to complete the current write process.
[0150] Thus, based on Figure 12 the embodiments shown, the multi-level cache architecture and the adaptive scheduling mechanism provided by the embodiments of the present disclosure are fully utilized, so as to realize efficient data writing and storage management.
[0151] The following introduces the device embodiments of the present disclosure, which can be used to execute the data processing method for block storage in the above embodiments of the present disclosure. For the details not disclosed in the device embodiments of the present disclosure, please refer to the embodiments of the data processing method for block storage in the above of the present disclosure.
[0152] Figure 13The block diagram of a data processing device for block storage according to an embodiment of the present disclosure is shown. For ease of description, certain steps of the above method are described corresponding to modules. It should be understood that the corresponding modules for executing one or several steps of the above method can be one or more hardware modules specifically configured to execute the corresponding steps, or implemented by a processor configured to execute the corresponding steps, or stored in a computer-readable medium for implementation by a processor, or implemented through a certain combination.
[0153] Referring to Figure 13 As shown, a data processing device for block storage according to an embodiment of the present disclosure includes a construction module 1310, a determination module 1320, a writing module 1330, and a processing module 1340.
[0154] Among them, the construction module 1310 is used to construct a mapping index tree, and the mapping index tree is used to uniformly manage the unified mapping relationship between the logical block address and the physical block address. The mapping index tree is a multi-level cache architecture including a cache layer, a segmented memory table, and a multi-level ordered table arranged in sequence; The determination module 1320 is used to determine the storage priority of the data to be written according to the received data writing request, and the storage priority is used to characterize the importance of the data to be written; The writing module 1330 is used to write the data to be written into the cache layer or the segmented memory table for storage according to the storage priority; and The processing module 1340 is used to merge and migrate the stored data to the next-level storage space when the stored data in the segmented memory table or the multi-level ordered table meets the merge migration rule.
[0155] In some embodiments of the present disclosure, writing the data to be written into the cache layer or the segmented memory table for storage according to the storage priority includes: when the storage priority is the first level, writing the data to be written into the cache layer for storage; when the storage priority is the second level, writing the data to be written into the segmented memory table for storage, and the importance of the second level is higher than that of the first level.
[0156] In some embodiments of the present disclosure, the processing module 1340 is used to: determine the first ratio of the stored data in the segmented memory table to its own total storage capacity; merge and migrate the stored data in the segmented memory table with the first ratio greater than or equal to the first capacity threshold to the multi-level ordered table.
[0157] In some embodiments of the present disclosure, the processing module 1340 is further used to: obtain system load information; determine the first capacity threshold according to the system load information.
[0158] In some embodiments of the present disclosure, the processing module 1340 is further configured to: for a multi-level ordered list, when the number of ordered lists at the L n level reaches a quantity threshold, and the storage data in the ordered lists at the L n level accounts for a second ratio of its own total storable capacity that is greater than or equal to a second capacity threshold, at least select some of the ordered lists from the L n level as the ordered lists to be merged, where n is an integer greater than or equal to 0; merge and sort the storage data of the ordered lists to be merged at the L n level and the storage data of the ordered lists to be merged at the L n+1 level in the order of logical block addresses to obtain a temporary ordered list, where the ordered lists to be merged at the L n+1 level are ordered lists whose key value ranges intersect with the key value ranges of the ordered lists to be merged at the L n level; write the temporary ordered list into the L n+1 level, and discard the ordered lists to be merged at the L n level and the L n+1 level.
[0159] In some embodiments of the present disclosure, the processing module 1340 is further configured to: determine the data failure ratio corresponding to the ordered list according to the number of failed key-value pairs and the total number of key-value pairs in the ordered list; when the data failure ratio reaches a failure ratio threshold, merge and migrate the storage data in the ordered list to the ordered list at the next lower level.
[0160] In some embodiments of the present disclosure, the mapping index tree is integrated into the kernel block layer to eliminate duplicate log records and mapping management.
[0161] In some embodiments of the present disclosure, the processing module 1340 is further configured to: determine the logical block address to be queried according to the received data query request; query in each storage space in the order of the cache layer, the segmented memory table, and the multi-level ordered list according to the logical block address; and feedback the physical block address corresponding to the logical block address according to the query result.
[0162] The present disclosure also provides an electronic device. Figure 14 A schematic diagram showing a hardware implementation manner using a processing system is shown.
[0163] As Figure 14As shown, the hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc. The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one connecting line is shown in this figure, but it does not mean that there is only one bus or one type of bus.
[0164] The present disclosure also provides a readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the above-mentioned method. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0165] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.
[0166] Computer programs or instructions can be stored in a readable storage medium or transmitted from one readable storage medium to another. For example, the computer programs or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0167] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, system, or computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0168] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0169] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks Figure 1 in the process or processes and / or blocks.
[0171] In the description of this specification, the descriptions with reference to the terms "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. mean that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described may be combined in any one or more embodiments / ways or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0172] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure, rather than limiting the scope of the present disclosure. For those skilled in the art, other changes or variations can be made on the basis of the above disclosure, and these changes or variations are still within the scope of the present disclosure.
Claims
1. A data processing method for block storage, characterized in that, Including: Constructing a mapping index tree for uniformly managing the mapping relationship between logical block addresses and physical block addresses. The mapping index tree is a multi-level cache architecture including a cache layer, a segmented memory table, and a multi-level ordered table arranged in sequence; Determining the storage priority of the data to be written according to the received data write request, where the storage priority is used to represent the importance of the data to be written; Writing the data to be written into the cache layer or the segmented memory table for storage according to the storage priority; And, When the stored data in the segmented memory table or the multi-level ordered table satisfies the merge migration rule, merging and migrating the stored data to the next lower-level storage space.
2. The method according to claim 1, wherein Writing the data to be written into the cache layer or the segmented memory table for storage according to the storage priority includes: When the storage priority is the first level, writing the data to be written into the cache layer for storage; When the storage priority is the second level, writing the data to be written into the segmented memory table for storage, where the importance of the second level is higher than that of the first level.
3. The method according to claim 1, characterized in that, When the stored data in the segmented memory table or the multi-level ordered table satisfies the merge migration rule and merging and migrating the stored data to the next lower-level storage space, the method includes: Determining the first ratio of the stored data in the segmented memory table to its own total storage capacity; Merging and migrating the stored data in the segmented memory table with the first ratio greater than or equal to the first capacity threshold to the multi-level ordered table.
4. The method according to claim 3, characterized in that, The method further includes: Obtaining system load information; Determining the first capacity threshold according to the system load information.
5. The method according to claim 3, characterized in that, When the stored data in the segmented memory table or the multi-level ordered table satisfies the merge migration rule and merging and migrating the stored data to the next lower-level storage space, the method further includes: For a multi-layer ordered list, when at L n the number of ordered lists at the hierarchical level reaches a quantity threshold, and at L n the proportion of the stored data in the ordered lists at the hierarchical level to their own total storable capacity is greater than or equal to a second capacity threshold, at least some of the ordered lists are selected from the L n hierarchical level as the ordered lists to be merged, where n is an integer greater than or equal to 0; Merge the stored data of the ordered list to be merged at level L n with the stored data of the ordered list to be merged at level L n+1 and perform merge sorting in the order of logical block addresses to obtain a temporary ordered list, where the ordered list to be merged at level L n+1 is an ordered list whose key value range intersects with the key value range of the ordered list to be merged at level L n ; Write the temporary ordered list to L n+1 in the hierarchy and discard the ordered list to be merged in the L n hierarchy and the L n+1 hierarchy.
6. The method according to claim 5, wherein The method further includes: Determining the data failure ratio of the ordered table according to the number of invalid key-value pairs and the total number of key-value pairs in the ordered table; When the data failure ratio reaches the failure ratio threshold, merging and migrating the stored data in the ordered table to the ordered table at the next lower level.
7. The method according to any one of claims 1-6, characterized in that, The mapping index tree is integrated into the kernel block layer to eliminate duplicate log records and mapping management.
8. The method according to any one of claims 1-6, characterized in that, The method further includes: Determining the logical block address to be queried according to the received data query request; Querying in each storage space according to the logical block address in the order of the cache layer, the segmented memory table, and the multi-level ordered table; Feedbacking the physical block address corresponding to the logical block address according to the query result.
9. An electronic device, characterized in that, Including: A memory that stores execution instructions; And, A processor that executes the execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Hybrid mapping operation method, device and equipment for storage unit and storage medium
CN110543435A
Data processing method and device and electronic equipment
CN113010455A
Metadata index storage method and device for distributed object storage system
CN118093592A
Tiered storage using storage class memory
US10126981B1
Multi-tile memory management mechanism
US20210263853A1
Cited By
Cache request scheduling method and artificial intelligence chip
CN121029427A