Distributed multi-level cache and storage method based on cold and hot data label identification

By using a two-dimensional tag generation and tag mirroring queue mechanism, combined with dynamic cold and hot tag values ​​and non-disruptive migration technology, the problems of resource allocation imbalance and service interruption during migration in multi-level caching and storage management are solved, achieving efficient and reliable data access and migration.

CN120973308AActive Publication Date: 2025-11-18WUHAN SPARK ZHONGDA INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511088780.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-18
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing multi-level caching and storage management methods lack dynamic identification and adaptability in terms of data access temporal locality, burstiness, and volatility, leading to resource allocation imbalances, service interruptions during migration, and data inconsistency issues.

Method used

A dual-dimensional tag generation mechanism is adopted, which combines access frequency and temporal locality to generate dynamic hot and cold tag values. A tag mirror queue is established to prepare for cross-level migration, and uninterrupted migration is achieved through sharded incremental replication, dual-write synchronization and atomic switching.

Benefits of technology

It enables adaptive resource allocation based on changes in system activity, reduces migration latency, improves cache hit rate, ensures high-concurrency business continuity and data consistency, and solves the problems of resource waste and service interruption in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973308A_ABST
    Figure CN120973308A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage, in particular to a distributed multi-level caching and storage method based on cold and hot data label identification, which comprises the following steps of: acquiring access frequency and access time locality of data blocks in real time, generating dynamic cold and hot label values, and storing the dynamic cold and hot label values into a database; wherein the access time locality is a standard deviation reciprocal of an access interval in the current time window; the data blocks are distributed to different storage hierarchies according to the dynamic cold and hot label values H, meanwhile, a cross-hierarchy label mirror image queue is established, and the label mirror image queue records a pre-distributed address of a target hierarchy; and in a source level keeping service state, writing the data block copy into a target level address specified by the label mirror image queue, and switching an access path in an atomized manner after writing is completed. According to the method, delay jitter caused by burst migration is avoided, a continuous physical space alignment strategy is adopted for the pre-allocated address, the fragment rate is reduced, the SSD sequential write-in performance is improved, migration operation is achieved, and resource switching is rapidly completed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data storage, and particularly relates to a distributed multi-level cache and storage method based on cold and hot data label identification. BACKGROUND

[0002] With the wide application of data centers, cloud storage platforms and distributed computing systems, the cold and hot distribution characteristics of data access in the storage system are increasingly prominent. In order to improve the storage performance and resource utilization efficiency, more and more systems adopt a multi-level storage architecture, such as a hierarchical cache and storage system constructed by combining memory (DRAM), solid state disk (SSD) and hard disk (HDD), to achieve the balance of different performance-cost ratios.

[0003] The existing multi-level cache and storage management method divides the data blocks into cold and hot based on static or single-dimensional features (such as access frequency), and determines the level according to the cold and hot. However, in actual business scenarios, data access has obvious time locality, burstiness and volatility. There are deficiencies in judging the cold and hot state only by the number of accesses: lack of modeling of access time distribution, leading to that the data blocks with high-frequency access in a short time are not identified as hot spots, affecting the cache hit rate; the layering strategy lacks adaptability, and the fixed threshold cannot be dynamically adjusted according to the overall heat change of the system, which easily leads to unbalanced resource allocation, such as memory overflow or long-term accumulation of active data in HDD; the traditional layer migration method mostly uses full replication and blocking switching, and cannot normally provide services during migration, affecting the continuity of high-concurrency business; if concurrent writing occurs during migration, there is a lack of efficient synchronization mechanism, which may lead to inconsistency between source and target copies, increasing the risk of data integrity. SUMMARY

[0004] The present application provides a distributed multi-level cache and storage method based on cold and hot data label identification, which supports dynamic cold and hot identification, has a prediction migration capability, and realizes smooth migration of data without interruption, to meet the high-performance access demand and reliability guarantee requirement in a large-scale data environment.

[0005] A distributed multi-level cache and storage method based on cold and hot data label identification, comprising the following steps:

[0006] S1, double-dimensional label generation: real-time collection of access frequency and access time locality of data blocks to generate a dynamic cold and hot label value H, wherein the access time locality is the inverse of the standard deviation of the access interval in the current time window;

[0007] S2, level cooperative decision: distributing the data blocks to the memory, SSD or HDD level according to the dynamic cold and hot label value H, and establishing a label mirroring queue across levels, which records the pre-allocated address of the target level;

[0008] S3, non-interrupt migration execution: when the data block needs to be migrated across the hierarchy, the data block copy is written to the target hierarchy address designated by the tag mirror queue under the service state of the source hierarchy, and the access path is switched atomically after the writing is completed.

[0009] Optionally, the collection of access frequency in S1 includes setting a sliding time window, counting the number of times the data block is accessed within the sliding time window, and calculating the access frequency per unit time as the ratio of the number of accesses to the sliding time window.

[0010] Optionally, the calculation of access time locality in S1 includes obtaining the timestamp sequence of all access events within the sliding event window, calculating the set of adjacent access time differences, performing standard deviation calculation on the set of adjacent access time differences, and defining the access time locality as the inverse of the standard deviation.

[0011] Optionally, the tag mirror queue is a data structure used to record and manage cross-hierarchy migration preparation information in advance, and its function is to prepare addresses and paths for future migration of data blocks without interrupting current data services, like a "migration reservation book" or "pre-migration list", recording information such as which data blocks are about to be migrated, where they are prepared to migrate to, and addresses have been reserved.

[0012] Supporting non-interrupt migration, the mirror queue has allocated addresses for the target hierarchy before the migration starts, so the data block can continue to serve in the source hierarchy while being copied to the target hierarchy silently, realizing "hot migration". Traditional migration operations often need to apply for addresses and establish paths after triggering migration, while the mirror queue completes these tasks in advance, so that the migration process only needs to switch the access direction, greatly reducing the delay. The queue records which data block tag values are continuously warming up or cooling down, providing a basis for the system to dynamically adjust the hierarchy distribution. The dynamic hot-cold tag value is generated according to the access frequency and time locality, and is represented as:

[0013] H = F·log2(1+L);

[0014] Where H is the dynamic hot-cold tag value, L is the access time locality, reflecting the degree of frequent access and time aggregation, F is the access frequency, and log2 represents the logarithmic function, used to suppress the non-linear amplification of L.

[0015] Optionally, S2 includes a double threshold setting, setting a first threshold θ1 allocated to the memory hierarchy and a second threshold θ2 allocated to the HDD hierarchy, and performing hierarchy allocation:

[0016] When H ≥ θ2, allocate to the memory hierarchy;

[0017] When θ2≤H<θ1, allocate to the SSD hierarchy;

[0018] When H < θ2, allocate to the HDD level.

[0019] Optionally, the S2 further comprises defining a tag value change slope k, which is determined by the change trend of the hot and cold tag values of the certain data block in the recent multiple periods, and the tag mirror queue comprises:

[0020] For the data block with continuously rising H value, pre-allocate the physical address at a higher level;

[0021] For the data block with continuously falling H value, pre-allocate the physical address at a lower level.

[0022] Optionally, the pre-allocated address comprises:

[0023] Allocate the continuous physical space in the free block of the target level;

[0024] Write the pre-allocated address into the tag mirror queue metadata, and associate the identifier of the source level data block.

[0025] Optionally, the S3 specifically comprises:

[0026] S31, incremental copy phase:

[0027] Start the asynchronous copy of the data block copy while continuing to process the read and write requests at the source level;

[0028] Fragment the data block into fixed-size migration units, and sequentially write them into the target level address specified by the tag mirror queue;

[0029] S32, write operation synchronization: for the write request occurring during the migration process, perform double write operation;

[0030] S33, atomic switching: when all migration units are copied and the redo log is empty, update the global routing table metadata; switch the data block access path to the target level address through atomic operation.

[0031] Optionally, the execution of the double write operation comprises:

[0032] Simultaneously update the source level data and the corresponding migration unit of the target level;

[0033] Record the double write operation of the non-migration unit to the redo log.

[0034] Optionally, the S3 further comprises, after the switching succeeds, releasing the service state of the source level data block, and releasing the source level storage space after delaying the heartbeat period.

[0035] Optionally, after the access path is switched, the following feedback is triggered:

[0036] Global routing table update -> new access request directs to new level;

[0037] New level access event -> becomes input source for S1 hot / cold tag value calculation;

[0038] a) Rewriting the new level physical address to the hot / cold tag database, replacing the original storage location record;

[0039] b) Reset the access statistics window of the data block, empty the historical access event queue;

[0040] c) Adjust the sliding window length W based on the new level media type, and then recalculate the H value:

[0041] Memory level: W = 10s;

[0042] SSD level: W = 30s;

[0043] HDD level: W = 90s.

[0044] Advantages of the present application:

[0045] The present application, by introducing a two-dimensional tag generation mechanism, comprehensively considers the access frequency and access time locality of the data block, dynamically calculates the hot / cold tag value, and on this basis, uses a dynamic threshold algorithm based on the global distribution mean and standard deviation to manage the data block in layers. Compared with the traditional static strategy, this mechanism can adaptively adjust the allocation boundary according to the system heat change, effectively alleviate the memory congestion in peak period and the resource waste in low peak period, reduce the SSD utilization rate fluctuation, and improve the overall access hit rate.

[0046] The present application, based on the trend change of the hot / cold tag value, constructs a tag mirror queue mechanism, realizes the advance recognition of the data block heat and the pre-allocation of the target level physical address, introduces the tag value slope judgment standard, automatically triggers the migration preparation combined with the continuous rising / falling trend, avoids the delay jitter caused by sudden migration, uses the continuous physical space alignment strategy for pre-allocation address, reduces the fragmentation rate, improves the SSD sequential write performance and realizes the fast completion of resource switching in migration operation.

[0047] The present application proposes a non-interrupt migration execution mechanism that supports uninterrupted read / write service during migration, uses a combination of sharding incremental replication, double-write synchronization and atomic switching to realize smooth transition of data blocks. Write conflicts are solved by vector clock, compensation writing of unsynchronized operations is ensured by redo log, and strong consistency of routing table switching is guaranteed by RAFT protocol. Finally, in the actual high-concurrency video stream editing scene, "frame rate zero drop" is realized, solving the problem of data hot migration process stuttering, frame loss or data inconsistency in existing systems. BRIEF DESCRIPTION OF DRAWINGS

[0048] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Fig. 1 This is a schematic diagram of the method flow according to an embodiment of the present invention;

[0050] Fig. 2 This is a schematic diagram of the method execution logic in an embodiment of the present invention. Detailed Implementation

[0051] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. For some well-known technologies, those skilled in the art may also use other alternative methods to implement the invention. Moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0052] like Figs. 1-2 As shown, a distributed multi-level caching and storage method based on hot and cold data tag identification includes the following steps:

[0053] S1, Two-dimensional label generation: Real-time collection of data block access frequency and access time locality to generate dynamic hot and cold label values ​​H, where access time locality is the reciprocal of the standard deviation of the access interval within the current time window;

[0054] S2, Hierarchical collaborative decision-making: Based on the dynamic cold and hot tag value H, data blocks are allocated to memory, SSD or HDD levels, and a cross-level tag mirroring queue is established. The tag mirroring queue records the pre-allocated address of the target level.

[0055] S3, Uninterrupted Migration Execution: When a data block needs to be migrated across levels, while maintaining service at the source level, a copy of the data block is written to the target level address specified by the tag mirror queue. After the writing is completed, the access path is switched atomically.

[0056] S1 specifically includes:

[0057] S11, Access Frequency F Collection: Set a time window W, count the total number of times the target data block is accessed N within the time window, and calculate the access frequency per unit time. Where F is the access frequency, N is the number of accesses within the window, and W is the sliding time window length, which is dynamically set according to the storage level.

[0058] Memory layer W = 10s, SSD layer W = 30s, HDD layer W = 120s.

[0059] S12, access time locality L calculation: obtain the timestamp sequence {t1, t2, …, t n} of all access events within the time window W, and calculate the adjacent access time difference set:

[0060] Δ = {Δt1 = t2 - t1, Δt2 = t3 - t2, …, Δt n-1 = t n -t n-1}; calculate the standard deviation of Δ using the Welford online iterative algorithm, and obtain:

[0061]

[0062] where Δ is the adjacent access time difference set, σ is the standard deviation, which measures the dispersion of time intervals, q is the number of elements in set Δ, and ε is a constant to prevent zero, set to 10 -6 seconds to ensure that σ Δ = 0 is meaningful, L represents access time locality, and the larger the value, the stronger the time concentration (i.e., local access is dense).

[0063] S13, dynamic hot and cold label value H generation: generate a dynamic hot and cold label value based on access frequency and time locality:

[0064] H = F · log2(1 + L);

[0065] where H is the dynamic hot and cold label value, which reflects the access frequency and time concentration, log2 represents the logarithm function with base 2, which is used to suppress the non-linear amplification of L to avoid label jumping caused by extreme fluctuations.

[0066] S2 specifically includes:

[0067] S21, hierarchical dynamic allocation: set dynamic thresholds based on the size of label value H and its distribution in the system: θ1 = μ + 2σ H , θ2 = μ - σ H ;

[0068] where μ represents the mean of all data block H values in the current system, σ H represents the standard deviation of all H values in the current system, θ1 is the lower threshold (first threshold) allocated to the memory hierarchy, and θ2 is the upper threshold (second threshold) allocated to the HDD hierarchy.

[0069] Based on the above thresholds, the data block allocation strategy is implemented:

[0070]

[0071] The mean and standard deviation of the label value H of all current data blocks can dynamically depict the hot and cold state distribution of the entire system, μ represents the current overall "average heat level", σ H represents the degree of hot and cold difference (the greater the fluctuation range, the greater σ H ).

[0072] Distinguish "hot", "warm", and "cold" data:

[0073] H≥μ+2σ H : far above the average heat, belongs to the extremely hot data block, should be put into the fastest memory response first;

[0074] H<μ-σ H : significantly lower than the average heat, belongs to the extremely cold data block, put into the lowest cost HDD;

[0075] μ-σ H ≤H<μ+2σ H : in the middle fluctuation interval, classified as warm data block, put into SSD to balance cost and performance.

[0076] System load and access pattern fluctuate greatly over time, fixed threshold can easily lead to: memory overflow or cold data misleaving high-speed layer in peak period, resource vacancy in trough period, excessive migration; dynamic threshold mechanism can adapt to system state in real time, improving the robustness of the strategy.

[0077] S22, label mirror queue generation:

[0078] The label mirror queue is a data structure used to record and manage cross-level migration preparation information in advance, and its role is to prepare addresses and paths for future migration of data blocks without interrupting current data services, like a "migration reservation book" or "pre-migration list", recording which data blocks are about to migrate, where they are preparing to migrate, and the address has been reserved, etc.

[0079] Supports non-interrupt migration, before migration starts, the mirror queue has already allocated addresses for the target level, so data blocks can continue to be served in the source level while being copied to the target level quietly, realizing "hot migration". Traditional migration operations often need to apply for addresses and establish paths after triggering migration, while the mirror queue completes these tasks in advance, so that the migration process only needs to switch access direction, greatly reducing the delay. The queue records which data block's label value is continuously warming up or cooling down, providing a basis for the system to dynamically adjust the level distribution.

[0080] Queue content example (structure):

[0081] Each label mirror queue record contains:

[0082] Data block identification (source location, such as the xth segment of the yth block of the SSD);

[0083] Pre-migration target address (specific location of the target tier memory or HDD);

[0084] Migration state (such as "pre-allocated", "completed", "invalid", etc.).

[0085] The tag mirror queue is generated as follows: when the hot and cold tag value H of the data block shows a continuous fluctuation trend, pre-allocate physical space for it in the target tier, and trigger the generation of the tag mirror queue:

[0086] Define the slope of the tag value change as: where H t is the H value of the current period, H t-3 is the H value three periods ago, and k is the average growth or decline rate;

[0087] The judgment criteria are as follows:

[0088]

[0089] The change trend (i.e. "slope") of the tag value in multiple consecutive periods is calculated in order to capture the persistent change in the hotness of the data block, rather than occasional fluctuations. If the hotness value of a data block continues to rise, it may soon reach the migration threshold and need to be placed in faster storage. If it continues to decline, it may soon be downgraded and moved to slower and cheaper storage. By observing the trend over multiple periods rather than a single value, false positives can be reduced and migration "back and forth jitter" can be avoided. The tag mirror queue pre-allocates target tier addresses, reserving space for data blocks in advance when the trend is obvious, avoiding delays caused by address application during actual migration. It supports non-disruptive migration, and when the data block actually meets the migration conditions, the recorded address in the queue can be used to switch the path, achieving smooth transition. The "slope trend judgment + mirror queue recording" method enables early identification of data hot and cold changes and pre-allocation of resources, which is a key foundation for efficient, non-intrusive, and multi-level data management.

[0090] S23, pre-allocate address management, pre-allocate target tier addresses for data blocks that need to be migrated:

[0091] Reserve space in the free contiguous blocks of the target tier:

[0092] Memory tier: aligned by 4KB;

[0093] SSD tier: aligned by 1MB;

[0094] HDD tier: aligned by 4MB.

[0095] Write the pre-allocated address into the tag mirror queue as key-value pairs:

[0096] Mapping format: <source level identifier, target level address>;

[0097] This queue storage structure is lightweight, with each mapping entry occupying only 16 bytes, effectively supporting efficient indexing and migration path switching under millions of data blocks.

[0098] S3 specifically includes:

[0099] S31, Incremental Replication Phase: While continuously providing read and write services at the source level, an asynchronous replica replication process is initiated, dividing the target data block into fixed-size migration units, denoted as: U = {u1, u2, ..., u...} c}; • Each migration unit u i Sequentially write to the target address range pre-allocated by the tag mirroring queue; where, u i Let be the i-th migration unit. The unit size depends on the migration path: 128KB for the memory layer and 1MB for the SSD / HDD layer. c is the number of units that need to be migrated in the data block, which is equal to the total size of the data block divided by the size of each migration unit.

[0100] S32, Write Operation Synchronization: During migration, if a migration unit is being written to, a double write operation is performed.

[0101] S321, Simultaneously write the user's write request:

[0102] Source level original position;

[0103] The corresponding migration unit address in the target level.

[0104] S322, If the write request is applied to a cell that has not yet been migrated, the operation is written to the redo log:

[0105] Where RedoLog represents the set of redo logs, u j This indicates a migration unit that has not yet been copied; op represents the write operation content; t represents the operation timestamp; U copied This represents the set of cells that have been copied.

[0106] S33, Atomization Switching:

[0107] When both of the following conditions are met:

[0108] Condition 1: All migration units have been replicated.

[0109] Condition 1: The redo log is empty (all delayed writes have been synchronized);

[0110] Then the atomic switch operation is performed to update the global route table:

[0111] Global Route[B id ]=TargetLayerAddress; means updating the access path of the data block in the system to the new target layer address. Global Route is a global route table that stores the current location of each data block; Bid is the unique identifier of the data block (Block ID); TargetLayerAddress is the pre-allocated physical address of the data block in the target storage layer; the expression means: update the access path of data block Bid to its new address, that is, "switch the access path". The update operation is based on the consensus protocol (RAFT) and ensures that more than half of the nodes agree.

[0112] S34, resource cleaning, after switching successfully:

[0113] The source layer data block is marked as read-only;

[0114] Delay releasing its physical space, release time is:

[0115] T release =T switch +3·T heartbeat ; where, T switch is the switching completion time, T heartbeat represents the current heartbeat period of the system, default 200ms, dynamically adjusted by network conditions.

[0116] Heartbeat period refers to the time interval of sending state signals, i.e. "heartbeat", between nodes or components in a distributed system, used to check whether the node is online, judge whether the communication is normal, and coordinate resource status or task execution; in the non-interrupt migration scheme, "heartbeat period" is used to control the timing of resource release. For example: when the data block is successfully migrated from SSD to memory, the copy on SSD will not be deleted immediately, but will wait for 3 heartbeat periods to confirm that the system is stable, there is no network jitter or exception, and then safely release the resource.

[0117] After switching the access path, the following feedback is triggered:

[0118] Global route table update → new access request directed to new layer;

[0119] New layer access event → become input source for S1 hot / cold tag value calculation;

[0120] a) Rewrite the new layer physical address to the hot / cold tag database, replacing the original storage location record;

[0121] b) Reset the access statistics window of the data block and clear the historical access event queue;

[0122] c) Adjust the sliding window length W based on the new tier media type and recalculate H value:

[0123] Memory tier: W = 10s;

[0124] SSD tier: W = 30s;

[0125] HDD tier: W = 90s.

[0126] The present application encompasses any alternatives, modifications, equivalent methods and solutions made to the essence and scope of the present application. In order to make the public have a thorough understanding of the present application, specific details are described in the following preferred embodiments of the present application, and the present application can also be fully understood without the description of these details to those skilled in the art. In addition, in order to avoid unnecessary confusion to the essence of the present application, well-known methods, processes, procedures, elements and circuits, etc. are not described in detail.

[0127] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can also be made, which should be considered as the protection scope of the present application.

Claims

1. A distributed multi-level cache and storage method based on cold-hot data tag identification, characterized in that, The method comprises the following steps: S1, two-dimensional label generation: collecting access frequency and access time locality of data blocks in real time, and generating dynamic hot and cold label value H, wherein the access time locality is the reciprocal of the standard deviation of access intervals in a current time window; S2, hierarchical collaborative decision: distributing data blocks to memory, SSD or HDD levels according to the dynamic hot and cold label value H, and establishing a label mirror queue across levels, which records the pre-allocated address of the target level; S3, non-interrupt migration execution: when the data block needs to be migrated across levels, the data block copy is written to the target level address specified by the label mirror queue while maintaining the service state in the source level, and the access path is switched atomically after the writing is completed.

2. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 1, characterized in that, The collection of access frequency in S1 comprises setting a sliding time window, counting the number of accesses of the data block in the sliding time window, and calculating the access frequency per unit time as the ratio of the number of accesses to the sliding time window.

3. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 2, characterized in that, The calculation of access time locality in S1 comprises obtaining the timestamp sequence of all access events in the sliding event window, calculating the set of adjacent access time differences, and calculating the standard deviation of the set of adjacent access time differences, and defining the access time locality as the reciprocal of the standard deviation.

4. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 3, characterized in that, The dynamic hot and cold label value is generated according to the access frequency and time locality, and is expressed as: H=F·log2(1+L); Wherein, H is the dynamic hot and cold label value, L is the access time locality, reflecting the access frequency and time aggregation, F is the access frequency, log2 represents the logarithmic function, and is used to suppress the nonlinear amplification of L.

5. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 1, characterized in that, S2 comprises double-threshold setting, setting the first threshold θ1 for allocation to the memory level and the second threshold θ2 for allocation to the HDD level, and performing hierarchical allocation: When H≥θ2, it is allocated to the memory level; When θ2≤H<θ1, it is allocated to the SSD level; When H<θ2, it is allocated to the HDD level.

6. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 5, characterized in that, S2 further comprises defining a label value change slope k, which is determined by the change trend of the cold and hot label value of a certain data block in the recent multiple periods, and the label mirror queue comprises: For data blocks with continuously rising H values, pre-allocate physical addresses in higher levels; For data blocks with continuously falling H values, pre-allocate physical addresses in lower levels.

7. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 6, characterized in that, The pre-allocated address comprises: Allocating continuous physical space in the free block of the target level; Write the pre-allocated address to the label mirror queue metadata, and associate the identifier of the source level data block.

8. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 1, characterized in that, S3 specifically comprises: S31, incremental replication phase: While the source level continues to process read and write requests, start asynchronous replication of data block copies; Divide the data block into fixed-size migration units and sequentially write them to the target level address specified by the label mirror queue; S32, write operation synchronization: for the write request occurring in the migration process, perform double write operation; S33, atomic switching: when all migration units are replicated and the redo log is empty, update the global routing table metadata; switch the data block access path to the target level address through atomic operation.

9. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 8, characterized in that, The execution of the double write operation comprises: Simultaneously update the source level data and the corresponding migration unit of the target level; The double write operation record of the non-migrated unit is recorded to the redo log.

10. The distributed multi-level cache and storage method based on cold and hot data tag identification according to claim 8, characterized in that, The S3 further comprises releasing the service state of the source hierarchical data block after the switching succeeds, and releasing the source hierarchical storage space after delaying the heartbeat period.

Citation Information

Patent Citations

  • Data migration method and device, terminal equipment and storage medium

    CN112256675A

  • Data storage method and device, data query method and device, equipment and storage medium

    CN115878513A

  • Data migration method and device

    CN116185993A

  • Data block hierarchical storage method and device, computer equipment and storage medium

    CN118034593A

  • Data fragmentation and table division autonomous extension system and method

    CN118606295A

Cited By

  • Memory control method and memory storage device

    CN121433586A