Dynamic management method and system for multi-layer data cache

Through hierarchical analysis and stability classification based on data access paths in multi-layer data cache management, the cache-level storage method is optimized, which solves the problem that traditional methods are difficult to finely adjust cache resources in complex access mode, and achieves more efficient cache resource utilization and data throughput capabilities.

CN119987685AActive Publication Date: 2025-05-13SHENZHEN DOMOD DIGITAL TECH

Patent Information

Application Number
CN202510459379.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Traditional multi-layer data cache management methods are difficult to make fine adjustments when facing complex access modes, resulting in wasted or overflowing cache resources, affecting cache utilization and data reading performance.

Method used

Through hierarchical analysis and stability classification based on data access paths, the data cache hierarchical storage method is optimized, the stable path data is stored in the cache area, and the non-stable path data is stored in the low-speed cache area, and the storage order and elimination strategy are adjusted according to the access cross frequency and data competition situation.

Benefits of technology

It improves the precise allocation efficiency of cache resources, reduces the waste of cache resources caused by access path jumps, improves the continuity of data access and cache hit rate, and enhances the overall data throughput capability and system load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987685A_ABST
    Figure CN119987685A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data cache management, in particular to a multilayer data cache dynamic management method and system, and the method comprises the following steps: based on a data access path, extracting a data access path hierarchy identifier, analyzing path hierarchy depth, recording access continuity, classifying the path to be stable or unstable, and obtaining path hierarchy distribution information. In the invention, through hierarchical analysis and stability classification of the data access paths, the cache resource allocation efficiency, the stable paths and the unstable paths are improved, the access delay is reduced, the access continuity is improved, the cache waste caused by hopping is reduced, the space utilization rate is improved, the cache pollution is avoided, the storage sequence is optimized, and the hit rate is improved; through dynamic elimination, resource competition is reduced, storage medium I / O load is monitored, medium load is balanced, throughput capacity is improved, data flow is optimized, cache coordination capacity is improved, cache levels can flexibly adapt to access mode changes, and data storage stability and load balance are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data cache management, and in particular to a multi-layer data cache dynamic management method and system. Background Art

[0002] The field of data cache management technology includes data storage, access optimization, cache replacement strategy and multi-layer cache structure design. Its core content involves how to reasonably allocate cache resources in computer systems or distributed environments, improve data reading speed, and reduce access latency. Currently, data cache management is widely used in databases, files, cloud computing platforms and embedded devices. Traditional data cache management methods mainly adopt static or dynamic allocation strategies, by setting a fixed cache size or adjusting it according to access frequency. Some methods combine machine learning or adaptive algorithms to optimize cache hit rate. Multi-layer data cache technology is an important part of cache management. It uses cache storage media at different levels, combined with cache consistency maintenance and replacement strategies, to achieve efficient data access.

[0003] Among them, the multi-layer data cache dynamic management method refers to dynamically adjusting the data distribution strategy between cache levels in a multi-layer cache system by real-time monitoring of data access patterns, cache hit rates and storage medium status. This method covers technical matters such as cache space allocation, data migration decisions and cache replacement rule settings. Specific methods include dynamic adjustment of cache space based on access frequency and timeliness, cache content migration based on data priority or access characteristics, and formulation of cache elimination strategies in combination with statistical analysis or rule matching methods. This method adjusts the data storage location by monitoring cache usage and optimizes the utilization of cache resources in combination with specific logic.

[0004] Traditional technologies mainly rely on static or dynamic allocation strategies based on access frequency, which makes it difficult to fine-tune cache resources when facing complex access patterns. The fixed cache size method is prone to storage space waste or cache overflow when the data access pattern changes greatly, affecting cache utilization. The adjustment method based on access frequency is easily affected by the suddenness of data access and cannot quickly respond to real-time changes in data access requirements, resulting in high-frequency access data not being stored in the cache area first, affecting data reading performance. Although traditional multi-layer cache management uses storage media at different levels for data storage, it lacks fine-grained analysis of access path stability and data cross-competition, resulting in a decrease in cache hit rate. Data migration strategies are difficult to adapt to complex access scenarios. Data elimination strategies rely on fixed rules and cannot be flexibly adjusted, resulting in high-value data being eliminated prematurely, affecting data reading efficiency. The storage medium has limited load balancing and control capabilities and cannot dynamically adjust the data storage ratio according to the I / O task load, causing storage bottlenecks, reducing overall throughput, and affecting storage stability under high-load environments. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings in the prior art and to propose a multi-layer data cache dynamic management method and system.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: a multi-layer data cache dynamic management method, comprising the following steps: S1: Based on the data access path, extract the data access path level identifier, analyze the path level depth, record the access continuity, classify the path as stable or unstable, and obtain the path level distribution information; S2: Based on the path level distribution information, analyze the data cache level storage ratio, store stable path data in the high-speed cache area, store unstable data in the low-speed cache area, adjust the storage area, and obtain the cache space allocation status; S3: extracting data block access records, detecting access crossover frequency, filtering frequently crossover data blocks, analyzing cache storage status, adjusting storage order and data elimination, and obtaining cache data priority according to the cache space allocation status; S4: calling the cache data priority, monitoring the storage medium I / O tasks, analyzing the task execution status, adjusting the cache data storage ratio according to the task load, adjusting the storage medium of part of the data, and obtaining the storage medium load change trend; S5: Based on the storage medium load change trend, analyze the data matching situation, call the cache allocation rule, analyze the data throughput situation, and output the storage status of the data at the cache level.

[0007] As a further solution of the present invention, the path level distribution information includes path level identification, path level depth, path access continuity, access stability parameter classification results, and path stability categories; the cache space allocation status includes data cache level storage ratio, cache area storage data, unstable path storage data, and storage area adjustment results; the cache area data priority includes data block access records, frequently accessed cross-data blocks, cache area storage status, data storage order, and competition elimination data; the storage medium load change trend includes storage medium I / O tasks, task execution status, task load, and storage medium adjustment results; the storage status of the data at the cache level includes data matching status, storage area allocation, cache allocation rules, and data throughput.

[0008] As a further solution of the present invention, the step of obtaining the path level distribution information is specifically as follows: S111: based on the data access path, extract the path level identifier, analyze the level depth, record the path access continuity, identify the level access times, and obtain the path level distribution characteristic value; S112: calling the path level distribution characteristic value, classifying the paths according to the access stability parameter, screening the stable paths and the unstable paths, calculating the proportion of the stable paths, and obtaining the path stability ratio; S113: calling the path stability ratio, combining the access trend of each level path, analyzing the access change range of the differentiated level paths, using the formula: ; Calculate the path level access fluctuation value and obtain the path level distribution information; in, Represents the path level access fluctuation value, Representative The number of visits to the level, Represents the number of visits to the previous level. Representative The total number of levels in the hierarchy, Representative The stability ratio of the hierarchical path, Represents the total number of path levels.

[0009] As a further solution of the present invention, the step of obtaining the cache space allocation status is specifically as follows: S211: Based on the path level distribution information, calculate the data storage amount of the path level, compare the storage capacity, filter the path levels whose storage ratio exceeds the threshold, and obtain the path level storage ratio screening result; S212: calling the path level storage ratio screening result, storing the data in differential cache areas according to data stability classification, comparing the storage ratios of high-speed and low-speed cache areas, and obtaining a cache area data storage deviation value; S213: Based on the data storage deviation value of the cache area, analyze the adjustment range of the storage area to the access jump degree, using the formula: ; Calculate the storage area adjustment range to obtain the cache space allocation status; in, Represents the storage area adjustment range, Representative The amount of low-speed cache storage at each level, Representative The cache storage capacity of the level, Representative The frequency of visits to the level, Representative Baseline visit frequency of the tier, Representative The storage scaling factor for the tier.

[0010] As a further solution of the present invention, the step of obtaining the priority of the cache data is specifically as follows: S311: extracting data block access records according to the cache space allocation status, calculating the number of data block accesses and time intervals, analyzing access intersection conditions, and obtaining access intersection metrics; S312: calling the access intersection metric value, filtering the intersection data blocks according to the intersection number threshold, calculating the cache occupancy ratio, analyzing the distribution trend, and obtaining the contention data block set; S313: Based on the set of competing data blocks, analyze the cache storage status and data block access priority, adjust the storage order and elimination strategy according to the priority, and use the formula: ; Get the cache data priority; in, Represents the cache data priority, Represents a data block The number of crossover visits, Represents a data block The time interval between the last visit, Represents a data block The cache occupancy ratio, Represents a data block The storage location index in the cache, Represents the mean value of the data block storage location index in the cache.

[0011] As a further solution of the present invention, the step of acquiring the storage medium load change trend is specifically: S411: calling the cache data priority, monitoring the storage medium I / O tasks, calculating the storage medium I / O response time ratio, screening the storage medium whose task ratio exceeds the task threshold, and obtaining I / O task ratio distribution data; S412: Analyze the load change rate of the storage medium based on the I / O task proportion distribution data, adjust the cache data storage ratio of the load storage medium, and obtain a cache allocation adjustment coefficient; S413: Adjust the data storage medium according to the cache allocation adjustment coefficient, analyze the load change trend after the data adjustment, and use the formula: ; Obtain storage media load change trends; in, Represents the storage medium load change trend value. represents the load level of the jth storage medium after adjustment, represents the load level of the jth storage medium before adjustment, represents the original task load stability of the jth storage medium, represents the data access fluctuation amplitude of the jth storage medium, Indicates the total number of storage media to be adjusted.

[0012] As a further solution of the present invention, the steps of obtaining the storage status of the data at the cache level are specifically as follows: S511: based on the storage medium load change trend, extract the data block access frequency, migration times, and storage duration, analyze the data matching degree, filter the matching conditions, and obtain the data matching status classification value; S512: Call the data matching status classification value to analyze the data throughput of the storage layer, using the formula: ; Calculate the data throughput state characteristic value, determine the storage area adjustment demand, and obtain the storage area adjustment demand; in, Represents the characteristic value of data throughput status, Represents the current storage level throughput, Represents the raw storage tier throughput, represents the throughput fluctuation range, represents the mean throughput, represents the data migration rate, Represents the data migration impact factor; S513: Based on the storage area adjustment demand, according to the cache allocation rule, analyzing the adjusted data storage status, and outputting the storage status of the data at the cache level.

[0013] The multi-layer data cache dynamic management system is used to execute the multi-layer data cache dynamic management method, and the system includes: The path access analysis module extracts the hierarchical identifier based on the data access path, identifies the hierarchical depth, records the path access sequence, selects stable paths and calculates the distribution ratio to obtain the path hierarchical distribution data; The cache layer allocation module calculates the storage ratio of the high-speed cache area and the low-speed cache area based on the path layer distribution data, stores the stable path set and the unstable path set, adjusts the jump path storage area, and obtains the cache space allocation status; The data competition detection module extracts data block access records based on the cache space allocation state, filters the cross-accessed data blocks in a short period of time, calculates the number of cross-accesses and classifies them, analyzes the storage proportion of the cross-accessed data blocks in the cache area, adjusts the storage order and filters the cross-accessed data blocks, and obtains the data competition evaluation result; The cache data priority adjustment module calls the cache data priority based on the data competition evaluation result, identifies the load of the storage medium I / O task, adjusts the storage ratio, selects part of the data for storage medium migration, and obtains the storage medium load status; The storage medium load balancing module screens low-frequency access data to adjust the storage area based on the storage medium load status, calls cache allocation rules, calculates cache data throughput, adjusts the data storage level, and obtains the storage status of data at the cache level.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, through hierarchical analysis and stability classification based on data access paths, the data cache hierarchical storage method is optimized, the precise allocation efficiency of cache resources is improved, and the stable path data is stored in the high-speed cache area, which reduces data access delay, improves the continuity of data access, and reduces the waste of cache resources caused by access path jumps. Unstable path data is stored in the low-speed cache area to improve cache space utilization, avoid high-frequency data from reducing reading efficiency due to cache pollution, optimize storage order based on access cross frequency and data competition, improve the effective hit rate of cache area data, and reduce resource competition through dynamic elimination strategy, improve cache stability, monitor storage medium I / O task load, dynamically adjust the storage ratio of cache data according to storage pressure, balance the load of different media, and improve overall data throughput. The data storage location is adjusted based on the data matching situation and cache allocation rules, the data flow efficiency of the cache level is optimized, and the collaborative ability of multi-layer data cache is improved. The overall optimization strategy improves the adaptability of cache resources, enables the cache hierarchy structure to flexibly respond to changes in data access patterns, improves data storage stability and access efficiency, reduces data access latency, and at the same time enhances the data collaborative management capabilities between cache levels, improving data throughput and system load balancing. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the workflow of the present invention; Figure 2 It is a flow chart for obtaining path level distribution information in the present invention; Figure 3 A flowchart for obtaining the cache space allocation status in the present invention; Figure 4 A flowchart of obtaining the priority of buffer data in the present invention; Figure 5 A flowchart for obtaining a storage medium load change trend in the present invention; Figure 6 This is a flow chart for obtaining the storage status of data at the cache level in the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0018] Example 1: Please refer to Figure 1 The present invention provides a technical solution: a multi-layer data cache dynamic management method, comprising the following steps: S1: Based on the data access path, extract the path level identifier, analyze the path level depth, record the path access continuity, classify the path according to the access stability parameter, filter the stable path and the unstable path, and obtain the path level distribution information; S2: Based on the path level distribution information, analyze the data cache level storage ratio, store stable path data in the high-speed cache area, and store unstable path data in the low-speed cache area. Adjust the storage area according to the access jump degree to obtain the cache space allocation status; S3: Extract data block access records based on cache space allocation, detect access crossovers in a short period of time, determine data competition based on the number of crossovers, filter data blocks with frequent access crossovers, analyze cache storage status, adjust data storage order, adjust data elimination based on competition, and obtain cache data priority; S4: Call the cache data priority, monitor the storage medium I / O tasks, analyze the task execution status, adjust the cache data storage ratio according to the task load, adjust the storage medium of some data, and obtain the storage medium load change trend; S5: Based on the change trend of storage medium load, analyze the data matching situation and adjust the storage area, call the cache allocation rules, analyze the data throughput, and output the storage status of the data at the cache level.

[0019] Path level distribution information includes path level identification, path level depth, path access continuity, access stability parameter classification results, and path stability categories. Cache space allocation status includes data cache level storage ratio, cache area storage data, unstable path storage data, and storage area adjustment results. Cache area data priority includes data block access records, frequently accessed cross-data blocks, cache area storage status, data storage order, and competition elimination data. Storage medium load change trends include storage medium I / O tasks, task execution status, task load, and storage medium adjustment results. The storage status of data at the cache level includes data matching status, storage area allocation, cache allocation rules, and data throughput.

[0020] See also Figure 2 , the specific steps for obtaining path level distribution information are: S111: based on the data access path, extract the path level identifier, analyze the level depth, record the path access continuity, identify the level access times, and obtain the path level distribution characteristic value; The access path of each file is extracted through the access records of the log file, and then the paths are identified and classified, and the access frequency of each path in different time periods is counted. This process helps managers understand data access patterns and storage efficiency. For example, in the server of an e-commerce website, by analyzing the access data of the storage path of product images, server resources can be adjusted and data call speed can be optimized. This method directly affects the structural optimization of the file system through specific path access records, and finally obtains the path level distribution characteristic value.

[0021] S112: calling the path level distribution characteristic value, classifying the paths according to the access stability parameter, screening the stable paths and the unstable paths, calculating the proportion of the stable paths, and obtaining the path stability ratio; Cloud service providers monitor and classify the data request paths of different customers and adjust resource allocation based on the stability and volatility of the paths. This includes collecting access logs within a specific period of time, calculating the access stability of each path, and allocating more cache resources to paths that show high stability. This analysis helps improve the service response speed and resource utilization efficiency, and obtains the path stability ratio.

[0022] S113: Call the path stability ratio, combine the access trend of each level path, analyze the access change range of the differentiated level paths, and use the formula: ; Calculate the path level access fluctuation value and obtain the path level distribution information; in, Represents the path level access fluctuation value, Representative The number of visits to the level, Represents the number of visits to the previous level. Representative The total number of levels in the hierarchy, Representative The stability ratio of the hierarchical path, Represents the total number of path levels; In big data analysis, the calculation of path-level access fluctuation values ​​is crucial for predicting data traffic and resource requirements. Assume that in a cloud storage system, it is necessary to analyze file access at different levels to optimize server load and extract the number of accesses at each level from log records. For example, within a certain period of time: Level 1 path ( ) Number of visits ; Level 2 Path ( ) Number of visits ; Level 3 Path ( ) Number of visits ; Level 4 Path ( ) Number of visits ; The access data comes from log statistics, which calculates the access fluctuations between levels; Among them, the stability ratio of each level Calculated by the visit frequency in the past 30 days, for example: , , , ; Substitute the data and calculate: ; ; ; ; Calculate the result path level access fluctuation value , this value indicates that the traffic fluctuation between access levels is large, which means that the access changes at different levels are significant. This data can be further used to predict future data access trends, and combined with historical data to determine whether it is necessary to optimize the cache strategy or adjust the server resource configuration, and finally obtain the path level distribution information.

[0023] See also Figure 3 , the specific steps for obtaining the cache space allocation status are: S211: Calculate the data storage amount of the path level based on the path level distribution information, compare the storage capacity, filter the path levels whose storage ratio exceeds the threshold, and obtain the path level storage ratio screening result; First, we monitor the access logs of different data centers, collect data access frequency and path information, and calculate the data storage capacity of each path level through the data. For example, in data center A, we monitor that path P1 is accessed 300 times in a week, while path P2 is accessed 50 times. Based on the number of accesses, we estimate the data storage requirements of P1 and P2, and compare the storage capacity of each path level to determine which levels have a storage share that exceeds the predetermined threshold. This threshold is set based on historical data and storage policies, such as 25%. If the storage share of level L1 reaches 30%, it is considered to exceed the threshold. Through detailed calculation and comparison, we filter out the path levels that need special attention and obtain the path level storage share screening results.

[0024] S212: Calling the path level storage ratio screening result, storing the data in the differentiated cache area according to the data stability classification, comparing the high-speed cache area storage ratio and the low-speed cache area storage ratio, and obtaining the cache area data storage deviation value; Data is classified according to its stability. Stable data (such as data that is frequently accessed and changes little) is stored in the high-speed cache area, and unstable data (such as data that is infrequently accessed or changes frequently) is stored in the low-speed cache area. For example, the data change frequency of path level L2 is once a month, while the data change frequency of L3 is once a week. Based on this, L2 data is classified as stable and L3 as unstable. Through this classification, the storage ratio of each type of data in the high-speed and low-speed cache areas is compared to determine whether there is a deviation. The deviation threshold is set to 10%. If the deviation between the high-speed cache ratio and the low-speed cache ratio of L3 reaches 15%, adjustments are made to finally obtain the cache area data storage deviation value.

[0025] S213: Based on the data storage deviation value of the cache area, analyze the adjustment range of the storage area to the access jump degree, using the formula: ; Calculate the storage area adjustment range to obtain the cache space allocation status; in, Represents the storage area adjustment range, Representative The amount of low-speed cache storage at each level, Representative The cache storage capacity of the level, Representative The frequency of visits to the level, Representative Baseline visit frequency of the tier, Representative The storage adjustment factor of the tier; According to the data storage deviation value of the cache area, the adjustment range of the storage area to the access jump degree is further calculated. , , , , Respectively represent The storage capacity of the low-speed cache area, the storage capacity of the high-speed cache area, the actual access frequency, the reference access frequency, and the storage adjustment factor of the level; Set specific parameter values, assuming there are three path levels ( ), and its monitoring data are as follows: Tier 1: GB, GB, times / week, times / week, ; Tier 2: GB, GB, times / week, times / week, ; Tier 3: GB, GB, times / week, times / week, ; Substitute the above values ​​into the formula to calculate: Compute the sum of the stored differences: ; Calculate the square root of the sum of the squares of the visit frequency deviations: ; Calculate the sum of the storage adjustment factors: ; Calculate the final adjustment: ; The calculation results show that the adjustment range of the storage area is 0.807, which means that under the current path level distribution, the storage ratio of high-speed and low-speed caches needs to be adjusted accordingly to reduce the mismatch between access frequency and storage deviation. Based on this adjustment range, storage reallocation is performed to migrate part of the low-speed cache data to the high-speed cache or vice versa, so as to finally obtain the cache space allocation status.

[0026] See also Figure 4 , the specific steps for obtaining the cache data priority are: S311: extracting data block access records according to the cache space allocation status, calculating the number of data block accesses and time intervals, analyzing the access intersection status, and obtaining the access intersection metric value; Monitor the access of each data block in the cache. Assume that in a cloud storage service, analyze which data blocks are frequently accessed based on log files. The information is used for subsequent optimization operations. Count the number of accesses and time intervals for each data block. This statistic is based on real-time data stream analysis. For example, in a database, data block A is accessed 100 times within an hour, with an average access interval of 5 minutes. Calculate the access intersection, which refers to the situation where different data blocks are accessed by multiple users or programs in the same period of time. By setting the cross-access threshold, for example, if two data blocks are accessed more than 50 times within 10 minutes, it is considered that there is a high cross-access, and finally the access cross-access metric value is obtained. This value is obtained through comparative analysis. For example, the above data blocks A and B are accessed 120 times in total within 10 minutes, so their access cross-access metric value is determined to be high.

[0027] S312: calling the access crossover metric value, filtering the crossover data blocks according to the crossover number threshold, calculating the cache occupancy ratio, analyzing the distribution trend, and obtaining the contention data block set; Retrieve the crossover metric value calculated in the previous step from the database. For example, extract the access crossover metric value of data blocks A and B from the cache management software, and filter the crossover data blocks according to the crossover number threshold. The screening is based on the preset crossover access threshold. For example, if the crossover number exceeds 100 times, it is considered to be a high-crossover data block. Calculate its cache occupancy ratio, that is, the ratio of the cache size occupied by the data block to the total cache size. For example, data blocks A and B occupy 30% of the cache together. Analyze the distribution trend. Determine whether the distribution of data blocks in the cache is concentrated or dispersed through statistical analysis. For example, a chart shows that data blocks A and B are mainly concentrated in the first 50% of the cache. Get a set of competing data blocks. The set reflects the data blocks that need to be managed first. For example, data blocks A and B are marked as high-competition data blocks because they are frequently accessed and occupy a large cache.

[0028] S313: Based on the set of competing data blocks, analyze the cache storage status and data block access priority, adjust the storage order and elimination strategy according to the priority, and use the formula: ; Get the cache data priority; in, Represents the cache data priority, Represents a data block The number of crossover visits, Represents a data block The time interval between the last visit, Represents a data block The cache occupancy ratio, Represents a data block The storage location index in the cache, Represents the mean value of the data block storage location index in the cache; First, extract the data blocks marked as high contention from the log or cache management module. For example, in the high-speed cache of a cloud computing platform, data blocks A and B have been identified as high-contention data blocks by the system because they are frequently accessed by multiple users. Read their relevant parameters from the database or cache monitoring, including access frequency, access interval, cache occupancy ratio and storage location index, and analyze the cache storage status. The analysis process involves calculating the utilization of the current cache storage and checking the distribution of high-contention data blocks in the cache. For example, the current total cache capacity is 200GB, and data blocks A and B occupy 15GB and 10GB respectively, occupying a total of 12.5% ​​of the cache space. At the same time, count the distribution of all data blocks, evaluate whether there is a local storage hotspot, and calculate the access priority of the data blocks. This process needs to be based on the number of access crossovers, the most recent access time interval, the cache occupancy ratio and the storage location index of the data blocks. A specific calculation is performed on data block A, where the number of access crossovers of data block A is , the time interval between the last visit Minutes, cache occupancy ratio (15GB / 200GB), storage location index , the mean value of the cache data block storage location index , substitute into the formula and calculate as follows: ; Similarly, if the data block B is calculated, its parameters are set as , minute, (10GB / 200GB), , and the same calculation is obtained: ; Adjust the storage order and elimination strategy according to the cache priority. For example, in cache management, data block B with a higher cache priority is moved to a higher level cache, while data block A with a lower priority will be moved to a slower storage area or considered for elimination. Finally, the cache data priority is obtained. The result shows that data block B will be stored first due to its higher access priority, while data block A will be replaced or have its storage level lowered due to its lower priority.

[0029] See also Figure 5 , the specific steps for obtaining the storage medium load change trend are: S411: calling the cache data priority, monitoring the storage medium I / O tasks, calculating the storage medium I / O response time ratio, screening the storage medium whose task ratio exceeds the task threshold, and obtaining the I / O task ratio distribution data; First, set different data priorities. For example, in actual application scenarios, such as financial transactions, transaction data has the highest priority, followed by user verification data. Calculate the proportion of each data type in the total task and set a task threshold, assuming it is 30%. When the task proportion of any data type exceeds this threshold, the alarm mechanism is automatically activated to adjust the priority of high-proportion tasks. Next, monitor the I / O response of the storage medium. If the I / O response time is found to be long, reallocate this part of the data to optimize the overall storage medium performance. Finally, obtain the I / O task proportion distribution data. For example, in a financial transaction, it is found that the I / O response time of transaction data is 0.01 seconds, and that of user verification data is 0.02 seconds.

[0030] S412: Analyze the load change rate of the storage medium based on the I / O task proportion distribution data, adjust the cache data storage ratio of the load storage medium, and obtain a cache allocation adjustment coefficient; The load change rate of each storage medium is analyzed. By real-time monitoring of data access frequency and data read and write times, the load change of each storage medium is calculated. If the load change rate of a storage medium exceeds the safety threshold, its data storage strategy is automatically adjusted. For example, data writing to the storage medium is reduced, and data backup operations on the storage medium are increased. The cache allocation adjustment coefficient is obtained. For example, if the adjustment coefficient is 1.5, it means that the data write rate needs to be increased by 50% to reduce the pressure on high-load storage media.

[0031] S413: Adjust the data storage medium according to the cache allocation adjustment coefficient, analyze the load change trend after the data adjustment, and use the formula: ; Obtain storage media load change trends; in, Represents the storage medium load change trend value. represents the load level of the jth storage medium after adjustment, represents the load level of the jth storage medium before adjustment, represents the original task load stability of the jth storage medium, represents the data access fluctuation amplitude of the jth storage medium, Represents the total amount of storage media to be adjusted; Adjust the data storage medium distribution and recalculate the storage medium load change trend. In actual applications, for example, a data storage contains 5 storage media, and the initial load level of each storage medium is different. After adjusting the storage strategy, the load level of each storage medium also changes accordingly. Assume that the initial load levels are: , , , , ; After adjustment, the storage medium load level changes to: , , , , ; The historical task load stability and data access fluctuation range are set as follows: , , , , ; , , , , ; Substitute the formula to gradually calculate the load change trend value of each storage medium: ; ; ; ; ; Add the load change trend values ​​of each storage medium to obtain the total load change trend value of the storage medium: ; The overall load adjustment of the storage medium is low, indicating that the data allocation strategy still maintains a certain balance after adjustment, and does not cause drastic changes in the load of some storage media. The load optimization threshold can be set based on this value. For example, if the threshold is set to 0.5, if the total load change trend value If it exceeds 0.5, it is necessary to further adjust the data storage strategy, reallocate some data to storage media with lower load, and establish a storage media load change trend.

[0032] See also Figure 6 , the steps to obtain the storage status of data at the cache level are as follows: S511: based on the storage medium load change trend, extract the data block access frequency, migration times, and storage duration, analyze the data matching degree, filter the matching conditions, and obtain the data matching status classification value; First, the access frequency of each data block is monitored. For example, in a cloud storage environment, the number of accesses to each data block per hour is obtained through log analysis tools. The access data is then processed through a series of numerical integration, including averaging and comparing historical data to determine the activity and usage frequency of the data block, and calculating the number of migrations of data blocks between different storage layers. This involves tracking storage operation logs, such as the number of records of data blocks moving from the cache layer to the low-speed storage layer. It is also necessary to calculate the storage duration of the data block, that is, the total storage time of the data block since its creation. The timestamp information of the data block needs to be extracted from the metadata management. According to the set storage adaptation threshold, if a data block has an access frequency of less than 10 times / day, it is marked as low-frequency access. The matching status of the data block is screened and classified, and finally the data matching status classification value is obtained. The value reflects the storage efficiency and access pattern of the data, which is calculated based on the actual storage usage.

[0033] S512: Call the data matching status classification value to analyze the data throughput of the storage layer, using the formula: ; Calculate the data throughput state characteristic value, determine the storage area adjustment demand, and obtain the storage area adjustment demand; in, Represents the characteristic value of data throughput status, Represents the current storage level throughput, Represents the raw storage tier throughput, represents the throughput fluctuation range, represents the mean throughput, represents the data migration rate, Represents the data migration impact factor; Monitor the data input and output rates of each storage layer and measure the data read and write speed. For example, in a distributed storage, the storage layers include cache (such as SSD), standard storage (such as HDD) and archive storage (such as tape storage). First, calculate the data throughput of each storage layer. Assume that the current throughput of the SSD layer is MB / s, while the historical throughput for the previous cycle was MB / s, the difference between the two can be used to evaluate the throughput change trend, calculate the average and fluctuation range of throughput, and assume that the average throughput data of this layer in the past 24 hours is MB / s, with a fluctuation range of MB / s, it can be used to determine whether the current throughput is in a stable state and calculate the data migration rate. This value can be obtained by counting the total number of migrations of statistical blocks between different storage tiers. Assuming that the current monitored data migration rate is , that is, 12% of the data blocks are migrated to different storage tiers every hour, when calculating the data throughput state characteristic value; Among them, the data migration influencing factor The value depends on the type of storage media to be migrated and the impact of the migration. Assume that the current setting is , substitute all the values ​​into: ; ; ; Calculated data throughput state characteristic value , used to determine whether the storage area needs to be adjusted. If the set storage adjustment threshold is 60, the current calculated value has exceeded this threshold, indicating that the throughput of the storage area needs to be optimized. The calculated result of the storage area adjustment demand is that adjustment is required. The result can be used to further adjust the data storage allocation strategy, such as migrating part of the data from the SSD layer to the HDD layer, or adding additional cache at the SSD layer.

[0034] S513: Based on the storage area adjustment demand, according to the cache allocation rule, analyzing the adjusted data storage status, and outputting the storage status of the data at the cache level; Adjust the position of data between different storage layers according to the data throughput status characteristic value. If the throughput status characteristic value of a certain layer shows data congestion, migrate some data to a lower or higher-speed storage layer to balance the load. The adjustment process involves the application of storage resource allocation algorithms, such as dynamically adjusting the storage location of data blocks based on preset thresholds and throughput models to improve the overall performance and response speed of storage, and outputting the adjusted data storage status. This is achieved by updating the storage metadata and user interface to ensure that users and administrators can clearly see the adjusted data distribution and output the storage status of data at the cache layer.

[0035] The multi-layer data cache dynamic management system is used to execute the multi-layer data cache dynamic management method, and the system includes: The path access analysis module extracts the hierarchical identifier based on the data access path, identifies the hierarchical depth, records the path access sequence, selects stable paths and calculates the distribution ratio to obtain the path hierarchical distribution data; The cache hierarchical allocation module calculates the storage ratio of the high-speed cache area and the low-speed cache area based on the path hierarchical distribution data, stores the stable path set and the unstable path set, adjusts the jump path storage area, and obtains the cache space allocation status; The data competition detection module extracts data block access records based on the cache space allocation status, filters the cross-accessed data blocks in a short period of time, calculates the number of cross-accesses and classifies them, analyzes the storage ratio of cross-accessed data blocks in the cache area, adjusts the storage order and filters the cross-accessed data blocks, and obtains the data competition evaluation results; The cache data priority adjustment module calls the cache data priority based on the data competition evaluation result, identifies the load of the storage medium I / O task, adjusts the storage ratio, selects part of the data for storage medium migration, and obtains the storage medium load status; The storage medium load balancing module screens low-frequency access data and adjusts the storage area based on the storage medium load status, calls the cache allocation rules, calculates the cache data throughput, adjusts the data storage level, and obtains the storage status of the data at the cache level.

[0036] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A multi-layer data cache dynamic management method, characterized in that: The following steps are involved: S1: Based on the data access path, extract the data access path level identifier, analyze the path level depth, record the access continuity, classify the path as stable or unstable, and obtain the path level distribution information; S2: Based on the path level distribution information, analyze the data cache level storage ratio, store stable path data in the high-speed cache area, store unstable data in the low-speed cache area, adjust the storage area, and obtain the cache space allocation status; S3: extracting data block access records, detecting access crossover frequency, filtering frequently crossover data blocks, analyzing cache storage status, adjusting storage order and data elimination, and obtaining cache data priority according to the cache space allocation status; S4: calling the cache data priority, monitoring the storage medium I / O tasks, analyzing the task execution status, adjusting the cache data storage ratio according to the task load, adjusting the storage medium of part of the data, and obtaining the storage medium load change trend; S5: Based on the storage medium load change trend, analyze the data matching situation, call the cache allocation rule, analyze the data throughput situation, and output the storage status of the data at the cache level.

2. The multi-layer data cache dynamic management method according to claim 1, characterized in that: The path level distribution information includes the path level identifier, the path level depth, the path access continuity, the access stability parameter classification result, and the path stability category; the cache space allocation status includes the data cache level storage ratio, the cache area storage data, the unstable path storage data, and the storage area adjustment result; the cache area data priority includes the data block access record, the frequently accessed cross-data block, the cache area storage status, the data storage order, and the competition elimination data; the storage medium load change trend includes the storage medium I / O task, the task execution status, the task load, and the storage medium adjustment result; the storage status of the data at the cache level includes the data matching status, the storage area allocation, the cache allocation rule, and the data throughput status.

3. The multi-layer data cache dynamic management method according to claim 1, characterized in that: The steps for obtaining the path level distribution information are specifically as follows: S111: based on the data access path, extract the path level identifier, analyze the level depth, record the path access continuity, identify the level access times, and obtain the path level distribution characteristic value; S112: calling the path level distribution characteristic value, classifying the paths according to the access stability parameter, screening the stable paths and the unstable paths, calculating the proportion of the stable paths, and obtaining the path stability ratio; S113: calling the path stability ratio, combining the access trend of each level path, analyzing the access change range of the differentiated level paths, using the formula: ; Calculate the path level access fluctuation value and obtain the path level distribution information; in, Represents the path level access fluctuation value, Representative The number of visits to the level, Represents the number of visits to the previous level. Representative The total number of levels in the hierarchy, Representative The stability ratio of the hierarchical path, Represents the total number of path levels.

4. The multi-layer data cache dynamic management method according to claim 3, characterized in that: The steps for obtaining the cache space allocation status are specifically as follows: S211: Based on the path level distribution information, calculate the data storage amount of the path level, compare the storage capacity, filter the path levels whose storage ratio exceeds the threshold, and obtain the path level storage ratio screening result; S212: calling the path level storage ratio screening result, storing the data in differential cache areas according to data stability classification, comparing the storage ratios of high-speed and low-speed cache areas, and obtaining a cache area data storage deviation value; S213: Based on the data storage deviation value of the cache area, analyze the adjustment range of the storage area to the access jump degree, using the formula: ; Calculate the storage area adjustment range to obtain the cache space allocation status; in, Represents the storage area adjustment range, Representative The amount of low-speed cache storage at each level, Representative The cache storage capacity of the level, Representative The frequency of visits to the level, Representative Baseline visit frequency of the tier, Representative The storage scaling factor for the tier.

5. The multi-layer data cache dynamic management method according to claim 4, characterized in that: The steps for obtaining the priority of the cache data are specifically as follows: S311: extracting data block access records according to the cache space allocation status, calculating the number of data block accesses and time intervals, analyzing access intersection conditions, and obtaining access intersection metrics; S312: calling the access intersection metric value, filtering the intersection data blocks according to the intersection number threshold, calculating the cache occupancy ratio, analyzing the distribution trend, and obtaining the contention data block set; S313: Based on the set of competing data blocks, analyze the cache storage status and data block access priority, adjust the storage order and elimination strategy according to the priority, and use the formula: ; Get the cache data priority; in, Represents the cache data priority, Represents a data block The number of crossover visits, Represents a data block The time interval between the last visit and the last visit of Represents a data block The cache occupancy ratio, Represents a data block The storage location index in the cache, Represents the mean value of the data block storage location index in the cache.

6. The multi-layer data cache dynamic management method according to claim 5, characterized in that: The steps of acquiring the storage medium load change trend are specifically as follows: S411: calling the cache data priority, monitoring the storage medium I / O tasks, calculating the storage medium I / O response time ratio, screening the storage medium whose task ratio exceeds the task threshold, and obtaining I / O task ratio distribution data; S412: Analyze the load change rate of the storage medium based on the I / O task proportion distribution data, adjust the cache data storage ratio of the load storage medium, and obtain a cache allocation adjustment coefficient; S413: Adjust the data storage medium according to the cache allocation adjustment coefficient, analyze the load change trend after the data adjustment, and use the formula: ; Obtain storage media load change trends; in, Represents the storage medium load change trend value. represents the load level of the jth storage medium after adjustment, represents the load level of the jth storage medium before adjustment, represents the original task load stability of the jth storage medium, represents the data access fluctuation amplitude of the jth storage medium, Indicates the total number of storage media to be adjusted.

7. The multi-layer data cache dynamic management method according to claim 6, characterized in that: The steps for obtaining the storage status of the data at the cache level are specifically as follows: S511: based on the storage medium load change trend, extract the data block access frequency, migration times, and storage duration, analyze the data matching degree, filter the matching conditions, and obtain the data matching status classification value; S512: Call the data matching status classification value to analyze the data throughput of the storage layer, using the formula: ; Calculate the data throughput state characteristic value, determine the storage area adjustment demand, and obtain the storage area adjustment demand; in, Represents the characteristic value of data throughput status, Represents the current storage level throughput, Represents the raw storage tier throughput, represents the throughput fluctuation range, represents the mean throughput, represents the data migration rate, represents the data migration influencing factor; S513: Based on the storage area adjustment demand, according to the cache allocation rule, analyzing the adjusted data storage status, and outputting the storage status of the data at the cache level.

8. A multi-layer data cache dynamic management system, characterized in that: According to the multi-layer data cache dynamic management method according to any one of claims 1 to 7, the system comprises: The path access analysis module extracts the hierarchical identifier based on the data access path, identifies the hierarchical depth, records the path access sequence, selects stable paths and calculates the distribution ratio to obtain the path hierarchical distribution data; The cache layer allocation module calculates the storage ratio of the high-speed cache area and the low-speed cache area based on the path layer distribution data, stores the stable path set and the unstable path set, adjusts the jump path storage area, and obtains the cache space allocation status; The data competition detection module extracts data block access records based on the cache space allocation state, filters the cross-accessed data blocks in a short period of time, calculates the number of cross-accesses and classifies them, analyzes the storage proportion of the cross-accessed data blocks in the cache area, adjusts the storage order and filters the cross-accessed data blocks, and obtains the data competition evaluation result; The cache data priority adjustment module calls the cache data priority based on the data competition evaluation result, identifies the load of the storage medium I / O task, adjusts the storage ratio, selects part of the data for storage medium migration, and obtains the storage medium load status; The storage medium load balancing module screens low-frequency access data to adjust the storage area based on the storage medium load status, calls cache allocation rules, calculates cache data throughput, adjusts the data storage level, and obtains the storage status of data at the cache level.

Citation Information

Patent Citations

  • Method and system for realizing consistency of distributed multi-level caches of industrial data

    CN118820133A

  • Shunting processing method, system and equipment for cache data and storage medium

    CN119025560A

  • Cache optimization method and device, equipment, storage medium and computer program product

    CN119484641A

  • Data interaction method, device, equipment and storage medium for multi-layer stacked memory

    CN119781696A

  • Use of differing granularity heat maps for caching and migration

    US20140207995A1

Cited By

  • Cloud multi-modal data dynamic archiving system for big data

    CN120179184A

  • CAD drawing loading method and system combined with cache optimization

    CN120315779A

  • A CAD drawing loading method and system combined with cache optimization

    CN120315779B

  • Self-adaptive dynamic priority caching method, system, equipment and medium

    CN120578349A