Intelligent hierarchical caching method and system based on hyper-converged architecture

By building a progressive integrated technology chain and a multi-dimensional prediction model, combined with a risk-sensitive reward-improved reinforcement learning method, the problems of design fragmentation and disconnection between prediction and scheduling in existing intelligent hierarchical caching methods are solved, dynamic adaptability and global optimization of cache management are achieved, the misplacement rate of hot and cold data and the risk of media aging are reduced, and the stability and performance of the system are improved.

CN120705077AActive Publication Date: 2025-09-26TIANJIN DEV ZONE ESINT NETWORK SYST

Patent Information

Application Number
CN202511207812.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-26
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

The existing intelligent hierarchical caching method based on hyper-converged architecture has a design fragmentation at the overall architectural level. The prediction and scheduling between each link are relatively disconnected, resulting in the inability to dynamically optimize the cache tiering strategy with the system operation status. It is difficult to balance long-term performance and media life. The initial tiering results are disconnected from the actual business load. The load prediction granularity is coarse and the neighborhood correlation is insufficiently captured. The scheduling strategy is unbalanced and the physical constraints of modeling capacity, bandwidth, temperature and life are not effectively unified, resulting in system instability or premature aging of the media under extreme load.

Method used

A progressive integrated technology chain consisting of 'initialization classification - multi-dimensional load prediction - constraint-aware scheduling - intelligent cache execution' is constructed. An initial cache classification method based on heat score calculation is adopted, and an adaptive load three-channel prediction model improved by timing alignment is introduced. Combined with the reinforcement learning method improved by risk-sensitive rewards, cache resource scheduling is carried out to achieve the full-process linkage of resource status perception, load trend prediction and risk constraint optimization.

Benefits of technology

It significantly improves the dynamic adaptability and global optimization capabilities of cache management, reduces the misplacement rate of hot and cold data and the initial scheduling cost, improves the time consistency and resource availability of prediction results, dynamically balances performance and resource security, and reduces the risk of cache system instability or premature media aging under high load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705077A_ABST
    Figure CN120705077A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent hierarchical caching method and system based on a hyper-converged architecture. The method comprises the steps of converged architecture collection, initial cache grading, three-way load prediction, cache resource scheduling and intelligent hierarchical caching. The invention relates to the technical field of intelligent grading of data caches, in particular to an intelligent grading caching method and system based on a hyper-converged architecture, and adopts an initial cache grading method based on popularity score calculation to realize a more dynamic, fine-grained and controllable three-layer cache initialization structure. The cold and hot data misplacement rate and the initial scheduling cost are effectively reduced; a self-adaptive load three-channel prediction model improved by time sequence alignment enhancement is adopted, a time sequence alignment enhancement and neighborhood enhancement mechanism is introduced, and a dual-channel gating fusion backbone network of a time channel and a resource channel is combined to carry out three-path load prediction; and performing cache resource scheduling by adopting a reinforcement learning method based on risk-sensitive reward improvement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent hierarchical data caching, and specifically to an intelligent hierarchical caching method and system based on a hyper-converged architecture. Background Art

[0002] An intelligent hierarchical caching method and system based on a hyper-converged architecture is a distributed technology that integrates computing, storage, and network resources. Through dynamic data tiering and intelligent caching strategies using machine learning algorithms, it prioritizes high-frequency data for storage in high-performance storage tiers and shifts low-frequency data to low-cost storage tiers. This significantly reduces I / O latency and improves resource utilization. While leveraging the horizontal scalability of hyper-convergence to ensure high system availability, it is suitable for scenarios such as cloud computing and edge computing that require real-time response and massive data processing.

[0003] However, existing intelligent approaches suffer from design fragmentation at the overall architectural level, with prediction and scheduling between various links being disconnected. This results in cache tiering strategies being unable to dynamically optimize based on system operating conditions, and makes it difficult to balance long-term performance with media lifespan.

[0004] The existing initial cache tiering process relies heavily on static media characteristics or simple access frequency, lacking comprehensive modeling of multiple factors such as access timing, read / write characteristics, and request size. This leads to a technical problem in which the initial tiering results are out of sync with actual business loads.

[0005] Existing load prediction methods have technical problems such as coarse prediction granularity in cache management, insufficient capture of sudden access patterns and neighborhood correlations, and a lack of constraint fusion between multi-channel prediction results, which can easily lead to imbalanced scheduling strategies.

[0006] In existing cache resource scheduling methods, there are simple reward function designs that only consider performance improvement or migration costs, without uniformly modeling the physical constraints of capacity, bandwidth, temperature and lifespan with delay risks. This leads to technical problems such as scheduling strategies easily violating hard limits or causing lifespan overdraft under extreme loads. Summary of the Invention

[0007] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides an intelligent hierarchical caching method and system based on a hyper-converged architecture. In view of the existing data hierarchical caching method based on a hyper-converged architecture, there are problems in the existing intelligent method that there is design fragmentation at the overall architecture level, and the prediction and scheduling between each link are relatively disconnected, resulting in the inability to dynamically optimize the cache grading strategy with the system operation status, and it is difficult to balance long-term performance and media life. Starting from the top-level design of the system, this solution constructs a progressive integrated technology chain consisting of "initialization grading-multi-dimensional load prediction-constraint-aware scheduling-intelligent cache execution", which realizes resource status perception, load trend prediction, risk constraint optimization and hierarchical execution under a unified architecture. The whole process linkage significantly improves the dynamic adaptability and global optimization ability of cache management, and solves the structural bottleneck of "static initial setting, step-by-step independence, and lack of feedback" of the traditional method; in view of the technical problem that the existing initial cache classification setting relies more on static media characteristics or simple access frequency, lacks comprehensive modeling of multiple factors such as access timing, read-write characteristics and request size, resulting in the initial classification results being out of touch with the actual business load, this solution creatively adopts an initial cache classification method based on heat score calculation, introduces weighted indicators such as number of visits, recency, read-write ratio, and request size, and combines hysteresis interval control and usage condition restrictions to achieve more dynamic, fine-grained and The controllable three-layer cache initialization structure effectively reduces the misplacement rate of hot and cold data and the initial scheduling cost. In view of the technical problems in the existing load prediction methods, such as the coarse prediction granularity of traditional load prediction methods in cache management, insufficient capture of sudden access patterns and neighborhood correlations, and lack of constraint fusion between multi-channel prediction results, which easily leads to imbalance in scheduling strategies, this solution creatively adopts the adaptive load three-channel prediction model improved by timing alignment enhancement to perform three-way load prediction. By introducing timing alignment enhancement and neighborhood enhancement mechanisms, combined with the dual-channel gating fusion backbone network of time channel and resource channel, prototype constraints, cross-time domain consistency constraints and single-channel constraints are embedded in the three-way prediction of access type, access heat and high latency risk. Tonality constraints achieve significant improvements in prediction results in terms of time consistency, category stability and resource availability, providing a more robust multi-dimensional prediction basis for cache scheduling; in existing cache resource scheduling methods, there is a simple reward function design that only considers performance improvement or migration cost, but does not uniformly model the physical constraints of capacity, bandwidth, temperature and lifespan with delay risk, resulting in the scheduling strategy easily violating hard limits or causing lifespan overdraft under extreme loads. This solution creatively adopts a reinforcement learning method based on risk-sensitive reward improvement for cache resource scheduling, so that scheduling decisions can dynamically balance between performance and resource security, significantly reducing the risk of cache system instability or premature media aging under high load.

[0008] The technical solution adopted by the present invention is as follows: The present invention provides an intelligent hierarchical caching method based on a hyper-converged architecture, the method comprising the following steps:

[0009] Step S1: Fusion architecture acquisition;

[0010] Step S2: initial cache hierarchy;

[0011] Step S3: three-way load prediction;

[0012] Step S4: cache resource scheduling;

[0013] Step S5: Intelligent hierarchical caching.

[0014] Furthermore, in step S1, the fusion architecture collection is used to collect and organize the operation status and data access status data of computing, storage and network devices, specifically by building a hyper-converged architecture to collect, integrate and store data resources to obtain the original data of system operation;

[0015] The specific content of the system operation original data includes equipment operation status data, storage medium attribute data and data access record data.

[0016] Furthermore, in step S2, the initial cache hierarchy is used to divide the cache space into initial levels according to the speed, capacity, and cost of different storage media, and set the usage conditions and restrictions of each level. Specifically, based on the original data of the system operation, the initial cache hierarchy is constructed by initializing the three-level structure of the cache level to obtain the cache level initialization structure, including the following steps:

[0017] Step S21: Cache tier division, specifically, dividing the cache space into a fast tier, a high-speed tier, and a capacity tier. The fast tier uses media with the lowest latency and the smallest capacity among the three tiers, the high-speed tier uses media with intermediate latency and capacity among the three tiers, and the capacity tier uses media with the highest latency and the highest capacity among the three tiers.

[0018] Step S22: Data block management, specifically using fixed-size data blocks as management units, and calculating a heat score for each data block by comprehensively weighting the access count factor, access recency factor, read-write ratio factor, and request size factor based on historical access records.

[0019] Step S23: Setting a file write strategy, specifically setting a hierarchical strategy for file writes in the storage system, including temporarily storing newly written data in a write buffer of the fast tier, and subsequently migrating the data to other storage tiers in priority order based on access feature determination results;

[0020] Prioritize writing large-volume sequential read data to the high-speed tier to reduce the fast tier's occupancy and optimize subsequent access performance.

[0021] Step S24: hysteresis interval control, used as a scheduling screening condition for cache resource scheduling in step S4, specifically setting a dual-threshold hysteresis control mechanism, setting independent heat score thresholds for upward and downward adjustments between the fast tier, high-speed tier, and capacity tier, and setting a data block regulation mechanism between tiers;

[0022] Step S25: Usage restrictions. Specifically, unified usage conditions and restrictions are set for different tiers, including limiting the capacity utilization of each tier to no higher than a preset capacity threshold; reducing the data migration rate and prioritizing data migration to the low-speed tier when the device temperature reaches a preset temperature threshold; setting a periodic write budget for solid-state media and limiting the write ratio to no more than the budget; placing data that exceeds a preset object capacity threshold in a single write by default in the high-speed tier or capacity tier; and reserving a minimum access bandwidth for foreground services in various operations to ensure business continuity.

[0023] Step S26: Initial cache hierarchy, specifically generating a cache hierarchy initialization structure, including three layers of media type and capacity upper limit, migration rate upper limit, temperature and life threshold, hysteresis interval, warm-up list and fixed list, and initial hierarchy mapping table for each data block;

[0024] The cache level initialization structure is used as input for subsequent steps, for analysis by three-way load prediction in step S3, and for planning and boundary control by cache resource scheduling in step S4.

[0025] Furthermore, in step S3, the three-way load prediction is used to predict the business type, access popularity, and high latency risk of each type of data. Specifically, based on the system operation original data and the cache level initialization structure, an adaptive load three-channel prediction model enhanced with timing alignment is used to perform three-way load prediction to obtain comprehensive load prediction reference data, including the following steps:

[0026] Step S31: time-aware alignment, specifically extracting access time series from the system operation raw data, and calculating the time-aware alignment factor of each data item to obtain time-aligned access feature data;

[0027] Step S32: Access domain enhancement, specifically, by calculating the neighborhood enhancement weights between adjacent data items, constructing access heat features between adjacent data items, and obtaining neighborhood enhancement access feature data;

[0028] Step S33: constructing an adaptive backbone network, specifically constructing a dual-channel gated fusion network as an adaptive backbone network for data feature extraction, including a time channel subnet, a resource channel subnet, and a feature fusion sub-block, and extracting features using the adaptive backbone network to obtain fused feature data;

[0029] The time channel subnet specifically performs time series convolution processing on the time feature vector and extracts access morphology features. The resource channel subnet specifically constructs a multi-layer perceptron and extracts resource status features from cache resource status parameters. The feature fusion subblock specifically constructs a gated fusion mechanism to perform gated fusion on the access morphology features and the resource status features to obtain fused feature data.

[0030] Step S34: constructing three prediction channels, specifically, constructing an access type prediction channel, an access popularity prediction channel, and a high latency risk prediction channel based on the adaptive backbone network, thereby obtaining three prediction channels, and performing data prediction for the access type prediction task, the access popularity prediction task, and the high latency risk prediction task, respectively, to obtain task branch prediction data;

[0031] Step S35: Soft budget constraint loss is improved, specifically by introducing a soft budget constraint loss. By calculating the data occupied capacity exceeding the threshold under the predicted popularity in the next 10 minutes, and comparing the data occupied capacity exceeding the threshold with the available budget capacity, the excess portion is included in the optimization target according to the square loss, limiting the prediction result to the range of system capacity and resource budget, and integrating the soft budget constraint loss with the prototype constraint loss, the cross-temporal consistency loss, the monotonicity constraint loss and the task prediction loss, a joint optimization loss function is obtained.

[0032] Step S36: Three-way load prediction, specifically, by constructing the adaptive backbone network and the three-way prediction channel, combining the joint optimization loss function, performing model training to obtain a three-way load prediction model, and obtaining comprehensive load prediction reference data by using the three-way load prediction model;

[0033] The load comprehensive prediction reference data specifically includes data identification, access prediction type, predicted access popularity, predicted access delay and high delay risk prediction probability.

[0034] Furthermore, in step S4, the cache resource scheduling is used to formulate an optimal data placement and migration plan between cache layers under the constraints of capacity, bandwidth, temperature, and lifespan. Specifically, based on the system operation original data, the cache layer initialization structure, and the load comprehensive prediction reference data, a reinforcement learning method based on risk-sensitive reward improvement is used to perform cache resource scheduling to obtain cache resource scheduling decision data, including the following steps:

[0035] Step S41: State constraint construction: Based on the system operation raw data, cache layer initialization structure, and load comprehensive prediction reference data, a scheduling state vector and constraint boundaries are constructed. The state vector includes the popularity, latency risk, current layer, and recent migration records of each data object. The constraint boundaries include layer capacity, available bandwidth, temperature, lifespan thresholds, and their budget parameters.

[0036] Step S42: Risk-sensitive reward improvement, which is used to construct a reinforcement learning problem with budget and bandwidth constraints. Specifically, a reward function is constructed that introduces a performance benefit term, a migration cost penalty, and a risk-sensitive constraint penalty to obtain a risk-sensitive reward improvement reward function.

[0037] Step S43: reinforcement learning training, specifically, performing reinforcement learning training through the state constraint construction and the risk-sensitive reward improvement, by defining an action vector, to obtain a cache resource scheduling model;

[0038] The action vector includes a layer migration action, a migration target layer, and a migration speed limit;

[0039] The level migration action specifically includes maintaining the level, raising the level, and sinking the level operations;

[0040] Step S44: Cache resource scheduling, specifically, performing cache resource scheduling based on the system operation original data, the cache level initialization structure, and the load comprehensive prediction reference data by using the cache resource scheduling model to obtain cache resource scheduling decision data;

[0041] The cache resource scheduling decision data specifically includes a cache resource scheduling migration object identifier, a cache resource scheduling operation type, a cache resource source level, a cache resource target level, a migration speed limit, and a decision time.

[0042] Furthermore, in step S5, the intelligent hierarchical cache is used to perform data migration and adjustment between different cache layers, specifically to control data migration and hierarchical adjustment between different cache layers based on the cache resource scheduling decision data, and to move up, down or adjust the target data across layers according to a preset migration strategy under the condition of meeting the operation constraints and the heat score threshold, and to dynamically update the data distribution of each layer in combination with the real-time business priority and access heat to form a cache scheduling structure that meets the balance between performance and lifespan, thereby obtaining an intelligent cache scheduling structure.

[0043] The present invention provides an intelligent hierarchical cache system based on a hyper-converged architecture, comprising a data acquisition module, an initial grading module, a load prediction module, a resource scheduling module and an intelligent grading module;

[0044] The data acquisition module is used for fusion architecture collection, obtains the original system operation data through fusion architecture collection, and sends the original system operation data to the initial classification module, the load prediction module and the resource scheduling module;

[0045] The initial grading module is used for initial cache grading, obtains a cache level initialization structure through the initial cache grading, and sends the cache level initialization structure to the load prediction module and the resource scheduling module;

[0046] The load prediction module is used for three-way load prediction, obtains load comprehensive prediction reference data through the three-way load prediction, and sends the load comprehensive prediction reference data to the resource scheduling module;

[0047] The resource scheduling module is used for cache resource scheduling, obtains cache resource scheduling decision data through cache resource scheduling, and sends the cache resource scheduling decision data to the intelligent classification module;

[0048] The intelligent grading module is used for intelligent grading cache, and an intelligent cache scheduling structure is obtained through intelligent grading cache.

[0049] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0050] (1) In view of the existing data hierarchical caching methods based on hyper-converged architecture, there are problems such as design fragmentation at the overall architecture level and disconnection between prediction and scheduling between various links, which leads to the inability to dynamically optimize the cache tiering strategy with the system operation status and difficulty in balancing long-term performance and media life. Starting from the top-level design of the system, this solution constructs a progressive integrated technology chain consisting of "initialization tiering - multi-dimensional load prediction - constraint-aware scheduling - intelligent cache execution". Under a unified architecture, it realizes the full process linkage of resource status perception, load trend prediction, risk constraint optimization and hierarchical execution, significantly improving the dynamic adaptability and global optimization capabilities of cache management, and solving the structural bottleneck of "static initialization, step-by-step independence, and lack of feedback" of the traditional method;

[0051] (2) In view of the technical problem that the existing initial cache classification setting process relies more on static media characteristics or simple access frequency, lacks comprehensive modeling of multiple factors such as access timing, read-write characteristics and request size, resulting in the initial classification results being out of touch with the actual business load, this solution creatively adopts an initial cache classification method based on heat score calculation, introduces weighted indicators such as number of accesses, recency, read-write ratio, and request size, and combines hysteresis interval control and usage condition restrictions to achieve a more dynamic, fine-grained and controllable three-layer cache initialization structure, effectively reducing the hot and cold data misplacement rate and initial scheduling costs;

[0052] (3) In view of the technical problems in existing load prediction methods, such as coarse prediction granularity in cache management, insufficient capture of sudden access patterns and neighborhood correlations, and lack of constraint fusion between multi-channel prediction results, which easily leads to imbalance in scheduling strategies, this solution creatively adopts an adaptive load three-channel prediction model improved by timing alignment enhancement to perform three-way load prediction. By introducing timing alignment enhancement and neighborhood enhancement mechanisms, combining a dual-channel gating fusion backbone network of time channel and resource channel, and embedding prototype constraints, cross-time domain consistency constraints and monotonicity constraints in the three-way prediction of access type, access popularity and high latency risk, it achieves significant improvements in the prediction results in terms of time consistency, category stability and resource availability, providing a more robust multi-dimensional prediction basis for cache scheduling;

[0053] (4) In the existing cache resource scheduling methods, there is a simple reward function design that only considers performance improvement or migration cost, but does not uniformly model the physical constraints of capacity, bandwidth, temperature and lifespan with delay risk, resulting in the scheduling strategy easily violating hard limits or causing lifespan overdraft under extreme loads. This solution creatively adopts a reinforcement learning method based on risk-sensitive reward improvement to perform cache resource scheduling, so that scheduling decisions can dynamically balance performance and resource security, significantly reducing the risk of cache system instability or premature media aging under high loads. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A schematic diagram of a process flow of an intelligent hierarchical caching method based on a hyper-converged architecture provided by the present invention;

[0055] Figure 2 A schematic diagram of an intelligent hierarchical caching system based on a hyper-converged architecture provided by the present invention;

[0056] Figure 3 This is a schematic diagram of the process of initial cache classification in step S2;

[0057] Figure 4 This is a schematic diagram of the process of three-way load prediction in step S3;

[0058] Figure 5 This is a flow chart of cache resource scheduling in step S4.

[0059] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION

[0060] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0061] In the description of the present invention, it should be understood that terms such as "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limiting the present invention.

[0062] Example 1, see Figure 1The present invention provides an intelligent hierarchical caching method based on a hyper-converged architecture, the method comprising the following steps:

[0063] Step S1: Fusion architecture acquisition;

[0064] Step S2: initial cache hierarchy;

[0065] Step S3: three-way load prediction;

[0066] Step S4: cache resource scheduling;

[0067] Step S5: Intelligent hierarchical caching.

[0068] By performing the above operations, we can address the problems in the existing data hierarchical caching methods based on hyper-converged architecture, such as the design fragmentation of the existing intelligent methods at the overall architectural level, the disconnection between prediction and scheduling between various links, which results in the inability to dynamically optimize the cache tiering strategy with the system operating status, and the difficulty in balancing long-term performance and media life. Starting from the top-level design of the system, this solution constructs a progressive integrated technology chain consisting of "initialization tiering - multi-dimensional load prediction - constraint-aware scheduling - intelligent cache execution", and realizes the full-process linkage of resource status perception, load trend prediction, risk constraint optimization and hierarchical execution under a unified architecture, significantly improving the dynamic adaptability and global optimization capabilities of cache management, and solving the structural bottlenecks of traditional methods of "static initialization, step-by-step independence, and lack of feedback".

[0069] Example 2, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S1, the converged architecture collection is used to collect and organize the operation status and data access status data of computing, storage and network devices. Specifically, the data resources are collected, integrated and stored by building a hyper-converged architecture to obtain the original data of system operation;

[0070] The specific content of the system operation raw data, including equipment operation status data, storage medium attribute data and data access record data;

[0071] The device operating status data specifically includes processor utilization, memory usage, storage read and write speed, access queue length, network transmit and receive speed, and round-trip latency;

[0072] The storage medium attribute data specifically includes the storage medium operating temperature, available capacity and expected lifespan;

[0073] The data access record data specifically includes access object identification, access type, access time, access size, access interval, return delay, cache hit record and timeout record;

[0074] Preferably, the construction of a hyper-converged architecture specifically includes hardware resource deployment, network configuration management, data acquisition program deployment, data integration storage and capacity expansion design;

[0075] In a preferred embodiment, the hardware resource deployment specifically deploys at least three physical nodes, each of which is configured with an x86 or ARM architecture server, including a processor, memory, high-speed solid-state storage media, large-capacity mechanical storage media, and a high-speed network interface; computing virtualization is achieved through server virtualization software (such as KVM or VMware ESXi), and the physical CPU and memory are abstracted into a virtual resource pool; distributed storage software is used to aggregate the storage media of each node into a unified storage pool to support block, file, and object storage services;

[0076] The network configuration management specifically implements logical isolation and traffic scheduling between virtual machines by building a virtualized logical network, allocates high-priority bandwidth to storage traffic by configuring service quality control policies, and achieves elastic expansion of network functions by embedding virtual firewalls and load balancers in nodes;

[0077] The data acquisition program deployment specifically involves deploying a lightweight data acquisition program on each node to obtain the node's operating information, including computing resource status, storage resource status, network resource status, and data access records. The data acquisition program samples the operating information at preset time intervals or trigger conditions, and transmits the information to a centralized processing area via an internal data channel.

[0078] The data integration and storage is specifically to receive the operation information by building a centralized processing area and storage components, and to aggregate, classify and standardize the collected information from each node to obtain the system operation raw data in a unified format and store it in the storage component;

[0079] The capacity expansion design specifically refers to deploying multiple redundant copies of the data acquisition program and the storage components, and balancing data distribution through consistent hashing when adding new nodes.

[0080] Example 3, see Figure 1 、 Figure 2 and Figure 3 This embodiment is based on the above embodiment. In step S2, the initial cache hierarchy is used to divide the cache space into initial levels according to the speed, capacity, and cost of different storage media, and set the usage conditions and restrictions of each level. Specifically, based on the original data of the system operation, the cache hierarchy is initialized by constructing a three-layer structure, performing initial cache hierarchy, and obtaining a cache hierarchy initialization structure, including the following steps:

[0081] Step S21: Cache tier division, specifically, dividing the cache space into a fast tier, a high-speed tier, and a capacity tier. The fast tier uses media with the lowest latency and the smallest capacity among the three tiers, the high-speed tier uses media with intermediate latency and capacity among the three tiers, and the capacity tier uses media with the highest latency and the highest capacity among the three tiers.

[0082] Preferably, the initial capacity ratio of the three layers is 10%:30%:60%, and a 10% safety margin is reserved for each layer;

[0083] Step S22: Data block management, specifically using fixed-size data blocks as management units, and calculating a heat score for each data block by comprehensively weighting the access count factor, access recency factor, read-write ratio factor, and request size factor based on historical access records.

[0084] The calculation formula of the heat score is:

[0085] ;

[0086] Where H is the popularity score, w1 is the access count weight, F is the access count factor, w2 is the access recency weight, R is the access recency factor, w3 is the read-write ratio weight, W is the read-write ratio factor, w4 is the request size weight, and S is the request size factor.

[0087] Preferably, the default recommended weight of the access count weight is set to 0.4, the default recommended weight of the access recency weight is set to 0.3, the default recommended weight of the read-write ratio weight is set to 0.2, and the default recommended weight of the request size weight is set to 0.1;

[0088] The specific calculation formula of the visit number factor is:

[0089] ;

[0090] Where F is the access factor, ln is the natural logarithm function, count is the total number of accesses to the data block, and C max The upper limit of the number of configurable visits is 10,000 by default;

[0091] The specific calculation formula of the access recency factor is:

[0092] ;

[0093] Where R is the access recentness factor, exp is the natural base function, It is the time between the current time and the last visit, the unit is set to hours. is the time decay factor, which is set to 6 by default and is expressed in hours;

[0094] The specific calculation formula of the read-write scale factor is:

[0095] ;

[0096] Where W is the read-write ratio factor, read ps is the number of reads, write ps is the number of writes;

[0097] The specific calculation formula of the request size factor is:

[0098] ;

[0099] Where S is the request size factor, min is the minimum value function, AvgReqSize is the average request size, S ref The default value is 256, and the unit is KB.

[0100] Preferably, the size of the management unit is specifically set to 1MB or 4MB; when the heat score is ≥80, the data block is initially placed in the fast tier; when the heat score is 50-79, the data block is initially placed in the high-speed tier; when the heat score is <50, the data block is initially placed in the capacity tier;

[0101] Step S23: Setting a file write strategy, specifically setting a hierarchical strategy for file writes in the storage system, including temporarily storing newly written data in a write buffer of the fast tier, and subsequently migrating the data to other storage tiers in priority order based on access feature determination results;

[0102] Prioritize writing large-volume sequential read data to the high-speed tier to reduce the fast tier's occupancy and optimize subsequent access performance.

[0103] Preferably, the large volume is determined according to a relative capacity threshold. When the size of a single file or data block is ≥ 5% of the total capacity of the fast layer, it is considered as large volume data.

[0104] Step S24: hysteresis interval control, used as a scheduling screening condition for cache resource scheduling in step S4, specifically setting a dual-threshold hysteresis control mechanism, setting independent heat score thresholds for upward and downward adjustments between the fast tier, high-speed tier, and capacity tier, and setting a data block regulation mechanism between tiers;

[0105] Preferably, the migration direction and heat score threshold for the three-tier cache are set as follows:

[0106] When a data block is moved from the capacity tier to the high-speed tier, its heat score must be at least 60 points. When it is moved from the high-speed tier to the capacity tier, the migration can be triggered as long as its heat score is less than 50 points.

[0107] When a data block is transferred from the high-speed tier to the fast tier, its heat score must be at least 80 points. When it is transferred from the fast tier to the high-speed tier, the migration can be triggered only if its heat score is less than 70 points.

[0108] When a data block is directly promoted from the capacity tier to the express tier, its heat score must be at least 90 points. When it is directly demoted from the express tier to the capacity tier, the migration is triggered only if its heat score is less than 60 points.

[0109] Step S25: Usage restrictions. Specifically, unified usage conditions and restrictions are set for different tiers, including limiting the capacity utilization of each tier to no higher than a preset capacity threshold; reducing the data migration rate and prioritizing data migration to the low-speed tier when the device temperature reaches a preset temperature threshold; setting a periodic write budget for solid-state media and limiting the write ratio to no more than the budget; placing data that exceeds a preset object capacity threshold in a single write by default in the high-speed tier or capacity tier; and reserving a minimum access bandwidth for foreground services in various operations to ensure business continuity.

[0110] Preferably, the preset capacity threshold is set to 90% of the capacity by default;

[0111] Step S26: Initial cache hierarchy, specifically generating a cache hierarchy initialization structure, including three layers of media type and capacity upper limit, migration rate upper limit, temperature and life threshold, hysteresis interval, warm-up list and fixed list, and initial hierarchy mapping table for each data block;

[0112] The cache level initialization structure is used as input for subsequent steps, for analysis by three-way load prediction in step S3, and for planning and boundary control by cache resource scheduling in step S4.

[0113] By performing the above operations, we can address the technical problem that in the existing initial cache grading process, the existing initial cache grading settings mostly rely on static media characteristics or simple access frequency, lack comprehensive modeling of multiple factors such as access timing, read-write characteristics and request size, resulting in the initial grading results being out of touch with the actual business load. This solution creatively adopts an initial cache grading method based on heat score calculation, introduces weighted indicators such as number of accesses, recency, read-write ratio, and request size, and combines hysteresis interval control and usage condition restrictions to achieve a more dynamic, fine-grained and controllable three-tier cache initialization structure, effectively reducing the misplacement rate of hot and cold data and the initial scheduling cost.

[0114] Example 4, see Figure 1 、 Figure 2 and Figure 4 This embodiment is based on the above embodiment. In step S3, the three-way load prediction is used to predict the business type, access popularity, and high latency risk of each type of data. Specifically, based on the original data of the system operation and the cache level initialization structure, an adaptive load three-channel prediction model enhanced with timing alignment is used to perform three-way load prediction to obtain comprehensive load prediction reference data, including the following steps:

[0115] Step S31: Time-aware alignment, specifically extracting access time series from the system operation raw data and calculating the time-aware alignment factor of each data item to obtain time-aligned access feature data. The calculation formula is:

[0116] ;

[0117] Where, F i is the time-aligned access feature of the i-th data item in the system operation raw data, i is the data item index, Is the access time index, used as the time index for access time series. is the access time weight, count i is the number of visits, j is the neighborhood data item index, used as a normalized index, count j is the normalized number of visits;

[0118] Preferably, the access time weight is calculated by introducing a time decay coefficient and a data level change time node, and the calculation formula is:

[0119] ;

[0120] Where t is the current time index, is the time decay factor, ranging from [0.05, 0.2], is the data level migration compensation parameter, and its value range is [0.2, 0.6]. The whole is a conditional value function. When the condition When the condition is true, the value is 1, and when the condition is false, the value is 0. m It is the time for data level migration. is the migration process time parameter, the value range is [5,15], the unit is minutes;

[0121] Step S32: Access domain enhancement, specifically, by calculating the neighborhood enhancement weights between adjacent data items, constructing access heat features between adjacent data items, and obtaining neighborhood enhancement access feature data. The calculation formula is:

[0122] ;

[0123] Where, is the neighborhood enhanced access feature data of the i-th data item, j is the neighborhood data item index, i is the data item index, L ij is the neighborhood enhancement weight, is the heat score calculated in step S22 corresponding to the j-th neighborhood data item, is the heat score attenuation factor, ranging from [0.5, 1.5], and d(i, j) is the distance between adjacent data items;

[0124] The calculation formula of the neighborhood enhancement weight is:

[0125] ;

[0126] Where, L ij is the neighborhood enhancement weight, Coaccess(i,j) is the number of simultaneous accesses, and access is the number of times the data item is accessed;

[0127] Step S33: constructing an adaptive backbone network, specifically constructing a dual-channel gated fusion network as an adaptive backbone network for data feature extraction, including a time channel subnet, a resource channel subnet, and a feature fusion sub-block, and extracting features using the adaptive backbone network to obtain fused feature data;

[0128] The time channel subnet specifically performs time series convolution processing on the time feature vector and extracts access morphology features. The resource channel subnet specifically constructs a multi-layer perceptron and extracts resource status features from cache resource status parameters. The feature fusion subblock specifically constructs a gated fusion mechanism to perform gated fusion on the access morphology features and the resource status features to obtain fused feature data.

[0129] Preferably, the temporal channel subnet specifically uses a set of 1D convolutional layers (with a convolution kernel size of 3, a stride of 1, and 64 channels) combined with a temporal attention mechanism to capture short-term fluctuations and long-term trends in access patterns and output access morphological features; the resource channel subnet specifically uses a two-layer multilayer perceptron (with hidden layer dimensions of 128 and 64, and an activation function of ReLU) to perform nonlinear mapping on resource states to obtain resource state feature vectors;

[0130] Step S34: constructing three prediction channels, specifically, constructing an access type prediction channel, an access popularity prediction channel, and a high latency risk prediction channel based on the adaptive backbone network, thereby obtaining three prediction channels, and performing data prediction for the access type prediction task, the access popularity prediction task, and the high latency risk prediction task, respectively, to obtain task branch prediction data;

[0131] Preferably, the access type prediction channel uses a 256-unit fully connected layer for feature compression and a ReLU activation function, followed by a Softmax output layer with an output dimension equal to the number of business type categories. Prototype constraint loss is introduced during training to dynamically maintain the feature center of each category and limit the Euclidean distance between the current batch feature and the category center. The prototype constraint loss is calculated as follows:

[0132] ;

[0133] Where, is the prototype constraint loss, i is the data item index, is the fusion feature data, x i is the original data input corresponding to the i-th data item, It is the yth i The feature center of the predicted category, y i is the predicted category corresponding to the i-th data item, is the feature center stability weight, the default value is 0.1, c is the category index, is the feature center of the cth access category, is the feature center of the cth access category in the last iteration during model training, and ||·||2 is the L2 norm operator;

[0134] The access popularity prediction channel uses a shared two-layer 128-unit fully connected network to extract time-invariant features. At the output, it is split into three independent linear regression heads, corresponding to the 1-minute, 10-minute, and 1-hour prediction scales, respectively. Each regression head consists of a 64-unit ReLU fully connected layer and a linear output layer. By introducing cross-temporal consistency loss, the exponential sliding average is used to constrain the prediction smoothness between different time scales. The calculation formula of the cross-temporal consistency loss is:

[0135] ;

[0136] Where, is the cross-temporal consistency loss, It is the predicted visit popularity at the 10-minute scale. is the exponential moving average operator, is the first smoothing coefficient, is the second smoothing coefficient, It is the predicted visit popularity at the 1-minute scale. It is the predicted visit popularity at the hourly scale;

[0137] The high-latency risk prediction channel first extracts risk-related features through a 128-unit fully connected layer and uses a Sigmoid activation function to limit the value range to (0, 1). A linear layer is then used to generate a delay prediction value. A combination of load rate, concurrency, and session cost parameters is introduced when calculating the delay value. A monotonicity constraint loss is used to ensure that the delay prediction value does not decrease as the load rate increases. The calculation formula for the monotonicity constraint loss is:

[0138] ;

[0139] Where, is the monotonicity constraint loss, is the comprehensive prediction score of the i-th data item, which is used to measure the comprehensive quality of the prediction results, where p is the latency percentile index, and the specific value range is {95,99}, which is used to represent p95 latency and p99 latency. is the proportion of the control variable of the i-th data item, and its value range is (0,1);

[0140] The calculation formula of the comprehensive prediction score is:

[0141] ;

[0142] Where, is the comprehensive prediction score, It is the prediction result of the input data, specifically used to represent the prediction output of the high latency risk prediction channel. is the scaling factor, is the proportion of the control variable, with a value range of (0,1), c s is the scenario weight coefficient, which is used to express the estimation of service time dispersion, and s is the medium service time reference parameter;

[0143] Step S35: Soft budget constraint loss is improved, specifically by introducing a soft budget constraint loss. By calculating the data occupied capacity exceeding the threshold under the predicted popularity in the next 10 minutes, and comparing the data occupied capacity exceeding the threshold with the available budget capacity, the excess portion is included in the optimization target according to the square loss, limiting the prediction result to the range of system capacity and resource budget, and integrating the soft budget constraint loss with the prototype constraint loss, the cross-temporal consistency loss, the monotonicity constraint loss and the task prediction loss, a joint optimization loss function is obtained.

[0144] The calculation formula of the soft budget constraint loss is:

[0145] ;

[0146] Where, is the soft budget constraint loss, is the sigmoid function, is the fast layer heat threshold, is the smoothing parameter, C L1 is the maximum budget capacity corresponding to the fast layer, reserve is the budget reservation ratio parameter, It is a positive square operation, which is used to indicate that only the part exceeding the budget will be penalized;

[0147] The calculation formula of the joint optimization loss function is:

[0148] ;

[0149] Where, is the joint optimization loss function, is the task prediction loss, which is used to represent the weighted average prediction loss of access type prediction, access popularity prediction, and high latency risk prediction. is the prototype constraint loss, is the cross-temporal consistency loss, is the monotonicity constraint loss, is the soft budget constraint loss, is the prototype constraint weight, the default value is 0.5, is the cross-time consistency weight, the default value is 0.3, Is the monotonicity constraint loss, the default value is 0.2, is the soft budget constraint weight, the default value is 0.4;

[0150] Step S36: Three-way load prediction, specifically, by constructing the adaptive backbone network and the three-way prediction channel, combining the joint optimization loss function, performing model training to obtain a three-way load prediction model, and obtaining comprehensive load prediction reference data by using the three-way load prediction model;

[0151] The load comprehensive prediction reference data specifically includes data identification, access prediction type, predicted access popularity, predicted access delay and high delay risk prediction probability.

[0152] By performing the above operations, in order to address the technical problems in existing load prediction methods, such as the coarse prediction granularity of traditional load prediction methods in cache management, insufficient capture of sudden access patterns and neighborhood correlations, and lack of constraint fusion between multi-channel prediction results, which easily leads to imbalance in scheduling strategies, this solution creatively adopts an adaptive load three-channel prediction model improved by timing alignment enhancement to perform three-way load prediction. By introducing timing alignment enhancement and neighborhood enhancement mechanisms, combining a dual-channel gating fusion backbone network of time channel and resource channel, and embedding prototype constraints, cross-time domain consistency constraints, and monotonicity constraints in the three-way prediction of access type, access popularity, and high latency risk, it achieves significant improvements in the prediction results in terms of time consistency, category stability, and resource availability, providing a more robust multi-dimensional prediction basis for cache scheduling.

[0153] Example 5, see Figure 1 、 Figure 2 and Figure 5 This embodiment is based on the above embodiment. In step S4, the cache resource scheduling is used to formulate an optimal data placement and migration plan between cache layers under the constraints of capacity, bandwidth, temperature, and lifespan. Specifically, based on the original system operation data, the cache layer initialization structure, and the load comprehensive prediction reference data, a reinforcement learning method based on risk-sensitive reward improvement is used to perform cache resource scheduling to obtain cache resource scheduling decision data, including the following steps:

[0154] Step S41: State constraint construction: Based on the system operation raw data, cache layer initialization structure, and load comprehensive prediction reference data, a scheduling state vector and constraint boundaries are constructed. The state vector includes the popularity, latency risk, current layer, and recent migration records of each data object. The constraint boundaries include layer capacity, available bandwidth, temperature, lifespan thresholds, and their budget parameters.

[0155] Step S42: Risk-sensitive reward improvement is used to construct a reinforcement learning problem with budget and bandwidth constraints. Specifically, a reward function is constructed that introduces performance benefit terms, migration cost penalties, and risk-sensitive constraint penalties to obtain a risk-sensitive reward improvement reward function. The calculation formula is:

[0156] ;

[0157] Where R is the risk-sensitive reward improvement function, i is the data item index, is the predicted visit popularity, is the predicted access latency of the i-th data item at the current cache level, is the predicted access latency of the i-th data item at the target cache level, C migis the migration cost penalty, which is calculated by dividing the data volume by the migration speed limit. U is the predicted probability of high latency risk. is the capacity constraint violation, is the bandwidth constraint violation, is the temperature constraint violation, is the lifetime constraint violation amount;

[0158] Step S43: reinforcement learning training, specifically, performing reinforcement learning training through the state constraint construction and the risk-sensitive reward improvement, by defining an action vector, to obtain a cache resource scheduling model;

[0159] The action vector includes a layer migration action, a migration target layer, and a migration speed limit;

[0160] The level migration action specifically includes maintaining the level, raising the level, and sinking the level operations;

[0161] Step S44: Cache resource scheduling, specifically, performing cache resource scheduling based on the system operation original data, the cache level initialization structure, and the load comprehensive prediction reference data by using the cache resource scheduling model to obtain cache resource scheduling decision data;

[0162] The cache resource scheduling decision data specifically includes a cache resource scheduling migration object identifier, a cache resource scheduling operation type, a cache resource source level, a cache resource target level, a migration speed limit, and a decision time.

[0163] By performing the above operations, we address the technical problem that in existing cache resource scheduling methods, there is a simple reward function design that only considers performance improvement or migration cost, but does not uniformly model the physical constraints and delay risks of capacity, bandwidth, temperature and lifespan, which leads to scheduling strategies easily violating hard limits or causing lifespan overdraft under extreme loads. This solution creatively adopts a reinforcement learning method based on risk-sensitive reward improvement to perform cache resource scheduling, so that scheduling decisions can dynamically balance performance and resource security, significantly reducing the risk of cache system instability or premature media aging under high load.

[0164] Example 6, see Figure 1 and Figure 2This embodiment is based on the above embodiment. In step S5, the intelligent hierarchical cache is used to perform data migration and adjustment between different cache layers. Specifically, based on the cache resource scheduling decision data, the data migration and hierarchical adjustment between different cache layers are controlled. Under the condition of meeting the operation constraints and the heat score threshold, the target data is moved up, down or adjusted across layers according to the preset migration strategy, and the data distribution of each layer is dynamically updated in combination with the real-time business priority and the access heat to form a cache scheduling structure that meets the balance between performance and life, thereby obtaining an intelligent cache scheduling structure.

[0165] Example 7, see Figure 1 and Figure 2 , this embodiment is based on the above embodiment, and the present invention provides an intelligent hierarchical cache system based on a hyper-converged architecture, including a data acquisition module, an initial grading module, a load prediction module, a resource scheduling module and an intelligent grading module;

[0166] The data acquisition module is used for fusion architecture collection, obtains the original system operation data through fusion architecture collection, and sends the original system operation data to the initial classification module, the load prediction module and the resource scheduling module;

[0167] The initial grading module is used for initial cache grading, obtains a cache level initialization structure through the initial cache grading, and sends the cache level initialization structure to the load prediction module and the resource scheduling module;

[0168] The load prediction module is used for three-way load prediction, obtains load comprehensive prediction reference data through the three-way load prediction, and sends the load comprehensive prediction reference data to the resource scheduling module;

[0169] The resource scheduling module is used for cache resource scheduling, obtains cache resource scheduling decision data through cache resource scheduling, and sends the cache resource scheduling decision data to the intelligent classification module;

[0170] The intelligent grading module is used for intelligent grading cache, and an intelligent cache scheduling structure is obtained through intelligent grading cache.

[0171] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0172] While the embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that various changes, modifications, substitutions, and alterations can be made to the embodiments without departing from the principles and spirit of the invention.

[0173] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. An intelligent hierarchical caching method based on a hyper-converged architecture, characterized by: The method comprises the following steps: Step S1: Converged architecture collection, by building a hyper-converged architecture to collect, integrate and store data resources to obtain the original data of system operation; Step S2: Initial cache grading, based on the original data of the system operation, initializing the three-layer structure of the cache hierarchy to perform initial cache grading and obtain the cache hierarchy initialization structure; Step S3: Three-way load prediction, based on the system operation original data and the cache level initialization structure, adopt the adaptive load three-channel prediction model improved by timing alignment enhancement to perform three-way load prediction and obtain load comprehensive prediction reference data, including the following steps: Step S31: Time-aware alignment; Step S32: Access domain enhancement; Step S33: Adaptive backbone network construction, specifically constructing a dual-channel gated fusion network, including a time channel subnet, a resource channel subnet and a feature fusion sub-block to obtain fusion feature data; Step S34: Three-way prediction channel construction, specifically on the basis of the adaptive backbone network, respectively constructing an access type prediction channel, an access heat prediction channel and a high delay risk prediction channel to obtain task branch prediction data; Step S35: Soft budget constraint loss improvement, specifically introducing soft budget constraint loss and integrating it to obtain a joint optimization loss function; Step S36: Three-way load prediction; Step S4: Cache resource scheduling: Based on the system operation raw data, the cache level initialization structure, and the load comprehensive prediction reference data, a reinforcement learning method based on risk-sensitive reward improvement is used to perform cache resource scheduling to obtain cache resource scheduling decision data; the risk-sensitive reward improvement constructs a reward function that introduces performance benefit terms, migration cost penalties, and risk-sensitive constraint penalties; Step S5: Intelligent hierarchical caching to obtain an intelligent cache scheduling structure.

2. The intelligent hierarchical caching method based on a hyper-converged architecture according to claim 1, characterized in that: In step S1, the fusion architecture collection is used to collect and organize the operation status and data access status data of computing, storage and network equipment, specifically by building a hyper-converged architecture to collect, integrate and store data resources to obtain the original data of system operation; the specific content of the original data of system operation includes device operation status data, storage medium attribute data and data access record data.

3. The intelligent hierarchical caching method based on a hyper-converged architecture according to claim 2, characterized in that: In step S2, the initial cache hierarchy is used to divide the cache space into initial levels according to the speed, capacity, and cost of different storage media, and set the usage conditions and restrictions of each level. Specifically, based on the original data of the system operation, the initial cache hierarchy is constructed by initializing the three-level structure of the cache level to obtain the cache level initialization structure, including the following steps: Step S21: Cache tier division, specifically, dividing the cache space into a fast tier, a high-speed tier, and a capacity tier. The fast tier uses media with the lowest latency and the smallest capacity among the three tiers, the high-speed tier uses media with intermediate latency and capacity among the three tiers, and the capacity tier uses media with the highest latency and the highest capacity among the three tiers. Step S22: Data block management, specifically using fixed-size data blocks as management units, and calculating a heat score for each data block by comprehensively weighting the access count factor, access recency factor, read-write ratio factor, and request size factor based on historical access records. Step S23: Setting a file write strategy, specifically setting a hierarchical strategy for file writes in the storage system, including temporarily storing newly written data in a write buffer of the fast tier, and subsequently migrating the data to other storage tiers in priority order based on access feature determination results; Prioritize writing large-volume sequential read data to the high-speed tier to reduce the fast tier's occupancy and optimize subsequent access performance. Step S24: hysteresis interval control, used as a scheduling screening condition for cache resource scheduling in step S4, specifically setting a dual-threshold hysteresis control mechanism, setting independent heat score thresholds for upward and downward adjustments between the fast tier, high-speed tier, and capacity tier, and setting a data block regulation mechanism between tiers; Step S25: Use condition restrictions, specifically setting unified use conditions and restrictions for different levels; Step S26: Initial cache hierarchy, specifically generating a cache hierarchy initialization structure.

4. The intelligent hierarchical caching method based on a hyper-converged architecture according to claim 3, characterized in that: In step S2, the cache hierarchy initialization structure includes three layers of media type and capacity upper limit, migration rate upper limit, temperature and life threshold, hysteresis interval, warm-up list and fixed list, and initial hierarchy mapping table for each data block; The cache level initialization structure is used as input for subsequent steps, for analysis by three-way load prediction in step S3, and for planning and boundary control by cache resource scheduling in step S4.

5. The intelligent hierarchical caching method based on a hyper-converged architecture according to claim 4, characterized in that: In step S3, the three-way load prediction is used to predict the business type, access popularity, and high latency risk of each type of data. Specifically, based on the system operation original data and the cache level initialization structure, an adaptive load three-channel prediction model enhanced with timing alignment is used to perform three-way load prediction to obtain comprehensive load prediction reference data, including the following steps: Step S31: time-aware alignment, specifically extracting access time series from the system operation raw data, and calculating the time-aware alignment factor of each data item to obtain time-aligned access feature data; Step S32: Access domain enhancement, specifically, by calculating the neighborhood enhancement weights between adjacent data items, constructing access heat features between adjacent data items, and obtaining neighborhood enhancement access feature data; Step S33: constructing an adaptive backbone network, specifically constructing a dual-channel gated fusion network as an adaptive backbone network for data feature extraction, including a time channel subnet, a resource channel subnet, and a feature fusion sub-block, and extracting features using the adaptive backbone network to obtain fused feature data; The time channel subnet specifically performs time series convolution processing on the time feature vector and extracts access morphology features. The resource channel subnet specifically constructs a multi-layer perceptron and extracts resource status features from cache resource status parameters. The feature fusion subblock specifically constructs a gated fusion mechanism to perform gated fusion on the access morphology features and the resource status features to obtain fused feature data. Step S34: constructing three prediction channels, specifically, constructing an access type prediction channel, an access popularity prediction channel, and a high latency risk prediction channel based on the adaptive backbone network, thereby obtaining three prediction channels, and performing data prediction for the access type prediction task, the access popularity prediction task, and the high latency risk prediction task, respectively, to obtain task branch prediction data; Step S35: Improvement of soft budget constraint loss, specifically, introducing soft budget constraint loss. By calculating the data occupied capacity exceeding the threshold under the predicted popularity in the next 10 minutes, and comparing the data occupied capacity exceeding the threshold with the available budget capacity, the excess portion is included in the optimization target according to the square loss, limiting the prediction result to the range of system capacity and resource budget, and integrating the soft budget constraint loss with the prototype constraint loss, cross-temporal consistency loss, monotonicity constraint loss and task prediction loss to obtain a joint optimization loss function; Step S36: Three-way load prediction, specifically, through the adaptive backbone network construction and the three-way prediction channel construction, combined with the joint optimization loss function, model training is performed to obtain a three-way load prediction model, and by using the three-way load prediction model, load comprehensive prediction reference data is obtained.

6. The intelligent hierarchical caching method based on a hyper-converged architecture according to claim 5, characterized in that: In step S3, the load comprehensive prediction reference data specifically includes data identification, access prediction type, predicted access popularity, predicted access delay and high delay risk prediction probability.

7. The intelligent hierarchical caching method based on a hyper-converged architecture according to claim 6, characterized in that: In step S4, the cache resource scheduling is used to formulate an optimal data placement and migration plan between cache layers under the constraints of capacity, bandwidth, temperature, and lifespan. Specifically, based on the system operation original data, the cache layer initialization structure, and the load comprehensive prediction reference data, a reinforcement learning method based on risk-sensitive reward improvement is used to perform cache resource scheduling to obtain cache resource scheduling decision data, including the following steps: Step S41: State constraint construction: Based on the system operation raw data, cache layer initialization structure, and load comprehensive prediction reference data, a scheduling state vector and constraint boundaries are constructed. The state vector includes the popularity, latency risk, current layer, and recent migration records of each data object. The constraint boundaries include layer capacity, available bandwidth, temperature, lifespan thresholds, and their budget parameters. Step S42: Risk-sensitive reward improvement, which is used to construct a reinforcement learning problem with budget and bandwidth constraints. Specifically, a reward function is constructed that introduces a performance benefit term, a migration cost penalty, and a risk-sensitive constraint penalty to obtain a risk-sensitive reward improvement reward function. Step S43: reinforcement learning training, specifically, performing reinforcement learning training through the state constraint construction and the risk-sensitive reward improvement, by defining an action vector, to obtain a cache resource scheduling model; The action vector includes a layer migration action, a migration target layer, and a migration speed limit; The level migration action specifically includes maintaining the level, raising the level, and sinking the level operations; Step S44: Cache resource scheduling, specifically, performing cache resource scheduling based on the system operation original data, the cache level initialization structure, and the load comprehensive prediction reference data by using the cache resource scheduling model to obtain cache resource scheduling decision data; The cache resource scheduling decision data specifically includes a cache resource scheduling migration object identifier, a cache resource scheduling operation type, a cache resource source level, a cache resource target level, a migration speed limit, and a decision time.

8. The intelligent hierarchical caching method based on a hyper-converged architecture according to claim 7, characterized in that: In step S5, the intelligent hierarchical cache is used to perform data migration and adjustment between different cache layers. Specifically, based on the cache resource scheduling decision data, the data migration and hierarchical adjustment between different cache layers are controlled. Under the condition of meeting the operation constraints and the heat score threshold, the target data is moved up, down or adjusted across layers according to the preset migration strategy, and the data distribution of each layer is dynamically updated in combination with the real-time business priority and access heat to form a cache scheduling structure that meets the balance between performance and lifespan, thereby obtaining an intelligent cache scheduling structure.

9. An intelligent hierarchical caching system based on a hyper-converged architecture, for implementing an intelligent hierarchical caching method based on a hyper-converged architecture as claimed in any one of claims 1 to 8, characterized in that: It includes data acquisition module, initial classification module, load prediction module, resource scheduling module and intelligent classification module.

10. The intelligent hierarchical cache system based on a hyper-converged architecture according to claim 9, characterized in that: The data acquisition module is used for fusion architecture collection, obtains the original system operation data through fusion architecture collection, and sends the original system operation data to the initial classification module, the load prediction module and the resource scheduling module; The initial grading module is used for initial cache grading, obtains a cache level initialization structure through the initial cache grading, and sends the cache level initialization structure to the load prediction module and the resource scheduling module; The load prediction module is used for three-way load prediction, obtains load comprehensive prediction reference data through the three-way load prediction, and sends the load comprehensive prediction reference data to the resource scheduling module; The resource scheduling module is used for cache resource scheduling, obtains cache resource scheduling decision data through cache resource scheduling, and sends the cache resource scheduling decision data to the intelligent classification module; The intelligent grading module is used for intelligent grading cache, and an intelligent cache scheduling structure is obtained through intelligent grading cache.

Citation Information

Patent Citations

  • Network resource allocation and scheduling method under hyper-converged architecture

    CN117240806A

  • Cache data management and control method and system based on distributed encrypted storage

    CN119025566A

  • Model reasoning framework of cloud drawing platform

    CN120011059A

  • Memory access optimization method based on intelligent cache management

    CN120295942A

  • Distributed metadata management method and storage system based on cloud computing

    CN120335723A

Cited By

  • Cold and hot data hierarchical processing method and system for industrial internet

    CN121833657A