A cold and hot data layered storage method and system for a smart computing scenario
By acquiring task metadata in real time and building a dynamic heat model, the migration of data between multiple storage layers is dynamically adjusted, solving the problems of lagging cold and hot data identification and resource waste in the layered storage of cold and hot data in intelligent computing scenarios, and realizing efficient data access and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-16
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, the cold and hot data tiered storage method suffers from problems such as delayed cold and hot data identification, lack of business collaboration, and conservative migration strategies in intelligent computing scenarios, resulting in low data access efficiency and resource waste.
The system employs a task scheduling interface module, a data acquisition module, a heat analysis module, and a migration decision module. By acquiring task metadata in real time, collecting access behavior logs, and building a dynamic heat model, it dynamically adjusts the migration of data between multiple storage layers, thereby achieving accurate determination of data heat and reasonable allocation of storage resources.
It improves data access efficiency and resource utilization in intelligent computing scenarios, takes into account both high-performance access and low-cost storage needs, and solves the problems of poor storage adaptability and resource waste.
Smart Images

Figure CN122488998A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of storage technology, and in particular to a method and system for tiered storage of hot and cold data for intelligent computing scenarios. Background Technology
[0002] In intelligent computing scenarios, to balance storage performance and cost, a tiered storage technology for hot and cold data is typically adopted: frequently accessed hot data is stored on high-speed media, while infrequently accessed cold data is migrated to low-cost media.
[0003] In related technologies, tiered storage methods are widely used in general object storage service scenarios, such as log archiving, video backup, and enterprise file storage. These tiered storage methods can usually be implemented based on access frequency statistics (such as AWS S3 Intelligent-Tiering), lifecycle strategies for automatically cooling data based on time rules (such as Alibaba Cloud Object Storage Service (OSS)), and hot / cold classification based on log analysis (such as related solutions for the Hadoop Distributed File System (HDFS)). However, these methods usually have drawbacks such as delayed hot / cold identification, lack of business collaboration, and conservative migration strategies. Summary of the Invention
[0004] This disclosure is made in view of the above-mentioned problems. This disclosure provides a method and system for hierarchical storage of hot and cold data for intelligent computing scenarios.
[0005] According to one aspect of this disclosure, a method for hierarchical storage of hot and cold data for intelligent computing scenarios is provided, applied to a hierarchical storage system for hot and cold data in intelligent computing scenarios. The system includes a task scheduling interface module, a data acquisition module, a heat analysis module, a migration decision module, and a hierarchical storage module. The hierarchical storage module includes multiple storage layers divided according to access frequency. The method includes: The task scheduling interface module acquires and parses the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time; wherein, the task metadata includes at least the task identifier, the task priority value, and the current storage path of multiple data to be accessed required to run the intelligent computing task. The data acquisition module collects access behavior logs for each of the data to be accessed in real time. The heat analysis module uses a pre-built dynamic heat model to determine the current access heat of each of the data to be accessed based on the task priority value, the access behavior log, and the time decay factor. The migration decision module determines the migration strategy for each piece of data to be accessed based on the task priority, the current access popularity, the current storage path, and the current bandwidth utilization of the multiple storage layers; and migrates each piece of data to be accessed between the multiple storage layers according to the migration strategy. The hierarchical storage module stores each piece of data to be accessed in its corresponding storage layer.
[0006] According to another aspect of this disclosure, a hot and cold data tiered storage system for intelligent computing scenarios is provided to implement the hot and cold data tiered storage method for intelligent computing scenarios described in the embodiments of this disclosure. The system includes a task scheduling interface module, a data acquisition module, a heat analysis module, a migration decision module, and a tiered storage module. The tiered storage module includes multiple storage layers divided according to access heat. The task scheduling interface module is configured to acquire and parse the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time; wherein, the task metadata includes at least the task identifier, the task priority value, and the current storage path of multiple data to be accessed required to run the intelligent computing task. The data acquisition module is configured to collect access behavior logs of each of the data to be accessed in real time. The heat analysis module is configured to determine the current access heat of each of the data to be accessed based on the task priority value, the access behavior log, and the time decay factor using a pre-built dynamic heat model. The migration decision module is configured to determine a migration strategy for each piece of data to be accessed based on the task priority, the current access popularity, the current storage path, and the current bandwidth utilization of the multiple storage layers; and to migrate each piece of data to be accessed between the multiple storage layers according to the migration strategy. The hierarchical storage module is configured to store each piece of data to be accessed in a corresponding storage layer.
[0007] As will be described in detail below, the cold and hot data hierarchical storage method for intelligent computing scenarios according to embodiments of this disclosure acquires and parses the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time through a task scheduling interface module; wherein, the task metadata includes at least a task identifier, a task priority value, and the current storage path of multiple data to be accessed required to run the intelligent computing task; a data acquisition module collects access behavior logs of each data to be accessed in real time; a heat analysis module uses a pre-built dynamic heat model based on task priority value, access behavior logs, and time decay factor to determine the current access heat of each data to be accessed; a migration decision module... Based on task priority, current access popularity, current storage path, and current bandwidth utilization of multiple storage layers, a migration strategy for each data to be accessed is determined. According to the migration strategy, the data to be accessed is migrated between multiple storage layers. The tiered storage module stores each data to be accessed in its corresponding storage layer, enabling dynamic and accurate determination of data popularity and reasonable allocation of storage resources. This balances the high-performance access requirements of intelligent computing tasks with storage cost control, improving storage resource utilization and the operational efficiency of intelligent computing tasks. It solves the problems of poor storage adaptability, resource waste, and high cost in existing technologies, making it suitable for massive data storage in various intelligent computing scenarios.
[0008] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description
[0009] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 A schematic diagram of the framework of a tiered storage system for hot and cold data in intelligent computing scenarios provided by an exemplary embodiment of this disclosure is shown. Figure 2 A flowchart illustrating a method for hierarchical storage of hot and cold data for intelligent computing scenarios, provided by an exemplary embodiment of this disclosure, is shown. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0012] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0013] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0014] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0015] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0016] In existing object storage systems, hot and cold data identification and tiered storage are common data management techniques, widely used in cloud storage, archive storage, big data platforms, and other scenarios. Typical solutions include: The first type is hierarchical storage based on access frequency statistics: statistically analyze the number of historical accesses, set thresholds to migrate data (such as AWS S3 Intelligent-Tiering), and store data on storage media with different performance and cost according to the access frequency and importance of the data to optimize the utilization of storage resources. However, there are the following technical disadvantages: (1) It does not have the ability to migrate data in advance: by analyzing the access frequency and importance of the data, the hot and cold attributes of the data are judged. It is a passive identification mechanism and cannot complete the necessary data loading and caching before task scheduling; (2) It does not support cross-system linkage migration: migration is carried out after detecting changes in data access frequency or importance. It may be that the migration has not been completed when the data needs to be accessed frequently, which affects the system performance. In other words, the system does not consider the diversity of data sources of the intelligent computing platform (such as object storage, distributed file system, etc.), and the data flow efficiency is low in the heterogeneous platform environment; (3) The static weight model is not suitable for intelligent computing scenarios: the method is suitable for general data storage environment and mainly focuses on changes in data access frequency and importance. However, the importance of data in AI scenarios has strong phased characteristics. The static judgment mechanism of this scheme is difficult to adapt to the differentiated dependence of data on different stages of model training.
[0017] The second type is a lifecycle strategy based on time-based rules to automatically cool down data (such as Alibaba Cloud OSS): This strategy uses static parameters such as data access frequency, time, and weight to achieve automatic hierarchical storage and policy configuration, suitable for data lifecycle management in general business systems. It can analyze data access frequency, time, and other characteristics to automatically classify data into cold and hot data and store them on storage media with different performance and cost levels, thereby improving the performance and efficiency of the storage system. However, the following technical drawbacks exist: (1) Coarse data access granularity: The determination of hot and cold data is mainly based on the access frequency and time characteristics of the data to determine the hot and cold attributes. The access frequency is mostly evaluated at the granularity of files or objects, and it fails to be refined to identify the data heat according to the model, task, and session dimensions in the intelligent computing scenario. It lacks consideration for business semantics and task scheduling, which may lead to misjudgment of data heat; (2) Insufficient timeliness of data migration, unable to achieve task-level data pre-scheduling: Migration is carried out after detecting changes in data access frequency or time characteristics, which may result in the migration not being completed when the data needs to be accessed frequently. It lacks a mechanism to link with the running status of upper-layer AI tasks and cannot perceive the real-time data requirements of training or inference tasks in advance; (3) System isolation and insufficient collaboration: The system is not integrated with the computing resource platform and mainly focuses on the management of hot and cold data within the storage system. It lacks collaboration with upper-layer business systems, does not have overall scheduling optimization capabilities, and is prone to resource idleness or data blockage.
[0018] The third approach involves log analysis to assist in hot and cold data classification: predicting data popularity by analyzing access logs (e.g., the hot and cold data classification scheme of the Hadoop Distributed File System (HDFS)). Based on HDFS's tag manager and tier manager, it can classify storage media and formulate corresponding storage strategies for different storage directories within HDFS, achieving automatic tiered storage of hot and cold data. Through the collaborative work of intelligent storage systems, a visualized big data platform can be built to achieve automatic tiering of hot and cold data in HDFS. This method primarily targets the HDFS file system within big data platforms, focusing on the tiered management of hot and cold data within HDFS. However, the following technical drawbacks exist: (1) Lagging identification mechanism: By collecting and statistically analyzing data access events in the HDFS storage directory to measure data temperature, it is a passive identification mechanism. This method relies on the historical access statistics of the data and requires a period of behavior collection to make a cold or hot judgment. There is a problem of untimely response to changes in data status; (2) Not combined with task scheduling information: The cold and hot identification of the HDFS layer cannot perceive the usage intention of the upper layer business, resulting in some data that will be frequently accessed being identified as cold data, and the migration is delayed; (3) Conservative migration strategy: The data is scanned periodically by a timer, and the cold and hot data are triggered by preset rules to migrate the data. The cold and hot switching cannot be completed in a short time, which makes it difficult to meet the real-time requirements of AI training, inference and other intelligent computing tasks.
[0019] It is evident that the hierarchical storage methods in related technologies typically suffer from drawbacks such as delayed hot / cold data identification, lack of business collaboration, and conservative migration strategies. Specifically, these drawbacks include: (1) delayed hot / cold data identification: relying on historical access data to passively determine the popularity, failing to respond promptly to sudden changes in data access patterns during AI tasks; (2) lack of business collaboration: being isolated from upper-layer task scheduling systems (such as AI training platforms), failing to perceive the datasets required by tasks about to be executed, resulting in data that needs to be accessed when a task starts still being located in the cold layer, causing I / O wait; (3) conservative migration strategies: using timed scanning or periodic migration, failing to dynamically execute in conjunction with system load, potentially seizing bandwidth during peak business periods and affecting performance.
[0020] Therefore, in order to solve the above problems, this disclosure provides a method for tiered storage of hot and cold data for intelligent computing scenarios, which is applied to a tiered storage system for hot and cold data in intelligent computing scenarios. Figure 1 A schematic diagram of the framework of a tiered cold and hot data storage system for intelligent computing scenarios provided by an exemplary embodiment of this disclosure is shown. Figure 1 As shown, the cold and hot data hierarchical storage system 100 for intelligent computing scenarios includes a task scheduling interface module 101, a data acquisition module 102, a heat analysis module 103, a migration decision module 104, and a hierarchical storage module 105.
[0021] For example, the tiered storage module 105 may include multiple storage tiers divided according to access frequency. Table 1 shows the division criteria for multiple storage tiers in the tiered storage module provided in this disclosure embodiment.
[0022] Table 1. Criteria for Dividing Multiple Storage Layers in a Tiered Storage Module
[0023] As shown in Table 1, multiple storage layers can include a hot layer, a warm layer, and a cold layer, each with a corresponding access frequency range. Specifically, the access frequency of data stored in the hot layer is greater than a first preset access frequency; the access frequency of data stored in the warm layer is greater than or equal to a second preset access frequency and less than or equal to the first preset access frequency; and the access frequency of data stored in the cold layer is less than the second preset access frequency. The storage media, performance indicators, and applicable data differ among the multiple storage layers divided according to access frequency.
[0024] Here, the specific values of the first preset access popularity and the second preset access popularity can be selected according to actual needs, and this embodiment of the disclosure does not impose specific limitations on them. In the method of this embodiment of the disclosure, the first preset access popularity is set to 85, and the second preset access popularity is set to 75.
[0025] This disclosure provides a method for tiered storage of hot and cold data for intelligent computing scenarios. It can work closely with intelligent computing platforms to perform tiered storage of hot and cold data with forward-looking perception and dynamic migration capabilities, thereby improving data access efficiency and resource utilization in intelligent computing scenarios.
[0026] Figure 2 A flowchart illustrating a hierarchical storage method for hot and cold data in intelligent computing scenarios, provided by an exemplary embodiment of this disclosure, is shown. Figure 2 As shown, this method for tiered storage of hot and cold data for intelligent computing scenarios includes: S201, the task scheduling interface module acquires and parses the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time; among which, the task metadata includes at least the task identifier, the task priority value, and the current storage path of multiple data to be accessed required to run the intelligent computing task; S202, The data acquisition module collects access behavior logs of each data to be accessed in real time; S203, the heat analysis module uses a pre-built dynamic heat model to determine the current access heat of each piece of data to be accessed based on task priority, access behavior logs and time decay factor; S204, the migration decision module determines the migration strategy for each piece of data to be accessed based on task priority, current access popularity, current storage path, and current bandwidth utilization of multiple storage layers; and migrates each piece of data to be accessed between multiple storage layers according to the migration strategy. S205, the hierarchical storage module stores each piece of data to be accessed in its corresponding storage layer.
[0027] Specifically, the task scheduling interface module can connect to the task submission interface of the intelligent computing platform in real time. When the intelligent computing platform submits a task to be run, the scheduling system generates a task description file (JSON format) to obtain the task data. The task scheduling interface module immediately obtains the task data packet through the task submission interface, parses the task data packet using a preset parsing algorithm, extracts the task metadata, and then uploads the task metadata to the heat analysis module.
[0028] Task metadata can include a task identifier, task priority value, and the current storage path of multiple data items to be accessed for running the intelligent computing task. The task identifier uniquely distinguishes different intelligent computing tasks, preventing conflicts in task scheduling and data migration. Task priority values can be set based on the task type, which may include training tasks, inference tasks, data preprocessing tasks, and log analysis, but is not limited to these. Task priority values can be set from 1 to 10, with core intelligent computing tasks such as neural network model training and inference tasks set at 8-10 points (high priority), and auxiliary intelligent computing tasks such as data preprocessing and log analysis set at 1-7 points (low priority). The current storage path of the data to be accessed includes the storage node address, storage directory, and data file name, used to accurately locate the current storage location of each piece of data to be accessed.
[0029] Step 201 enables precise parsing of the task metadata of the intelligent computing tasks submitted by the intelligent computing platform, realizing a one-to-one binding between intelligent computing tasks and corresponding data to be accessed. Task identifiers avoid scheduling conflicts between tasks and data, and task priority values provide priority basis for subsequent heat analysis and migration decisions, ensuring that the core data of high-priority intelligent computing tasks can be processed first, laying the foundation for the entire hierarchical storage process.
[0030] The data acquisition module is deployed as a resident process on each intelligent computing node or object storage proxy node. IOMonitor continuously collects client read and write behavior data and collects access behavior logs for all data to be accessed in real time. These logs include access timestamps, operation types (GET (read operation) / PUT (write operation) / DELETE (delete operation)), operation frequency, and data size. The data acquisition module can collect access behavior logs for each piece of data to be accessed and upload them to the heat analysis module in real time, providing accurate and continuous behavioral data for determining data hotness / coldness. Simultaneously, this collection process uses an asynchronous acquisition mode, which does not consume the intelligent computing task's runtime resources, ensuring the normal operation of the intelligent computing task.
[0031] The popularity analysis module can load a pre-built dynamic popularity model, using task priority, access behavior logs, and time decay factor as input parameters to calculate the current access popularity of each piece of data to be accessed. Here, the time decay factor is negatively correlated with the most recent access time; the more recent the access time, the larger the time decay factor. For example, the time decay factor can be calculated based on the reciprocal of the last (i.e., the previous) access time of the data to be accessed; the closer to the current time, the larger the time decay factor (that is, the higher the access frequency of the data to be accessed, the larger the time decay factor). Introducing the time decay factor when calculating the current access popularity avoids the interference of historical access behavior on the current access popularity in static judgment, making the popularity calculation results more consistent with the real-time access characteristics of the data to be accessed in intelligent computing scenarios, and providing a reliable basis for subsequent migration decisions.
[0032] After the intelligent computing platform submits the intelligent computing task to be run, the migration decision module immediately marks the data to be accessed under the current storage path ( / data / imagenet path) as "about to be hot" and sets the migration priority to the highest.
[0033] The migration decision module obtains the current bandwidth utilization of each storage layer in real time, and makes a comprehensive decision based on task priority, current access popularity, current storage path, and current bandwidth utilization of multiple storage layers to determine the migration strategy for each data to be accessed, triggering asynchronous fine-grained migration, replacing the traditional timed full scan mechanism.
[0034] For example, if the current access popularity of the data to be accessed is greater than 85 (the data to be accessed is hot data) and the task priority value is greater than or equal to 8 points, it will be migrated to the hot layer regardless of its current storage path; if the data to be accessed is already stored in the hot layer and the current bandwidth utilization of the hot layer is less than 70%, it will remain unchanged; if the current bandwidth utilization of the hot layer is greater than or equal to 70%, the data corresponding to the high task priority value will be migrated first to ensure the access performance of core task data.
[0035] If the current access popularity of the data to be accessed is greater than or equal to 75 and less than or equal to 85 (the data to be accessed is warm data), there is no need to split the storage according to the task priority value. This avoids warm data occupying hot layer resources and also prevents warm data from being stored in cold layer, which would cause excessive access latency, thus achieving adaptive storage of warm data.
[0036] If the current access popularity of the data to be accessed is less than 75 (the data to be accessed is cold data), regardless of the task priority, it will be migrated to the cold layer to release the storage resources of the hot and warm layers and reduce storage costs; if the cold data is already stored in the cold layer, it will remain unchanged.
[0037] The migration decision module generates data migration instructions based on migration decisions and controls the storage cluster to execute migration operations. This module incorporates a standardized storage adapter, internally encapsulating migration interfaces for various heterogeneous storage systems (such as Ceph, object storage, and HDFS). It provides a unified set of read, write, and migration instructions, supporting seamless integration with these heterogeneous storage systems and enabling cross-platform data scheduling. For example, assuming data to be accessed is being migrated from a cold layer to a hot layer, the migration decision module can call the ReliableAutonomic Distributed Object Store (RADOS) interface in the distributed storage system (Ceph) to migrate the data from a Hard Disk Drive (HDD) pool to a Solid State Drive (SSD) pool. For cross-storage system migrations (such as HDFS → object storage), data block transmission and verification mechanisms are enabled.
[0038] In addition, during data migration, a token bucket algorithm can be used to control the migration rate, ensuring that it does not exceed the bandwidth limit. If migration fails (e.g., due to network interruption), it is rolled back to the original storage layer and an error log is recorded. Based on this, while ensuring high-performance access to high-priority, high-frequency data, the bandwidth load of each storage layer can be balanced, avoiding bandwidth congestion at a single layer, while reducing invalid data migration and improving migration efficiency.
[0039] Based on migration decisions, the tiered storage module stores each piece of data to be accessed in its corresponding storage layer. Hot data is stored in the hot layer to meet the low latency and high concurrency read requirements of intelligent computing tasks; warm data is stored in the warm layer to meet the needs of warm data that has a "moderate access frequency, does not require extreme performance but needs to avoid high latency"; and cold data is stored in the cold layer to reduce overall storage costs. This achieves physical isolation and tiered management of hot and cold data, taking into account both the high-performance access and low-cost storage requirements of intelligent computing scenarios, improving the overall utilization of storage resources, and ensuring data path traceability so as not to affect the normal read and operation of intelligent computing tasks.
[0040] According to the technical solution of the exemplary embodiments of this disclosure, the task scheduling interface module acquires and parses the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time. The task metadata includes at least a task identifier, a task priority value, and the current storage path of multiple data items to be accessed required to run the intelligent computing task. The data acquisition module collects access behavior logs of each data item in real time. The heat analysis module uses a pre-built dynamic heat model based on the task priority value, access behavior logs, and time decay factor to determine the current access heat of each data item. The migration decision module determines the migration strategy for each data item based on the task priority value, current access heat, current storage path, and the current bandwidth utilization of multiple storage layers. The data items are migrated between multiple storage layers according to the migration strategy. The hierarchical storage module stores each data item in its corresponding storage layer, enabling dynamic and accurate determination of data heat and reasonable allocation of storage resources. This balances the high-performance access requirements of intelligent computing tasks with storage cost control, improves storage resource utilization and intelligent computing task operation efficiency, and solves the problems of poor storage adaptability, resource waste, and high cost in existing technologies. It is suitable for massive data storage in various intelligent computing scenarios.
[0041] In some embodiments, the access behavior log includes at least read and write operations recorded by access timestamp; Using a pre-built dynamic popularity model based on task priority, access behavior logs, and time decay factors, the current access popularity of each piece of data to be accessed is determined, including: Count the number of read operations and the number of write operations for each piece of data to be accessed within a preset historical time window, and calculate the historical read frequency and historical write frequency for each piece of data to be accessed. Using a pre-built access frequency prediction model, based on historical read frequency and historical write frequency, the predicted read frequency and predicted write frequency of each data to be accessed within a preset future time window are predicted. Calculate the time decay factor based on the most recent access time of each piece of data to be accessed; The task priority, predicted read frequency, predicted write frequency, and time decay factor are input into a pre-built dynamic heat model to calculate the current access heat of each piece of data to be accessed.
[0042] Specifically, the access behavior logs mentioned above may include access timestamps, operation types (GET (read operation) / PUT (write operation) / DELETE (delete operation)), operation frequency, data size, etc.
[0043] The aforementioned access behavior logs, indexed by access timestamps, fully record the system's read and write operations on each piece of data to be accessed. This provides original and reliable data for subsequent access frequency statistics and popularity calculations, ensuring the authenticity and accuracy of popularity assessments.
[0044] The aforementioned preset historical time window can be a sliding time window. By limiting the preset historical time window, invalid historical data that has not been accessed for a long time can be filtered out. The number of read operations and write operations can be counted separately, and the historical read frequency and historical write frequency of each data to be accessed can be calculated. This allows for the differentiation of the contribution of different operation types to data popularity, making the characterization of historical access features more refined. Here, the size of the preset historical time window can be set according to actual needs, and this embodiment does not specifically limit it. In the method of this embodiment, the preset historical time window can be set to the past 24 hours.
[0045] The aforementioned access frequency prediction model can be a pre-built Long Short-Term Memory (LSTM) network. Based on historical read and write frequencies, the model predicts the predicted read and write frequencies of each piece of data to be accessed within a preset future time window. The size of the preset future time window can be set according to actual needs, and this embodiment does not impose specific limitations on it. In the method of this embodiment, the preset future time window can be set to the next hour.
[0046] This disclosure embodiment predicts future frequencies based on historical frequencies, which can extend the historical access patterns of data to future time periods to correct access popularity and achieve forward-looking prediction of the subsequent access popularity of data, avoiding the problem that the popularity assessment lags behind the actual access trend due to relying solely on historical statistical values.
[0047] The time decay factor is negatively correlated with the most recent access time. It can be the reciprocal of the most recent access time of the data to be accessed. The more recent the access time, the larger the time decay factor. This can weaken the influence of the historical popularity of data that has not been accessed for a long time, strengthen the weight of recently active data, and make the access popularity assessment more in line with the actual access status of the current system.
[0048] The task priority, predicted read frequency, predicted write frequency, and time decay factor are input into a pre-built dynamic popularity model to calculate the current access popularity of each piece of data to be accessed. Through multi-dimensional feature fusion calculation, the task priority reflects the importance at the business level, the predicted read frequency and predicted write frequency reflect the access trend, and the time decay factor reflects the access timeliness. The four factors work together to dynamically and accurately obtain the current access popularity of each piece of data to be accessed, improving the reliability and applicability of the popularity assessment results.
[0049] For example, the dynamic heat model is represented as: Heat obj )= α × f read + β × f write + c ×T -1 recent + d ×T P Among them, Heat( obj This indicates the current access popularity of each piece of data to be accessed; f read Indicates the predicted read frequency; f write Indicates the predicted write frequency; T -1 recent T represents the time decay factor; P Indicates task priority; α The weights represent the predicted read frequency; β The weights represent the predicted write frequency; c Indicates the weight of the time decay factor; d This indicates the weight of the task priority value.
[0050] Here, α , β , c and d The value can be dynamically adjusted based on the task type of the intelligent computing task to be run. For example, if the task type of the intelligent computing task to be run is a training task, α Set to 0.4 (focus on reading). β Set to 0.3 (emphasis on write); if the task type of the intelligent computing task to be run is a reasoning task, α Set to 0.6, β Set it to 0.1.
[0051] For example, in the early stages of training (epoch < 10), because checkpoint files are generated frequently, the write weight is increased. β Set to 0.4; during the inference phase, increase the read weight, and... α Set it to 0.7 to focus on real-time response.
[0052] For example, high-priority tasks (such as urgent tasks) d The value is 0.2, for a normal task. d The value is 0.1.
[0053] In some embodiments, the migration strategy includes the target storage path and the timing of the migration; Based on task priority, current access frequency, current storage path, and current bandwidth utilization of multiple storage layers, a migration strategy for each piece of data to be accessed is determined, including: Based on the current access popularity and the current storage path, determine the target storage path for each piece of data to be accessed; Based on task priority and the current bandwidth utilization of the storage layer corresponding to the target storage path in multiple storage layers, the migration timing of each data to be accessed is determined.
[0054] Specifically, the embodiments of this disclosure can decompose the migration strategy into two dimensions: the target storage path and the migration timing, so as to achieve fine-grained control over the data storage location and the migration execution time, thereby avoiding problems such as unreasonable allocation of storage resources and migration conflicts caused by a single strategy configuration, and improving the flexibility and controllability of data migration.
[0055] Current access popularity reflects the frequency and importance of accessing data, while the current storage path indicates the current storage layer of the data. Each storage layer has a corresponding access popularity. Therefore, combining current access popularity and current storage path can accurately determine the target storage path for each piece of data to be accessed, providing a basis for tiered storage of the data. This ensures efficient access to the data based on its current access popularity while optimizing overall storage resource utilization and reducing storage costs.
[0056] Task priority values ensure that high-priority business data is migrated first, avoiding impact on critical business operations due to migration delays. The current bandwidth utilization rate of the storage layer corresponding to the target storage path reflects the real-time load of that storage layer. Therefore, by determining the migration timing of each data to be accessed based on task priority values and the current bandwidth utilization rate of the storage layer corresponding to the target storage path, and selecting a period of bandwidth idle time to perform the migration, the bandwidth contention and performance interference of the migration operation on normal business access can be reduced, thereby improving the overall stability of the system and the migration success rate.
[0057] In some embodiments, the target storage path for each piece of data to be accessed is determined based on the current access frequency and the current storage path, including: If the current access popularity does not match the access popularity of the storage layer corresponding to the current storage path, the current storage path is updated based on the current access popularity and the access popularity of multiple storage layers to obtain the target storage path for each data to be accessed. If the current access popularity matches the access popularity of the storage layer corresponding to the current storage path, the current storage path will be determined as the target storage path for each piece of data to be accessed.
[0058] Specifically, as described above, this embodiment pre-configures corresponding access popularity intervals for different storage layers. It matches the current access popularity of each piece of data to be accessed with the access popularity of the storage layer corresponding to the current storage path to determine whether the current storage path of the data to be accessed needs to be adjusted. By matching the access popularity with the access popularity corresponding to the storage layer, data to be accessed with unreasonable storage locations can be quickly filtered out, avoiding redundant scheduling operations on data to be accessed that has already been adapted to the storage layer, thus improving the targeting and execution efficiency of tiered storage management.
[0059] When the current access popularity does not match the access popularity of the storage layer corresponding to the current storage path, the current storage path is updated based on the current access popularity and the access popularity of multiple storage layers to obtain the target storage path for each piece of data to be accessed. When the access popularity does not match the storage layer, it indicates that the actual access demand of the data to be accessed does not match the performance specifications of the current storage layer. In this case, a new storage layer is matched according to the current access popularity. This allows hot data to fall into the high-performance, low-latency hot layer to ensure access response speed, while low-intensity data falls into the high-capacity, low-cost cold layer, optimizing the allocation of storage resources and achieving a balance between data access performance and storage cost.
[0060] If the current access popularity matches the access popularity of the storage layer corresponding to the current storage path, the current storage path is determined as the target storage path for each piece of data to be accessed. For data whose access popularity matches the storage layer, the original storage path remains unchanged, and there is no need to trigger a migration process. This can effectively reduce bandwidth consumption, IO overhead, and system load caused by data migration, ensure that the storage scheduling logic is simple and efficient, and improve the stability and response speed of the overall tiered storage system.
[0061] In some embodiments, the migration timing of each piece of data to be accessed is determined based on task priority values and the current bandwidth utilization of the storage layer corresponding to the target storage path in multiple storage layers, including: If the current bandwidth utilization of the storage layer corresponding to the target storage path is greater than or equal to the preset bandwidth utilization and the task priority value is less than the preset priority value in multiple storage layers, the migration timing of the data to be accessed will be delayed.
[0062] Specifically, embodiments of this disclosure can set a preset bandwidth utilization rate and a preset priority value to determine whether data migration will cause bandwidth pressure on the storage layer corresponding to the target storage path and the importance of the service to which the data belongs. By setting dual judgment conditions, the storage layer load status and service priority can be taken into account, avoiding bandwidth congestion or obstruction of high-priority tasks due to blind migration. Here, both the preset bandwidth utilization rate and the preset priority value can be set according to actual needs, and embodiments of this disclosure do not impose specific limitations on them. In the method of embodiments of this disclosure, the preset bandwidth utilization rate is set to 70%, and the preset priority value is set to 8 points.
[0063] When the current bandwidth utilization of the storage layer corresponding to the target storage path is greater than or equal to 70%, it indicates that the bandwidth of that storage layer is already under high load. Immediately performing migration at this time would further consume bandwidth resources and affect the access response speed of normal services on that storage layer. Meanwhile, for data with lower task priority, the urgency and importance of the business are relatively low, allowing for a suitable delay in migration. By delaying the migration timing in this scenario, the bandwidth peak of the storage layer corresponding to the target storage path can be avoided, preventing low-priority tasks from consuming valuable bandwidth resources and ensuring the stable operation of high-priority services and critical access operations.
[0064] By controlling the timing of the aforementioned delayed migration, it is possible to achieve staggered scheduling of migration traffic and business traffic. Under the premise of ensuring that the overall access performance of the system is not affected, the migration scheduling of low-priority data can be smoothly completed, thereby improving the stability and resource utilization of the storage system in high-concurrency and high-load scenarios.
[0065] In some embodiments, the system may further include a storage adaptation module; after migrating the data to be accessed between multiple storage tiers according to a migration strategy, the method further includes: Based on the migration strategy, the storage adaptation module updates the current storage path of each data to be accessed in the task metadata and records the migration time.
[0066] Specifically, the storage adaptation module serves as the connecting unit between the migration decision module and the intelligent computing platform. It is specifically responsible for updating the path and recording information after data migration, which can decouple migration execution from metadata maintenance, thereby improving the clarity of system module responsibilities and overall maintainability.
[0067] After data migration is complete, the storage adaptation module synchronously updates the "current storage path" field in the task metadata of the intelligent computing tasks to be run in the metadata database according to the executed migration strategy. This ensures that the task metadata is consistent with the actual data storage location, avoiding subsequent access failures and addressing anomalies caused by outdated or incorrect storage path information. This ensures that the upper-layer business has an accurate and reliable perception of the data storage location. Simultaneously, recording the migration time in the task metadata creates a complete data migration history, facilitating subsequent tracing, statistics, and analysis of changes in data access frequency, storage layer scheduling frequency, and migration traffic timing. This provides a valid basis for subsequent dynamic heat model optimization, storage strategy adjustment, and system maintenance troubleshooting.
[0068] In some embodiments of this disclosure, the migration decision module is further provided with a hot and cold oscillation suppression mechanism to avoid frequent data migration between different storage layers and to ensure the stability and scheduling reliability of the hierarchical storage system.
[0069] The migration decision module counts the number of migrations of the same data to be accessed within a preset time period. If the number of migrations of the same data to be accessed within the preset time period is greater than or equal to the preset number of migrations, the data object is marked as "oscillating data". By counting the number of migrations and determining the threshold, oscillating data that repeatedly migrates between multiple storage layers due to frequent fluctuations in access popularity can be accurately identified, providing an accurate basis for subsequent cooling suppression and avoiding misjudging normally scheduled migration data as oscillating data. Here, the preset time period and the preset number of migrations can be set according to actual needs, and this embodiment does not specifically limit them. In the method of this embodiment, the preset time period can be set to 24 hours, and the preset number of migrations can be set to 5 times.
[0070] After identifying oscillating data, the migration decision module automatically triggers a cooling strategy. For data marked as oscillating, the cooling strategy is implemented, pausing the migration process for a preset cooling time window. Once the access popularity of the data stabilizes, it is re-evaluated at the storage layer and migration is re-determined. By pausing migration and setting a cooling time window, continuous migration of data to be accessed due to small fluctuations in popularity within a short period can be effectively avoided, reducing the ineffective consumption of storage layer I / O and bandwidth resources. Scheduling is only performed after the data access popularity has fully converged and stabilized, thus suppressing the oscillating migration of hot and cold data from a mechanism perspective. Here, the preset cooling time window can be set according to actual needs, and this embodiment does not specifically limit it. In the method of this embodiment, the preset cooling time window can be set to 12 hours.
[0071] To facilitate operation and maintenance management, this embodiment also includes a manual review process. When oscillation data is identified and a cooling strategy is triggered, the migration decision module automatically generates a corresponding oscillation data report and pushes it to the administrator console. Automatic generation and reporting of oscillation reports allows administrators to promptly detect abnormal oscillation data and scheduling anomalies in the system, improving the real-time performance and observability of system operation and maintenance.
[0072] Based on the oscillation data report, administrators can manually adjust parameters such as the preset migration count and preset cooldown time window, or directly bind the oscillation data to a fixed storage layer. By providing a manual intervention and configuration adjustment entry point, manual optimization methods can be added on top of the mechanism suppression, further strengthening the control over oscillation data, adapting to the storage scheduling needs of different business scenarios, and improving system flexibility and applicability.
[0073] Table 2 shows a comparison of the advantages of the cold and hot data hierarchical storage method for intelligent computing scenarios provided by the embodiments of this disclosure with traditional solutions.
[0074] Table 2. Comparison of Advantages between Cold and Hot Data Tiered Storage Methods and Traditional Solutions for Intelligent Computing Scenarios
[0075] The beneficial effects of the cold and hot data hierarchical storage method for intelligent computing scenarios provided in this disclosure are as follows: (1) Significantly improved real-time performance and foresight: Deeply integrated with the intelligent computing platform scheduling system, it parses task metadata in real time and triggers migration, predicts data access needs through task metadata, and completes data preheating before task execution, avoiding access delays caused by delayed identification in traditional solutions.
[0076] (2) Optimization of the accuracy of hot and cold judgment: Introduce semantic features such as task type and access behavior logs into the dynamic heat model to enhance the generalization ability of the model. Multi-dimensional feature modeling can distinguish training / inference tasks and stage data dependencies (such as checkpoint files) and reduce the misjudgment rate.
[0077] (3) Improved resource utilization and system efficiency: The on-demand migration strategy combined with the load awareness mechanism reduces invalid I / O operations and reduces storage layer bandwidth pressure.
[0078] (4) Robustness enhancement: Based on migration history and cooling time window, abnormal migration behavior is intelligently suppressed, and the hot and cold shock suppression mechanism avoids metadata conflicts and performance jitter caused by frequent data migration.
[0079] (5) Improved compatibility and scalability: Standardized adapters encapsulate underlying storage differences, provide a unified management interface, support heterogeneous storage platforms, and adapt to diverse data source requirements in intelligent computing scenarios.
[0080] This disclosure also provides a tiered cold and hot data storage system for intelligent computing scenarios, used to implement the aforementioned tiered cold and hot data storage method for intelligent computing scenarios, such as... Figure 1 As shown, the system 100 includes a task scheduling interface module 101, a data acquisition module 102, a heat analysis module 103, a migration decision module 104, and a hierarchical storage module 105. The hierarchical storage module 105 includes multiple storage layers divided according to access heat. The task scheduling interface module 101 is configured to acquire and parse the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time; wherein, the task metadata includes at least the task identifier, the task priority value, and the current storage path of multiple data to be accessed required to run the intelligent computing task. The data acquisition module 102 is configured to collect access behavior logs of each of the data to be accessed in real time. The heat analysis module 103 is configured to use a pre-built dynamic heat model to determine the current access heat of each of the data to be accessed based on the task priority value, the access behavior log and the time decay factor. The migration decision module 104 is configured to determine a migration strategy for each piece of data to be accessed based on the task priority value, the current access popularity, the current storage path, and the current bandwidth utilization of the multiple storage layers; and to migrate each piece of data to be accessed between the multiple storage layers according to the migration strategy. The hierarchical storage module 105 is configured to store each of the data to be accessed in a corresponding storage layer.
[0081] For details, please refer to the previous section on the hierarchical storage method for hot and cold data in intelligent computing scenarios; it will not be repeated here.
[0082] In some embodiments, the access behavior log includes at least read and write operations recorded by access timestamp; The heat analysis module 103 is also configured to count the number of read operations and the number of write operations of each of the data to be accessed within a preset historical time window, and to calculate the historical read frequency and historical write frequency of each of the data to be accessed. Using a pre-built access frequency prediction model, based on the historical read frequency and the historical write frequency, the predicted read frequency and predicted write frequency of each of the data to be accessed within a preset future time window are predicted; Calculate the time decay factor based on the most recent access time of each of the data to be accessed; The task priority value, the predicted read frequency, the predicted write frequency, and the time decay factor are input into a pre-built dynamic heat model to calculate the current access heat of each of the data to be accessed. The dynamic heat model is expressed as: Heat obj )= α × f read + β × f write + c ×T -1 recent + d ×T P Among them, Heat( obj The number () indicates the current access popularity of each of the data items to be accessed; f read This indicates the predicted read frequency; f write T represents the predicted write frequency; -1 recent T represents the time decay factor; P This indicates the priority value of the task; α The weights represent the predicted read frequencies; β The weights representing the predicted write frequency; c This represents the weight of the time decay factor; d The weight represents the priority value of the task. For details, please refer to the previous section on the hierarchical storage method for hot and cold data in intelligent computing scenarios; it will not be repeated here.
[0083] In some embodiments, the migration strategy includes the target storage path and the migration timing; The migration decision module 104 is further configured to determine the target storage path for each of the data to be accessed based on the current access popularity and the current storage path; Based on the task priority value and the current bandwidth utilization of the storage layer corresponding to the target storage path in the multiple storage layers, the migration timing of each of the data to be accessed is determined.
[0084] For details, please refer to the previous section on the hierarchical storage method for hot and cold data in intelligent computing scenarios; it will not be repeated here.
[0085] In some embodiments, the migration decision module 104 is further configured to update the current storage path based on the current access popularity and the access popularity corresponding to the plurality of storage layers when the current access popularity does not match the access popularity of the storage layer corresponding to the current storage path, so as to obtain the target storage path for each of the data to be accessed. If the current access popularity matches the access popularity of the storage layer corresponding to the current storage path, the current storage path is determined as the target storage path for each of the data to be accessed.
[0086] For details, please refer to the previous section on the hierarchical storage method for hot and cold data in intelligent computing scenarios; it will not be repeated here.
[0087] In some embodiments, the migration decision module 104 is further configured to delay the migration timing of the data to be accessed when the current bandwidth utilization rate of the storage layer corresponding to the target storage path in the plurality of storage layers is greater than or equal to a preset bandwidth utilization rate and the task priority value is less than a preset priority value.
[0088] For details, please refer to the previous section on the hierarchical storage method for hot and cold data in intelligent computing scenarios; it will not be repeated here.
[0089] In some embodiments, the system further includes a storage adapter module ( Figure 1 (Not shown in the image), configured to update the current storage path of each of the data to be accessed in the task metadata based on the migration strategy, and record the migration time.
[0090] For details, please refer to the previous section on the hierarchical storage method for hot and cold data in intelligent computing scenarios; it will not be repeated here.
[0091] The above description is merely an illustration of some embodiments of this disclosure and the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0092] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A method for cold-hot data hierarchical storage in an intelligent computing scenario, characterized in that, A tiered storage system for hot and cold data, applicable to intelligent computing scenarios, comprises a task scheduling interface module, a data acquisition module, a heat analysis module, a migration decision module, and a tiered storage module. The tiered storage module includes multiple storage layers divided according to access frequency. The method includes: The task scheduling interface module acquires and parses the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time; wherein, the task metadata includes at least the task identifier, the task priority value, and the current storage path of multiple data to be accessed required to run the intelligent computing task. The data acquisition module collects access behavior logs for each of the data to be accessed in real time. The heat analysis module uses a pre-built dynamic heat model to determine the current access heat of each of the data to be accessed based on the task priority value, the access behavior log, and the time decay factor. The migration decision module determines the migration strategy for each piece of data to be accessed based on the task priority, the current access popularity, the current storage path, and the current bandwidth utilization of the multiple storage layers; and migrates each piece of data to be accessed between the multiple storage layers according to the migration strategy. The hierarchical storage module stores each piece of data to be accessed in its corresponding storage layer.
2. The method of claim 1, wherein, The access behavior log includes at least read and write operations recorded by access timestamp; The step of using a pre-built dynamic popularity model to determine the current access popularity of each piece of data to be accessed based on the task priority value, the access behavior log, and the time decay factor includes: The number of read operations and the number of write operations for each of the data to be accessed within a preset historical time window are counted, and the historical read frequency and historical write frequency of each of the data to be accessed are calculated. Using a pre-built access frequency prediction model, based on the historical read frequency and the historical write frequency, the predicted read frequency and predicted write frequency of each of the data to be accessed within a preset future time window are predicted; Calculate the time decay factor based on the most recent access time of each of the data to be accessed; The task priority value, the predicted read frequency, the predicted write frequency, and the time decay factor are input into a pre-built dynamic heat model to calculate the current access heat of each of the data to be accessed. The dynamic heat model is expressed as: Heat( obj )= α × f read + β × f write + γ ×T -1 recent + δ ×T P wherein Heat( obj ) represents a current access heat degree of each of the to-be-accessed data; f read represents the predicted read frequency; f write represents the predicted write frequency;T -1 recent represents the time decay factor;T P represents the task priority value; α represents a weight of the predicted read frequency; β represents a weight of the predicted write frequency; γ represents a weight of the time decay factor; δ represents a weight of the task priority value.
3. The method of claim 1, wherein, The migration strategy includes the target storage path and the timing of the migration. The process of determining the migration strategy for each piece of data to be accessed based on the task priority, the current access popularity, the current storage path, and the current bandwidth utilization of the multiple storage layers includes: Based on the current access popularity and the current storage path, determine the target storage path for each of the data to be accessed; Based on the task priority value and the current bandwidth utilization of the storage layer corresponding to the target storage path in the multiple storage layers, the migration timing of each of the data to be accessed is determined.
4. The method of claim 3, wherein, The step of determining the target storage path for each piece of data to be accessed based on the current access popularity and the current storage path includes: If the current access popularity does not match the access popularity of the storage layer corresponding to the current storage path, the current storage path is updated based on the current access popularity and the access popularity corresponding to the multiple storage layers to obtain the target storage path for each of the data to be accessed. If the current access popularity matches the access popularity of the storage layer corresponding to the current storage path, the current storage path is determined as the target storage path for each of the data to be accessed; and / or, The step of determining the migration timing of each piece of data to be accessed based on the task priority value and the current bandwidth utilization of the storage layer corresponding to the target storage path among the multiple storage layers includes: If the current bandwidth utilization rate of the storage layer corresponding to the target storage path in the plurality of storage layers is greater than or equal to the preset bandwidth utilization rate, and the task priority value is less than the preset priority value, the migration timing of the data to be accessed is delayed.
5. The method according to any one of claims 1 to 4, wherein The system further includes a storage adaptation module; after migrating the data to be accessed among the multiple storage layers according to the migration strategy, the method further includes: Based on the migration strategy, the storage adaptation module updates the current storage path of each piece of data to be accessed in the task metadata and records the migration time.
6. A cold and hot data hierarchical storage system for intelligent computing scenarios, used to implement the cold and hot data hierarchical storage method for intelligent computing scenarios as described in any one of claims 1 to 5, characterized in that, The system includes a task scheduling interface module, a data acquisition module, a heat analysis module, a migration decision module, and a hierarchical storage module. The hierarchical storage module includes multiple storage layers divided according to access popularity. The task scheduling interface module is configured to acquire and parse the task metadata of the intelligent computing tasks submitted by the intelligent computing platform in real time; wherein, the task metadata includes at least the task identifier, the task priority value, and the current storage path of multiple data to be accessed required to run the intelligent computing task. The data acquisition module is configured to collect access behavior logs of each of the data to be accessed in real time. The heat analysis module is configured to determine the current access heat of each of the data to be accessed based on the task priority value, the access behavior log, and the time decay factor using a pre-built dynamic heat model. The migration decision module is configured to determine a migration strategy for each piece of data to be accessed based on the task priority, the current access popularity, the current storage path, and the current bandwidth utilization of the multiple storage layers; and to migrate each piece of data to be accessed between the multiple storage layers according to the migration strategy. The hierarchical storage module is configured to store each piece of data to be accessed in a corresponding storage layer.
7. The system as described in claim 6, characterized in that, The access behavior log includes at least read and write operations recorded by access timestamp; The heat analysis module is also configured to count the number of read operations and the number of write operations of each of the data to be accessed within a preset historical time window, and to calculate the historical read frequency and historical write frequency of each of the data to be accessed. Using a pre-built access frequency prediction model, based on the historical read frequency and the historical write frequency, the predicted read frequency and predicted write frequency of each of the data to be accessed within a preset future time window are predicted; Calculate the time decay factor based on the most recent access time of each of the data to be accessed; The task priority value, the predicted read frequency, the predicted write frequency, and the time decay factor are input into a pre-built dynamic heat model to calculate the current access heat of each of the data to be accessed. The dynamic heat model is expressed as: Heat( obj )= α × f read + β × f write + γ ×T -1 recent + δ ×T P Among them, Heat( obj The number () indicates the current access popularity of each of the data items to be accessed; f read This indicates the predicted read frequency; f write T represents the predicted write frequency; -1 recent This represents the time decay factor; T P This indicates the priority value of the task; α The weights represent the predicted read frequencies; β The weights representing the predicted write frequency; γ This represents the weight of the time decay factor; δ The weight represents the priority value of the task.
8. The system as described in claim 6, characterized in that, The migration strategy includes the target storage path and the timing of the migration. The migration decision module is also configured to determine the target storage path for each of the data to be accessed based on the current access popularity and the current storage path; Based on the task priority value and the current bandwidth utilization of the storage layer corresponding to the target storage path in the multiple storage layers, the migration timing of each of the data to be accessed is determined.
9. The system as described in claim 8, characterized in that, The migration decision module is further configured to update the current storage path based on the current access popularity and the access popularity corresponding to the multiple storage layers when the current access popularity does not match the access popularity of the storage layer corresponding to the current storage path, so as to obtain the target storage path for each of the data to be accessed. If the current access popularity matches the access popularity of the storage layer corresponding to the current storage path, the current storage path is determined as the target storage path for each of the data to be accessed. And / or, The migration decision module is also configured to delay the migration of the data to be accessed when the current bandwidth utilization of the storage layer corresponding to the target storage path is greater than or equal to a preset bandwidth utilization and the task priority value is less than a preset priority value.
10. The system as described in any one of claims 6 to 9, characterized in that, The system also includes a storage adaptation module, configured to update the current storage path of each piece of data to be accessed in the task metadata based on the migration strategy, and record the migration time.