Hot and cold data exchange method, system and storage medium based on optical storage
By collecting access feature information and evaluating the heat of logical intervals in the optical-magnetic hybrid storage system, the misjudgment problem of hot and cold data identification and migration in the existing technology is solved, adaptive adjustment and efficient exchange of data storage locations are achieved, and the stability of the storage system and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202510943024.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing data center storage systems lack in-depth modeling of overall data behavior patterns in the identification and migration of hot and cold data. This leads to migration strategies being prone to misjudgment, failing to dynamically respond to changes in storage layer load, and lacking batch collaborative migration capabilities, affecting system efficiency and stability.
By collecting access characteristic information of data in the optical-magnetic hybrid storage system, calculating the access life cycle factor, dividing the data blocks into multiple logical intervals and performing heat assessment, generating data temperature classification results, formulating data placement strategies, allocating hot data to the magnetic storage layer and cold data to the optical storage layer, and performing batch data migration by prioritizing migration tasks.
It achieves precise separation and efficient exchange of hot and cold data, reduces metadata storage space usage, reduces storage overhead for heat calculations, reduces frequent fluctuations in data between different temperature levels, improves storage resource utilization, and reduces system resource consumption and the impact of migration operations on normal business.
Smart Images

Figure CN120428929B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, system and storage medium for exchanging hot and cold data based on optical storage. Background Art
[0002] Currently, data centers widely use multi-level tiered storage architectures to manage massive amounts of data. A common structure is a combination of cache, magnetic storage, and optical storage. In this system, hot data is typically stored in DRAM, SSDs, or HDDs, which offer higher read and write performance, while cold data is transferred to storage media such as optical disk arrays or magnetic tapes, which have lower power consumption, longer lifespans, but slower access speeds. Most existing methods for identifying and migrating hot and cold data rely on a single access frequency or access time as the basis for classification. These methods trigger data migration only when access patterns suddenly change, and they lack structural analysis of logical data. Traditional methods for determining heat are often based on the behavior of a single block of data and lack in-depth modeling of the overall data behavior patterns, making migration strategies prone to misjudgment. Data migration scheduling often uses static rules for execution, failing to dynamically respond to changes in the storage tier's load and lacking the ability for coordinated batch migration, which in turn impacts the overall efficiency and stability of the system.
[0003] However, in actual applications, data access behavior is characterized by periodicity, locality, and suddenness. A single indicator cannot fully reflect the temperature change trend of the data. In addition, the lifecycles and access patterns of different types of data vary significantly. For example, the temperature change paths of logs, images, and intermediate data for model training are not consistent. If a unified threshold strategy is used for division, it will cause problems such as frequent migration of large amounts of boundary data and hot and cold switching fluctuations. More importantly, most current systems lack the ability to dynamically schedule based on the status of storage resources when performing migration tasks. They fail to evaluate the cost and benefit of tasks, and cannot realize the combination and execution control of batch tasks. This leads to frequent problems such as uneven utilization of storage resources, increased energy consumption, and reduced response performance. Summary of the Invention
[0004] The present application provides a method, system and storage medium for hot and cold data exchange based on optical storage, which is used to solve the problem of how to achieve accurate identification of logical interval temperature, dynamic grading of temperature levels, optimal matching of storage mapping and efficient organization of batch migration tasks under an optical-magnetic hybrid architecture.
[0005] In a first aspect, the present application provides a method for exchanging hot and cold data based on optical storage, which includes: collecting access feature information of data in an optical-magnetic hybrid storage system to obtain a data access feature information set; calculating an access life cycle factor based on the data access feature information set, dividing the data block into multiple logical intervals and calculating the heat value of the logical interval to obtain a logical interval heat evaluation result; generating a data temperature grading result based on the logical interval heat evaluation result; formulating a data placement strategy based on the data temperature grading result, allocating hot data to the magnetic storage layer, and allocating cold data to the optical storage layer, and generating a data placement mapping table; performing data migration operations according to the data placement mapping table, processing migration tasks according to priority, performing batch data migration, and updating data location mapping records.
[0006] In a second aspect, the present application provides a hot and cold data exchange system based on optical storage, the hot and cold data exchange system based on optical storage comprising:
[0007] An acquisition module is used to acquire access characteristic information of data in the optical-magnetic hybrid storage system to obtain a data access characteristic information set;
[0008] An access module, configured to calculate an access lifecycle factor based on the data access feature information set, divide the data block into a plurality of logical intervals, calculate the heat values of the logical intervals, and obtain a logical interval heat evaluation result;
[0009] A generating module, configured to generate a data temperature classification result based on the logic interval heat evaluation result;
[0010] an allocation module, configured to formulate a data placement strategy based on the data temperature classification result, allocate hot data to the magnetic storage layer, allocate cold data to the optical storage layer, and generate a data placement mapping table;
[0011] The execution module is used to execute data migration operations according to the data placement mapping table, process migration tasks according to priority, perform batch data migration, and update data location mapping records.
[0012] In a third aspect, a hot and cold data exchange device based on optical storage is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory to enable the hot and cold data exchange device based on optical storage to execute the above-mentioned hot and cold data exchange method based on optical storage.
[0013] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is run on the computer, the computer executes the above-mentioned hot and cold data exchange method based on optical storage.
[0014] In the technical solution provided by the present application, by collecting access feature information of data in the optical-magnetic hybrid storage system, comprehensive information on data access patterns is obtained, which provides an accurate data basis for the subsequent separation of hot and cold data, and avoids data classification errors caused by incomplete collection of access feature information in traditional data management methods; by introducing an access life cycle factor calculation mechanism and dividing data blocks into logical intervals for heat evaluation, the storage overhead of heat calculation is reduced while maintaining the accuracy of heat evaluation. This method calculates the heat of logical intervals rather than individual data blocks, which greatly reduces the metadata storage space occupied; the data temperature classification results generated based on the logical interval heat evaluation results are more stable, reducing the frequent fluctuations of data between different temperature levels, and providing data storage space for accurate classification. The proposed data placement strategy provides a reliable basis for the formulation of storage strategies; the formulated data placement strategy fully utilizes the characteristics of optical storage media suitable for storing cold data and magnetic storage media suitable for storing hot data, realizing efficient utilization of storage resources; during the execution of data migration operations, priority sorting and batch migration processing are used to reduce system resource consumption and the impact of migration operations on normal business. For the case where hot data is cooled to cold data, a delayed write strategy is adopted to avoid unnecessary migration caused by short-term fluctuations in data temperature; the entire method forms a closed-loop data management mechanism, which can adaptively adjust the data storage location according to changes in data access patterns, realize the precise separation and efficient exchange of hot and cold data, and meet the multiple requirements of storage performance, capacity and cost in the big data environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 This is a schematic diagram of an embodiment of a method for exchanging hot and cold data based on optical storage in an embodiment of the present application;
[0017] Figure 2 This is a schematic diagram of an embodiment of a hot and cold data exchange system based on optical storage in an embodiment of the present application;
[0018] Figure 3 It is a schematic block diagram of the structure of a hot and cold data exchange device based on optical storage in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The embodiments of the present application provide a method, system, and storage medium for exchanging hot and cold data based on optical storage.
[0020] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, a method for exchanging hot and cold data based on optical storage includes:
[0021] Step S101: collecting access characteristic information of data in the optical-magnetic hybrid storage system to obtain a data access characteristic information set;
[0022] Step S102: Calculate the access lifecycle factor based on the data access feature information set, divide the data block into multiple logical intervals, and calculate the heat value of the logical interval to obtain a logical interval heat evaluation result;
[0023] Step S103: Generate a data temperature classification result based on the logic interval heat evaluation result;
[0024] Step S104: Formulate a data placement strategy based on the data temperature classification result, allocate hot data to the magnetic storage layer, allocate cold data to the optical storage layer, and generate a data placement mapping table;
[0025] Step S105: Execute data migration operations according to the data placement mapping table, process migration tasks according to priority, perform batch data migration, and update data location mapping records.
[0026] It is understandable that the execution subject of the present application can be a hot and cold data exchange system based on optical storage, or a terminal or a server, which is not limited here. The embodiment of the present application is described by taking the server as the execution subject as an example.
[0027] Specifically, a data access log system is established to record access operations for each data block in the optical-magnetic hybrid storage system in real time, including access timestamps, access types, access frequency, data size, and creation time. Standardized access log data is generated by preprocessing the raw access log data, removing outliers and supplementing missing values. Aggregate analysis is then performed at preset time intervals (e.g., 24 hours), calculating the access frequency of each data block in different time periods to generate time-series access frequency data. This time-series access frequency data is then correlated with the data block creation time to calculate lifecycle characteristic values and record the migration history of data blocks between different storage tiers. Ultimately, this data is integrated into a complete set of data access characteristic information.
[0028] Characteristic parameters such as the data block's most recent access timestamp, access frequency, creation time, and migration count are extracted from the data access feature information set. The difference between the most recent access timestamp and the current timestamp is calculated and divided by the data block's expected lifecycle to obtain the temporal locality value. The ratio of the access frequency to the maximum access frequency in the system is calculated to obtain the frequency normalization value. The ratio of the number of migrations to the maximum migration threshold is calculated and subtracted from 1 to obtain the migration cost value. Quantum statistical entropy analysis is used to calculate the quantum entropy value of the data access pattern, quantifying the degree of uncertainty in data access behavior and obtaining the access entropy weight factor. Based on these calculation results, the data blocks are divided into logical intervals according to their creation time, application, and data type. The comprehensive heat value of each logical interval is calculated through a weighted combination of the temporal locality value, frequency normalization value, migration cost value, and access entropy weight factor.
[0029] An initial temperature classification threshold is set for the heat assessment result of the logical interval. Heat values above the first threshold are classified as hot data, those between the first and second thresholds are classified as warm data, and those below the second threshold are classified as cold data, thus obtaining the initial data temperature classification result. Through data volume balance analysis, the proportion of data volume at each temperature level is calculated, and the temperature classification threshold is adjusted according to the difference between the data distribution ratio and the system's preset target distribution ratio. The historical trend of the heat assessment results of the logical interval is analyzed to identify the boundary intervals with frequent heat value fluctuations, set the state transition hysteresis period, and control the frequency of temperature level changes. After performing the temperature classification calculation, the capacity utilization of each storage layer is taken into consideration. When the capacity of the magnetic storage layer is close to saturation, more boundary data will be classified as cold data, and when the read and write bandwidth of the optical storage layer is close to saturation, some boundary data will be classified as warm data.
[0030] The storage requirements for hot, warm, and cold data in the data temperature grading results are analyzed, and the access latency requirements, read / write ratio, and data size for each temperature tier are calculated. The magnetic storage layer is divided into high-performance and standard areas, and the optical storage layer is divided into active and archive areas. Data correlation analysis is performed, and a data correlation network diagram is constructed to identify data groups that are frequently accessed simultaneously. Hot data is allocated to the high-performance area of the magnetic storage layer, warm data to the standard area of the magnetic storage layer, and cold data with an access frequency exceeding a preset threshold (0.5 times per day or 3 times per week) is allocated to the active area of the optical storage layer. Long-term inaccessible cold data is allocated to the archive area of the optical storage layer, ensuring that highly correlated data groups are physically close. The data placement plan is load-balanced, leveraging the parallel access characteristics of the optical-magnetic hybrid storage system to evenly distribute data across physical devices, avoiding excessive load on any single point. Finally, a data placement mapping table is generated.
[0031] Compare the current location of the data with the target location, identify the data that needs to be migrated, and generate a list of migration tasks. Perform a priority analysis on the migration tasks, assigning the highest priority to data with significant temperature level changes, assigning a higher priority to data with large differences between the current location and the target location, and assigning a medium priority to data that is predicted to be frequently accessed in the near future. Perform a preliminary assessment of the migration task queue, calculate the ratio of resource consumption to expected benefits, and filter out tasks where benefits outweigh costs. Combine small, similar tasks into batch migration tasks, develop a batch execution plan, and schedule execution during time periods with low system load. When performing data migration, for the migration of hot data to cold data, adopt a delayed write strategy, marking the data but not moving it temporarily until resources are sufficient or the data has not been accessed for a long time, and then perform the actual migration. After the migration is completed, update the data location mapping record to record the new data physical location information and migration history information.
[0032] For example, after deploying a hybrid optical-magnetic storage system, a large data center collected access characteristics for historical data and found that user profile data for a business system was accessed more frequently during the day on weekdays and less frequently at night and on weekends. When calculating the access lifecycle factor, the temporal locality of this data was 0.75 (weekdays) and 0.25 (non-weekdays), the frequency normalization value was 0.6, the migration cost was 0.8, and the access entropy weighting factor was 0.7. After partitioning this data into logical intervals, the overall heat value during weekday daytime was 0.72, classifying it as hot data; the overall heat value during non-weekday periods was 0.45, classifying it as warm data. The data placement strategy allocated hot data during the daytime to the high-performance area of the magnetic storage tier, and warm data during non-weekdays to the standard area of the magnetic storage tier. Every night, when system load is lower, data migration is performed to move warm data to an appropriate storage location. The hot data is then moved back to the high-performance area before work begins the next morning, enabling efficient data exchange between different storage tiers.
[0033] In the embodiment of the present application, by collecting access feature information of data in the optical-magnetic hybrid storage system, comprehensive information on data access patterns is obtained, which provides an accurate data basis for the subsequent separation of hot and cold data, and avoids data classification errors caused by incomplete collection of access feature information in traditional data management methods; by introducing an access life cycle factor calculation mechanism and dividing data blocks into logical intervals for heat evaluation, the storage overhead of heat calculation is reduced while maintaining the accuracy of heat evaluation. This method calculates the heat of logical intervals rather than individual data blocks, which greatly reduces the metadata storage space occupied; the data temperature classification results generated based on the logical interval heat evaluation results are more stable, reducing the frequent fluctuations of data between different temperature levels, and providing data placement The data placement strategy fully utilizes the characteristics of optical storage media suitable for storing cold data and magnetic storage media suitable for storing hot data, thus realizing efficient utilization of storage resources. During the execution of data migration operations, priority sorting and batch migration processing are used to reduce system resource consumption and the impact of migration operations on normal business. For the case where hot data is cooled to cold data, a delayed write strategy is adopted to avoid unnecessary migration caused by short-term fluctuations in data temperature. The entire method forms a closed-loop data management mechanism, which can adaptively adjust the data storage location according to changes in data access patterns, realize the precise separation and efficient exchange of hot and cold data, and meet the multiple requirements of storage performance, capacity and cost in the big data environment.
[0034] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0035] The access operation of each data block in the optical-magnetic hybrid storage system is recorded in real time to obtain the original access log data;
[0036] Preprocess the original access log data to obtain standardized access log data;
[0037] Aggregate and analyze the standardized access log data at preset time intervals, calculate the access frequency of each data block in different time periods, and obtain time series access frequency data;
[0038] Perform correlation analysis on the time series access frequency data and the data block creation time, calculate the life cycle characteristic value of the data block, and obtain the data life cycle characteristic data;
[0039] Perform statistical analysis on historical data block migration records, record the number of migrations and reasons between different storage layers, and obtain historical data migration data.
[0040] Standardized access log data, time-series access frequency data, data lifecycle feature data, and data migration history data are integrated to establish an index structure and obtain a data access feature information set.
[0041] Specifically, a hybrid optical-magnetic storage system refers to a multi-layer heterogeneous storage architecture that combines magnetic storage media (such as HDDs or SSDs) with optical storage media (such as Blu-ray disc arrays). This architecture has the characteristics of both high-speed response and low-energy long-term storage. In this method, in order to achieve accurate identification and hierarchical storage of hot and cold data, it is necessary to obtain its access feature information at the data block level and perform multi-dimensional feature extraction and logical structuring processing to generate a data access feature information set with structured semantics. This process relies on a set of strictly defined access log collection and processing procedures. It is necessary to follow specific steps to clean, regularize, convert and construct feature values for access data to ensure that the constructed data structure can directly support subsequent hierarchical judgments.
[0042] During the real-time recording of access operations, a lightweight log tracking component needs to be embedded in the storage system. This component tracks every access operation to each data block (a data block is the smallest unit of read or write operations in the storage layer, typically 4KB or 8KB). This component records information such as the access timestamp, access operation type (read or write), access thread or call source, operation duration, and the size of the data block involved. A data block is typically uniquely identified by its logical volume number, offset within the block, and storage layer type (magnetic or optical). Log records are asynchronously cached in the master control unit's log buffer and flushed in batches every minute to reduce system I / O load. During continuous operation, this module generates a high density of raw access log data. This raw data is an unstructured, multi-field event sequence that requires further preprocessing before analysis. This preprocessing step involves normalizing the raw log fields, converting to a uniform unit, and performing field cleansing. For example, access timestamps are uniformly converted to a standard UNIX time format (seconds), access types are encoded using binary values (read as 0, write as 1), and non-business access records such as system background scans or health checks are filtered out. Each access log record is processed as a fixed-structure JSON object or structured table entry, forming a standardized access log data set. This standardized access log data is then partitioned and aggregated according to time windows, determining the number of accesses to each data block during different time periods (e.g., hourly or daily). Statistics are generated by creating an aggregated hash table for each data block identifier, accumulating access counts using time segments as keys, and outputting time-series access frequency data. Records are recorded as a set of key-value pairs: {data block identifier: [T1: f1, T2: f2,..., Tn: fn]}, where Tn is the time window number and fn is the access frequency for the corresponding time period.
[0043] To further extract data lifecycle characteristics, the aforementioned time-series access frequency data needs to be aligned and analyzed with the data block's creation time. The data block creation time is recorded during the first write operation and permanently bound to its metadata structure, forming a timestamp, Tc. By performing a sliding window test on the distance between each data block's frequency curve and Tc, we determine whether its peak access period is close to or far from the creation time. If the access peak is concentrated near the creation time, the lifecycle tends to be short-term. If access fluctuates periodically or remains active later in the process, the lifecycle tends to be long-term. Based on this, a lifecycle characteristic is defined as a range label (short-term / medium-term / long-term) and a numerical density metric (such as average access frequency / first peak time offset), generating structured data lifecycle characteristic data. The migration history of data blocks is summarized and analyzed. In the optical-magnetic hybrid system, each cross-tier migration of a data block triggers a migration event, including the source storage tier ID, target storage tier ID, migration time, migration trigger reason (such as temperature level change, space load balancing, etc.), and whether it is part of a batch migration. This information is aggregated to calculate the number of migrations for each data block and cluster the migration reasons to identify frequent migrations due to misjudgment of temperature levels or short-term bursts of access. This information forms a migration statistics table, which includes fields such as the total number of migrations for each data block, a migration periodicity analysis label (frequent migration / stable), and a migration direction trend (heating / cooling), forming historical data migration data.
[0044] The above four types of data - standardized access log data, time-series access frequency data, data lifecycle feature data, and data migration history data - are jointly integrated to establish an efficient index structure and build a data access feature information set. In the actual design, this index structure uses an inverted index combined with a B+ tree index mechanism to establish a primary index item for each data block, with the data block ID as the key. The corresponding value structure is a set of field reference pointers pointing to the location of the above four types of structured data. Through this structure, access feature queries for any data block can be completed with O(logN) complexity, and batch retrieval based on time, access behavior, or storage layer status as filtering conditions is supported, meeting the requirements of subsequent logical interval division and heat evaluation processes for data density and accessibility.
[0045] For example, in one application scenario, a data block within a system with video streaming distribution characteristics was created at 2:00 AM one day. Within the first 12 hours of that day, the data block was accessed 600 times. Access frequency then plummeted, with only five accesses over the next two days. Reads predominated, accounting for over 95% of accesses. Within 24 hours of creation, the data block was marked as hot and migrated from optical storage to magnetic storage. 24 hours after the sudden drop in access, it was migrated back to optical storage. Frequency distribution sequences derived from access log aggregation revealed high concentration of hotness, short-term concentration of lifecycle characteristics, and a strong cooling trend indicated by migration records. By mapping the four data structures to primary index entries, subsequent hotness calculations and temperature grading can directly access frequency, migration, and lifecycle parameters based on the index, eliminating the need for a full data traversal or repeated calculations. This case demonstrates how access data, through structured processing and indexing, can be logically integrated and how data behavior characteristics can be reflected through the combined use of multiple dimensions.
[0046] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0047] Extract the most recent access timestamp, access frequency, creation time, and migration times of the data block from the data access feature information set to obtain data block feature parameters;
[0048] The difference between the most recently accessed timestamp and the current timestamp in the data block characteristic parameters is calculated and divided by the expected life cycle of the data block to obtain the temporal locality value;
[0049] Calculate the ratio of the access frequency in the data block characteristic parameters to the maximum access frequency observed in the system to obtain a frequency normalization value;
[0050] Calculate the ratio of the number of migrations in the data block characteristic parameters to the maximum migration number threshold set by the system, and subtract the ratio from 1 to obtain the migration cost value;
[0051] Perform quantum statistical entropy analysis on data blocks, calculate the quantum entropy value of data access patterns, quantify the degree of uncertainty and chaos in data access behavior, identify data with deterministic access patterns and random access characteristics, and obtain access entropy weight factors;
[0052] Based on the creation time, application and data type of the data blocks, data blocks with similar attributes are divided into the same logical interval. Combined with the time locality value, frequency normalization value, migration cost value and access entropy weight factor, the comprehensive heat value of each logical interval is calculated through weighted combination to obtain the logical interval heat evaluation result.
[0053] Specifically, obtaining and processing the heat information of the logical interval must rely on the core feature parameters extracted from the data access feature information set, and transforming these parameters through clear mathematical operations to generate quantifiable heat evaluation data. The data block feature parameter extraction operation first locates the field and retrieves the value from the established data access feature information set. This set has been standardized in the previous step and contains fields such as the most recent access timestamp, access frequency, data creation time, and migration count of each data block. The access timestamp refers to the system time when the data block was last accessed, the access frequency refers to the cumulative number of times the data block is accessed within a unit time window (for example, 24 hours), the data creation time refers to the time point of the first write operation of the data block, and the migration count refers to the total number of operations in which the data block is migrated from one storage layer to another storage layer in the past time period.
[0054] After obtaining these fields, the first step is to calculate the temporal locality value, which measures the activity of the data block within its lifetime. The calculation logic involves subtracting the current system time from the data block's most recent access time to obtain the access interval value. This value is then divided by the data block's expected lifetime value to obtain the normalized temporal locality value. The expected lifetime value is determined based on the data block's data type and business logic. For example, the expected lifetime of some video surveillance data may be 30 days, while that of online log data may be 7 days. A temporal locality value closer to 0 indicates that the data block has been recently accessed and has a high degree of timeliness. A temporal locality value closer to 1 indicates that the data has not been accessed for a long time and is in a cooling-off phase.
[0055] Access frequencies are normalized. To eliminate outliers caused by peak service loads or unusual access behavior, the maximum access frequency over a period of time across all data blocks in the system is first calculated as a standard upper limit. The access frequency of the current data block is divided by this maximum value to obtain the normalized frequency value. This normalization ensures that frequency values across different service types are comparable when calculating data block popularity, thus avoiding misjudgments of data popularity due to inherent differences in service access. The migration cost is then calculated. This value penalizes data blocks that are frequently migrated in the system, preventing excessive system resource consumption due to unstable popularity fluctuations. The calculation logic is to compare the number of migrations a data block has undergone with the system's maximum migration threshold and then subtract the ratio from 1. The maximum migration threshold is typically set based on the system's maximum migration load, such as 3. If a data block has undergone two migrations and the threshold is 3, the ratio is 2 / 3, and the migration cost is 1-2 / 3 = 1 / 3. A larger value indicates a lower data migration burden and a higher suitability for further migration.
[0056] Next, we conduct quantum statistical entropy analysis on data access behavior, with the goal of measuring the degree of certainty and randomness of data block access behavior. The specific approach is to perform probability distribution modeling on the access time series data of the data block, use the number of accesses in each time slice as a sample, construct an access frequency distribution function, and then calculate the statistical entropy of the distribution. If the access behavior is concentrated at certain fixed time points or fluctuates periodically, the entropy value is low, indicating that the access regularity is strong; if the distribution is uniform or fluctuates frequently, the entropy value is high, indicating that the access uncertainty is strong. In order for the entropy value to be able to participate in the subsequent heat calculation, it is necessary to further normalize the entropy value according to the maximum entropy value range to form an access entropy weight factor, which is used to characterize the degree of influence of the uncertainty of the data block in the heat determination.
[0057] After completing the above parameter extraction and conversion, multiple data blocks with similar characteristics need to be grouped into the same logical interval based on three types of metadata: creation time, application, and data type. A logical interval is a collection of data blocks formed according to certain semantic aggregation rules. For example, the classification criteria include log data generated by the same application within the same time period or video files of the same type. This grouping method facilitates subsequent unified evaluation and batch processing.
[0058] After logical interval division is completed, the temporal locality value, frequency normalization value, migration cost value, and access entropy weight factor of all data blocks within the same interval are collected separately, and a weighted summation method is used to generate the interval's comprehensive heat value. The weight coefficient must be determined in advance through system learning or policy setting to balance the contribution of each characteristic parameter to heat. For example, for business types that are more sensitive to access frequency, the weight of the frequency normalization value can be increased, and for systems with higher migration costs, the impact coefficient of the migration cost value can be increased. The comprehensive heat value of each logical interval serves as an evaluation indicator of the hot and cold status of the interval and is involved in subsequent temperature grading and storage strategy formulation.
[0059] For example, assume that a hybrid optical-magnetic storage system contains a set of model cache data generated by an image recognition service. These data blocks are all created at 12:00 PM on a certain day. They are structured parameter fragments with sizes ranging from 16KB to 64KB. Access frequency is very concentrated within the first eight hours after creation, with an average of 30 accesses per hour, then drops sharply to less than once per day. The data block was last accessed 48 hours before the current time, and the current system time is 72 hours after its creation. The system has a default lifetime of 96 hours for such data blocks. According to data processing logic, the temporal locality value is equal to the access interval of 48 hours divided by the lifetime of 96 hours, which is 0.5. This data block ranks 10th in frequency systemwide, with a maximum access frequency of 80 times per hour. The normalized frequency value is 30 / 80 = 0.375. Migration statistics show that it has been migrated once. The maximum migration threshold is 3, resulting in a migration cost of 1-1 / 3 = 0.666. Access to this data block is highly concentrated in the early stages and becomes sparser later. Entropy analysis yields a medium entropy value, normalized to 0.45. Based on the consistency of the application and data type, this data block is grouped with other similar blocks into a single logical interval. Weights are set as follows: temporal locality 0.3, frequency normalization 0.3, migration cost 0.2, and access entropy weight 0.2. Substituting all these parameters, the logical interval heat is calculated as: 0.3 × 0.5 + 0.3 × 0.375 + 0.2 × 0.666 + 0.2 × 0.45 = 0.15 + 0.1125 + 0.1332 + 0.09 = 0.4857. This heat value indicates that this logical interval is in a warm state, and its storage location can be further delineated based on the heat threshold.
[0060] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0061] An initial temperature classification threshold is set for the heat evaluation result of the logical interval. The logical interval with a heat value higher than the first threshold is classified as hot data, the logical interval with a heat value between the second threshold and the first threshold is classified as warm data, and the logical interval with a heat value lower than the second threshold is classified as cold data, thereby obtaining the initial data temperature classification result.
[0062] Perform data balance analysis on the initial data temperature classification results, calculate the data volume proportion of each temperature level, and obtain the data distribution ratio;
[0063] According to the difference between the data distribution ratio and the target distribution ratio preset by the system, the temperature classification threshold is adjusted. When the proportion of hot data exceeds the target value, the first threshold is increased, and when the proportion of cold data is lower than the target value, the second threshold is lowered to obtain the revised temperature classification threshold;
[0064] Analyze the historical trend of the heat evaluation results of the logical intervals, identify the boundary intervals where the heat value fluctuates frequently, set the state transition hysteresis period for these intervals, and obtain the temperature stability control parameters;
[0065] The logic interval heat evaluation result, the corrected temperature classification threshold and the temperature stability control parameter are input into the temperature classification processor, and the temperature classification calculation is performed to obtain the optimized data temperature classification result;
[0066] The optimized data temperature grading results are analyzed for storage load impact. Taking into account the current capacity utilization of each storage layer, more boundary data is classified as cold data when the capacity of the magnetic storage layer is close to saturation, and some boundary data is classified as warm data when the read and write bandwidth of the optical storage layer is close to saturation, thus obtaining the data temperature grading results.
[0067] Specifically, using the logical interval heat assessment results as input, an initial temperature grading threshold is set, delineating the heat ranges corresponding to each level of data. A logical interval is a unit consisting of a collection of data blocks with consistent access characteristics. Heat is a quantitative metric calculated by weighting the temporal locality value, frequency-normalized value, migration cost value, and access entropy weight factor of the data block, and has a clear numerical form. When setting the initial temperature grading threshold, two boundary heat values are set to divide the heat space into three intervals: intervals above the upper threshold are classified as hot data, those between the upper and lower thresholds are classified as warm data, and those below the lower threshold are classified as cold data. This division method has a clear numerical basis and serves as the basis for subsequent data migration and placement strategies. After completing the initial division, a data balance analysis is performed on the results. The essence of this analysis is to calculate the proportion of logical intervals in the three temperature levels within the total number of intervals in the system. The specific process involves traversing all logical intervals, mapping their temperature levels to category labels based on the interval in which the heat value falls, summing the counts for each category, and normalizing the total count to obtain the proportion of hot, warm, and cold data. The statistics here need to be completed using parallel computing in a high-concurrency data access environment, usually using distributed MapReduce to perform counting and accumulation.
[0068] After determining the percentage of each category, the statistical result needs to be subtracted from the system's preset target ratios. These target ratios are reference ratios set during system design based on storage resource allocation and access latency requirements. For example, hot data should be controlled at 20%, warm data at 50%, and cold data at 30%. By subtracting the actual ratios from the target ratios, we can determine whether the current classification is appropriate. If the hot data exceeds the set threshold, the hotness determination is too broad, and the upper threshold should be raised to reclassify some logical intervals that were previously hot data as warm data. If the cold data ratio is too low, the cold zone classification is too strict, and the lower threshold should be lowered appropriately to allow more boundary logical intervals to enter the cold data range. The threshold adjustment logic fine-tunes the current threshold based on the difference ratio to ensure that the system temperature level remains dynamically balanced. After threshold adjustment, historical trends in logical interval heat levels should be analyzed to identify boundary intervals with frequent heat fluctuations. This analysis involves accessing the history of the logical interval's heat changes over multiple cycles. Using a sliding window, the stability of the heat sequence is determined. If a logical interval's heat value repeatedly crosses warm or cold thresholds over multiple cycles, it is considered to have frequent heat fluctuations. Such intervals are particularly prone to migration oscillations, requiring a state transition hysteresis. This means that if the interval's heat value crosses the threshold again, its temperature level is not immediately adjusted. Instead, the transition is reassessed after the delay expires. The length of the hysteresis is determined based on the average interval of the interval's past heat change cycles and stored in the temperature stability control parameters.
[0069] After obtaining the logical interval heat assessment results, the revised temperature classification thresholds, and the temperature stability control parameters, all three are input into the temperature classification processor for calculation. Based on the input heat values and the revised threshold mapping logic, the processor, combined with the state hysteresis parameters, determines whether to allow temperature classification adjustments in the current cycle and outputs a new logical interval temperature label. This process incorporates hysteresis control as a final classification condition, helping to buffer boundary intervals and reduce frequent temperature fluctuations caused by unstable data. After obtaining the optimized data temperature classification results, a storage load impact analysis is performed to calibrate the final classification results. This analysis requires real-time monitoring of the capacity utilization and read / write bandwidth usage of the magnetic and optical storage layers. If the magnetic storage layer utilization exceeds a certain warning value, such as 85%, it is considered to be space-constrained. To alleviate the load, data near the boundary of the warm and hot thresholds should be relocated toward warm data. If the optical storage layer's read / write bandwidth is nearing full capacity, some logical intervals at the cold-warm boundary should be relocated toward warm data to reduce I / O congestion caused by frequent writes. This adjustment operation is implemented through secondary classification judgment. The rule logic is to re-screen resource-sensitive data in the boundary interval set, reverse the classification, and update it to the final temperature label.
[0070] For example, in a hybrid optical-magnetic storage system, the heat values of a batch of intermediate cached data blocks generated by the model inference service fluctuated significantly over the past week. Some logical intervals dropped from 0.81 to 0.76 within 72 hours, then rebounded to 0.83 within the next 24 hours. According to the initial tiering logic, these heat values remained within the hot data range and were classified as hot. After counting all logical intervals, the proportion of hot data reached 27%, exceeding the system's target of 20%. Therefore, the upper threshold was adjusted, lowering the heat value from the original boundary to 0.85. This resulted in the interval with a heat value of 0.83 being classified as warm data in the next tiering. Furthermore, this logical interval, having crossed the threshold three times in a row, was marked as having frequently fluctuating heat. Even though its heat value remained above the current hot data threshold of 0.85 during this tiering cycle, it was temporarily retained as warm data due to a hysteresis mechanism. Subsequent storage layer load analysis detected an increase in write latency in the magnetic storage layer, with the current space occupied at 92%. Consequently, several logical intervals on the warm-hot boundary were reclassified as warm data to alleviate the magnetic storage load. The entire process is driven by the heat value of the logical interval, combined with threshold adjustment strategy, heat stability control and storage resource occupancy feedback to form a complete set of dynamically adjustable temperature grading mechanisms, ensuring that the data classification results have real-time and resource-aware capabilities, and supporting the efficient scheduling of hot and cold data under the optical-magnetic hybrid architecture.
[0071] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0072] Perform storage demand analysis on hot data, warm data, and cold data in the data temperature grading results, calculate the access latency requirements, read-write ratio, and data size of each temperature grade data, and obtain storage demand parameters;
[0073] Partitioning the magnetic storage layer into a high-performance area and a standard area, and dividing the optical storage layer into an active area and an archive area according to storage demand parameters, thereby obtaining a storage layer partition structure;
[0074] Perform correlation analysis on the data in the temperature classification results, construct a data correlation network diagram, and obtain data correlation group information;
[0075] Based on the storage layer partition structure and data association group information, hot data is allocated to the high-performance area of the magnetic storage layer, warm data is allocated to the standard area of the magnetic storage layer, cold data with access frequency higher than the preset threshold is allocated to the active area of the optical storage layer, and cold data that has not been accessed for a long time is allocated to the archive area of the optical storage layer. At the same time, it ensures that highly associated data groups are allocated to similar physical locations, thus obtaining a preliminary data placement plan;
[0076] The initial data placement plan was optimized for load balancing. Leveraging the parallel access characteristics of the optical-magnetic hybrid storage system, data was evenly distributed across physical devices to avoid excessive load at any single point, resulting in a balanced data placement plan.
[0077] The balanced data placement scheme is converted into a data location mapping table, which records the physical location information of the data in each logical interval in the storage system, including the storage layer type, device identifier and physical address, to obtain the data placement mapping table.
[0078] Specifically, in the hot-cold data exchange method based on optical storage, logical intervals are temperature-classified to generate hot, warm, and cold data. Further storage resource adaptation and precise positioning are then performed based on the actual access requirements of each data type. This process first uses a storage requirements analysis operation to perform statistical analysis on the data characteristics of each temperature level, focusing on three key dimensions: access latency requirements, read-write ratio, and average data block size. Access latency requirements are determined by pairing historical access behavior with request response time. The access latency distribution for each logical interval type is calculated, and its 95th percentile value is extracted as the typical response requirement. The read-write ratio is derived from the access type field in the raw access logs, calculating the ratio of write to read operations within each logical interval type. Data size is directly derived from the physical length field of each data block, and the average value is calculated after aggregation. After the logical interval labels are determined, these three statistics are aggregated and grouped in memory. This results in a high-frequency, low-latency, and high-write ratio pattern for hot data, a medium-latency read-write balance for warm data, and a high-latency, low-access, and low-write pattern for cold data. Once the storage requirements are determined, the magnetic and optical storage layers are structurally partitioned. The magnetic storage layer is sub-partitioned using performance-based logic, divided into high-performance and standard zones. This partitioning logic is based on the device's I / O throughput, queue depth, and concurrent scheduling capabilities. More responsive storage devices (such as high-speed HDDs or high-performance SSDs) are assigned to the high-performance zone, while slower-responding but larger-capacity devices are assigned to the standard zone. The optical storage layer is divided into active and archive zones based on data activity requirements. The active zone is preferentially allocated to optical disk arrays or multi-laser channel systems with relatively high read bandwidth. Cold data with access frequencies exceeding a preset threshold (0.5 times per day or three times per week) is allocated to the active zone of the optical storage layer. The preset threshold is 0.5 times per day or three times per week. The archive zone is allocated to single-laser readers or low-speed, low-power devices. This partitioning operation is based on the device's actual performance indicators and current system resource information. It is not statically set; instead, labels are dynamically assigned based on the storage layer's operating status and device performance model.
[0079] Perform correlation analysis on hierarchical data and construct a data correlation network graph. This graph structure uses logical intervals as graph nodes and co-access relationships, co-write batches, or frequent cross-access writes in adjacent time periods as edge weights to establish directed or undirected edges. Co-access relationships are obtained through clustering analysis of access logs. If two logical intervals are frequently accessed by the same thread within the same time window, they are considered to be correlated; co-write batches are derived by clustering metadata-annotated time windows; and adjacent write times determine whether access behavior linkage occurs within a sliding window. The graph structure constructed in this way reflects the logical closeness between data. Graph partitioning algorithms such as Louvain modularity clustering or label propagation are used to extract associated data groups and generate data correlation group information. This information structure is a collection of logical interval sets, indicating which logical intervals should be placed close together when distributing data.
[0080] Based on the formed storage layer partition structure and data association group information, preliminary data placement mapping operations are performed. In this process, hot data is placed in the high-performance area of the magnetic storage layer first, and the principle of priority proximity of the association group is followed during the allocation process, that is, if multiple logical intervals in the same group are hot data, their physical locations are arranged on the same disk or the same RAID group; warm data is allocated to the standard area of the magnetic storage layer, and is sorted and allocated according to the size of the logical interval and the remaining space of the device; cold data is divided into active and inactive categories according to the activity level, and is respectively placed in the active area and archive area of the optical storage layer, and the storage order gives priority to the sequential write optimization of the current writable segment of the optical disc. In all the above categories, the physical location assignment operation of the logical interval must take into account the current load status of the device and the space fragmentation to avoid waste of resources caused by scattered writing.
[0081] After the initial placement is complete, load balancing optimization is performed on the resulting data layout. This operation iterates data redistribution based on the current resource utilization of each storage device. The total logical extent capacity, current I / O task queue depth, and predicted future access popularity are calculated for each device. Data migration simulations are performed on devices with excessively high load indicators, attempting lateral migration without changing the storage tier. The target devices involved in the migration simulation must meet similar performance metrics to ensure that data access performance is not degraded due to the migration. Data migration paths are generated, selecting devices based on the shortest distance first principle, ultimately forming a balanced physical placement mapping plan. This plan is converted into a standardized data location mapping table. The mapping table contains fields such as the logical extent number, storage tier type label (magnetic or optical), device identifiers (such as disk serial number, optical disc number), physical address offset, and storage zone tags (such as high-performance zone, archive zone). The table is stored in a queryable index structure. This table serves as the core address lookup basis for the storage system, used for address resolution of subsequent access requests and generation of migration task plans.
[0082] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0083] Compare the current data location with the target location based on the data placement mapping table, identify the data that needs to be migrated, and generate a migration task list;
[0084] Prioritize the tasks in the migration task list to obtain a prioritized migration task queue;
[0085] Pre-evaluate the prioritized migration task queue, calculate the ratio of resource consumption to expected benefit for each migration task, select tasks with greater benefits than costs, and obtain the optimized migration task set.
[0086] Combine the small, similar tasks in the optimized migration task set into batch migration tasks, formulate a batch execution plan, and schedule them to be executed during periods of low system load to obtain a batch migration execution plan.
[0087] Perform data migration according to the batch migration execution plan. For migrations where hot data is cooled down to cold data, a delayed write strategy is used. This strategy marks the data but does not move it until sufficient resources are available or the data has not been accessed for a long time. The actual migration is then performed to obtain the migration execution results.
[0088] Based on the migration execution result, the data location mapping record is updated to record the new data physical location information and the migration history information, including the migration time, source location, target location and migration reason, to obtain an updated data location mapping record.
[0089] Specifically, the data placement mapping table details the current physical storage location and expected target location of each logical interval. Migration task identification uses the data placement mapping table as input and compares the current location field against the target location field. If the two fields point to different storage tiers or device identifiers, the logical interval is deemed to require migration and is added to the migration task list. This process uses the logical interval identifier as the index primary key, and the current location and target location as comparison fields. Identification is accomplished through a full table scan and conditional filtering, resulting in a structured migration task list containing fields such as the logical interval number, source location, target location, data size, and migration reason code. After the migration task list is generated, tasks are prioritized. The priority calculation model is based on a multi-dimensional scoring mechanism, with the scoring dimensions including the change in the logical interval's temperature level, the current storage tier load status, and the predicted access intensity. The temperature level change is calculated as the absolute difference between the previous and next temperature values, reflecting the urgency of data temperature adjustment. The current storage tier load status is derived from a comprehensive calculation of real-time I / O utilization and capacity utilization. The predicted access intensity is determined by the recent trend in access frequency growth. The weighted sum of these three factors forms a priority score, with higher values indicating greater migration urgency. This ranking creates a prioritized migration task queue, which serves as the basis for subsequent scheduling.
[0090] A pre-assessment of migration tasks is performed to measure the ratio of each task's resource consumption to its expected benefit, preventing inefficient migration tasks from disrupting system resources. Resource consumption is calculated based on the data size, the bandwidth between the source and target devices, and the system load level during the expected migration period. Expected benefit is determined by the improvement in access latency of the logical interval at the new storage location after migration, changes in I / O concurrency efficiency, and the load balancing effect of the target storage layer. The two values are then calculated and their ratio is calculated. If the ratio falls below the set threshold, the task is eliminated. Ultimately, tasks with a reasonable ratio of resource investment to migration value remain, forming the optimized migration task set.
[0091] After the migration task set is determined, small migration tasks with similar characteristics are aggregated to form batch migration tasks. The criteria for determining similar tasks are that the target location is in the same storage device or physical area, the data size is below the threshold, and the migration direction is consistent (for example, all are magnetic-to-optical migration). The aggregation operation uses the same target device as the clustering condition. After the tasks are merged, a unified migration batch number is generated, and a batch execution plan is formulated. The time schedule of the batch execution plan must be based on the load forecast results provided by the system resource monitoring module. The execution time is scheduled during the I / O idle period (such as the early morning period or the business off-peak period). The plan structure records information such as the migration batch number, task list, expected start time, resource allocation quota, etc.
[0092] During batch migration tasks, a delayed write strategy is used for data migrations from hot data to cold data. This strategy sets a pending migration mark for the temperature drop interval, rather than immediately performing the physical move. The actual migration is triggered only after a certain period of time passes with no access events and resource status conditions are met. This strategy can prevent frequent hot-cold switching caused by short-term fluctuations, leading to repeated migrations. It uses both a delay threshold (e.g., 24 hours of no access) and a resource availability threshold (e.g., when the write load on the magnetic storage layer falls below a set value) as the migration trigger.
[0093] After the migration operation is complete, the data location mapping record is updated based on the migration execution results. This update replaces the logical interval current location information field in the original mapping table with the target location field. The updated information structure must retain the version identifier to support subsequent migration traceability analysis. Simultaneously, a new migration history record is added to each migration record. Fields include the logical interval number, migration timestamp, source device ID, target device ID, migration reason code, and execution batch number. This data is stored in the migration history log for subsequent optimization analysis and policy adjustments.
[0094] For example, suppose a batch of hot data generated by the model inference service has seen its heat value drop from 0.83 to 0.58 over the past eight hours, shifting from hot to warm. At the same time, the standard area utilization of the current magnetic storage layer has reached 91%. According to the data placement mapping table, the target location for this batch of logical intervals should be the standard area of the magnetic storage layer. However, due to further stabilization of the cooling trend, the target location of some data needs to be adjusted to the active area of the optical storage layer. A location comparison detected an inconsistency between the target location field and the current location field, marking the data for migration. Priority analysis indicates a 0.25 temperature drop, exceeding the storage layer load threshold. Access prediction indicates a decreasing access probability in the future, resulting in a high score ranking. Pre-migration assessment calculations indicate that this migration will reduce average I / O response time by 35ms and free up 14% of the write bandwidth of high-performance disks. Resource consumption is manageable, and the benefits significantly outweigh the costs. The data size of these logical intervals is all under 64KB, with the same target location. They are aggregated into a batch migration task and scheduled for execution at 3:00 AM, during system idle time. A delayed migration strategy is set for hot-to-cold data, marking it for migration. The migration is triggered only after no access events are confirmed for the next 24 hours. After the migration is complete, the target location field in the data location mapping table is written with the new value. A new migration history record is created, including the migration reason code "Temperature reduction." The execution batch number and timestamp fields are also recorded, enabling closed-loop management of the entire migration process.
[0095] In a specific embodiment, the process of performing the step of identifying data that needs to be migrated may specifically include the following steps:
[0096] Parse the data placement mapping table, extract the current storage location information and target storage location information of each logical interval, and obtain a location information comparison table;
[0097] Analyze storage layer location differences based on the location information comparison table, identify logical intervals where storage layers have changed, mark these intervals as objects to be migrated, and obtain a preliminary migration object set.
[0098] Perform data volume statistics for each logical interval in the initial migration object set, calculate the data size, number of data blocks involved in the migration, and estimated migration time, and obtain migration resource requirement data;
[0099] Detect changes in continuous storage space based on the location information comparison table, combine data that is physically continuous but will be dispersed after migration into a whole migration task, and combine data that is physically dispersed but will be integrated after migration into a whole migration task, to obtain a migration task combination plan;
[0100] Integrate the migration resource demand data and the migration task combination plan, group the logical intervals according to the migration direction, generate a cooling migration task set from the magnetic storage layer to the optical storage layer and a heating migration task set from the optical storage layer to the magnetic storage layer, and obtain a directional migration task set;
[0101] Convert the directional migration task set into a migration task description in a standard format, including task identifier, data interval range, source location, target location, data size, and migration type, to generate a migration task list.
[0102] Specifically, the data entries for each logical interval recorded in the mapping table are read sequentially, and the current location field and target location field are extracted. These fields represent the physical storage device where the current data resides and the recommended new storage area based on temperature perception, respectively. In this solution, a logical interval refers to a set of data blocks that are logically related, have consistent access characteristics, and have the same temperature level. The extraction process is based on a bidirectional mapping index structure. While ensuring the order consistency of the location mapping table, the two sets of location information are stored in a comparison buffer as key-value pairs, resulting in a location information comparison table. The core structure of this table consists of three items: the logical interval number, the current location path, and the target location path, which provide support for subsequent difference analysis and migration logic judgment. After constructing the location information comparison table, a difference comparison is performed to identify logical intervals whose current and target physical locations are in different storage tiers. The judgment criteria are that the tier fields marked by the current and target locations are inconsistent, for example, the tier mark changes from a magnetic storage tier mark to an optical storage tier mark. The matching result is a Boolean vector, where a true value indicates that migration is required and a false value indicates that migration is not required. After the traversal is completed, all logical interval number sets that meet the migration conditions are recorded as the preliminary migration object set, which becomes the original input source for subsequent resource scheduling and migration task combination.
[0103] For each logical interval in the initial migration object set, it is necessary to extract the number of data blocks it contains and the size of each data block. This is then used to cumulatively calculate the overall data size and number of data blocks for the logical interval. This is combined with the maximum available bandwidth and average I / O speed of a single device set by the current system settings to calculate the estimated time required for the migration. This estimated time is calculated by dividing the total data volume of the logical interval by the allocatable bandwidth of the current target storage layer device, and multiplying it by the redundancy check and synchronization timing factors to form the migration resource requirement data. The migration resource requirement data includes fields such as the logical interval number, number of data blocks, total data size, estimated migration time, target device number, and allocatable resource level, which are used for sorting and prioritizing subsequent migration tasks during scheduling.
[0104] The next step is to further analyze the arrangement of logical intervals in physical space based on the location information comparison table. If it is found that some logical intervals currently have continuous physical block addresses in the magnetic storage layer (that is, the starting and ending physical addresses can form an arithmetic sequence), but their target locations are multiple non-contiguous disc segments in the optical storage layer, then this group of logical intervals must be split into multiple subtasks to avoid spatial fragmentation and increased positioning complexity caused by the overall migration. Conversely, if multiple physically dispersed logical intervals correspond to continuous writable segments on the disc at the target location, they can be merged and combined into a single migration task, performing continuous write operations to optimize data writing performance. This judgment requires combining the block-level physical location distribution map and comparing the intervals with the target partition block table through linear address mapping.
[0105] After migration tasks are grouped, they are categorized by migration direction. This determines whether the logical interval is migrating from the magnetic storage tier to the optical storage tier, or from the optical storage tier back to the magnetic storage tier. The former is referred to as the cooling migration task set, while the latter is referred to as the warming migration task set. This grouping is based on the source and target storage tier type fields, with each set further subcategorized based on the target device. Migrating logical intervals in the warming direction typically carries a high risk of concurrent access, while migrating in the cooling direction is often used for archiving and optimizing space allocation. This migration direction classification allows for the development of different migration strategies and resource allocation models.
[0106] The above migration task set is converted into a unified migration task description format to form a standard migration task list. Each task description includes fields such as task number, migration direction (cooling or heating), logical interval number, starting physical location, target physical location, data size, target device number, task batch number, and estimated duration. This task list serves as input to the subsequent batch migration scheduling module, which, in conjunction with the system load prediction scheduler, performs resource allocation, scheduling window selection, and concurrent migration control.
[0107] For example, after a periodic heat adjustment analysis, it was detected that the temperature level of the five logical intervals dropped from warm to cold, and their target positions were pointed from the standard area of magnetic storage to the active area of the optical storage layer. The current location information extracted from the data location mapping table is the adjacent sector of disk group B, and the target position is the five free optical disc logical blocks of optical disc array A. After constructing the location information comparison table, it was found that the storage level has changed and it is marked as an object to be migrated. Based on the block size of 32KB and a total of 160 data blocks, the total amount of data to be migrated is calculated to be 5MB, the current write bandwidth of optical storage is 50MB / s, and the redundancy coefficient is 1.2. The estimated migration time is calculated to be 0.12 seconds. After analyzing the target optical disc partition table, it was found that the five target logical blocks are physically continuous areas, and the five logical intervals can be merged into a single batch migration task. This task is included in the cooling migration task set and is numbered MIG-C001. The task description records the task direction as magnetic to optical, the logical interval ID is SEQ1001 to SEQ1005, the source device is disk group B, the target device is optical disk array A, and the estimated execution time is 1:30 a.m. on the same day. The task list is completed and enters the scheduling phase.
[0108] The above describes the hot and cold data exchange method based on optical storage in the embodiment of the present application. The following describes the hot and cold data exchange system based on optical storage in the embodiment of the present application. Figure 2 In the embodiment of the present application, an embodiment of a hot and cold data exchange system based on optical storage includes:
[0109] The acquisition module 201 is used to collect access characteristic information of data in the optical-magnetic hybrid storage system to obtain a data access characteristic information set;
[0110] An access module 202 is configured to calculate an access lifecycle factor based on the data access feature information set, divide the data block into multiple logical intervals, and calculate the heat values of the logical intervals to obtain a logical interval heat evaluation result;
[0111] A generating module 203 is configured to generate a data temperature classification result based on the logic interval heat evaluation result;
[0112] an allocation module 204 for formulating a data placement strategy based on the data temperature classification result, allocating hot data to the magnetic storage layer and cold data to the optical storage layer, and generating a data placement mapping table;
[0113] The execution module 205 is configured to execute data migration operations according to the data placement mapping table, process migration tasks according to priority, perform batch data migration, and update data location mapping records.
[0114] Through the coordinated cooperation of the above components, by collecting access feature information of data in the optical-magnetic hybrid storage system, comprehensive information on data access patterns is obtained, providing an accurate data basis for the subsequent separation of hot and cold data, and avoiding data classification errors caused by incomplete collection of access feature information in traditional data management methods; by introducing an access life cycle factor calculation mechanism and dividing data blocks into logical intervals for heat evaluation, the storage overhead of heat calculation is reduced while maintaining the accuracy of heat evaluation. This method calculates the heat of logical intervals rather than individual data blocks, greatly reducing the storage space occupied by metadata; the data temperature grading results generated based on the logical interval heat evaluation results are more stable, reducing the frequent fluctuations of data between different temperature levels, and providing data The proposed method provides a reliable basis for the formulation of data placement strategies; the formulated data placement strategies fully utilize the characteristics of optical storage media suitable for storing cold data and magnetic storage media suitable for storing hot data, thereby realizing efficient utilization of storage resources; during the execution of data migration operations, priority sorting and batch migration processing are used to reduce system resource consumption and the impact of migration operations on normal business. For the case where hot data is cooled to cold data, a delayed write strategy is adopted to avoid unnecessary migration caused by short-term fluctuations in data temperature; the entire method forms a closed-loop data management mechanism, which can adaptively adjust the data storage location according to changes in data access patterns, realize the precise separation and efficient exchange of hot and cold data, and meet the multiple requirements of storage performance, capacity and cost in the big data environment.
[0115] above Figure 2The hot and cold data exchange system based on optical storage in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The hot and cold data exchange device based on optical storage in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0116] Figure 3 This is a schematic diagram of the structure of a hot and cold data exchange device based on optical storage, provided in an embodiment of the present invention. This hot and cold data exchange device 300 based on optical storage may vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors), memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) storing applications 333 or data 332. The memory 320 and storage media 330 may be either ephemeral or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each of which may include a series of instructions and operations within the hot and cold data exchange device 300 based on optical storage. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, executing the series of instructions and operations stored in the storage medium 330 on the hot and cold data exchange device 300 based on optical storage, thereby implementing the steps of the above-described method for hot and cold data exchange based on optical storage.
[0117] The optical storage-based hot and cold data exchange device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The illustrated structure of the optical storage-based hot and cold data exchange device does not limit the optical storage-based hot and cold data exchange device provided by the present invention and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0118] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to execute the steps of the optical storage-based hot and cold data exchange method.
[0119] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0120] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a hot and cold data exchange device based on optical storage (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0121] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for exchanging hot and cold data based on optical storage, characterized in that: The method comprises: Access characteristic information of data in the optical-magnetic hybrid storage system is collected to obtain a data access characteristic information set; The access life cycle factor is calculated according to the data access feature information set, the data block is divided into multiple logical intervals and the heat value of the logical interval is calculated to obtain the logical interval heat evaluation result, including: extracting the most recent access timestamp, access frequency, creation time and migration times of the data block from the data access feature information set to obtain the data block characteristic parameters; performing a difference calculation between the most recent access timestamp and the current timestamp in the data block characteristic parameters, and dividing it by the expected life cycle of the data block to obtain the time locality value; performing a ratio calculation between the access frequency in the data block characteristic parameters and the maximum access frequency observed in the system to obtain the frequency normalization value; performing a ratio calculation between the migration timestamp in the data block characteristic parameters and the maximum access frequency observed in the system to obtain the frequency normalization value; performing a ratio calculation between the migration timestamp in the data block characteristic parameters and the maximum access frequency observed in the system to obtain the frequency normalization value. The number of migrations is calculated by comparing it with the maximum migration number threshold set by the system, and the ratio is subtracted from 1 to obtain the migration cost value; quantum statistical entropy analysis is performed on the data block to calculate the quantum entropy value of the data access pattern, quantify the degree of uncertainty and chaos of the data access behavior, identify data with deterministic access rules and random access characteristics, and obtain an access entropy weight factor; based on the creation time, application and data type of the data block, data blocks with similar attributes are divided into the same logical interval, and the comprehensive heat value of each logical interval is calculated by combining the time locality value, the frequency normalization value, the migration cost value and the access entropy weight factor through a weighted combination method to obtain a logical interval heat evaluation result; generating a data temperature classification result based on the heat evaluation result of the logical interval; formulating a data placement strategy based on the data temperature classification result, allocating hot data to the magnetic storage layer and cold data to the optical storage layer, and generating a data placement mapping table; Execute data migration operations according to the data placement mapping table, process migration tasks according to priority, perform batch data migration, and update data location mapping records.
2. The method for exchanging hot and cold data based on optical storage according to claim 1, characterized in that: The data access characteristic information collection of the data in the optical-magnetic hybrid storage system to obtain a data access characteristic information set includes: Recording access operations to each data block in the optical-magnetic hybrid storage system in real time to obtain original access log data; Preprocessing the original access log data to obtain standardized access log data; Performing aggregate analysis on the standardized access log data at preset time intervals, calculating the access frequency of each data block in different time periods, and obtaining time series access frequency data; Performing correlation analysis on the time series access frequency data and the data block creation time, calculating the life cycle characteristic value of the data block, and obtaining data life cycle characteristic data; Performing statistical analysis on the historical migration records of the data blocks, recording the number of migrations and reasons of the data blocks between different storage layers, and obtaining historical data migration data; The standardized access log data, the time series access frequency data, the data life cycle characteristic data and the data migration history data are integrated to establish an index structure to obtain a data access characteristic information set.
3. The method for exchanging hot and cold data based on optical storage according to claim 1, characterized in that: Generating a data temperature classification result based on the logic interval heat evaluation result includes: An initial temperature classification threshold is set for the heat evaluation result of the logical interval, and a logical interval with a heat value higher than a first threshold is classified as hot data, a logical interval with a heat value between a second threshold and the first threshold is classified as warm data, and a logical interval with a heat value lower than the second threshold is classified as cold data, thereby obtaining an initial data temperature classification result; Performing a data volume balance analysis on the initial data temperature classification results, calculating the data volume proportion of each temperature level, and obtaining the data distribution ratio; Adjust the temperature classification threshold according to the difference between the data distribution ratio and the target distribution ratio preset by the system. When the proportion of hot data exceeds the target value, increase the first threshold. When the proportion of cold data is lower than the target value, lower the second threshold. Then, a revised temperature classification threshold is obtained. Analyze the historical trend of the heat evaluation results of the logical intervals, identify the boundary intervals where the heat value fluctuates frequently, set state transition hysteresis periods for these intervals, and obtain temperature stabilization control parameters; Inputting the logic interval heat evaluation result, the modified temperature classification threshold and the temperature stability control parameter into a temperature classification processor, performing temperature classification calculation, and obtaining an optimized data temperature classification result; A storage load impact analysis is performed on the optimized data temperature grading results. Taking into account the current capacity utilization of each storage layer, more boundary data is classified as cold data when the capacity of the magnetic storage layer is close to saturation, and some boundary data is classified as warm data when the read and write bandwidth of the optical storage layer is close to saturation, to obtain the data temperature grading results.
4. The method for exchanging hot and cold data based on optical storage according to claim 1, characterized in that: The step of formulating a data placement strategy based on the data temperature classification result, allocating hot data to the magnetic storage layer and cold data to the optical storage layer, and generating a data placement mapping table includes: Performing storage demand analysis on the hot data, warm data, and cold data in the data temperature classification results, calculating the access delay requirements, read-write ratio, and data size of the data at each temperature level, and obtaining storage demand parameters; Partitioning the magnetic storage layer according to the storage demand parameters, dividing the magnetic storage layer into a high-performance area and a standard area, and dividing the optical storage layer into an active area and an archive area, to obtain a storage layer partition structure; Performing correlation analysis on the data in the temperature classification results, constructing a data correlation network diagram, and obtaining data correlation group information; Based on the storage layer partition structure and the data association group information, hot data is allocated to the high-performance area of the magnetic storage layer, warm data is allocated to the standard area of the magnetic storage layer, cold data with an access frequency higher than a preset threshold is allocated to the active area of the optical storage layer, and cold data that has not been accessed for a long time is allocated to the archive area of the optical storage layer. At the same time, it is ensured that data groups with high association are allocated to similar physical locations, thereby obtaining a preliminary data placement plan; Performing load balancing optimization on the preliminary data placement plan, utilizing the parallel access characteristics of the optical-magnetic hybrid storage system to evenly distribute data to each physical device, avoiding excessive load at a single point, and obtaining a balanced data placement plan; The balanced data placement scheme is converted into a data location mapping table, and the physical location information of the data of each logical interval in the storage system is recorded, including the storage layer type, device identification and physical address, to obtain a data placement mapping table.
5. The method for exchanging hot and cold data based on optical storage according to claim 1, characterized in that: The performing of data migration operations according to the data placement mapping table, processing migration tasks according to priority, performing batch data migration, and updating data location mapping records include: Comparing the current location of the data with the target location according to the data placement mapping table, identifying the data that needs to be migrated, and generating a migration task list; Performing priority analysis on the tasks in the migration task list to obtain a migration task queue with priority; Pre-evaluate the prioritized migration task queue, calculate the ratio of resource consumption to expected benefit of each migration task, select tasks whose benefits outweigh their costs, and obtain an optimized migration task set; Merging small similar tasks in the optimized migration task set into batch migration tasks, formulating a batch execution plan, and arranging execution during a time period with low system load to obtain a batch migration execution plan; Execute data migration according to the batch migration execution plan. For the migration of hot data to cold data, adopt a delayed write strategy to mark the data but not move it temporarily until sufficient resources are available or the data has not been accessed for a long time, and then perform the actual migration to obtain the migration execution result; The data location mapping record is updated based on the migration execution result, recording new data physical location information and migration history information, including migration time, source location, target location and migration reason, to obtain an updated data location mapping record.
6. The method for exchanging hot and cold data based on optical storage according to claim 5, characterized in that: The step of comparing the current location of the data with the target location according to the data placement mapping table, identifying the data to be migrated, and generating a migration task list includes: Parsing the data placement mapping table, extracting current storage location information and target storage location information of each logical interval, and obtaining a location information comparison table; Performing storage layer location difference analysis based on the location information comparison table, identifying logical intervals where storage layers have changed, marking these intervals as objects to be migrated, and obtaining a preliminary migration object set; Performing data volume statistics on each logical interval in the preliminary migration object set, calculating the data size, number of data blocks, and estimated migration time involved in the migration, and obtaining migration resource requirement data; Detecting changes in continuous storage space based on the location information comparison table, combining data that is physically continuous but will be dispersed after migration into an overall migration task, and combining data that is physically dispersed but will be integrated after migration into an overall migration task, to obtain a migration task combination plan; Integrating the migration resource demand data with the migration task combination plan, grouping the logical intervals according to the migration direction, generating a cooling migration task set from the magnetic storage layer to the optical storage layer and a heating migration task set from the optical storage layer to the magnetic storage layer, and obtaining a directional migration task set; The directional migration task set is converted into a migration task description in a standard format, including a task identifier, a data interval range, a source location, a target location, a data size, and a migration type, and a migration task list is generated.
7. A hot and cold data exchange system based on optical storage, characterized in that: For implementing the hot and cold data exchange method based on optical storage according to any one of claims 1 to 6, the hot and cold data exchange system based on optical storage comprises: An acquisition module is used to acquire access characteristic information of data in the optical-magnetic hybrid storage system to obtain a data access characteristic information set; The access module is used to calculate the access life cycle factor according to the data access feature information set, divide the data block into multiple logical intervals and calculate the heat value of the logical interval to obtain the logical interval heat evaluation result, including: extracting the most recent access timestamp, access frequency, creation time and migration times of the data block from the data access feature information set to obtain the data block characteristic parameters; performing difference calculation on the most recent access timestamp and the current timestamp in the data block characteristic parameters, and dividing it by the expected life cycle of the data block to obtain the time locality value; performing ratio calculation on the access frequency in the data block characteristic parameters and the maximum access frequency observed in the system to obtain the frequency normalization value; performing ratio calculation on the access frequency in the data block characteristic parameters and the maximum access frequency observed in the system to obtain the frequency normalization value; performing ratio calculation on the access frequency in the data block characteristic parameters and the maximum access frequency observed in the system to obtain the frequency normalization value; performing ratio calculation on the access frequency in the data block characteristic parameters and the maximum access frequency observed in the system to obtain the frequency normalization value. The migration times of the data blocks are calculated by comparing them with the maximum migration times threshold set by the system, and the ratio is subtracted from 1 to obtain the migration cost value; quantum statistical entropy analysis is performed on the data blocks to calculate the quantum entropy value of the data access pattern, quantify the degree of uncertainty and chaos of the data access behavior, identify data with deterministic access rules and random access characteristics, and obtain the access entropy weight factor; based on the creation time, application and data type of the data blocks, data blocks with similar attributes are divided into the same logical interval, and the comprehensive heat value of each logical interval is calculated by weighted combination in combination with the time locality value, the frequency normalization value, the migration cost value and the access entropy weight factor to obtain the logical interval heat evaluation result; A generating module, configured to generate a data temperature classification result based on the logic interval heat evaluation result; an allocation module, configured to formulate a data placement strategy based on the data temperature classification result, allocate hot data to the magnetic storage layer, allocate cold data to the optical storage layer, and generate a data placement mapping table; The execution module is used to execute data migration operations according to the data placement mapping table, process migration tasks according to priority, perform batch data migration, and update data location mapping records.
8. A hot and cold data exchange device based on optical storage, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the method for exchanging hot and cold data based on optical storage according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is enabled to perform the hot and cold data exchange method based on optical storage according to any one of claims 1 to 6.
Citation Information
Patent Citations
Dynamic data distribution storage method and system for RAID (Redundant Array of Independent Disks)
CN119668506A
Data management method and system based on optomagnetic fusion storage
CN120233956A