Cold and hot data object life cycle feature extraction method based on deep learning
Through deep learning-based methods, extract the life cycle characteristics of data objects, predict their future popularity changes, and optimize data storage strategies based on this information, solving the technical problems of data popularity identification and storage strategy optimization, and achieving efficient and economical data storage management.
Patent Information
- Application Number
- CN202510220185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
How to use the evolutionary law of historical data popularity to establish an effective model to predict future data access trends, thereby guiding the optimal configuration of data storage and solving the technical difficulties faced by data popularity identification and storage strategy optimization.
The life cycle feature extraction method of hot and cold data objects based on deep learning is adopted. By obtaining the access frequency, access interval, access time distribution of the data objects, the feature vector describing data viscosity is constructed, and data clusters of different levels of viscosity are divided using clustering algorithms, their popularity types and life cycle stages are analyzed, and future heat change trends are predicted, and the optimal storage location and migration method are determined based on this information.
It realizes accurate identification of data popularity and dynamic storage strategy optimization, improves the performance and cost efficiency of the storage system, and ensures data security and fast accessibility.
Smart Images

Figure CN120216971A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a method for extracting lifecycle features of hot and cold data objects based on deep learning. Background Art
[0002] In the complex scenario of contemporary big data storage management, data access patterns show a high degree of heterogeneity and dynamism. The high frequency of hot data requests directly affects the response efficiency of the system, while cold data, although sparsely accessed, occupies a considerable proportion of storage resources. How to achieve in-depth analysis of the time locality and frequency locality of data access patterns to instantly identify the data heat state and dynamically adjust the data storage location and storage medium selection accordingly is a major technical challenge currently faced. Data heat is not static and unchanged, it fluctuates over time, hot data can turn cold, and cold data can also regain attention due to business changes, so accurate tracking of heat changes becomes the core. The introduction of the concept of data stickiness, that is, defining its value level by evaluating the degree of data reuse, provides a new perspective for distinguishing data activity. Data stickiness refers to the degree to which data is repeatedly accessed and used by users in the storage system. By analyzing data stickiness, the temperature of different data objects can be determined. Data with high data stickiness is usually frequently accessed and used, and is regarded as a hot object. Their frequent access patterns indicate high activity and value. Based on this, adopting differentiated storage strategies, ensuring that hot data is stored in high-speed media, cold data is placed in low-speed media with better cost-effectiveness, and realizing intelligent migration of data between media at all levels are key measures to improve system performance. However, how to use the evolution law of historical data popularity to establish an effective model to predict future data access trends, thereby further guiding the optimal configuration of data storage, is still a technical problem that needs to be overcome. Summary of the invention
[0003] The present invention provides a method for extracting life cycle features of hot and cold data objects based on deep learning, which mainly includes:
[0004] Obtain the access frequency, access interval, access time distribution, and number of repeated accesses of data objects as access pattern features, as well as data creation time, last access time, and access times as lifecycle features, and construct a feature vector that describes data stickiness;
[0005] A clustering algorithm is used to cluster the feature vectors. According to the clustering results, the data is divided into data clusters with different stickiness levels. The data in each stickiness level data cluster has similar access patterns and life cycle characteristics.
[0006] For each stickiness level data cluster, analyze its access time distribution and life cycle stage to determine its popularity type, including continuous popularity, periodic popularity, temporary popularity, and cold popularity.
[0007] Predict the heat change trend of the sticky level data cluster in the next period of time according to its current heat type and historical access pattern, and obtain the heat change curve of each cluster;
[0008] Combine the sticky level, heat type and heat change trend of the sticky level data cluster to determine its optimal storage location and migration method. Store the hot data with high stickiness on the hot storage device with fast access speed and high cost, and migrate the cold data with low stickiness to the cold storage device with slow access speed and low cost;
[0009] For data clusters with different sticky levels and heat types, adopt corresponding caching and prefetching strategies, including increasing the cache allocation and prefetching frequency for continuously hot data with high stickiness, prefetching in advance during the heat rise period for periodic hot data, dynamically adjusting the cache space for temporary hot data, and reducing the cache allocation and lowering the prefetching frequency for cold data;
[0010] By real-time monitoring the access situation and heat change of the data, dynamically adjust the migration and replication methods of the data between different storage layers. When the actual access pattern of the data deviates from the prediction, trigger data reclustering and storage optimization;
[0011] For data with different life cycle stages and business values, perform corresponding data protection and backup, including using a short backup cycle and a low recovery time method for data protection and backup for data in the active period with high business value, and reducing the backup frequency and extending the recovery time to save backup storage space and cost for data in the archival period with low business value.
[0012] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:
[0013] The present invention discloses a method for extracting the life cycle characteristics of hot and cold data objects based on deep learning. For each data cluster with a certain stickiness level, analyze its access time distribution and life cycle stage, and determine its heat type, including continuous heat, periodic heat, temporary heat, and cold. Based on the current heat type and historical access patterns of the data cluster, predict its heat change trend in the next period of time, so as to obtain the heat change curve of each cluster. Combining the stickiness level, heat type, and heat change trend of the data cluster with stickiness level, determine its optimal storage location and migration strategy in the storage system: store high-stickiness hot data on hot storage devices with fast access speed and high cost, and migrate low-stickiness cold data to cold storage devices with slow access speed and low cost; for data clusters with different stickiness levels and heat types, adopt corresponding caching and prefetching strategies; by real-time monitoring the access situation and heat change of the data, dynamically adjust the migration and replication methods of the data between different storage layers. When the actual access pattern of the data deviates from the prediction, trigger data reclustering and storage optimization; for data in different life cycle stages and with different business values, perform corresponding data protection and backup. Generally speaking, the present invention intelligently adjusts the data storage location and processing strategy by comprehensively analyzing the access pattern and life cycle characteristics of the data, effectively improving the performance and cost efficiency of the storage system, while ensuring the security and fast accessibility of the data. This method is particularly suitable for modern enterprises and cloud storage environments with large amounts of data and complex access patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flowchart of a method for extracting the life cycle characteristics of hot and cold data objects based on deep learning according to the present invention.
[0015] Figure 2 It is a schematic diagram of a method for extracting the life cycle characteristics of hot and cold data objects based on deep learning according to the present invention.
[0016] Figure 3 It is another schematic diagram of a method for extracting the life cycle characteristics of hot and cold data objects based on deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Such as Figures 1-3 , a method for extracting the life cycle characteristics of hot and cold data objects based on deep learning in this embodiment may specifically include:
[0019] S101. Obtain the access frequency, access interval, access time distribution of the data object, the number of repeated accesses as the access pattern feature, and the data creation time, most recent access time, and access count as the life cycle features, and construct a feature vector describing data stickiness.
[0020] Obtain the access log records of the data object, and extract the access frequency, access interval, and access time distribution attribute information. Count the number of repeated accesses of the data object and use it as an important indicator to measure the access pattern feature. Obtain the life cycle attributes of the data object, such as the creation time, most recent access time, and cumulative access count. According to the access pattern feature and life cycle attributes, construct a feature vector for describing data stickiness.
[0021] Specifically, to obtain the access log records of the data object, an access log collection module can be deployed in the data storage system to record information such as the timestamp, visitor IP, and access type of each access in real time. Through statistical analysis of the access log, it can be obtained that the average access frequency of the data object is 10 times per day, the average access interval is 2 hours, and the peak access period is concentrated from 10 am to 3 pm. At the same time, by counting the number of repeated accesses of the data object, it is found that 20% of the data objects are repeatedly accessed more than 50 times, indicating that these data objects have high access stickiness. When obtaining the life cycle attributes of the data object, information such as the creation time, most recent access time, and cumulative access count of the object can be extracted from the meta-database and constructed into a 50-dimensional feature vector together with the access pattern feature.
[0022] S102. Use a clustering algorithm to cluster the feature vector, and divide the data into different data clusters of stickiness levels according to the clustering result. The data within each data cluster of stickiness level has similar access patterns and life cycle features.
[0023] Use the feature vector to represent the access pattern and life cycle features of the data object, and judge the similarity degree between different data objects through the similarity measurement of the feature vector. Select a clustering algorithm for clustering according to the dimension and numerical distribution of the feature vector; judge the quality of the clustering result by calculating the tightness within the cluster and the separation degree between clusters. Among them, the tightness within the cluster is measured by the average distance, and the separation degree between clusters is measured by the inter-cluster distance. According to the clustering result, divide the data objects into different data clusters of stickiness levels, where the data objects within each data cluster have similarity.
[0024] Specifically, the access patterns and lifecycle characteristics of data objects are represented by feature vectors. By measuring the similarity of feature vectors, the similarity degree between different data objects is judged, providing a data basis for the clustering analysis algorithm. According to the dimension and numerical distribution of the feature vectors, a suitable clustering algorithm is selected, such as K-means clustering, hierarchical clustering, or density-based DBSCAN clustering, etc. By setting appropriate clustering parameters, the clustering result of the data is obtained. If the feature vector is high-dimensional and sparse, the density-based DBSCAN clustering algorithm is combined; if the feature vector is low-dimensional and dense, the K-means clustering algorithm is combined. The clustering result is evaluated and optimized. By calculating the tightness within the cluster and the separation between clusters, the quality of the clustering result is judged. The tightness within the cluster can be measured by the average distance, that is, calculating the average distance between each data object and the center of its belonging cluster. The smaller the distance, the higher the similarity of the data objects within the cluster, and the better the clustering effect. The separation between clusters can be measured by the inter-cluster distance, that is, calculating the distance between the centers of different clusters. The larger the distance, the greater the difference between different clusters, and the better the clustering effect. The Silhouette Coefficient can be used to comprehensively evaluate the clustering effect. The Silhouette Coefficient takes into account the similarity of the data object to its belonging cluster and the dissimilarity to other clusters. The value range of the Silhouette Coefficient is [-1, 1], and the larger the value, the better the clustering effect. According to the characteristics of the clustering algorithm, other evaluation metrics can also be used, such as the DBI index, Calinski-Harabasz index, etc., to evaluate the quality of the clustering result. If the clustering result is not satisfactory, the parameters of the clustering algorithm are adjusted or other more suitable clustering algorithms are selected until a satisfactory clustering effect is obtained. According to the clustering result, the data objects are divided into data clusters of different stickiness levels. The data objects within each data cluster have a certain degree of similarity, but there may be certain differences. By analyzing the characteristics of each data cluster, its belonging stickiness level is determined. Each data cluster of the stickiness level is analyzed in detail. By calculating statistical metrics such as the access frequency, access interval, and lifecycle length within the data cluster, the behavioral characteristics of different data clusters of the stickiness level are characterized. The access frequency can be obtained by calculating the number of accesses of each data object within a certain time window. The access interval can be obtained by calculating the average or median of the time difference between two consecutive accesses of each data object. The lifecycle length can be obtained by calculating the time span between the first access time and the last access time of the data object. For data clusters of different stickiness levels, the differences in their average lifecycle lengths are compared. According to the characteristic differences of the data clusters of the stickiness level, targeted data management strategies are formulated. For data clusters of high stickiness levels, they are stored on high-performance storage devices, such as SSDs or in-memory databases, to provide fast data access and retrieval capabilities.Meanwhile, these data can be backed up and disaster recovery processed more frequently to ensure data reliability and availability. For data clusters with a low stickiness level, they can be stored on low-cost storage devices such as SATA hard drives or tape libraries to save storage costs. For data that is rarely accessed or not accessed for a long time, data archiving or cold storage can be considered, and it can be moved to a cheaper offline storage medium. For data clusters with a high stickiness level, prefetching and caching mechanisms can be adopted to load frequently accessed data into memory or cache in advance to speed up data reading. For data clusters with a low stickiness level, lazy loading or on-demand reading methods can be used to reduce unnecessary data reading and processing overhead. At the same time, different data backup and recovery strategies, data security and access control strategies, etc. can also be set according to the characteristics of data clusters with different stickiness levels to meet the needs of different data management. Regularly re-cluster the data and divide the stickiness levels. According to the dynamic changes of the data, adjust the clustering results and stickiness levels of the data in a timely manner. This step can be used as a separate regular maintenance and update step, separated from the specific data management strategy. Through regular data re-clustering and stickiness level division, ensure that the data management strategy adapts to the actual access patterns and life cycle characteristics of the data, and improve the accuracy and effectiveness of data management. When constructing the data feature vector, indicators such as access frequency, the mean and variance of access intervals, and the entropy value of the access time distribution can be selected as features. For a data object, it is statistically obtained that its access times within a week are 100 times, the average access interval is 2 hours, the variance of the access interval is 0.8, and the entropy value of the access time distribution is 1.5. Combine these indicators into a feature vector [100, 2, 0.8, 1.5] to represent the access pattern features of this data object. At the same time, the life cycle length of the data object can also be used as a dimension of the feature vector, such as 30 days. By calculating the Euclidean distance or cosine similarity between different data objects, the similarity measure between them can be obtained. When selecting a clustering algorithm, factors such as data scale, feature dimension, and clustering shape can be considered. For large-scale data sets, the partition-based K-means clustering algorithm can be selected to divide the data into K clusters through iterative optimization. For high-dimensional data, the density-based DBSCAN clustering algorithm can be selected to discover clusters of any shape by setting density thresholds and minimum points. When determining the clustering parameters, the elbow method or silhouette coefficient can be used to select the optimal number of clusters K. By calculating the sum of squared clustering errors for different K values, it is found that when K = 5, the rate of decrease in the sum of squared clustering errors slows down, forming an elbow point, then K = 5 can be selected as the optimal number of clusters. When evaluating the clustering results, the average distance between each data object and the center of its cluster can be calculated. If the average distance is small, such as 0.5, it means that the tightness of the data objects within the cluster is high.Meanwhile, the minimum distance between different cluster centers can also be calculated. If the minimum distance is large, such as 2.0, it indicates that the separation degree between different clusters is good. For the silhouette coefficient, the average distance a between each data object and other objects within its cluster can be calculated, as well as the average distance b to the objects in the nearest cluster. Then, (b - a) / max(a, b) is calculated to obtain the silhouette coefficient of the object. The average value of the silhouette coefficients of all data objects is taken to get the silhouette coefficient of the entire clustering result. If this value is greater than 0.5, it indicates a good clustering effect. When dividing the stickiness levels, different stickiness level thresholds can be set according to the characteristics of each data cluster, such as access frequency, access interval, and life cycle length. Data clusters with an average access frequency greater than 100 times / week, an average access interval less than 1 day, and an average life cycle length greater than 30 days are defined as high stickiness level. Data clusters with an average access frequency less than 10 times / week, an average access interval greater than 7 days, and an average life cycle length less than 7 days are defined as low stickiness level. The remaining data clusters are defined as medium stickiness level. Different storage hierarchies and access mechanisms are adopted for data management. For data clusters with high stickiness level, they can be stored in SSD or in-memory databases, and the LRU cache policy is adopted to cache the recently accessed data in memory to improve the data reading speed. For data clusters with low stickiness level, they can be stored on SATA hard disks, and the lazy loading policy is adopted to load the data from the disk to memory only when needed, reducing unnecessary I / O overhead. Meanwhile, different backup strategies can also be set according to the access frequency and timeliness of the data. For example, for data with high stickiness level, incremental backup is performed daily and full backup is performed weekly; for data with low stickiness level, full backup is performed once a month to reduce the backup cost. When performing regular data reclustering, the sliding window method can be adopted. Every certain period, such as one month, the feature vectors of data objects are recalculated and compared with the previous clustering results. If it is found that the feature vectors of some data objects have changed significantly, such as a sudden decrease in access frequency, they are removed from the original cluster and reclustered. Meanwhile, the number and boundaries of clusters can be dynamically adjusted by setting the upper and lower limits of the number of data objects within a cluster to make it adapt to the data change trend. When the number of data objects in a cluster exceeds 100, it is considered to be split into two sub-clusters; when the number of data objects in a cluster is less than 10, it is considered to be merged into the most similar cluster. Through regular data reclustering, changes in data access patterns can be detected in a timely manner, data management strategies can be adjusted, and the adaptability and performance of the system can be improved.
[0025] S103. For each data cluster with a certain stickiness level, analyze its access time distribution and life cycle stage to determine its heat type, including continuous heat, periodic heat, temporary heat, and cold.
[0026] Statistically analyze the access time distribution of each data cluster with different stickiness levels, obtain the time series of access counts at different time granularities, and form the access count distribution curves at multiple time scales. Based on the smoothed access count distribution curves, judge the access patterns and popularity characteristics of the data clusters; by analyzing the number of peaks, peak intervals, and slope change morphological characteristics of the curves, determine whether the data clusters have periodic access rules, whether there are long-duration popularity persistence periods or sudden popularity; combined with the information on the life cycle stages of the data clusters, and considering factors such as the continuous access duration, access frequency, and access stability of the data clusters, judge the popularity types of the data clusters and classify them into different popularity categories such as continuously popular, periodically popular, temporarily popular, and cold.
[0027] Specifically, perform statistical analysis on the access time distribution of each data cluster with a certain stickiness level to obtain the time series of access counts at different time granularities, such as aggregating by hour, day, week, etc., to form the access count distribution curves at multiple time scales. Smooth the time series of access counts to remove accidental fluctuations and outliers, and extract the overall change trend of access counts. Methods such as simple moving average (SMA), weighted moving average (WMA), or exponential smoothing (EMA) can be used to smooth the curve. Among them, SMA is the arithmetic mean, WMA gives different weights according to the distance of data points, and EMA assigns weights according to the exponential decay function. The size of the smoothing window determines the degree of smoothing, and it needs to be selected and adjusted according to the characteristics of the data and the analysis requirements. According to the smoothed access count distribution curve, judge the access pattern and popularity characteristics of the data cluster. By analyzing the morphological characteristics of the curve, such as the number of peaks, peak intervals, slope changes, etc., determine whether the data cluster has a periodic access pattern, whether there are obvious periods of continuous popularity or sudden popularity. Combine the information on the life cycle stage of the data cluster, and comprehensively consider factors such as the continuous access duration, access frequency, and access stability of the data cluster to judge and classify the popularity type of the data cluster, and divide it into different popularity categories such as continuously hot, periodically hot, temporarily hot, cold, etc. For the data clusters determined to be continuously hot, set the access count threshold and the continuous duration threshold to screen out the continuously hot data clusters that meet the conditions; for the data clusters determined to be periodically hot, by analyzing the autocorrelation coefficient and partial autocorrelation coefficient of the access count time series, judge whether there is a significant periodicity, and the period length can be estimated by the lag order corresponding to the peak; for the data clusters determined to be temporarily hot, set the threshold for the sudden increase in access count and the upper limit of the continuous duration to determine the temporarily hot event and identify the temporarily hot event. For the data clusters determined to be continuously hot, their access counts are stable and continuously at a high level. The access count threshold and the continuous duration threshold can be set to screen out the continuously hot data clusters that meet the conditions. For example, if the access count of a data cluster exceeds 1000 times per day for 30 consecutive days, it can be determined as a continuously hot data cluster. Analyze the formation reasons and application value in combination with the business scenario, and formulate targeted data management. For the data clusters determined to be periodically hot, their access counts show an obvious periodic pattern. By analyzing the autocorrelation coefficient and partial autocorrelation coefficient of the access count time series, judge whether there is a significant periodicity. If the autocorrelation coefficient at a certain time lag shows a sharp peak, and the partial autocorrelation coefficient decays rapidly after this lag order, there may be a periodicity, and the period length can be estimated by the lag order corresponding to the peak. Use frequency domain analysis methods such as discrete Fourier transform (DFT) or fast Fourier transform (FFT) to extract the periodic components of the access count time series and explore the business rules and driving factors behind the periodic access behavior of the data cluster. For the data clusters determined to be temporarily hot, their access counts suddenly increase in a short period of time and then quickly decrease.It can be determined by setting the threshold for sudden increase in access times and the upper limit of the duration. If the daily access times of a data cluster suddenly increase to more than 10 times the average within 3 consecutive days and quickly return to the normal level within the next 5 days, it can be determined as a temporary popularity event. An anomaly detection method based on thresholds is adopted, such as marking data points outside the range of mean plus or minus three standard deviations as anomaly points to identify temporary popularity events. For the confirmed temporary popularity events, analyze their formation reasons, evaluate their impacts on system performance and resource utilization, and take corresponding measures such as caching and traffic limiting. For data clusters determined to be cold data clusters, it can be determined by setting the lower limit of access frequency and the upper limit of access interval. If the weekly access frequency of a data cluster is less than 10 times and the average access interval exceeds 30 days, it can be determined as a cold data cluster. To statistically analyze the access time distribution of data clusters at each stickiness level, the resample() function of the Pandas library can be used to resample and aggregate the access log data, such as aggregating df by hour. The aggregated data can be plotted into a distribution curve of access times at multiple time scales using the Matplotlib library. For the smoothing of access times, functions such as SMA, WMA, and EMA provided by the talib library can be used, and different window sizes such as 5, 10, 20, etc. can be set. The smoothed curve can more clearly reflect the overall trend of access times. The peaks and their attributes of the curve are detected using the find_peaks() function of the signal library, and the number of peaks, peak intervals, changes in curve slope, etc. are judged to further judge the periodic characteristics of the data cluster. If it is detected that a significant peak appears every 7 days on the curve and the slope between the peaks is gentle, it can be initially determined that the data cluster has a periodic pattern of weekly access peaks. Further, the seasonal_decompose() function of the statsmodels library can be used to perform seasonal decomposition on the access time series of access times to obtain the trend term, seasonal term, and residual term. If the seasonal term fluctuates significantly and regularly, the periodic characteristics are further verified. Combining the analysis of the business scenario, it is found that the access times of a certain data cluster will surge to more than 5 times the usual level every Friday from 6 pm to 10 pm, presumably related to the business peak after work, and it is recommended to focus on analyzing the access data during this period. For the determination of temporary popularity events, the mean and standard deviation of access times can be calculated first, and then the mean plus or minus three standard deviations are used as the threshold to mark the data points outside the threshold range. The average hourly access times of a certain data cluster is 100 and the standard deviation is 30, then the data points with access times exceeding 190 or less than 10 are marked as anomaly points. If the proportion of anomaly points in a certain time period exceeds 50% and the duration is within 24 hours, it can be determined as a temporary popularity event. By analyzing the access logs before and after the temporary popularity event, it is found that it may be related to a certain hot news event, resulting in a sudden increase in access volume in a short period of time.
[0028] S104. Predict the heat change trend of each sticky-level data cluster in the next period of time according to its current heat type and historical access pattern, and obtain the heat change curve of each cluster.
[0029] Obtain the historical access pattern data of each data cluster, where the historical access pattern data includes the distribution time series characteristics of the number of accesses, access intervals, and access times; by analyzing the historical access pattern data, extract the key features of the data cluster, where the key features include the average number of accesses, average access interval, and peak access time period; analyze whether the historical access pattern data has periodicity and trend, and select a time series prediction model according to the analysis results for training to obtain a heat change trend prediction model; use the optimized feature space and the heat change trend prediction model to predict the number of accesses of each data cluster in the next period of time to obtain the predicted heat change curve; perform trend analysis and outlier detection on the predicted heat change curve to judge the heat change trend of the data cluster in the next period of time and identify the abnormal fluctuation points or mutation points in the prediction curve.
[0030] Specifically, obtain the current popularity type and historical access pattern data for each data cluster with different stickiness levels, including time series features such as the number of accesses, access intervals, and access time distributions, and construct a feature vector for predicting the popularity change trend. Start from aspects such as time features, statistical features, trend features, periodic features, and adjacent features to extract key features that reflect the access behavior patterns. Use feature selection methods such as the filter method, wrapper method, and embedding method to select the most discriminative and predictive feature subset and optimize the feature space. According to the historical access pattern features of the data cluster, determine whether it has time series features such as periodicity and trend, and select a suitable time series prediction model, such as the ARIMA model, Prophet model, LSTM model, GRU model, Holt-Winters exponential smoothing model, etc. Obtain the popularity change trend prediction model through model training. Use methods such as the holdout method, cross-validation method, and sliding window method to evaluate the trained popularity change trend prediction model, and use evaluation metrics such as mean absolute error, mean squared error, mean absolute percentage error, and coefficient of determination to measure the prediction effect of the model. Use the trained popularity change trend prediction model to predict the number of accesses for each data cluster in the future for a period of time to obtain the predicted popularity change curve. Conduct trend analysis on the predicted popularity change curve. By calculating indicators such as the first-order difference or second-order difference, judge the popularity change trend of the data cluster in the future for a period of time, such as continuous increase, continuous decrease, or stable fluctuation, to provide a basis for the dynamic adjustment of data management strategies. Combine the results of outlier detection and trend analysis to conduct a comprehensive analysis of the predicted popularity change curve. By setting an outlier threshold and using methods such as the 3σ principle, quantile method, adaptive method, and business rules, identify the abnormal fluctuation points or mutation points in the predicted curve, analyze their possible causes, such as holiday effects and marketing activities, and evaluate their impact on prediction accuracy and strategy formulation. According to the comprehensive analysis results, dynamically adjust the popularity type and management strategy of the data cluster. For example, pre-classify the data clusters that are predicted to continuously increase and have no abnormalities as hot data for key management, and downgrade the data clusters that are predicted to continuously decrease and have no abnormalities to cold data to save resources. Regularly retrain and evaluate the popularity change trend prediction model, update and optimize the model using new historical access data to improve the accuracy and adaptability of the prediction. At the same time, monitor the prediction deviation and abnormal situations, and timely adjust the prediction model and parameter settings, such as feature selection and outlier threshold, and continuously iterate and optimize to ensure the reliability and effectiveness of the prediction results, providing reliable data support for the dynamic optimization of data management strategies. Through continuous model optimization and strategy adjustment, realize the real-time perception and prediction of the popularity change of data clusters, guide the intelligent and adaptive optimization of data management, and improve the performance and resource utilization efficiency of the system. When constructing the feature vector for predicting the popularity change trend, multi-dimensional time series features of the data cluster can be extracted.For a certain data cluster, the sequence of daily access times is statistically obtained as [1000, 1200, 1500, 1800, 2000], the sequence of access intervals is [20, 18, 16, 14, 12], and the sequence of access time distributions is as follows.
[0031] [0.2, 0.3, 0.5, 0.6, 0.8], where the access time distribution represents the proportion of the number of accesses in different time periods within a day. Align these sequences by time to obtain the feature vector at each time point. For example, [1000, 20, 0.2] represents the number of accesses, access interval, and access time distribution on the first day. Using filtering methods such as chi-square test and information gain, calculate the correlation between each feature and the popularity type, and select the top K features with the highest correlation. Then use wrapper methods such as recursive feature elimination to perform feature subset search to find the feature combination with the best prediction effect. When selecting a time series prediction model, it is possible to judge whether it has obvious periodicity and trend according to the historical access patterns of the data cluster. For a certain data cluster, by calculating the autocorrelation coefficient of its access number sequence, it is found that there is an obvious correlation peak at a lag of 7 days, indicating that it may have an access pattern with a weekly cycle; then by calculating the mean of the first-order difference sequence, if the mean is significantly not 0, it indicates that it may have an upward or downward trend. For data clusters with periodicity and trend, the SARIMA model can be selected for prediction. The SARIMA(p, d, q)(P, D, Q)_m model includes components such as seasonal autoregression (SAR), seasonal differencing (SI), and seasonal moving average (SMA), which can effectively capture the periodic and trend characteristics of the time series. Through methods such as grid search, optimize the hyperparameters of the model, such as p, d, q, P, D, Q, etc., and train to obtain the optimal SARIMA model. When evaluating the prediction effect, the method of sliding window cross-validation can be used. For example, set the size of the sliding window to 7 days and the step size to 1 day. Each time, use the data of the most recent 7 days as the test set, and the data before it as the training set, and repeat multiple times to obtain a series of prediction results. For each test set, calculate the error between the predicted value and the true value, such as the mean absolute percentage error (MAPE). If the average value of MAPE is less than 10%, it indicates that the prediction effect of the model is good. If the average value of MAPE is greater than 20%, it is necessary to further optimize the model or feature selection. When performing trend analysis on the prediction results, the first-order difference sequence of the prediction curve can be calculated to judge its change direction and rate. For the predicted access number sequence [1000, 1200, 1500, 1800, 2000] of a certain data cluster, calculate its first-order difference sequence [200, 300, 300, 200]. If the mean of the first-order difference sequence is significantly greater than 0, it indicates that the prediction curve shows an upward trend, and the upward rate is the mean of the first-order difference sequence. If the mean of the first-order difference sequence is significantly less than 0, it indicates that the prediction curve shows a downward trend. If the mean of the first-order difference sequence is close to 0, it indicates that the prediction curve is relatively stable. When performing outlier detection, the 3σ principle can be used to set the outlier threshold.For the prediction error sequence [-10%, 5%, 20%, -15%, 8%] of a certain data cluster, calculate its mean and standard deviation, obtaining a mean of 1.6% and a standard deviation of 13.4%. According to the 3σ principle, set the upper limit of the anomaly threshold to 41.8% and the lower limit to -38.6%. For each value in the prediction error sequence, if it exceeds the anomaly threshold range, it is marked as an anomaly point. For example, the point with a prediction error of 20% is marked as an anomaly point. Compare the anomaly points with the holiday calendar, marketing activity schedule, etc., and analyze the reasons for the anomalies. If the anomaly points are concentrated during holidays, it may be due to a sudden increase in traffic caused by the holiday effect. Dynamically adjust data management in combination with the prediction trend and anomalies. For data clusters with a continuously rising prediction and no anomalies, mark them as hot data in advance, and increase their storage and computing resource allocation. For example, store them in SSDs and increase the allocation ratio of CPU and memory to cope with future increases in traffic. For data clusters with a continuously declining prediction and no anomalies, gradually downgrade them to cold data and reduce their resource allocation. For example, store them in HDDs and lower the allocation ratio of CPU and memory to save system costs. For data clusters with an unclear prediction trend but anomalies, further analyze the reasons for the anomalies. If the anomalies are caused by a sudden increase in traffic due to marketing activities, temporarily adjust their resource allocation, such as increasing caching and load balancing, to cope with short-term traffic peaks. Continuously improve the accuracy and application effect of heat prediction through regular model optimization and strategy adjustment. Retrain the heat change trend prediction model every 30 days, using the actual access data of the most recent 30 days to update the model parameters and feature selection, enabling the model to learn the latest access patterns. Evaluate the prediction effect every 7 days, calculate metrics such as MAPE. If the MAPE of three consecutive evaluations exceeds 15%, trigger an early warning and conduct manual intervention and analysis. Dynamically adjust the anomaly threshold and data management strategy according to the evaluation results and business feedback. For example, adjust the anomaly threshold from 3σ to 2.5σ to increase the sensitivity of anomaly points.
[0032] S105. Combine the stickiness level, heat type, and heat change trend of the sticky-level data cluster to determine its optimal storage location and migration method. Store high-stickiness hot data on high-speed and high-cost hot storage devices, and migrate low-stickiness cold data to low-speed and low-cost cold storage devices.
[0033] Obtain the key attribute information of the stickiness level, heat type, and heat change trend of each data cluster with a certain stickiness level, which serves as the basis for determining its optimal storage location and migration method; combine factors such as storage capacity, storage unit price, data access volume, and access unit price to calculate the total cost under different storage schemes, and calculate the total benefits under different storage schemes to obtain the cost-benefit ratio; for data clusters with a stable heat type, predict the heat type in a future period according to its historical access pattern and heat change trend. If the prediction result is inconsistent with the current heat type, trigger data migration; for data clusters with a high frequency of heat type changes, monitor its heat metrics and access behaviors in real time. When it continuously exceeds the threshold range of the original heat type for multiple times, trigger the adjustment of the heat type and perform data migration.
[0034] Specifically, key attribute information such as the stickiness level, heat type, and heat change trend of each data cluster with a certain stickiness level is obtained, which serves as the basis for determining its optimal storage location and migration strategy. By designing quantization functions for each attribute, the attribute values are mapped to numerical values between 0 and 1, and different weight coefficients are assigned to calculate the comprehensive evaluation index of the data cluster. For example, the weight of the stickiness level attribute is 0.4, the weight of the heat type attribute is 0.3, and the weight of the heat change trend attribute is 0.3. For a data cluster with a high stickiness level, a hot heat type, and an upward heat change trend, its comprehensive evaluation index is 0.4×1 + 0.3×1 + 0.3×1 = 1. Different levels are divided according to the comprehensive evaluation index, and differentiated storage strategies are formulated. According to the comprehensive evaluation index of the data cluster, differentiated data storage and migration strategies are formulated. For high-stickiness hot data clusters with high index values, they are preferentially stored on hot storage devices with faster access speeds, such as solid-state drives or memory, etc., to improve the read / write performance and access experience of the data; for low-stickiness cold data clusters with low index values, they are considered to be migrated to cold storage devices with slower access speeds but lower storage costs, such as SATA hard drives, etc., to save storage resources and cost expenditures. When determining the optimal storage location of the data cluster, a multi-dimensional cost-benefit analysis method is adopted. Factors such as storage capacity, storage unit price, data access volume, and access unit price are comprehensively considered to calculate the total cost under different storage schemes. Factors such as data access performance, access latency, and user experience are comprehensively considered to calculate the total benefit under different storage schemes. For each storage scheme, according to the predicted access volume and access pattern of the data cluster, its total cost and total benefit are calculated to obtain the cost-benefit ratio. The storage scheme with the highest cost-benefit ratio is selected as the optimal storage location of the data cluster. The cost-benefit ratio is regularly re-evaluated, and the storage scheme is dynamically adjusted according to the actual access situation and cost changes of the data cluster. For data clusters with stable and changing heat types, different processing strategies are adopted. For data clusters with stable heat types, the heat type in the next period of time is predicted based on their historical access patterns and heat change trends. If the prediction result is consistent with the current heat type, the existing storage location remains unchanged; if the prediction result is inconsistent with the current heat type, such as changing from hot data to cold data, then data migration is triggered, and it is migrated from the hot storage device to the cold storage device. For data clusters with changing heat types, their heat indicators and access behaviors are monitored in real time. When they continuously exceed the threshold range of the original heat type for multiple times, automatic adjustment of the heat type is triggered. According to the new heat type, its optimal storage location is re-determined. If the new optimal storage location is inconsistent with the current location, the data migration process is started, and the data is migrated to the new optimal storage location. When performing data migration, incremental migration and batch migration strategies are adopted, that is, only the incremental data that has changed since the last migration is migrated, and at the same time, the data to be migrated is dispersed into multiple time windows to avoid having a greater impact on system performance and user access during the migration process.During the migration process, a data double-copy mechanism is adopted. For each data block to be migrated, a copy is first created on the target storage device, while the original data block is retained on the source storage device. In the double-copy state, read operations access the source storage device, and write operations update both the source storage device and the target storage device simultaneously, ensuring consistency through mechanisms such as two-phase commit. After the data synchronization on the target storage device is completed, the read / write access path is switched to the target storage device, and the source storage device serves as a backup. After the migration is completed, the backup data is deleted. Through the data copy mechanism and read / write redirection technology, the continuity and consistency of data access are ensured, and the impact of migration on the business is minimized. Regularly review and evaluate the data storage and migration strategies, and continuously optimize and improve the data storage and migration solutions in combination with the business development needs and technological development trends to enhance the performance, reliability, and cost-effectiveness of the system. When obtaining the attribute information of data clusters, a multi-dimensional attribute extraction method can be adopted. For example, for the stickiness level attribute, by analyzing the number of repeated accesses and the access time interval of users to the data cluster, if the average number of repeated accesses exceeds 10 times per month and the average access time interval is less than 7 days, it is determined to be highly sticky, and the quantization value is 1; if the average number of repeated accesses is 5-10 times per month and the average access time interval is 7-14 days, it is determined to be moderately sticky, and the quantization value is 0.6; in other cases, it is determined to be low sticky, and the quantization value is 0.3. For the heat type attribute, by analyzing the access frequency and access heat ranking of the data cluster, if the daily average access frequency exceeds 1000 times and the heat ranking is in the top 10%, it is determined to be hot data, and the quantization value is 1; if the daily average access frequency is 100-1000 times and the heat ranking is in the top 10%-30%, it is determined to be warm data, and the quantization value is 0.6; in other cases, it is determined to be cold data, and the quantization value is 0.3. For the heat change trend attribute, by analyzing the change in the access frequency of the data cluster in the most recent month, if it shows an upward trend for more than 3 consecutive days and the increase amplitude exceeds 50%, it is determined to be an upward trend, and the quantization value is 1; if it shows a downward trend for more than 3 consecutive days and the decrease amplitude exceeds 50%, it is determined to be a downward trend, and the quantization value is 0.3; in other cases, it is determined to be a stable trend, and the quantization value is 0.6. By synthesizing the quantization values and weights of the above three attributes, the comprehensive evaluation index of the data cluster can be obtained, which serves as the decision-making basis for data storage optimization. When conducting a cost-benefit analysis, a linear programming model can be used to solve the optimal storage solution.
[0035] S106. For data clusters with different stickiness levels and heat types, adopt corresponding caching and prefetching strategies, including for continuously hot data with high stickiness, increasing the cache allocation and prefetching frequency; for periodic hot data, prefetching in advance during its heat-up period; for temporary hot data, dynamically adjusting the cache space; and for cold data, reducing the cache allocation and lowering the prefetching frequency.
[0036] Obtain the stickiness level and popularity type information of each data cluster, and formulate a differentiated caching strategy; for the prefetching strategy of different data clusters, adopt a differentiated configuration that matches the caching strategy; record the cache hit rate and prefetch accuracy of each data cluster, and use them as the key evaluation metrics of the data cluster; periodically re-evaluate the popularity type of the data cluster to determine whether the popularity type of the data cluster has changed; if the popularity type of the data cluster has changed, adjust its caching strategy to achieve adaptive optimization of the cache.
[0037] Specifically, obtain the stickiness level and heat type information of each data cluster, and establish a data cluster attribute table. By analyzing the historical access logs of the data clusters, statistics such as access frequency, access interval, and number of repeated accesses are calculated to determine whether its stickiness level is high stickiness, medium stickiness, or low stickiness. At the same time, according to the time distribution law of the access logs, determine whether its heat type is continuously hot, periodically hot, temporarily hot, or cold data. Based on the stickiness level and heat type of the data clusters, formulate differentiated caching strategies. For continuously hot data clusters with high stickiness, use a larger cache space. For example, the cache space allocated for them is twice that of ordinary data clusters. At the same time, adopt a more aggressive cache update strategy, such as extending the cache expiration time to 2 days to ensure that hot data can continuously reside in the cache. For periodically hot data clusters, according to the periodic law of their historical accesses, predict the start time of the next heat up period, and start preloading their data into the cache 1 day before the heat up period arrives to reduce cache misses during the heat up period. For temporarily hot data clusters, monitor the change in their access volume in real time, calculate the ratio of the access volume in the current hour to the average access volume in the same hour in the past week. When this ratio exceeds 3 times continuously for 3 hours, trigger the dynamic expansion of the cache space, temporarily expanding its cache space to 2 times until the access peak ends and then releasing the expanded space. For cold data clusters, reduce the cache allocation amount. For example, adjust its cache space to 50% of that of ordinary data clusters. At the same time, increase the cache expiration speed to accelerate the elimination of cold data from the cache. For the prefetching strategies for different data clusters, adopt differentiated configurations that match the caching strategies. For continuously hot data clusters with high stickiness, adopt data prefetching 1 day in advance, that is, at 3 am every day, automatically preload the data blocks that may be accessed in the next 1 day into the cache to reduce the I / O latency during real-time access. For periodically hot data clusters, predict the heat curve in the next 7 days according to their access rules, and start the data prefetching task 1 day before each heat up period to cache their associated data in advance. For temporarily hot data clusters, start a short-term prefetching task within 1 hour after the sudden increase in their access volume to prefetch the data blocks accessed in the last 1 hour to cope with possible subsequent hot accesses. For cold data clusters, adopt a demand-loading strategy. When an access request arrives, first use a Bloom filter to determine whether the data block is in the cache. If not, load the data block from the disk. At the same time, according to the access frequency and temporal locality of the data block, decide whether to add it to the cache and its residence time in the cache. Record the cache hit rate and prefetch accuracy rate of each data cluster as the key evaluation indicators of the data cluster. At the same time, establish a global monitoring view of the cache resources, and real-time statistics of overall cache usage rate, number of prefetching tasks, cache hit rate, cache eviction rate and other indicators, and display the operating status and resource usage of the cache through a visual dashboard to timely detect and locate cache anomalies.For data clusters with a cache hit rate lower than 80% or a prefetch accuracy rate lower than 60%, trigger an alarm prompt to prompt relevant personnel to optimize their cache prefetch effect. The cache eviction method can also be optimized, and factors such as the access frequency, access time, and data size of the data should be comprehensively considered when evicting data. For each data block, record the timestamp T of its most recent access and the access count C, define a weight factor W and a decay factor D. For a data block with the current timestamp t, its comprehensive score S = C * W / (t - T + 1)^D, where T is the timestamp of the most recent access of the data block, and the weight factor W is used to adjust the impact of the access count on the score, that is, the higher the access count, the closer the access time, and the smaller the data block, the higher its score. When the cache space is insufficient, evict the data block with the lowest score until enough space is released. Regularly update the scores of all data blocks, and eliminate data blocks that have not been accessed for a long time to ensure the activity of the cache. Periodically re-evaluate the heat type of the data clusters, using a sliding window method. Set a time window of a fixed size, such as 24 hours, and slide the window forward for update at a fixed time interval, such as 1 hour. For each data cluster, count metrics such as the access count, the number of accessing users, and the access time distribution within the current window to obtain its current heat characteristics. Compare the heat characteristics of the current window with the historical window, calculate its change trend and anomaly degree, and determine whether the heat type of the data cluster has changed. If it has changed, adjust its cache policy in a timely manner to achieve adaptive optimization of the cache. When establishing a data cluster attribute table, association rule mining algorithms such as the Apriori algorithm or the FP-growth algorithm can be used to discover the association patterns between the access behaviors of the data clusters. By analyzing the access logs of data clusters A, B, and C, it is found that 80% of the users will access data cluster B within 1 hour after accessing data cluster A, and 60% of the users will access data cluster C within 2 hours after accessing data cluster B. Then it can be inferred that data clusters A, B, and C have a high degree of association and may belong to the same business scenario or user group. Therefore, they can be classified as the same highly sticky data cluster and similar cache policies can be adopted. At the same time, by calculating the kurtosis and skewness of the access frequency distribution of the data cluster, it can be judged whether it has an obvious heat periodicity. For data cluster D, it is found that the access frequency peak on Monday of each week is more than 2 times higher than that on other dates, and similar access peaks appear on three consecutive Mondays. Then it can be judged that data cluster D has a weekly heat pattern, and its periodic characteristics need to be focused on when formulating the cache policy. When implementing a differential cache space allocation strategy, comprehensively consider factors such as the access frequency, access heat, and data volume of the data clusters, and quantitatively evaluate the cache space requirements of each data cluster.For data cluster E, its daily access volume is 10,000 times, which is 2 times the average level, and its data volume is 50 GB. Then its cache space score = 10,000 / 5,000 * 2 + 50 / 100 = 4.5, that is, its cache space requirement is 5 times that of ordinary data clusters. According to the scoring results, the limited cache space is allocated proportionally to each data cluster. Data clusters with high scores get more cache space, and data clusters with low scores get less cache space, thus realizing the on-demand allocation and optimization of cache resources. When determining the cache invalidation strategy, a state transition model based on Markov chain can be adopted. According to the transition probability of the access state change of the data cluster, predict its access popularity in the next period of time and dynamically adjust the cache invalidation time. For data cluster F, based on its access logs in the recent 1 month, the transition probability from the high popularity state to the low popularity state is statistically obtained as 0.1, and the transition probability from the low popularity state to the high popularity state is 0.2. Then it can be predicted that the probability of maintaining the high popularity state in the next 1 week is 0.8. Therefore, its cache invalidation time can be extended to 1 week to reduce the cache invalidation and repeated loading of high popularity data. When implementing the prefetching strategy, a time series prediction model such as the ARIMA model or the Prophet model can be introduced. According to the historical access trend of the data cluster, predict the change in its access volume in the next period of time and start the data prefetching task in advance to reduce the cache miss rate. For data cluster G, by analyzing its hourly access volume data in the recent 1 month, use the Prophet model to fit the long-term trend, periodic changes, and holiday effects of its access volume, predict the access volume per hour per day in the next 1 week, and trigger the prefetching task 1 hour in advance for time periods with predicted access volume exceeding 100 times, and load the predicted high access data blocks into the cache in advance to cope with possible access peaks. When dealing with temporary hot data, monitor the difference between the current access volume of each data cluster and the historical access volume in the same period in real time, and quickly identify and locate explosive hot events through abnormal scoring indicators such as Z-score. For data cluster H, at 10 am, it is monitored that its access volume in the recent 5 minutes is 500 times, while the average access volume in this time period in history is 50 times, and the standard deviation is 20 times. Then its abnormal score Z = (500 - 50) / 20 = 22.5, which exceeds the preset abnormal threshold of 10. Therefore, it is determined as an abnormal hot event, and immediately trigger cache expansion and data prefetching to increase the cache capacity and processing ability to cope with possible continuously increasing access pressure. When optimizing cache eviction, a multi-level feedback LRU linked list structure can be used to dynamically adjust the position of data blocks in the linked list according to their access frequencies. Data blocks with higher access frequencies are promoted to the head of the linked list, and data blocks with lower access frequencies are demoted to the tail of the linked list. When evicting, preferentially select the data blocks at the tail.The cache space is divided into three levels, and each level maintains an LRU linked list. For a newly accessed data block, it is inserted into the head of the first-level linked list; for a data block that is accessed again, it is removed from the linked list of the current level and inserted into the head of the upper-level linked list; when the cache space is insufficient, the data block at the tail of the third-level linked list is evicted. Through this multi-level feedback mechanism, hot data and cold data can be more accurately distinguished, improving the cache hit rate and utilization rate. When evaluating the cache prefetch effect, the A / B test method can be used. The data clusters are randomly divided into two groups, one group applies the cache prefetch strategy and the other does not. By comparing indicators such as the cache hit rate, access latency, and resource consumption of the two groups, the benefits and costs of the cache prefetch strategy can be scientifically evaluated. For data clusters I and J, they are randomly divided into an experimental group and a control group. The cache prefetch is enabled in the experimental group and disabled in the control group. After continuous observation for 1 week, it is found that the average cache hit rate of the experimental group has increased by 10%, the average access latency has decreased by 20%, while the cache space utilization rate has only increased by 5%. Therefore, it can be considered that the cache prefetch strategy is effective and the benefits are greater than the costs, and it can be extended to other similar data clusters. When re-evaluating the heat type of data clusters, incremental clustering algorithms such as the BIRCH algorithm or the CluStream algorithm can be introduced to perform real-time clustering on the newly collected data cluster access metrics to identify new hot data clusters or data clusters with changed heat types, and update their cache policies in a timely manner. An incremental clustering is performed on the data cluster access logs of the most recent 24 hours every day at midnight. It is found that data cluster K originally belonged to periodic hot data, but the clustering results of the most recent 3 days show that it has changed to persistent hot data. Then its cache policy is immediately adjusted to increase the cache space and expiration time to avoid the loss of the cache of hot data. Through continuous incremental clustering and policy optimization, the adaptability and automation of cache management can be achieved, reducing the cost of manual operation and maintenance.
[0038] S107. By real-time monitoring the access situation and heat changes of data, dynamically adjust the migration and replication methods of data between different storage layers. When the actual access pattern of data deviates from the prediction, trigger data re-clustering and storage optimization.
[0039] Collect data access logs, extract key metrics such as the access time, access frequency, and access traffic of data objects; calculate in real time the current access heat and heat change trend of each data object, and compare with the historical access patterns to identify abnormal data objects with a prediction deviation greater than a preset threshold; if abnormal data objects are identified, trigger the retraining and verification of the heat change trend prediction model; obtain the access log data within a time window as a new training dataset, and use the new dataset to retrain the original heat prediction model; combine the access heat, data volume, and migration cost factors of data objects to quantitatively evaluate different storage levels, and select the level with the highest comprehensive score as the data migration target; re-cluster all data objects, select a clustering algorithm according to the business scenario and data characteristics, extract and standardize the key quantitative metrics of the data object access pattern to obtain new clustering clusters and optimization methods.
[0040] Specifically, collect data access logs, extract key metrics such as the access time, access frequency, and access traffic of data objects, and use a data stream processing engine such as Flink or Spark Streaming to calculate in real time the current access popularity and the trend of popularity change of each data object, and compare it with the historical access patterns to identify abnormal data objects with a large deviation from the prediction. The method for identifying abnormal data can be to construct a time series model of the access metrics of data objects, use techniques such as moving average and exponential smoothing for noise reduction and smoothing processing, and extract long-term trends and periodic patterns; then calculate the deviation degree between the current access metrics and the predicted values of the time series model, set a deviation degree threshold, and if it exceeds the threshold, it is determined to be abnormal; further distinguish temporary occasional anomalies and persistent trend anomalies through indicators such as the duration of the abnormal state and the trend of the change in the degree of abnormality. For the identified abnormal data objects, trigger the retraining and verification of the popularity change trend prediction model. Set the trigger conditions for model retraining, such as when the newly collected data volume reaches a certain scale, the prediction deviation continuously exceeds the threshold for a certain period of time, or model updates are performed regularly. Select an appropriate time window, collect the access log data within this window as the new training data set, and after preprocessing such as outlier filtering, missing value filling, and data normalization, use the new data set to retrain the original popularity change trend prediction model such as linear regression, time series decomposition, neural network, etc., and update the model parameters. Evaluate the prediction performance of the retrained model on the new test data set, such as metrics like mean squared error and mean absolute percentage error, optimize the model until a better model is found, and then deploy it to the online environment to replace the original model. If the prediction results show that the current storage level of the abnormal data object is no longer optimal, then initiate the storage level adjustment process. Considering factors such as the volume of data objects, access frequency, and resource utilization, formulate a data migration strategy, adopt strategies such as incremental migration, batch migration, and hot data first, generate migration tasks, migrate the data objects from the current storage layer to a higher-performance storage layer, and reconfigure the cache and backup strategies. After the migration is completed, promptly update the data directory and metadata information, including data version, storage location, access permissions, etc., and use technical means such as transaction mechanisms and consistency protocols to ensure the consistency of the metadata with the underlying storage and guarantee the correctness of subsequent access requests. Continuously monitor the changes in the access patterns of abnormal data objects, and judge whether the abnormality has returned to normal through indicators such as the duration of the abnormal state and the trend of the change in the degree of abnormality.If the access metrics match the predicted curve well for a continuous period, such as 7 days, it can be considered that the cause of the anomaly may be temporary. The triggering mechanism will gradually reduce the anomaly detection frequency and sensitivity of the data object, avoid frequent storage layer adjustments, and reduce system overhead. If the access metrics of the abnormal data object still deviate significantly from the predicted curve for a continuous period, such as 14 days, it can be considered that the previous heat change trend prediction model is no longer applicable, and it is necessary to re-cluster the data object and adjust its optimization strategy. Re-cluster all data objects, and select a suitable clustering algorithm according to the business scenario and data characteristics, such as K-Means, DBSCAN, GMM, etc. Extract key quantitative metrics of the data object access pattern, such as access frequency, access duration, access curve shape characteristics, etc., and standardize the metrics. Optimize the hyperparameters of the clustering algorithm, such as the number of clusters K, density radius Eps, etc., through methods such as the elbow method and silhouette coefficient. Use the optimized hyperparameters to cluster all data objects to obtain several clusters. For each cluster, statistically analyze the access metric distribution of the internal data objects, extract key central tendency and dispersion metrics, and assign a semantic label to the cluster. For abnormal data objects, calculate the similarity between their access metrics and each cluster, select the cluster with the highest similarity as its category, and apply the storage and cache optimization strategy of the cluster. When calculating the access heat of data objects in real time, a sliding window algorithm can be used. Set a fixed-size time window, such as 1 hour, and count metrics such as the number of accesses and access traffic within the window. Compare them with the metrics in the historical same-period window and calculate the change rate. For example, if the number of accesses to a data object within the current 1 hour is 1000, while the average number of accesses in the historical same-period window is 500, then its change rate is (1000 - 500) / 500 = 100%, indicating a significant increase in its heat. By setting a threshold for the change rate, such as 50%, quickly identify abnormal data objects with a sharp change in heat for key monitoring and optimization. At the same time, use algorithms such as EWMA (Exponentially Weighted Moving Average) to smooth the historical access metrics, fuse the metrics of the current window with a lower weight, and obtain a stable long-term trend curve less affected by the current heat as the benchmark for anomaly judgment. When identifying abnormal data, the Local Outlier Factor (LOF) algorithm can be used to map the access metrics of each data object into a multi-dimensional space, such as a three-dimensional space composed of (number of accesses, access traffic, access interval). In this space, calculate the average distance between each data object and its k nearest neighbor objects to obtain its local density. Then, calculate the ratio of the local density of the data object to the local density of its k nearest neighbor objects to obtain its LOF value. The larger the LOF value, the greater the difference between the data object and the surrounding objects, and the more likely it is to be an outlier point.The access metrics for a certain data object are (1000, 10MB, 10min), and the average metrics for its recent 5 data objects are (100, 1MB, 60min). Then the local density of this data object is approximately 10 times that of its k-nearest neighbors, and its LOF value is also approximately 10, which is much higher than the normal value of 1. It is very likely to be a hot spot anomaly. By setting a threshold for LOF, such as 5, abnormal data objects can be automatically identified, and subsequent processing strategies can be determined according to their degree of abnormality (the magnitude of the LOF value). When migrating data objects to different storage layers, the Belady algorithm can be used to optimize the migration order and strategy to minimize performance loss during the migration process. Through this cache replacement algorithm, the Belady algorithm, the data block that will not be accessed for the longest time in the future is replaced. By analyzing the historical access patterns and future access probabilities of each abnormal data object, its access cost at different storage layers can be estimated. For example, the access latency of SSD is 1ms, and the access latency of HDD is 10ms. Then, using the Belady algorithm with the goal of minimizing the total access latency, the optimal data object migration plan is calculated. For three abnormal data objects A, B, and C, their access probabilities within the next hour are 0.2, 0.6, and 0.5 respectively, and they are currently all located in the HDD hard disk drive layer. If the SSD layer can only accommodate 2 objects, the Belady algorithm will choose to migrate B and C to the SSD because they have the highest future access probabilities. The total expected latency after migration is 0.2×1 + 0.6×1 + 0.5×1 = 3.1ms, which is much less than the case without migration, which is 10ms. The algorithm will adjust the data object migration strategy dynamically according to the latest access statistics every hour. When using a shape-based time series clustering algorithm, such as the K-Shape algorithm, to identify the typical access patterns of data objects, first normalize the access metric time series of each data object to the same length and amplitude range, and then calculate the Shape-Based Distance (SBD) between different sequences as a similarity measure. SBD takes into account the differences in the local shape features of the sequences and is stable under transformations such as translation, stretching, and distortion. Based on SBD, clustering algorithms such as K-Medoids are used to aggregate time series with similar shapes together to obtain K typical access patterns, and each pattern corresponds to a cluster center sequence. For the weekly access volume sequences of a group of data objects, 3 patterns are obtained through K-Shape clustering. The first category has a high access volume on weekdays and a low access volume on weekends, showing a "5+2" rhythm. The second category increases linearly from Monday to Sunday, showing a "ramp-up" feature. The third category has overall random fluctuations without obvious patterns. By matching the access sequence of the new abnormal data object with these 3 patterns, it can be quickly determined which category it belongs to and optimized accordingly.When calculating the similarity of multi-dimensional access metrics, the Mahalanobis distance can be used. It takes into account the correlation and scale differences between different metrics and can more accurately measure the similarity between data objects. For data objects A and B, their access frequencies are 100 times per hour and 200 times per hour respectively, and their access traffic is 10 MB per hour and 15 MB per hour respectively. If the Euclidean distance is used, the distance between them is sqrt((100 - 200)^2+(10 - 15)^2)=100. However, if there is a strong positive correlation between the access frequency and the access traffic, such as a correlation coefficient of 0.8, then their difference should be amplified; on the contrary, if there is a strong negative correlation between them, such as a correlation coefficient of -0.8, then the difference should be reduced. The Mahalanobis distance can automatically learn the covariance matrix between access metrics, assign different weights to different metric combinations, and obtain a more reasonable similarity measure. For example, after calculation, the Mahalanobis distance between A and B is 120, which is higher than the Euclidean distance, indicating that the difference in their access patterns has been amplified and they may belong to different clustering clusters.
[0041] S108. For data with different life cycle stages and business values, perform corresponding data protection and backup, including for data that is in the active period and has a high business value, using a short backup cycle and a low recovery time method for data protection and backup; for data that is in the archival period and has a low business value, reduce the backup frequency and extend the recovery time to save backup storage space and costs.
[0042] Obtain the metadata information of the data, where the metadata information includes the creation time, last access time, access frequency, and importance level of the data. According to the metadata information, judge the life cycle stage and business value level of the data. For data with different life cycle stages and business value levels, match the corresponding policies from the predefined data protection and backup policy templates. If the data belongs to critical data, set redundant local copies and off-site copies, and store them dispersedly in different physical locations to avoid single point of failure.
[0043] Specifically, obtain the metadata information of the data, including the creation time, last access time, access frequency, importance level, etc. of the data. Through metadata analysis, judge the life cycle stage of the data, such as the active period, inactive period, archival period, etc. At the same time, evaluate the importance of the data to determine the business value level of the data. According to the life cycle stage and business value level of the data, match the predefined data protection and backup strategy templates. For the data in the active period and with high business value, select a higher-level protection and backup strategy template, which defines a shorter backup cycle such as daily backup, a lower recovery time objective such as within 1 hour recovery, and a larger number of backup copies. For the data in the archival period and with low business value, select a lower-level protection and backup strategy template, which defines a longer backup cycle such as monthly backup, a higher recovery time objective such as within 24 hours recovery, and a smaller number of backup copies. The storage location strategy for backup copies is to use more local copies such as 2 for critical data, stored on different physical disks or storage pools, and more off-site copies such as 2, stored in different cities, regions, and cloud service providers; for non-critical data, use fewer local copies such as 1 and off-site copies such as 1, and store them as scattered as possible to avoid single point of failure. After determining the data protection and backup strategy, automatically generate data backup and replication tasks. The backup task formulates a detailed execution plan according to the backup cycle and backup time window of the data, including the time points of full backup and incremental backup, the target backup storage location, backup method, etc.; the replication task formulates a detailed execution plan according to the number of copies and copy location requirements of the data, including the source and destination of replication, the network channel of replication, the bandwidth limit of replication, etc. The policy engine issues the backup and replication tasks to the corresponding storage nodes and data pipeline components, which are responsible for the specific execution and monitoring. The policy engine, through a unified scheduling framework, real-time grasps the location, status, load and other information of the storage nodes and data pipeline components, and realizes the automatic distribution, coordination, monitoring and exception handling of backup and replication tasks. During the execution of the backup task, continuously collect the metadata information of the backup data, including indicators such as the size of the backup data, backup time consumption, backup throughput, etc., and combine the runtime indicators such as the change rate of the data, access frequency, recovery time objective (RTO), data loss tolerance (RPO), etc., to dynamically adjust the execution plan of the backup task, such as adjusting the backup frequency, backup period, backup concurrency, etc., and minimize the impact of backup on the business as much as possible under the premise of meeting the data protection and recovery requirements, such as avoiding backup during the business peak period, off-peak backup, etc.The adjusted quantization is based on the following: if the data change rate exceeds 20% for three consecutive days and the access frequency exceeds 1000 times per hour for three consecutive days, the backup frequency will be adjusted from daily to every 12 hours; if the RTO of the data is less than 1 hour and the RPO is less than 5 minutes, the backup frequency will be adjusted to every 1 hour and the backup concurrency will be increased to 10; if the system resource utilization rate exceeds the threshold for three consecutive days, an alarm will be triggered and manual intervention will be required to analyze and adjust the backup strategy. After the backup is completed, automatic verification and integrity checks will be performed on the backup data, including comparison of the consistency between the backup metadata and the source data, read and write tests of the backup data, etc., to ensure the accuracy and recoverability of the backup data; at the same time, the metadata information of the source data will be updated, and information such as the location, version, and timestamp of the backup data will be recorded for subsequent recovery and management. For cross-domain and cross-region data replication, the optimal data transmission path will be selected through intelligent routing algorithms and network topology awareness. An adaptive path selection method based on reinforcement learning is adopted, abstracting the network topology as the state space, the optional paths of the replication task as the action space, defining the reward function according to the completion quality of the replication task, and continuously trial-and-error and learning through algorithms such as Q-learning to obtain the optimal path selection strategy. In the offline training stage, a reinforcement learning model is constructed using historical data; according to the real-time network status, the optimal transmission path is dynamically selected; the link quality is continuously monitored during the transmission process, and the path is adjusted in a timely manner to cope with network failures; the reinforcement learning model is iteratively updated regularly to continuously optimize the routing strategy. Considering factors such as network latency, bandwidth utilization, and operating costs, the optimal transmission path for cross-domain data replication is adaptively selected. Based on the backup metadata information, the backup storage space is monitored and managed in real time, tracking the backup volume and storage occupancy of data in different lifecycle stages. When the utilization rate of the backup storage space exceeds the preset threshold, such as 80%, an alarm mechanism is triggered to prompt the administrator to expand the backup storage resource pool; when abnormal situations such as backup failures and verification inconsistencies occur, a fault diagnosis process is started to analyze the cause of the failure and attempt automatic retry or manual intervention for recovery, continuously optimizing the data backup process and strategy to improve the automation and intelligence level of data protection. By extracting and analyzing various metadata attributes of the data, such as creation time, access time, access frequency, data size, etc., the automatic division of the data lifecycle can be realized. For a certain dataset, through statistical analysis, it is found that its access frequency in the last month is 10 times per day, while the average access frequency in the past year is only 1 time per month, and the data volume reaches 1TB. Then it can be basically judged that this dataset is currently in the active period and has high business value, and a higher level of protection and backup strategy should be adopted. For another dataset, there is no access record in the last year, the last modification time has exceeded 3 years, and the data volume is small, only 100GB. Then it can be considered that it has entered the archival period and has low business value. A simplified backup strategy can be adopted, and consideration can be given to migrating it to a low-cost cold storage medium.By judging the data lifecycle and business value through quantitative metrics and then mapping them to predefined policy templates, the automatic matching and application of data protection policies can be achieved. After determining the data protection and backup policies, the corresponding backup and replication tasks are automatically generated and executed. For backup tasks, a combination of incremental backup and synthetic full backup can be adopted, that is, incremental backup is performed daily to capture the amount of data changes, while a full backup is performed once a week or month to merge multiple incremental backups into a complete data copy. The backup time window can be dynamically adjusted according to the access heat of the data. For example, the backup is executed from 1 am to 3 am every day to avoid the business peak period. For cross-domain replication tasks, in order to make full use of the network bandwidth, a dynamic traffic control algorithm such as the TCPBBR congestion control algorithm can be adopted. According to the network latency and bandwidth estimation, the transmission rate and concurrency of the replication task are automatically adjusted to improve the cross-domain replication efficiency as much as possible on the premise of ensuring the security and reliability of the replicated data. As the core component of data protection, the policy engine realizes end-to-end automation by centrally managing all aspects such as backup and replication policies, tasks, scheduling, and execution. Based on the microservices architecture, the policy engine decouples different functional modules into independent services, such as policy management services, task scheduling services, log monitoring services, etc., and communicates through a unified RestAPI interface. Service registration and discovery use distributed middleware such as Eureka or Consul, and message queues such as Kafka or RabbitMQ are used for communication between services to achieve service decoupling, dynamic scaling, and high availability. During the execution of backup tasks, continuous monitoring and optimization are required, which requires real-time collection of various performance metrics of backup tasks, such as backup throughput, backup latency, CPU / memory / network / disk I / O utilization, etc., to form performance time-series data of task execution. When it is found that the backup latency is continuously higher than 100 ms for 5 minutes and the disk I / O utilization is continuously higher than 70% for 10 minutes, it is determined that the backup performance is abnormal. At this time, it is necessary to comprehensively evaluate whether the current backup policy matches in combination with dimensions such as data change rate, access frequency, and data volume. If the backup frequency is set too high, such as backing up once an hour, it may cause system resource overload. It is necessary to trigger the dynamic adjustment process of the backup policy, reduce the backup frequency to once a day until the system load returns to normal. The adjustment process is based on the MAPE closed-loop control model and realizes the adaptive optimization of the backup policy through steps such as monitoring, analyzing, planning, and executing. If the backup performance problem still cannot be solved after multiple adjustments, it is necessary to further analyze whether there are bottlenecks at the system architecture level, such as insufficient backup network bandwidth, insufficient CPU or memory of the backup server, etc., and trigger the resource expansion process to improve the processing capacity of the backup infrastructure.Backup data verification is a crucial part of ensuring data security and reliability. By using hash algorithms such as CRC32, MD5, and SHA256, the hash values of each data block in the backup file are calculated one by one and compared with the expected hash values recorded in the backup metadata. If inconsistencies are found, it indicates that the backup file has been damaged during transmission or storage, and it needs to be retransmitted or repaired to ensure the integrity and consistency of the backup data. At the same time, it is also necessary to regularly conduct recoverability tests on the backup data. Select some backup files and perform recovery drills in the test environment to check whether the recovered data is complete and usable and whether it is consistent with the production data to verify the effectiveness of the backup data. For data replication across domains and regions, the selection of the network transmission path is crucial, which directly affects the RTO and cost of the replication task. The adaptive routing algorithm based on reinforcement learning continuously optimizes the routing strategy through trial-and-error learning of learning-by-doing. For example, for a data replication task on a path, initially choose any path randomly. For example, the transmission time of a path is 120 minutes and the transmission cost is 1000 yuan. By continuously trying different path combinations, statistically analyzing the transmission performance and cost, and updating the reward value of the routing strategy. If it is found that the transmission time of this path is 100 minutes and the cost is 800 yuan, then increase the reward value of this path combination and gradually increase its selection probability. At the same time, an exploration mechanism is introduced, using the ε-greedy strategy, to randomly select a new path combination with a probability of ε to avoid premature convergence to the local optimum. After multiple rounds of iterative learning, the optimal transmission path set for different replication tasks is finally obtained to achieve global optimization of network routing. By collecting performance statistics data of components such as storage nodes, backup servers, and network links, a global monitoring and alerting and capacity management mechanism is constructed. Set the thresholds for key metrics, such as the disk space utilization rate exceeding 85%, the network packet loss rate exceeding 1%, the backup failure rate exceeding 0.1%, etc., to trigger different levels of alerts and notify the operation and maintenance personnel to handle them in a timely manner. Combining with the prediction of data growth trends, advance the expansion plan of backup storage resources. For example, if the disk space grows by 10% per month and will reach the design capacity of the storage cluster, 80TB, in half a year, then it is necessary to start the procurement process 2 months in advance, purchase new storage devices, and expand smoothly to avoid business interruption caused by the exhaustion of storage space.
[0044] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A method for extracting life cycle features of hot and cold data objects based on deep learning, characterized in that: The method comprises: Obtain the access frequency, access interval, access time distribution, and number of repeated accesses of data objects as access pattern features, as well as data creation time, last access time, and access times as lifecycle features, and construct a feature vector that describes data stickiness; A clustering algorithm is used to cluster the feature vectors. According to the clustering results, the data is divided into data clusters with different stickiness levels. The data in each stickiness level data cluster has similar access patterns and life cycle characteristics. For each stickiness level data cluster, analyze its access time distribution and life cycle stage to determine its popularity type, including continuous popularity, periodic popularity, temporary popularity, and cold popularity. According to the current popularity type and historical access pattern of the stickiness-level data cluster, its popularity change trend in the future is predicted to obtain the popularity change curve of each cluster; Combine the stickiness level, heat type and heat change trend of the stickiness level data cluster to determine its optimal storage location and migration method, store high-stickiness hot data on hot storage devices with fast access speed and high cost, and migrate low-stickiness cold data to cold storage devices with slow access speed and low cost; For data clusters with different viscosity levels and heat types, corresponding cache and prefetch strategies are adopted, including increasing cache allocation and prefetch frequency for high-viscosity continuous hot data, prefetching periodic hot data in advance during its heat rise period, dynamically adjusting cache space for temporary hot data, and reducing cache allocation and prefetch frequency for cold data; By monitoring data access and popularity changes in real time, the migration and replication of data between different storage layers can be dynamically adjusted. When the actual access pattern of data deviates from the prediction, data re-clustering and storage optimization are triggered. For data at different life cycle stages and business values, corresponding data protection and backup are carried out, including using short backup cycles and low recovery times to protect and backup data in the active period with high business value, and reducing the backup frequency and extending the recovery time for data in the archiving period with low business value to save backup storage space and costs.
2. The method according to claim 1, wherein: The access frequency, access interval, access time distribution, and number of repeated accesses of the data object are obtained as access pattern features, and the data creation time, last access time, and number of accesses are obtained as life cycle features to construct a feature vector describing data stickiness, including: Obtain access log records of data objects and extract access frequency, access interval and access time distribution attribute information; Count the number of repeated visits to statistical objects and use it as an important indicator to measure the characteristics of access patterns; Get the lifecycle attributes of the data object, including the creation time, the last access time, and the cumulative number of accesses; Based on the access pattern characteristics and lifecycle attributes, a feature vector is constructed to describe data stickiness.
3. The method according to claim 1, wherein: The clustering algorithm is used to cluster the feature vectors, and the data is divided into different stickiness level data clusters according to the clustering results. The data in each stickiness level data cluster has similar access patterns and life cycle characteristics, including: The access mode and life cycle characteristics of data objects are represented by feature vectors, and the similarity between different data objects is determined by measuring the similarity of feature vectors. Select a clustering algorithm for clustering based on the dimension and numerical distribution of the feature vector; The quality of clustering results is judged by calculating the compactness within the clusters and the separation between the clusters. The compactness within the clusters is measured by the average distance, and the separation between the clusters is measured by the inter-cluster distance. According to the clustering results, the data objects are divided into data clusters with different stickiness levels, wherein the data objects in each data cluster have similarities.
4. The method according to claim 1, wherein: For each stickiness level data cluster, the access time distribution and life cycle stage are analyzed to determine the heat type to which it belongs, including continuous heat, periodic heat, temporary heat, and cold, including: Perform statistical analysis on the access time distribution of each stickiness level data cluster, obtain the access time series at different time granularities, and form the access distribution curves at multiple time scales; According to the smoothed access frequency distribution curve, the access mode and heat characteristics of the data cluster are determined; By analyzing the number of peaks, peak intervals, and slope change characteristics of the curve, determine whether the data cluster has a periodic access pattern, whether there is a long period of continuous heat or sudden heat; Combined with the life cycle stage information of the data cluster, the continuous access duration, access frequency and access stability factors of the data cluster, the heat type of the data cluster is judged and divided into different heat categories: continuous hot, periodic hot, temporary hot and cold.
5. The method according to claim 1, wherein: The method predicts the heat change trend of the stickiness level data cluster in the future based on the current heat type and historical access mode of the cluster, and obtains the heat change curve of each cluster, including: Acquire historical access pattern data of each data cluster, wherein the historical access pattern data includes distribution time series characteristics of the number of accesses, access intervals, and access times; By analyzing the historical access pattern data, extracting key features of the data cluster, the key features include average number of accesses, average access interval and access peak time period; Analyze whether the historical access pattern data has periodicity and trend, and select a time series prediction model based on the analysis results, and perform training to obtain a popularity change trend prediction model; Using the optimized feature space and the popularity change trend prediction model, the number of visits to each data cluster in a future period of time is predicted to obtain a predicted popularity change curve; The predicted heat change curve is subjected to trend analysis and abnormal point detection to determine the heat change trend of the data cluster in the future period of time, and to identify abnormal fluctuation points or mutation points in the predicted curve.
6. The method according to claim 1, wherein: The method combines the viscosity level, heat type and heat change trend of the viscosity level data cluster to determine its optimal storage location and migration method, stores high-viscosity hot data on a hot storage device with fast access speed and high cost, and migrates low-viscosity cold data to a cold storage device with slow access speed and low cost, including: Obtain the key attribute information of each stickiness level data cluster, such as stickiness level, popularity type and popularity change trend, as a basis for determining its optimal storage location and migration method; Combine storage capacity, storage unit price, data access volume and access unit price to calculate the total cost under different storage solutions, and calculate the total benefits under different storage solutions to obtain the cost-effectiveness ratio; For data clusters with stable popularity types, the popularity type in the future is predicted based on its historical access pattern and popularity change trend. If the predicted result is inconsistent with the current popularity type, data migration is triggered; For data clusters with high frequency of heat type changes, their heat indicators and access behaviors are monitored in real time. When they exceed the threshold range of the original heat type for multiple consecutive times, the heat type adjustment is triggered and data migration is performed.
7. The method according to claim 1, wherein: The corresponding cache and pre-fetch strategies are adopted for data clusters with different viscosity levels and heat types, including increasing cache allocation and pre-fetch frequency for high-viscosity continuous hot data, pre-fetching periodically hot data in advance during its heat rise period, dynamically adjusting cache space for temporary hot data, and reducing cache allocation and pre-fetch frequency for cold data, including: Obtain information about the stickiness level and popularity type of each data cluster and formulate differentiated caching strategies; For pre-fetching strategies of different data clusters, differentiated configurations matching the cache strategy are adopted; Record the cache hit rate and prefetch accuracy of each data cluster and use them as key evaluation indicators of the data cluster; Re-evaluate the heat type of the data cluster periodically to determine whether the heat type of the data cluster has changed; If the heat type of the data cluster changes, the cache strategy is adjusted to achieve adaptive optimization of the cache.
8. The method according to claim 1, wherein: The method of dynamically adjusting the migration and replication of data between different storage layers by monitoring the access and popularity of data in real time, and triggering data re-clustering and storage optimization when the actual access mode of data deviates from the prediction, includes: Collect data access logs and extract key indicators of data object access time, access frequency and access traffic; Calculate the current access popularity and popularity change trend of each data object in real time, compare it with the historical access pattern, and identify abnormal data objects whose deviation from the prediction is greater than the preset threshold; If an abnormal data object is identified, it triggers the retraining and verification of the popularity change trend prediction model; Get access log data within a time window as a new training data set, and use the new data set to retrain the original popularity prediction model; Combined with the access popularity, data volume, and migration cost of data objects, quantitative evaluation is performed on different storage tiers, and the tier with the highest comprehensive score is selected as the data migration target; All data objects are re-clustered, clustering algorithms are selected according to business scenarios and data characteristics, key quantitative indicators of data object access patterns are extracted and standardized, and new clusters and optimization methods are obtained.
9. The method according to claim 1, wherein: The data protection and backup are performed accordingly for data in different life cycle stages and business values, including the use of short backup cycles and low recovery time methods for data in the active period and with high business value, and the use of reduced backup frequency and extended recovery time methods for data in the archiving period and with low business value to save backup storage space and costs, including: Acquire metadata information of the data, wherein the metadata information includes creation time, last access time, access frequency, and importance level of the data; Determine the life cycle stage and business value level of the data based on the metadata information; Match corresponding policies from predefined data protection and backup policy templates for data at different lifecycle stages and business value levels; If the data is critical, set up redundant local copies and off-site copies, and store them in different physical locations to avoid single point failures.
Citation Information
Cited By
Cold and hot data exchange method and system based on optical storage and storage medium
CN120428929A
Data dynamic caching method and device
CN120973829A
Cold and hot data distribution management method and device for nonvolatile storage device
CN121029093A
A non-volatile storage device hot and cold data shunting management method and device
CN121029093B
Game data processing method and device based on calorific value estimation
CN121051224A