Intelligent cold data migration method and system

Through time series analysis and clustering algorithms, and combined with data deduplication and correlation analysis, multi-level migration is performed using edge-cloud collaborative hierarchical method, which solves the misjudgment risk and resource waste of hot and cold data division and migration schemes in the existing technology, and achieves more efficient storage resource utilization and performance optimization.

CN119938282AInactive Publication Date: 2025-05-06BEIJING LEXUN TECH CO LTD

Patent Information

Application Number
CN202510437122.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has the risk of misjudgment and waste of resources in hot and cold data division and migration schemes, and cannot dynamically adapt to changes in data access patterns, resulting in poor storage performance.

Method used

The time series analysis algorithm is used to predict the state space of the data block, and automatically classify the hot and cold degree of the data block based on the clustering algorithm. Repeated data blocks are identified in clustering, combined with data deduplication technology, correlation analysis is performed through the Louvain algorithm, and multi-level migration processing is performed using the edge-cloud collaborative hierarchical method.

Benefits of technology

Automatic classification of the degree of hot and cold of data blocks and deduplication of duplicate data, dynamically adjust storage node allocation, improve the balance of storage resources and overall efficiency, and avoid the problem of excessive single-point storage pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938282A_ABST
    Figure CN119938282A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent cold data migration method and system, and relates to the technical field of data migration, and the method comprises the steps: predicting the state space of a data block through a time sequence analysis algorithm, automatically classifying the cold and hot degrees of the data block based on a clustering algorithm, and optimizing the migration data size in combination with a data deduplication technology after recognizing repeated data blocks in clustering. And carrying out association analysis on the data blocks needing to be migrated, carrying out secondary optimization on the migrated data volume, and carrying out multi-level migration processing on the data blocks needing to be migrated by utilizing an edge-cloud collaborative grading method in combination with a data block optimization result and a cold and hot degree clustering result. After the migration system automatically classifies the cold and hot degrees of the data blocks, repeated data blocks can be further identified in the clustering process, and by combining a data deduplication technology, invalid data migration is reduced, the storage utilization rate is improved, the distribution mode of storage nodes is dynamically adjusted, so that data storage resources are more balanced, and the problem that single-point storage pressure is too high is avoided; and the overall storage efficiency of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data migration, and in particular to an intelligent cold data migration method and system. Background Technology

[0002] With the continuous growth of data storage demand, distributed storage systems have become an important means to cope with large-scale data storage. However, among a large amount of data, only a small part of the data will be frequently accessed (hot data), while other data has not been accessed for a long time and becomes cold data. The ratio of hot and cold data is usually unbalanced. Cold data occupies high-performance storage media, resulting in resource waste and unnecessary operation and maintenance expenses. Therefore, it is necessary to use a migration system to intelligently migrate and manage cold data.

[0003] The existing technology has the following deficiencies:

[0004] 1. Traditional methods for partitioning hot and cold data usually rely on simple time thresholds or access frequency statistics, such as LRU (least recently used) or LFU (least frequently used) strategies. This approach has the risk of misjudgment, especially when the data access pattern has sudden or periodic changes.

[0005] 2. Traditional migration solutions are usually based on static rules (such as using simple polling strategies or fixed storage node selection strategies), which may cause some storage nodes to be overloaded, while other storage node resources are not fully utilized. This method cannot dynamically adapt to changes in data access patterns, which may lead to unreasonable storage of hot and cold data and affect storage performance.

[0006] Based on this, the present invention proposes an intelligent cold data migration method and system. After automatically classifying the hotness and coldness of data blocks, duplicate data blocks can be further identified in the clustering process, and combined with data deduplication technology, invalid data migration can be reduced, storage utilization can be improved, and the allocation method of storage nodes can be dynamically adjusted to make data storage resources more balanced, avoid the problem of excessive storage pressure at a single point, and improve the overall storage efficiency of the system. SUMMARY OF THE INVENTION

[0007] The purpose of the present invention is to provide an intelligent cold data migration method and system to solve the deficiencies in the background technology.

[0008] In order to achieve the above purpose, the present invention provides the following technical solution: an intelligent cold data migration method, the migration method comprising the following steps:

[0009] The migration system uses a time series analysis algorithm to predict the state space of data blocks, and automatically classifies the hotness and coldness of data blocks based on a clustering algorithm;

[0010] After identifying duplicate data blocks in the cluster, the amount of data to be migrated is optimized by combining data deduplication technology, and the Louvain algorithm is used to perform association analysis on the data blocks to be migrated, and the amount of data to be migrated is optimized again;

[0011] Combining the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated.

[0012] In a preferred embodiment, combining the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated, including the following steps:

[0013] Get all clusters that need data migration, and perform duplicate data processing and association analysis on the data blocks in the clusters that need data migration, recalculate the new cluster center value of each cluster that needs data migration, and divide all data blocks in all clusters into medium-frequency fluctuation data or low-frequency cold data based on the new cluster center value;

[0014] Multi-level migration processing includes:

[0015] S1: Edge Storage Cache

[0016] Store the intermediate frequency data block in the edge server as a temporary storage buffer;

[0017] S2: Batch migration to the cloud

[0018] Batch transfer of low-frequency cold data to the cloud;

[0019] S3: Access pattern adjustment

[0020] If cold data becomes hot, the edge node resumes access;

[0021] S4: Access mode adjustment

[0022] After the migration is complete, clean up duplicate or redundant data in the source storage.

[0023] In a preferred embodiment, all data blocks in all clusters are divided into medium-frequency fluctuation data or low-frequency cold data based on the new cluster center value, including the following steps:

[0024] By comparing the new cluster center value with the first gradient threshold, determine whether the data block in the cluster belongs to cold data or hot data. If the new cluster center value is greater than or equal to the first gradient threshold, the data block in the cluster is determined to be hot data. If the new cluster center value is less than the first gradient threshold, the data block in the cluster is determined to be cold data.

[0025] ​Compare the new cluster center value with the second gradient threshold, the second gradient threshold is less than the first gradient threshold, and the second gradient threshold is used to perform secondary division of cold data blocks. If the new cluster center value of a cluster is less than the first gradient threshold, and the new cluster center value is greater than or equal to the second gradient threshold, all data blocks in the cluster are divided into medium-frequency fluctuation data. If the new cluster center value of a cluster is less than the second gradient threshold, all data blocks in the cluster are divided into low-frequency cold data.

[0026] In a preferred embodiment, automatically classifying the hotness and coldness of data blocks based on a clustering algorithm includes the following steps:

[0027] After obtaining the state factors of all data blocks, the elbow method is first used to calculate the sum of squares of intra-cluster errors under different K values, and the inflection point is selected as the optimal K value;

[0028] Randomly select the state factors of K data blocks as the initial cluster centers of K clusters, calculate the difference between the state factors of the data blocks and the initial cluster centers of K clusters, and assign the data blocks to the cluster with the smallest difference. When all data blocks are assigned, calculate the mean of the state factors of the clusters as the new cluster centers;

[0029] Repeat the steps of data block allocation and new cluster center calculation until the new cluster centers of the K clusters converge and the clustering process is completed. Sort the K clusters from large to small according to the final new cluster centers to obtain the cluster list.

[0030] In a preferred embodiment, the migration system uses a time series analysis algorithm to predict the state space of the data block, including the following steps:

[0031] Obtain the time series data of each data block from the storage system. The time series data includes access time, access times, access interval, and read-write ratio. Remove abnormal access records and fill in missing data, and normalize the time series data.

[0032] Divide the access records of data blocks according to fixed time intervals, construct time series, calculate the time series data change trend of each data block in different time windows, and use Fourier transform to identify the periodic pattern of data access;

[0033] Use long short-term memory network to train deep neural network, learn the temporal relationship of data block access, and generate state factors for data blocks after predicting future access trends.

[0034] In a preferred embodiment, after identifying duplicate data blocks in the cluster, the amount of migrated data is optimized by combining data deduplication technology, including the following steps:

[0035] In the cluster list, mark the cluster whose new cluster center value is less than the first gradient threshold as the optimized cluster;

[0036] In all optimized clusters, obtain the duplication factor between two data blocks, and mark the two data blocks whose duplication factor is greater than the preset duplication threshold as duplicate data blocks;

[0037] Use the incremental storage method to deduplicate duplicate data blocks and only store the differences between data blocks. If the hash values ​​of multiple data blocks are consistent, only one copy of the data block is stored.

[0038] In a preferred embodiment, the calculation logic of the repetition factor is: calculate a fixed-length hash value for each data block in the optimized cluster, use LSH to perform similarity matching to identify the content similarity of the data block, obtain the hash value difference and content similarity between the two data blocks, normalize the hash value difference and content similarity, map the value range of the hash value difference and content similarity to [0,1], and subtract the hash value difference from the normalized content similarity to obtain the repetition factor.

[0039] In a preferred implementation, the Louvain algorithm is used to perform correlation analysis on the data blocks to be migrated, and the amount of migrated data is optimized secondary, including the following steps:

[0040] Define all stored data blocks as data blocks that need to be migrated, and first determine whether it is necessary to convert the data blocks that need to be migrated into data blocks that do not need to be migrated based on the correlation between the data blocks that need to be migrated and the data blocks that do not need to be migrated;

[0041] Through the Louvain algorithm, each data block that needs to be migrated is regarded as a community. The data blocks are traversed, and the modularity gain after joining the adjacent community is calculated. The community with the largest modularity gain is selected for merging, and the process is iterated until the modularity gain no longer increases;

[0042] The communities formed in the first stage are regarded as new nodes, the new association network is recalculated, and community merging is continued until the global optimal partition is reached.

[0043] In a preferred embodiment, the clustering process is completed until the new cluster centers of the K clusters converge, including the following steps:

[0044] If the change of the new cluster centers of K clusters in two consecutive iterations is less than the set change threshold, the new cluster centers are considered to converge, and the expression is: , where represents the new cluster center of cluster i in the tth iteration, represents the new cluster center of cluster i in the t-1th iteration, is the change threshold, is the number of clusters.

[0045] Intelligent cold data migration system, including data block hot and cold clustering module, data deduplication module, secondary optimization module, and multi-level migration module;

[0046] Data block hot and cold clustering module: uses time series analysis algorithm to predict the state space of data blocks, and automatically classifies the hot and cold degrees of data blocks based on clustering algorithm;

[0047] Data deduplication module: After identifying duplicate data blocks in the cluster, it combines data deduplication technology to optimize the amount of migrated data;

[0048] Secondary optimization module: Use the Louvain algorithm to perform correlation analysis on the data blocks that need to be migrated, and secondary optimize the amount of migrated data;

[0049] Multi-level migration module: Combined with the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated.

[0050] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0051] 1. The present invention predicts the state space of data blocks by using a time series analysis algorithm, and automatically classifies the hotness and coldness of data blocks based on a clustering algorithm. After identifying duplicate data blocks in clustering, the data deduplication technology is combined to optimize the amount of migrated data, and the Louvain algorithm is used to perform association analysis on the data blocks to be migrated, and the amount of migrated data is optimized twice. The edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks to be migrated, combining the data block optimization results and the hotness and coldness clustering results. After the migration system automatically classifies the hotness and coldness of data blocks, it can further identify duplicate data blocks in the clustering process, and combine data deduplication technology to reduce invalid data migration, improve storage utilization, and dynamically adjust the allocation method of storage nodes to make data storage resources more balanced, avoid the problem of excessive storage pressure at a single point, and improve the overall storage efficiency of the system.

[0052] 2. The present invention optimizes the amount of data to be migrated by combining data deduplication technology, and performs correlation analysis on the data blocks to be migrated through the Louvain algorithm to optimize the amount of data to be migrated. It can mine data sets with strong correlation, thereby ensuring the integrity of related data blocks during migration, reducing data access overhead, and optimizing subsequent data query performance. Brief Description of the Figures

[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0054] Figure 1 is a flow chart of the method of the present invention.

[0055] Figure 2 This is the framework structure diagram of the present invention. Specific implementation method

[0056] To make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be described clearly and completely in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] Example 1: Please refer to Figure 1 and Figure 2 As shown in the figure, the intelligent cold data migration method of this embodiment includes the following steps:

[0058] The migration system uses a time series analysis algorithm to predict the state space of a data block, including the following steps:

[0059] Obtain the access time, access count, access interval, read-write ratio and other time series data of each data block from the storage system, remove abnormal access records, such as instantaneous high-frequency access or invalid access requests, and fill in the missing data, normalize the access time, access count, access interval, read-write ratio and other features to make the access patterns of different data blocks comparable. After the normalization is completed, add the access time to the access count minus the access interval plus the read-write ratio to obtain the access index of the data block.

[0060] Divide the access records of data blocks according to fixed time intervals (such as hours, days, weeks), construct time series, calculate the change trend of time series data of each data block in different time windows, including the maximum, minimum, mean and variance of time series data, and use Fourier transform (FFT) to identify the periodic pattern of data access.

[0061] For each data block access sequence within a time window, the following features are calculated:

[0062] Maximum value (indicates peak): , where Represents the access index of the data block in time window i.

[0063] Minimum value (indicates low point): , n represents the number of time windows.

[0064] Mean (Reflecting the overall trend): .

[0065] Variance (Measure volatility): , the larger the variance, the more volatile the access pattern.

[0066] Assume the access index sequence is: , where represents the access index of the data block in the time window i, and the fast Fourier transform (FFT) of the time series X is performed, and the expression is:

[0067] , where is the Fourier transform result (spectral coefficient) of the gth frequency component, is the number of visits to the nth data point in the time series, is the total length of the time series, that is, the number of time windows of data block access records, is the frequency index, the value range is , n is the time index, corresponding to the data point number in the time series, the value range is , is the kernel function of the Fourier transform, where j is the imaginary unit (j 2 =-1), which represents a complex exponential function, used to convert time signals into the frequency domain.

[0068] Then the power spectrum calculation formula is: , where is the power spectrum. The frequency corresponding to the peak of the power spectrum reflects the main period of the access mode. The period Calculation formula: , where is the frequency corresponding to the peak of the power spectrum. If the power spectrum of a certain frequency component is significantly higher than other components, it indicates that the access pattern of the data block is periodic. Based on the results of Fourier transform (FFT), it can be determined whether there is a stable access cycle (such as daily, weekly, and monthly access patterns), identify high-frequency access data blocks, optimize data storage strategies (such as loading hot data in advance), distinguish long-term stable access patterns from burst access patterns, and adjust storage management strategies.

[0069] For example: After Fourier transform analysis, it is found that data access is mainly concentrated in working hours (9:00-18:00), and the corresponding frequency component shows a significant 24-hour periodic peak. In addition, it is found that there is a high number of visits every hour, indicating that the data block belongs to high-frequency access data.

[0070] Storage optimization strategy: Place the data block in memory cache (such as Redis, Memcached) or high-performance SSD to reduce disk I / O burden and improve access speed. During peak access times, dynamically allocate multiple storage nodes to provide services to prevent overload of a single storage device.

[0071] For the case where access patterns have long-term dependencies, a long short-term memory network is used to train a deep neural network to learn the temporal relationship of data block access and predict future access trends. The trend of temporal data changes for each data block in the future is calculated to generate a state factor for the data block.

[0072] Since the access patterns of data blocks have long-term dependencies (i.e., access behaviors in the past will affect future access trends), we can use long short-term memory networks (LSTMs) to train deep neural networks to learn the temporal relationships of data blocks. LSTMs solve the gradient vanishing problem in traditional RNN training through a memory gating mechanism, allowing the network to effectively capture long-span dependencies.

[0073] Build an LSTM model, including:

[0074] Input layer: accepts historical access data sequence.

[0075] Hidden layer: Multi-layer LSTM structure extracts long-term dependency features.

[0076] Output layer: predict the data block access trend in the future time window.

[0077] Use the data block access records of the past few days as input to predict the access patterns of the next few days. The training of the LSTM model belongs to the prior art and will not be described in detail in this application.

[0078] After training, use the LSTM model to predict the time series data change trend of each data block in the future:

[0079] Access time series: predict the time points of future accesses and determine whether there are periodic or sudden changes.

[0080] Access times series: predicts future access frequency and is used to evaluate changes in the hotness and coldness of data blocks.

[0081] Access interval sequence: Determines whether the access pattern becomes more concentrated or dispersed, and decides whether the data block needs to be cached.

[0082] Read-write ratio sequence: predict the read-write mode changes of future data blocks and optimize storage strategies (for example, high read-write ratio data blocks can be allocated to high-performance SSDs).

[0083] Thus, the access time growth rate, access number growth rate, access interval growth rate and read-write ratio growth rate are output, and the access time growth rate, access number growth rate, access interval growth rate and read-write ratio growth rate are summed to obtain the state factor.

[0084] The access time growth rate, access number growth rate, access interval growth rate and read-write ratio growth rate are calculated using the general growth rate calculation formula, and the expression is:

[0085] , where is the parameter growth rate, is the value of the parameter at time t2, is the value of the parameter at time t1.

[0086] Automatically classify the hotness and coldness of data blocks based on clustering algorithms, including the following steps:

[0087] The smaller the value of the status factor, the more likely the data block should be classified as a cold data block.

[0088] After obtaining the state factors of all data blocks, the elbow method is first used to calculate the intra-cluster error sum of squares (WSS) under different K values, and the inflection point is selected as the optimal K value. The state factors of K data blocks are randomly selected as the initial cluster centers of K clusters, and the difference between the state factors of the data blocks and the initial cluster centers of K clusters is calculated, and the data blocks are assigned to the cluster with the smallest difference. When all data blocks are assigned, the mean state factor of the cluster is calculated as the new cluster center, and the data block assignment and new cluster center calculation steps are repeated until the new cluster centers of K clusters converge and the clustering process is completed.

[0089] If the change of the cluster centers of K clusters in two consecutive iterations is less than the set change threshold, the algorithm is considered to have converged. The expression is: , where represents the new cluster center of cluster i in the tth iteration, represents the new cluster center of cluster i in the t-1th iteration, is the change threshold, is the number of clusters.

[0090] Sort the K clusters from large to small according to the final new cluster center to obtain a cluster list. In the cluster list, the lower the cluster is ranked, the colder the data block in the cluster is.

[0091] After identifying duplicate data blocks in the cluster, combine data deduplication technology to optimize the amount of migrated data, including the following steps:

[0092] In practical applications, cold data is rarely accessed (i.e., the smaller the new cluster center value of a cluster, the more likely the data blocks in the cluster are low-frequency cold data), and the migration ratio of duplicate data may be high, which is suitable as an optimization target. Therefore, in the cluster list, the cluster whose new cluster center value is less than the first gradient threshold is marked as the optimized cluster. The first gradient threshold is used to determine whether the data blocks in the cluster are cold data. When the new cluster center value is greater than or equal to the first gradient threshold, the data blocks in the cluster are determined not to be cold data.

[0093] Calculate a fixed-length hash value for each data block in the optimized cluster, such as MD5, SHA-256, to quickly compare whether the data content is the same. If multiple data blocks have the same hash value, they may be duplicate data. Use LSH for similarity matching to identify the content similarity of the data blocks. After obtaining the hash value difference and content similarity between the two data blocks, normalize the hash value difference and content similarity so that the value range of the hash value difference and content similarity is mapped to [0,1]. Subtract the hash value difference from the normalized content similarity to obtain the repetition factor. Mark the two data blocks with a repetition factor greater than the preset repetition threshold as duplicate data blocks. Use the incremental storage method (Delta-Encoding) to deduplicate the duplicate data blocks and only store the difference between the data blocks. If the hash values ​​of multiple data blocks are consistent, deduplication can be performed directly and only one data block is stored.

[0094] In this application, the difference in hash values ​​is usually calculated using the Hamming distance, which indicates the number of bits that differ between two hash values. This calculation method belongs to the prior art and will not be described in detail in this application.

[0095] Content similarity reflects the similarity of the actual content of two data blocks, not just the similarity of hash values. The goal of LSH is to approximate the content similarity through hash values, but the actual content similarity can be calculated by comparing the actual content of the data blocks. Therefore, after defining the content feature vector of the data block, the content similarity of the two data blocks is calculated. The expression is:

[0096] Content similarity = (D1·D2) / (||D1||*||D2||), where D1·D2 is the dot product of the content feature vector of data block 1 and the content feature vector of data block 2, and ||D1|| and ||D2|| are the norms of D1 and D2 (i.e., their respective lengths).

[0097] Select appropriate features based on the type and content of the data block. For example:

[0098] For text data blocks, you can extract word frequency, TF-IDF (term frequency-inverse document frequency), etc.

[0099] For image data blocks, the color histogram, texture features, edge information, etc. of the image can be extracted.

[0100] For audio data blocks, Mel-frequency cepstral coefficients (MFCC) can be extracted, etc.

[0101] The above-mentioned feature vector extraction method belongs to the prior art and will not be introduced one by one in this application.

[0102] Before data migration, detect and remove duplicate data blocks, migrate only unique data, improve migration efficiency, and use metadata to record mapping relationships for deduplicated data blocks to ensure that the original data can still be correctly found when the application accesses it.

[0103] Use the Louvain algorithm to perform correlation analysis on the data blocks that need to be migrated, and optimize the amount of migrated data twice, including the following steps:

[0104] Define all stored data blocks as data blocks that need to be migrated. First, determine whether the data blocks that need to be migrated need to be converted into data blocks that do not need to be migrated based on the correlation between the data blocks that need to be migrated and the data blocks that do not need to be migrated. Then, use the Louvain algorithm to treat each data block that needs to be migrated as an independent community (that is, each data block is stored separately). Traverse the data blocks and calculate the modularity gain after joining the adjacent community. Select the community with the largest modularity gain to merge. Iterate until the modularity gain no longer increases. Treat the community formed in the first stage as a new node, recalculate the new association network, and continue to merge the communities until the global optimal division is reached. Optimize data batch migration based on community structure, reduce unnecessary scattered data movement, and enhance the storage proximity of highly associated data blocks, effectively improving access efficiency.

[0105] Get the mutual information and covariance between the two data blocks, normalize the mutual information and covariance, and then sum the normalized mutual information and covariance to get the correlation coefficient between the two data blocks. The larger the correlation coefficient, the higher the correlation between the data blocks, that is, the use of one data block may be inseparable from another data block (even if the data block is cold data).

[0106] The calculation expression of mutual information is: , where is the mutual information, which indicates the amount of information shared between random variables X and Y, represents the joint probability distribution when X is x and Y is y, represents the marginal probability distribution of variable X taking the value x, represents the marginal probability distribution of variable Y taking the value y, Measures the degree of deviation from independence between X and Y. If X and Y are independent, then , mutual information is 0. Mutual information measures the amount of information shared between two variables. The larger the value, the stronger the correlation between the two variables.

[0107] The calculation expression of covariance is: , where is the covariance between the random variables X and Y, indicating how the two vary together, is the i-th data point in data block X, is the i-th data point in data block Y, is the mean of data block X, is the mean of the data block Y. Covariance is used to measure the strength of the relationship between two data blocks, indicating how the two variables change together. If the covariance is positive, it means that the two data blocks grow or decrease synchronously; if it is negative, it means that the two data blocks have opposite change trends.

[0108] Suppose we have 5 data blocks, their popularity (frequency of use), relevance and other characteristics are as follows:

[0109] Data block A: cold data, stored in a certain storage location.

[0110] Data block B: cold data, stored in another storage location.

[0111] Data block C: cold data, stored in a different location.

[0112] Data block D: hot data, stored in a storage location with high access frequency.

[0113] Data block E: hot data, stored in another storage location with high access frequency.

[0114] We perform similar calculations on the mutual information and covariance between other data blocks, and finally obtain Table 1:

[0115] Table 1 of correlation coefficients between data blocks

[0116] The correlation coefficient is the sum of the normalized mutual information and covariance, which indicates the correlation between two data blocks. The larger the correlation coefficient, the stronger the correlation between them. For example, the correlation coefficient between B and D is 1.3, indicating that there is a strong correlation between the two data blocks; while the correlation coefficient between A and D is 0.28, indicating that there is a weak correlation between them.

[0117] Based on the correlation coefficient, we can decide how to optimize the migration strategy of data blocks. Specifically:

[0118] It is determined that the data block pairs with correlation coefficients greater than or equal to the correlation threshold are strongly correlated, and the data block pairs with correlation coefficients less than the correlation threshold are weakly correlated. Assuming the correlation threshold is 0.9, the data block pairs with strong correlation are: A and B, B and D, B and E, and D and E. Since data blocks D and E are hot data, and B and D, B and E are strongly correlated, data block B is converted to a data block that does not need to be migrated. In addition, since A and B are highly correlated, in order to avoid data loss when reading hot data, data block A is also converted to a data block that does not need to be migrated.

[0119] Data blocks with a correlation coefficient less than 0.9 have a weak correlation and should be stored separately. Assuming the correlation threshold is 0.7, if we only look at cold data, B and C have a strong correlation, so we can migrate cold data blocks B and C to the same storage location.

[0120] In addition to the above examples, this application also has a variety of migration situations. In actual applications, there are a large number of data blocks, so we will not introduce them one by one.

[0121] Combining the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated, including the following steps:

[0122] In this application, the comparison result between the new cluster center value and the first gradient threshold is used to determine whether the data block in the cluster belongs to cold data or hot data. If the new cluster center value is greater than or equal to the first gradient threshold, the data block in the cluster is determined to be hot data. If the new cluster center value is less than the first gradient threshold, the data block in the cluster is determined to be cold data.

[0123] When it is determined that the data blocks in a cluster belong to cold data, it indicates that data migration is required for all data blocks in the cluster. After repeated data processing and association analysis are performed on the data blocks in the cluster that needs data migration, the new cluster center value of each cluster that needs data migration is recalculated, and then the new cluster center value is compared with the second gradient threshold. The second gradient threshold is less than the first gradient threshold. The second gradient threshold is used to perform secondary division of the cold data blocks. If the new cluster center value of the cluster is less than the first gradient threshold, and the new cluster center value is greater than or equal to the second gradient threshold, all data blocks in the cluster are divided into medium-frequency fluctuation data. If the new cluster center value of the cluster is less than the second gradient threshold, all data blocks in the cluster are divided into low-frequency cold data. After all clusters that need data migration are divided, two major categories are obtained, namely medium-frequency fluctuation data clusters and low-frequency cold data clusters.

[0124] The cloud adjusts the storage location according to the hotness or coldness of the data (for example, storing cold data in long-term storage such as S3 Glacier), sets data access policies, and if a data block becomes hot, it can be migrated from low-speed storage back to high-speed storage.

[0125] Phase 1: Edge Storage Caching

[0126] First, store the intermediate frequency data blocks in the edge server as a temporary storage buffer to reduce the direct writing pressure on the cloud.

[0127] Phase 2: Mass migration to the cloud

[0128] Batch transfer of low-frequency cold data to avoid network overhead caused by frequent migration of small files.

[0129] Phase 3: Access Pattern Adjustment

[0130] If some cold data becomes hot, the edge node can quickly restore access without having to re-download the entire data from the cloud.

[0131] Phase 4: Access Pattern Adjustment

[0132] After the migration is complete, clean up duplicate or redundant data in the source storage to optimize storage usage.

[0133] This application predicts the state space of data blocks by using a time series analysis algorithm, and automatically classifies the hotness and coldness of data blocks based on a clustering algorithm. After identifying duplicate data blocks in clustering, it optimizes the amount of data to be migrated by combining data deduplication technology, performs association analysis on the data blocks to be migrated by the Louvain algorithm, optimizes the amount of data to be migrated again, and combines the data block optimization results and the hotness and coldness clustering results to perform multi-level migration processing on the data blocks to be migrated using an edge-cloud collaborative grading method. After the migration system automatically classifies the hotness and coldness of data blocks, it can further identify duplicate data blocks during the clustering process, and combine data deduplication technology to reduce invalid data migration, improve storage utilization, and dynamically adjust the allocation method of storage nodes to make data storage resources more balanced, avoid the problem of excessive storage pressure at a single point, and improve the overall storage efficiency of the system.

[0134] This application optimizes the amount of data to be migrated by combining data deduplication technology, and performs correlation analysis on the data blocks to be migrated through the Louvain algorithm to optimize the amount of data to be migrated. It can mine data sets with strong correlations, thereby ensuring the integrity of related data blocks during migration, reducing data access overhead, and optimizing subsequent data query performance.

[0135] Example 2: The intelligent cold data migration system described in this embodiment includes a data block hot and cold clustering module, a data deduplication module, a secondary optimization module, and a multi-level migration module;

[0136] Data block hot and cold clustering module: uses time series analysis algorithm to predict the state space of data blocks, and automatically classifies the hot and cold degrees of data blocks based on clustering algorithm. The clustering results are sent to the data deduplication module and multi-level migration module;

[0137] Data deduplication module: After identifying duplicate data blocks in the cluster, the amount of migrated data is optimized by combining data deduplication technology, and the primary optimization result is sent to the secondary optimization module;

[0138] Secondary optimization module: Use the Louvain algorithm to perform correlation analysis on the data blocks that need to be migrated, secondary optimize the amount of migrated data, and send the secondary optimization results to the multi-level migration module;

[0139] Multi-level migration module: Combined with the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated.

[0140] The above formulas are dimensionless and numerical calculations. The formula is a formula obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0141] It should be understood that the term "and / or" in this article is only a description of the association relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the related objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0142] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0143] Ordinary technicians in this field can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0144] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. Intelligent cold data migration method, characterized by: The migration method comprises the following steps: The migration system uses a time series analysis algorithm to predict the state space of data blocks and automatically classifies the hotness and coldness of data blocks based on a clustering algorithm; After identifying duplicate data blocks in the cluster, the amount of data to be migrated is optimized by combining data deduplication technology, and the Louvain algorithm is used to perform association analysis on the data blocks to be migrated, and the amount of data to be migrated is optimized again; Combining the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated.

2. The intelligent cold data migration method according to claim 1, characterized in that: Combining the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated, including the following steps: Obtain all clusters that require data migration, perform duplicate data processing and association analysis on the data blocks in the clusters that require data migration, recalculate the new cluster center value of each cluster that requires data migration, and divide all data blocks in all clusters into medium-frequency fluctuation data or low-frequency cold data based on the new cluster center value; The multi-level migration process includes: S1: Edge Storage Cache The IF data blocks are stored in the edge server as a temporary storage buffer; S2: Batch migration to the cloud Batch transfer of low-frequency cold data to the cloud; S3: Access pattern tuning If the cold data becomes hot, the edge node resumes access; S4: Access pattern adjustment After the migration is complete, clean up duplicate or redundant data in the source storage.

3. The intelligent cold data migration method according to claim 2, characterized in that: Based on the new cluster center value, all data blocks in all clusters are divided into medium-frequency fluctuation data or low-frequency cold data, including the following steps: By comparing the new cluster center value with the first gradient threshold, it is determined whether the data block in the cluster belongs to cold data or hot data. If the new cluster center value is greater than or equal to the first gradient threshold, it is determined that the data block in the cluster belongs to hot data. If the new cluster center value is less than the first gradient threshold, it is determined that the data block in the cluster belongs to cold data. The new cluster center value is compared with the second gradient threshold, the second gradient threshold is less than the first gradient threshold, and the second gradient threshold is used to perform secondary division on the cold data blocks. If the new cluster center value of the cluster is less than the first gradient threshold, and the new cluster center value is greater than or equal to the second gradient threshold, all data blocks in the cluster are divided into medium-frequency fluctuation data. If the new cluster center value of the cluster is less than the second gradient threshold, all data blocks in the cluster are divided into low-frequency cold data.

4. The intelligent cold data migration method according to claim 3, characterized in that: Automatically classify the hotness and coldness of data blocks based on clustering algorithms, including the following steps: After obtaining the state factors of all data blocks, the elbow method is first used to calculate the sum of squares of intra-cluster errors under different K values, and the inflection point is selected as the optimal K value; Randomly select the state factors of K data blocks as the initial cluster centers of K clusters, calculate the difference between the state factors of the data blocks and the initial cluster centers of K clusters, and assign the data blocks to the cluster with the smallest difference. When all data blocks are assigned, calculate the mean state factor of the cluster as the new cluster center. Repeat the steps of data block allocation and new cluster center calculation until the new cluster centers of the K clusters converge and the clustering process is completed. Sort the K clusters from large to small according to the final new cluster centers to obtain a cluster list.

5. The intelligent cold data migration method according to claim 4, characterized in that: The migration system uses a time series analysis algorithm to predict the state space of a data block, including the following steps: Obtain the time series data of each data block from the storage system. The time series data includes access time, access count, access interval, and read-write ratio. Remove abnormal access records and fill in missing data, and normalize the time series data. Divide the access records of data blocks according to fixed time intervals, construct time series, calculate the time series data change trend of each data block in different time windows, and use Fourier transform to identify the periodic pattern of data access; The deep neural network is trained using the long short-term memory network to learn the temporal relationship of data block access and generate state factors for the data blocks after predicting future access trends.

6. The intelligent cold data migration method according to claim 5, characterized in that: After identifying duplicate data blocks in the cluster, we use data deduplication technology to optimize the amount of data to be migrated, including the following steps: In the cluster list, the cluster whose new cluster center value is less than the first gradient threshold is marked as the optimized cluster; In all optimized clusters, a repetition factor between two data blocks is obtained, and two data blocks whose repetition factor is greater than a preset repetition threshold are marked as repetitive data blocks; An incremental storage method is used to deduplicate duplicate data blocks, and only the differences between data blocks are stored. If the hash values ​​of multiple data blocks are consistent, only one copy of the data block is stored.

7. The intelligent cold data migration method according to claim 6, characterized in that: The calculation logic of the repetition factor is as follows: a fixed-length hash value is calculated for each data block in the optimized cluster, LSH is used for similarity matching to identify the content similarity of the data block, and after obtaining the hash value difference and content similarity between two data blocks, the hash value difference and content similarity are normalized so that the value range of the hash value difference and content similarity is mapped to [0, 1], and the repetition factor is obtained by subtracting the hash value difference from the normalized content similarity.

8. The intelligent cold data migration method according to claim 7, characterized in that: The Louvain algorithm is used to perform correlation analysis on the data blocks to be migrated, and the amount of data to be migrated is optimized again, including the following steps: All stored data blocks are defined as data blocks that need to be migrated, and first, based on the correlation between the data blocks that need to be migrated and the data blocks that do not need to be migrated, it is determined whether the data blocks that need to be migrated need to be converted into data blocks that do not need to be migrated; The Louvain algorithm treats each data block that needs to be migrated as a community, traverses the data blocks, calculates the modularity gain after joining the adjacent community, selects the community with the largest modularity gain to merge, and iterates until the modularity gain no longer increases; The communities formed in the first stage are regarded as new nodes, the new association network is recalculated, and community merging is continued until the global optimal partition is reached.

9. The intelligent cold data migration method according to claim 8, characterized in that: The clustering process is completed until the new cluster centers of the K clusters converge, including the following steps: If the change of the new cluster centers of K clusters in two consecutive iterations is less than the set change threshold, the new cluster centers are considered to converge, and the expression is: , where represents the new cluster center of cluster i in the tth iteration, represents the new cluster center of cluster i in the t-1th iteration, is the change threshold, is the number of clusters.

10. An intelligent cold data migration system, used to implement the migration method according to any one of claims 1 to 9, characterized in that: Including data block hot and cold clustering module, data deduplication module, secondary optimization module, and multi-level migration module; Data block hot and cold clustering module: uses time series analysis algorithm to predict the state space of data blocks, and automatically classifies the hot and cold degrees of data blocks based on clustering algorithm; Data deduplication module: After identifying duplicate data blocks in the cluster, it combines data deduplication technology to optimize the amount of migrated data; Secondary optimization module: performs correlation analysis on the data blocks to be migrated through the Louvain algorithm, and secondary optimizes the amount of migrated data; Multi-level migration module: Combined with the data block optimization results and the hot and cold degree clustering results, the edge-cloud collaborative grading method is used to perform multi-level migration processing on the data blocks that need to be migrated.

Citation Information

Patent Citations

  • Layered distributed cloud computing architecture and service delivery method

    CN101977242A

  • Image clustering method, device, computer equipment and readable storage medium

    CN113963221A

  • Data migration method and device, equipment and storage medium

    CN116974470A

  • Data standardization management method and system

    CN117171243A

  • Super-multi-objective optimization method and system based on knowledge migration

    CN118036329A

Cited By

  • Data storage method and system and electronic equipment

    CN120508260A

  • Data storage method, system and electronic device

    CN120508260B

  • Mass monitoring data-oriented edge computing storage optimization algorithm and system

    CN120610657A