Artificial intelligence-based data hot and cold layering adaptive storage optimization system and method

CN119127066BActive Publication Date: 2026-08-11北京雅弘商业管理有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]传统的数据存储系统在应对这一挑战时暴露出了诸多不足,首先,往往片面地将数据访问频率作为唯一的冷热划分标准,忽视了数据本身其他重要属性如新颖性、隐私敏感性等,导致冷热分层存储缺乏精准性,这种"一刀切"的存储策略缺乏灵活性和针对性,无法充分适应不同数据特征带来的差异化需求

Benefits of technology

[0054] By constructing an advanced cold/hot status prediction model, the cold/hot status of stored data can be accurately predicted in the future, laying the foundation for precise cold/hot tiered storage of data and thus significantly improving the utilization efficiency of storage resources. More ingeniously, a new evaluation index, the data novelty index, is introduced, which comprehensively considers multiple dimensions such as data collection time, similarity with other data, and data diversity. This allows for a more accurate and comprehensive assessment of the importance and value of the data, providing a more precise reference for judging the cold/hot status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119127066B_ABST
    Figure CN119127066B_ABST
Patent Text Reader

Abstract

This invention discloses an artificial intelligence-based data hot / cold tiered adaptive storage optimization system and method. The method includes: collecting storage time-series information of each piece of stored data in a storage device; inputting the storage time-series information into a hot / cold state prediction model to predict the hot / cold state of the stored data in the future; collecting the access frequency of the stored data within a unit period; establishing an access frequency analysis set based on the access frequency within the unit period and extracting access frequency features; inputting the access frequency features into a staged data identification model to obtain staged judgment results, which include yes or no; determining whether to migrate the stored data based on the staged judgment results and the predicted hot / cold state of the stored data in the future, and if not to migrate, determining whether to back up the data; and rationally formulating data migration or non-migration decisions to minimize frequent and unnecessary data migrations, thereby reducing storage operation and maintenance costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data storage technology, and more specifically to an artificial intelligence-based adaptive storage optimization system and method for hot and cold data tiering. Background Technology

[0002] With the rapid development of information technology and the deepening of digital transformation, the amount of data is experiencing an unprecedented explosive growth. The storage and efficient management of massive amounts of data has become a major challenge that needs to be addressed in the field of information technology.

[0003] Traditional data storage systems have revealed many shortcomings in addressing this challenge. First, they often unilaterally use data access frequency as the sole criterion for classifying data as hot or cold, ignoring other important attributes of the data itself, such as novelty and privacy sensitivity. This results in a lack of precision in hot and cold tiered storage. This "one-size-fits-all" storage strategy lacks flexibility and specificity, and cannot fully adapt to the differentiated needs brought about by different data characteristics.

[0004] Secondly, existing technologies often overlook the potential for periodic fluctuations in data access patterns when making data migration decisions. The access frequency of certain data may fluctuate periodically over time, but traditional methods are weak in identifying this, making it easy to make inappropriate and frequent migration decisions for such data, resulting in wasted storage resources and low efficiency.

[0005] More importantly, existing methods exhibit significant blindness in data migration, lacking a comprehensive assessment and weighing of the negative impacts of frequent and unnecessary migrations. Frequent data migrations not only consume substantial storage and network resources, impacting overall storage performance, but also significantly increase operational costs. Simultaneously, frequent data transmission increases the risk of data theft or tampering during transmission, posing a threat to data security and integrity. However, existing technologies do not adequately address these negative impacts, easily falling into the trap of blindly and frequently migrating data.

[0006] In summary, traditional data storage systems have significant shortcomings in addressing the challenges of the big data era, including cold / hot stratification criteria, identification of periodic fluctuations, and migration decisions. They also lack intelligent management capabilities that can adaptively adjust storage strategies based on data characteristics, posing numerous challenges to storage efficiency, cost, and security. New and innovative solutions are urgently needed to optimize and improve these systems.

[0007] Therefore, proposing an adaptive storage optimization system and method based on artificial intelligence for hot and cold data tiering to solve the difficulties of existing technologies is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0008] The purpose of this invention is to provide an adaptive storage optimization system and method for hot and cold data tiering based on artificial intelligence, in order to address the shortcomings of the prior art.

[0009] To achieve the above objectives, the present invention provides the following technical solution: an artificial intelligence-based adaptive storage optimization method for hot and cold data tiering, comprising:

[0010] The storage time sequence information of each piece of stored data in the storage device is collected. The storage time sequence information includes access frequency, privacy type, data novelty index and hot / cold status, and the hot / cold status includes cold status and hot status.

[0011] The storage time series information is input into the hot and cold status prediction model to predict the hot and cold status of the stored data in the future.

[0012] The frequency of access to collected and stored data within a unit period;

[0013] Establish an access frequency analysis set within a unit period and extract access frequency features;

[0014] The access frequency feature is input into the phased data identification model to obtain the phased judgment result, which includes yes or no.

[0015] Based on the phased judgment results of the stored data and the predicted hot / cold status of the stored data in the future, it is determined whether to migrate the stored data; if not to migrate, it is determined whether to back it up.

[0016] Furthermore, the method for determining whether to migrate the stored data includes:

[0017] If the stored data is currently stored on a hot data storage medium, and the thermal state of the stored data is predicted in the future, and the phased judgment result is yes or no, then the stored data will not be migrated.

[0018] If the stored data is currently stored on a hot data storage medium, and the predicted cold state of the stored data in the future is determined in a phased manner, then the stored data will not be migrated.

[0019] If the stored data is currently stored on a hot data storage medium, and the predicted cold state of the stored data in the future is determined in stages, and the result is negative, then the stored data will be migrated to a cold data storage medium.

[0020] If the stored data is currently stored on a cold data storage medium, and the future cold state of the stored data is predicted, if the phased judgment result is no or yes, then the stored data will not be migrated.

[0021] If the stored data is currently stored on a cold data storage medium, and the predicted hot state of the stored data in the future is determined in a phased manner, then the stored data will be migrated to a hot data storage medium.

[0022] Furthermore, methods for determining whether to perform a backup if migration is not performed include:

[0023] If the stored data is currently stored on a cold data storage medium, and the future hot state of the stored data is predicted, if the phased judgment result is yes, then the stored data is backed up to a hot data storage medium, and the access interface for the stored data is switched to the hot data storage medium; when the next prediction of the future cold state of the stored data is made, the backup of the stored data on the hot data storage medium is deleted.

[0024] Furthermore, the data novelty index is generated based on a comprehensive analysis of the reciprocal of the difference between the data collection time and the current time, the mean similarity between the stored data and other data, and the data type diversity index of the stored data.

[0025] Furthermore, the storage data type diversity index uses entropy to measure the diversity of categories in stored data.

[0026] Furthermore, the data novelty index is calculated as follows:

[0027] NS=ω1×TL+ω2×DT+ω2×VY;

[0028] TL is the reciprocal of the difference between the data acquisition time and the current time; DT is the mean similarity between the stored data and other data; VY is the data type diversity index; ω1, ω2, and ω3 are weight coefficients, which are non-negative numbers, and the sum of the weight coefficients is 1.

[0029] The TL formula is expressed as follows:

[0030] T1 represents the data acquisition time, and T2 represents the current time.

[0031] The DT formula is expressed as follows:

[0032] sx(A, B) i () represents the similarity score between the stored data and the i-th data among other data, where n is the number of other data;

[0033] In the formula a r Let r be the r-th element of data A. For other data B i The r-th element in the array, sx(A, B) i() represents the Euclidean distance between stored data and other data;

[0034] Furthermore, the formula for calculating VY is as follows:

[0035]

[0036] In the formula, p g It represents the probability of category g appearing in the stored data, where G is the total number of categories contained in the stored data.

[0037] Furthermore, the method for constructing the cold and hot state prediction model includes:

[0038] The prediction time step T, sliding step P2, and sliding window length P1 are preset; the stored time series information is transformed into multiple training samples using the sliding window method. Each training sample corresponds to a label and constitutes a set of training data. P1 and P2 are both integers greater than 1.

[0039] The training data is used as the input to the hot and cold state prediction model, and the hot and cold state of the stored data after the predicted time step T is used as the output. The hot and cold state of the actual stored data at the future time after each time step T is used as the prediction target, and the prediction accuracy is used as the training target. The hot and cold state prediction model is trained to obtain a hot and cold state prediction model that meets the prediction accuracy. The hot and cold state prediction model is a temporal convolutional network, a long short-term memory network, or a gated recurrent unit.

[0040] Furthermore, the training methods for the staged data recognition model include:

[0041] Collect V sets of identification data, where V is an integer greater than 1. The identification data includes access frequency features and the corresponding stage judgment results of the access frequency features. The access frequency features include access frequency change time series data, frequency mean and frequency standard deviation.

[0042] The vector of each set of identification data is used as the input of the stage data identification model. The stage data identification model outputs the stage judgment result corresponding to each set of access frequency features and uses the actual stage judgment result corresponding to each set of access frequency features as the prediction target. The training objective is to minimize the sum of prediction errors of all stage judgment results corresponding to all access frequency features. The stage data identification model is trained until the sum of prediction errors converges and then training stops. The stage data identification model is a regression prediction model.

[0043] Furthermore, the method for obtaining the access frequency change time-series data includes:

[0044] The difference between adjacent access frequencies in the access frequency analysis set is divided by the access frequency with the earlier acquisition time to obtain the access frequency change. This process is repeated to calculate the access frequency changes of adjacent access frequencies in turn. All access frequency changes constitute the access frequency change time series data.

[0045] Furthermore, the frequency mean is the mean of the access frequencies within the access frequency analysis set; the frequency standard deviation is the standard deviation of the access frequencies within the access frequency analysis set.

[0046] An AI-based adaptive storage optimization system for hot and cold data tiering, implementing the AI-based adaptive storage optimization method for hot and cold data tiering, includes the following components:

[0047] The first acquisition module is used to acquire the storage time sequence information of each stored data in the storage device. The storage time sequence information includes access frequency, privacy type, data novelty index and hot / cold status, and the hot / cold status includes cold status and hot status.

[0048] The first analysis module is used to input the storage time series information into the hot and cold status prediction model to predict the hot and cold status of the stored data in the future.

[0049] The second acquisition module is used to collect the access frequency of the stored data within a unit period.

[0050] The feature extraction module is used to establish an access frequency analysis set based on the access frequency within a unit period and extract access frequency features.

[0051] The second analysis module inputs the access frequency characteristics into the phased data identification model to obtain the phased judgment results, which include yes or no.

[0052] The comprehensive analysis module determines whether to migrate the stored data based on the phased judgment results of the stored data and the predicted hot / cold status of the stored data in the future. If migration is not required, it determines whether to back up the data.

[0053] The technical effects and advantages of the data cold and hot tiered adaptive storage optimization system and method based on artificial intelligence provided by this invention are as follows:

[0054] By constructing an advanced cold / hot status prediction model, the cold / hot status of stored data can be accurately predicted in the future, laying the foundation for precise cold / hot tiered storage of data and thus significantly improving the utilization efficiency of storage resources. More ingeniously, a new evaluation index, the data novelty index, is introduced, which comprehensively considers multiple dimensions such as data collection time, similarity with other data, and data diversity. This allows for a more accurate and comprehensive assessment of the importance and value of the data, providing a more precise reference for judging the cold / hot status.

[0055] Secondly, by analyzing the temporal characteristics of access frequency, an efficient staged data identification model is constructed, which can accurately identify the staged fluctuation characteristics in data access patterns, thereby avoiding making inappropriate frequent migration decisions for data with periodic access frequency fluctuations.

[0056] By comprehensively predicting future hot and cold data conditions and making phased judgments, we can make reasonable decisions on whether or not to migrate data, minimizing frequent and unnecessary data migrations, thereby reducing storage operation and maintenance costs and ensuring data transmission security and consistency.

[0057] In summary, the innovative solution proposed in this invention breaks through the traditional "one-size-fits-all" storage strategy, realizes intelligent data storage management, and can adaptively adjust the storage strategy to adapt to different data characteristics, thereby maximizing the utilization efficiency of storage resources. It is an important innovative achievement in data storage optimization in the era of big data. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0059] Figure 1 This is a schematic diagram of the AI-based adaptive cold and hot data tiered storage optimization system of the present invention;

[0060] Figure 2 This is a flowchart of the AI-based adaptive cold and hot data tiered storage optimization method of the present invention. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] It should be noted that when a component is said to be "fixed to" another component, it can be directly attached to the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component.

[0063] Example 1

[0064] like Figure 1 As shown, this embodiment of the data cold and hot tiered adaptive storage optimization system based on artificial intelligence includes a wired and / or wireless connection.

[0065] The first acquisition module is used to acquire the storage time sequence information of each piece of stored data in the storage device. The storage time sequence information includes access frequency, privacy type, data novelty index and hot / cold status, and the hot / cold status includes cold status and hot status.

[0066] Hot data corresponds to high access frequency, and vice versa. Privacy types include privacy categories (non-public data) and non-privacy categories (public data). Privacy category data is accessed less frequently because it requires additional security measures and access control, and is therefore classified as cold data. Non-privacy category data is widely shared and used, and its status as hot data needs to be determined based on access frequency, privacy type, data growth rate, and data novelty.

[0067] Data sets with a high novelty index, i.e., recently generated or updated datasets, are more likely to become hot data because they may contain new information that is valuable to decision-makers. Whether a dataset is hot data should be judged based on access frequency, privacy type, and data growth rate. Relatively old or outdated datasets may be more likely to become cold data because they have lost their timeliness.

[0068] The data novelty index is generated based on a comprehensive analysis of the reciprocal of the difference between the data collection time and the current time, the mean similarity between the stored data and other data, and the data type diversity index. The calculation method for the data novelty index is as follows:

[0069] NS=ω1×TL+ω2×DT+ω2×VY;

[0070] TL is the reciprocal of the difference between the data acquisition time and the current time. The TL value is positively correlated with the NS value. DT is the mean similarity between the stored data and other data (referring to stored data other than the current stored data). The DT value is positively correlated with the NS value. VY is the data type diversity index of the stored data. The VY value is positively correlated with the NS value. ω1, ω2, and ω3 are weight coefficients, which are non-negative and the sum of the weight coefficients is 1. They are used to balance the influence of each parameter on the novelty score.

[0071] The TL formula is expressed as follows:

[0072] T1 represents the data acquisition time, and T2 represents the current time.

[0073] The DT formula is expressed as follows:

[0074] sx(A, B) iThe similarity score between the stored data and the i-th data in other data is represented by , where n is the number of other data. When calculating the similarity score, the stored data and other data need to be converted into vectors, i.e., vectorization. This can be done through natural language processing (NLP), such as bag-of-words model, TF-IDF (Term Frequency-Inverse Document Frequency), one-hot encoding, etc.

[0075] In the formula a r Let r be the r-th element of data A. For other data B i The r-th element in the array. sx(A, B) i The value represents the Euclidean distance between stored data and other data. The smaller the value, the closer the stored data is to other data, and vice versa. The similarity between stored data and other data can also be evaluated by using Jacarbe similarity and cosine similarity.

[0076] VY uses entropy to measure the diversity of categories in stored data, calculated as follows:

[0077]

[0078] In the formula, p g VY (entropy) is the probability of category g appearing in the stored data, where G is the total number of categories in the stored data. The higher the VY (entropy), the greater the diversity and the stronger the data novelty. If all elements in the stored data belong to the same category, the entropy will be 0, indicating no diversity. If the category distribution is very uneven, such as most elements belonging to a few categories, the entropy value will be low, indicating low diversity, and vice versa.

[0079] Examples of categories are as follows:

[0080] Natural language processing technology can be used to classify stored data into categories such as biological classification (kingdom, phylum, class, order, family, genus, species), software application classification (operating system, office software, image editing, video editing, audio production, games), data type classification (quantitative data, qualitative data, text data, image data, audio data, video data), and economic industry classification (agriculture and manufacturing, service industry, finance, information technology), etc.

[0081] The first analysis module is used to input the storage time series information into the hot and cold status prediction model to predict the hot and cold status of the stored data in the future.

[0082] Methods for constructing cold and hot state prediction models include:

[0083] The prediction time step T, sliding step P2, and sliding window length P1 are preset; the stored time series information is transformed into multiple training samples using the sliding window method. Each training sample corresponds to a label and constitutes a set of training data. P1 and P2 are both integers greater than 1.

[0084] The first training sample is X1, X1 = [T1,PV1,YS1,NS1], and the hot / cold state corresponding to X1 is Y1; T1 is the timestamp in the first training sample, PV1 is the access frequency in the first training sample, YS1 is the privacy type in the first training sample, and NS1 is the data novelty index in the first training sample.

[0085] The training data is used as the input to the hot and cold state prediction model, and the hot and cold state of the stored data after the predicted time step T is used as the output. The hot and cold state of the actual stored data at the future time after each time step T is used as the prediction target, and the prediction accuracy is used as the training target. The hot and cold state prediction model is trained to obtain a hot and cold state prediction model that meets the prediction accuracy, that is, the hot and cold state prediction model is built. The hot and cold state prediction model can be an autoregressive integral moving average model, a temporal convolutional network (TCN), a long short-term memory network (LSTM), or a gated recurrent unit (GRU).

[0086] The second acquisition module is used to collect the access frequency of the stored data within a unit period.

[0087] The feature extraction module is used to establish an access frequency analysis set for access frequencies within a unit period and extract access frequency features, which include access frequency change time series data, frequency mean and frequency standard deviation.

[0088] Methods for obtaining access frequency variation time-series data include:

[0089] Dividing the difference between adjacent access frequencies in the access frequency analysis set by the access frequency collected earlier yields the access frequency change. This process is repeated to calculate the access frequency changes of adjacent access frequencies in turn. All access frequency changes constitute access frequency change time series data. Access frequency change time series data reflects the volatility or periodicity of data access. Stage-specific access changes mean that the access frequency is significantly different in different time periods, which can be identified and quantified through access frequency change time series data.

[0090] The frequency mean is the average access frequency within the access frequency analysis set; the frequency standard deviation is the standard deviation of the access frequency within the access frequency analysis set. The frequency mean provides the average level of access frequency and is a measure of the population, but it does not provide information about the temporal distribution of access frequency changes. The frequency standard deviation measures the fluctuation or dispersion of access frequency and reflects the consistency of data access. A high standard deviation indicates that the access frequency has large fluctuations and may have different access phases or cycles. The standard deviation is directly related to the data of phased access changes and can reveal the magnitude and frequency of access frequency changes.

[0091] Access frequency change time series data, frequency mean, and frequency standard deviation are complementary measures that together provide a comprehensive understanding of data access patterns. They help identify and analyze phased access changes, and through these measures, we can better understand the dynamic characteristics of data access and make more accurate data migration decisions.

[0092] The second analysis module inputs the access frequency characteristics into the phased data identification model to obtain phased judgment results, which include yes or no.

[0093] Training methods for phased data recognition models include:

[0094] Collect V sets of identification data, where V is an integer greater than 1. The identification data includes access frequency features and the corresponding stage-by-stage judgment results for the access frequency features.

[0095] The vector of each set of identification data is used as the input of the stage data identification model. The stage data identification model outputs the stage judgment result corresponding to each set of access frequency features and uses the actual stage judgment result corresponding to each set of access frequency features as the prediction target. The training objective is to minimize the sum of prediction errors of all stage judgment results corresponding to all access frequency features. The stage data identification model is trained until the sum of prediction errors converges and then training stops. The stage data identification model can be a regression prediction model or other suitable machine learning model, which is not specifically limited here.

[0096] The formula for calculating the prediction error is Z. V =(δ V -μ V ) 2 Z V The prediction error is represented by V, where V is the group number of the feature vector of the access frequency feature, and δ is the value of δ. V For the actual stage judgment result corresponding to the access frequency characteristics of the Vth group, μ V The interim judgment result is predicted for the access frequency characteristics of group V.

[0097] The comprehensive analysis module determines whether to migrate the stored data based on the phased judgment results of the stored data and the predicted hot / cold status of the stored data in the future. If migration is not required, it determines whether to back up the data.

[0098] The method for determining whether to migrate the stored data includes:

[0099] If the stored data is currently stored on a hot data storage medium, and the future thermal state of the stored data is predicted, with a phased judgment result of yes or no, then the stored data will not be migrated. Cold data storage media include HDDs, and hot data storage media include SSDs.

[0100] If the stored data is currently stored on a hot data storage medium, and the predicted future cold state of the stored data is determined by a phased judgment result, it indicates that the stored data is undergoing phased changes, meaning that the access frequency varies significantly over time. In the next predicted future time, the data is highly likely to become hot again. Therefore, the stored data will not be migrated. This avoids the following risks associated with frequent data migration: performance degradation of storage devices (because data needs to be copied and transmitted during migration, consuming network bandwidth and storage resources), increased storage costs (data migration involves the use of storage devices, network bandwidth, and migration tools; frequent migration does not bring the expected cost savings but instead increases overall expenses due to the cost of migration itself), risks to data consistency and integrity (data faces consistency and integrity risks during migration, especially when failures or interruptions occur), data security issues (data is exposed to insecure network environments during migration, increasing the risk of data leakage or tampering), and increased management complexity (frequent data migration increases the complexity of storage management, requiring more management work to track data location and status), etc.

[0101] If the stored data is currently stored on a hot data storage medium, and the prediction of the cold state of the stored data in the future is negative in the phased judgment, it means that the stored data is not data that changes periodically, that is, the access frequency changes with the period and the change is small. In this case, the stored data will be migrated to a cold data storage medium.

[0102] If the stored data is currently stored on a cold data storage medium, and the predicted cold state of the stored data in the future is determined in stages (either yes or no), then the stored data will not be migrated.

[0103] If the stored data is currently stored on a cold data storage medium, and the predicted hot state of the stored data in the future is determined in a phased manner, then the stored data will be migrated to a hot data storage medium.

[0104] If the stored data is currently stored on a cold data storage medium, and the predicted hot state of the stored data in the future is determined in a phased manner, then the stored data is backed up to a hot data storage medium, and the storage data access interface is switched to the hot data storage medium. When the next predicted cold state of the stored data in the future is reached, the backup of the stored data on the hot data storage medium is deleted, and the storage data access interface is switched back to the cold data storage medium. This avoids frequent migration of such stored data whose access frequency varies greatly with the period, thus preventing adverse effects, and also meets the user's access needs for the stored data.

[0105] This embodiment proposes an innovative AI-based data hot and cold tiered adaptive storage optimization system, which aims to address the challenges faced by traditional data storage systems.

[0106] This system, by constructing an advanced cold / hot status prediction model, can accurately predict the cold / hot status of stored data within a future time period, laying the foundation for precise cold / hot tiered storage of data and thus significantly improving the utilization efficiency of storage resources. More ingeniously, it introduces a brand-new evaluation indicator, the data novelty index, which comprehensively considers multiple dimensions such as data collection time, similarity with other data, and data diversity. This allows for a more accurate and comprehensive assessment of the importance and value of the data, providing a more precise reference for judging the cold / hot status.

[0107] Secondly, by analyzing the temporal characteristics of access frequency, an efficient staged data identification model is constructed, which can accurately identify the staged fluctuation characteristics in data access patterns, thereby avoiding making inappropriate frequent migration decisions for data with periodic access frequency fluctuations.

[0108] By comprehensively predicting future hot and cold data conditions and making phased judgments, we can make reasonable decisions on whether or not to migrate data, minimizing frequent and unnecessary data migrations, thereby reducing storage operation and maintenance costs and ensuring data transmission security and consistency.

[0109] In summary, the innovative solution proposed in this embodiment breaks through the traditional "one-size-fits-all" storage strategy, realizes intelligent data storage management, and can adaptively adjust the storage strategy to adapt to different data characteristics, thereby maximizing the utilization efficiency of storage resources. It is an important innovative achievement in data storage optimization in the era of big data.

[0110] Example 2

[0111] Please see Figure 2 As shown, this embodiment provides an artificial intelligence-based adaptive storage optimization method for hot and cold data tiering, including:

[0112] The storage time sequence information of each piece of stored data in the storage device is collected. The storage time sequence information includes access frequency, privacy type, data novelty index and hot / cold status, and the hot / cold status includes cold status and hot status.

[0113] The storage time series information is input into the hot and cold status prediction model to predict the hot and cold status of the stored data in the future.

[0114] The frequency of access to collected and stored data within a unit period;

[0115] Establish an access frequency analysis set within a unit period and extract access frequency features;

[0116] The access frequency characteristics are input into the phased data identification model to obtain the phased judgment results, which include yes or no.

[0117] Based on the phased judgment results of the stored data and the predicted hot / cold status of the stored data in the future, it is determined whether to migrate the stored data; if not to migrate, it is determined whether to back it up.

[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An artificial intelligence-based data hot and cold layering adaptive storage optimization method, characterized in that, include: The storage time sequence information of each piece of stored data in the storage device is collected. The storage time sequence information includes access frequency, privacy type, data novelty index and hot / cold status. Hot / cold status includes cold status and hot status. Privacy type includes privacy category and non-privacy category. The data novelty index is generated by comprehensive analysis based on the reciprocal of the difference between the storage data collection time and the current time, the average similarity between the stored data and other data, and the storage data type diversity index. The storage time series information is input into the hot and cold status prediction model to predict the hot and cold status of the stored data in the future. The frequency of access to collected and stored data within a unit period; Establish an access frequency analysis set within a unit period and extract access frequency features; The access frequency characteristics are input into the phased data identification model to obtain phased judgment results, which include yes or no, and are used to characterize the phased fluctuation and change characteristics. Based on the phased judgment results of the stored data and the predicted hot / cold status of the stored data in the future, it is determined whether to migrate the stored data; if not to migrate, it is determined whether to back it up.

2. The artificial intelligence-based data hot and cold tiering adaptive storage optimization method according to claim 1, characterized in that, The method for determining whether to migrate the stored data includes: If the stored data is currently stored on a hot data storage medium, and the thermal state of the stored data is predicted in the future, and the phased judgment result is yes or no, then the stored data will not be migrated. If the stored data is currently stored on a hot data storage medium, and the predicted cold state of the stored data in the future is determined in a phased manner, then the stored data will not be migrated. If the stored data is currently stored on a hot data storage medium, and the predicted cold state of the stored data in the future is determined in stages, and the result is negative, then the stored data will be migrated to a cold data storage medium. If the stored data is currently stored on a cold data storage medium, and the future cold state of the stored data is predicted, if the phased judgment result is no or yes, then the stored data will not be migrated. If the stored data is currently stored on a cold data storage medium, and the predicted hot state of the stored data in the future is determined in stages, and the result is negative, then the stored data will be migrated to a hot data storage medium.

3. The artificial intelligence-based data hot and cold tiering adaptive storage optimization method according to claim 2, characterized in that, If a migration is not performed, the methods for determining whether to perform a backup include: If the stored data is currently stored on a cold data storage medium, and the future hot state of the stored data is predicted, if the phased judgment result is yes, then the stored data is backed up to a hot data storage medium, and the storage data access interface is switched to the hot data storage medium; when the next prediction of the future stored data being in a cold state occurs, the backup of the stored data on the hot data storage medium is deleted, and the storage data access interface is switched to the cold data storage medium.

4. The artificial intelligence-based data hot and cold tiering adaptive storage optimization method according to claim 1, characterized in that, The storage data type diversity index uses entropy to measure the diversity of categories in stored data. 5.The AI-based data hot and cold tiering adaptive storage optimization method of claim 2, wherein, The method for constructing the cold and hot state prediction model includes: The prediction time step T, sliding step P2, and sliding window length P1 are preset; the stored time series information is transformed into multiple training samples using the sliding window method. Each training sample corresponds to a label and constitutes a set of training data. P1 and P2 are both integers greater than 1. The training data is used as the input to the hot and cold state prediction model, and the hot and cold state of the stored data after the predicted time step T is used as the output. The hot and cold state of the actual stored data at the future time after each time step T is used as the prediction target, and the prediction accuracy is used as the training target. The hot and cold state prediction model is trained to obtain a hot and cold state prediction model that meets the prediction accuracy. The hot and cold state prediction model is a temporal convolutional network, a long short-term memory network, or a gated recurrent unit.

6. The artificial intelligence-based data hot and cold tiering adaptive storage optimization method according to claim 5, characterized in that, Training methods for phased data identification models include: Collect V sets of identification data, where V is an integer greater than 1. The identification data includes access frequency features and the corresponding stage judgment results of the access frequency features. The access frequency features include access frequency change time series data, frequency mean and frequency standard deviation. The vector of each set of identification data is used as the input of the stage data identification model. The stage data identification model outputs the stage judgment result corresponding to each set of access frequency features and uses the actual stage judgment result corresponding to each set of access frequency features as the prediction target. The training objective is to minimize the sum of prediction errors of all stage judgment results corresponding to all access frequency features. The stage data identification model is trained until the sum of prediction errors converges and then training stops. The stage data identification model is a regression prediction model.

7. The artificial intelligence-based data hot and cold tiering adaptive storage optimization method according to claim 6, characterized in that, The method for obtaining the access frequency change time-series data includes: The difference between adjacent access frequencies in the access frequency analysis set is divided by the access frequency with the earlier acquisition time to obtain the access frequency change. This process is repeated to calculate the access frequency changes of adjacent access frequencies in turn. All access frequency changes constitute the access frequency change time series data. 8.The method of claim 6, wherein, The frequency mean is the mean of the access frequencies within the access frequency analysis set; the frequency standard deviation is the standard deviation of the access frequencies within the access frequency analysis set.

9. The system for data hot and cold layering and adaptive storage optimization based on artificial intelligence, characterized in that, The system for implementing the AI-based adaptive storage optimization method for hot and cold data tiering as described in any one of claims 1-8 comprises: The first acquisition module is used to acquire the storage time sequence information of each piece of stored data in the storage device. The storage time sequence information includes access frequency, privacy type, data novelty index and hot / cold status. The hot / cold status includes cold status and hot status. The privacy type includes privacy category and non-privacy category. The data novelty index is generated by comprehensive analysis based on the reciprocal of the difference between the data acquisition time and the current time, the average similarity between the stored data and other data, and the data type diversity index of the stored data. The first analysis module is used to input the storage time series information into the cold and hot state prediction model to predict the cold and hot state of the stored data in the future. The second acquisition module is used to collect the access frequency of the stored data within a unit period. The feature extraction module is used to establish an access frequency analysis set based on the access frequency within a unit period and extract access frequency features. The second analysis module inputs the access frequency characteristics into the phased data identification model to obtain phased judgment results, which include yes or no, and are used to characterize the phased fluctuation and change characteristics. The comprehensive analysis module determines whether to migrate the stored data based on the phased judgment results of the stored data and the predicted hot / cold status of the stored data in the future. If migration is not required, it determines whether to back up the data.

Citation Information

Patent Citations

  • Kafka cluster data consistency guarantee method based on message heat

    CN107666516A