Data popularity prediction method and device, electronic equipment and storage medium

By obtaining the access log of business data, and dynamically selecting target data characteristics using decision trees and timing models, the problem of difficulty in predicting the change in business data popularity in the existing technology is solved, and flexible and accurate data popularity prediction and storage resource optimization are achieved.

CN120354112APending Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510575323.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to predict the change in popularity of business data flexibly and effectively, and it is difficult to adapt to complex and changeable data access methods, making it difficult to balance storage systems between high performance and low cost.

Method used

By obtaining the access log of business data, extracting data access characteristics, attribute characteristics and storage device characteristics, selecting target data characteristics using the decision tree model, and combining the timing model to predict the popularity tags, dynamically adjusting the migration of data at different storage layers.

Benefits of technology

It realizes flexible and accurate prediction of the popularity of business data, improves storage resource utilization efficiency, reduces latency and system costs, and adapts to complex and changeable data access modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354112A_ABST
    Figure CN120354112A_ABST
Patent Text Reader

Abstract

The invention discloses a data popularity prediction method and device, electronic equipment and a storage medium, and relates to the technical field of storage, at least two decision tree models can be trained by using data features and popularity labels of business data in a preset time period, and target data features used for classifying the popularity labels are selected by using the trained decision tree models; therefore, the data features suitable for distinguishing the business data popularity labels can be dynamically selected. In addition, the time sequence model can be trained by using the target data features of the business data at each time point and the popularity label, so that the trained time sequence model is used for carrying out popularity label prediction on the business data. Therefore, the target data features capable of distinguishing the data popularity in the preset time period can be dynamically selected, meanwhile, the time sequence model can be trained by utilizing the target data features, and the popularity label of the business data can be predicted by utilizing the time sequence model, so that a relatively good data popularity prediction effect can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage technology, and in particular, to a method, device, electronic device and storage medium for predicting data heat. Background Art

[0002] With the rapid growth of data volume, the storage system faces the challenge of how to balance between high performance and low cost. For this purpose, a hierarchical structure is usually set in the storage system according to the performance of the storage medium, and the business data is migrated to different performance storage layers according to the heat of the business data. However, in the related art, the detection of the heat of business data is usually based only on static rules, which not only makes it difficult to predict the heat change of business data, but also difficult to adapt to complex and changeable data access methods. Summary of the Invention

[0003] The present application provides a method, device, electronic device and storage medium for predicting data heat, so as to at least solve the technical problem in the related art that it is difficult to predict the heat of business data.

[0004] To solve the above technical problem, the present application provides a method for predicting data heat, including:[[]]

[0005] Obtaining the access log of business data in a preset time period, and obtaining the data features of business data at each time point from the access log; wherein, the data features include data access features, data attribute features, storage device features and storage device performance features;

[0006] Determining the heat label of business data at each time point according to the data access feature;

[0007] Training at least two decision tree models by using the data features and heat labels, and selecting target data features for classifying the heat labels by using the trained decision tree models;

[0008] Training a time series model by using the target data features and heat labels of business data at each time point;

[0009] Obtaining the latest access log of business data, obtaining the latest value of the target data feature of business data from the latest access log, and inputting the latest value into the trained time series model to obtain the heat label prediction result of business data.

[0010] The present application also provides a device for predicting data heat, including:[[]]

[0011] A feature acquisition module, configured to obtain the access log of business data in a preset time period, and obtain the data features of business data at each time point from the access log; wherein, the data features include data access features, data attribute features, storage device features and storage device performance features;

[0012] A labeling module, configured to determine the popularity labels of business data at each time point according to data access characteristics;

[0013] A data feature selection module, configured to train at least two decision tree models by using data features and popularity labels, and select target data features for classifying popularity labels by using the trained decision tree models;

[0014] A prediction model training module, configured to train a time series model by using the target data features and popularity labels of business data at each time point;

[0015] A prediction module, configured to obtain the latest access log of business data, obtain the latest values of the target data features of business data from the latest access log, and input the latest values into the trained time series model to obtain the prediction result of the popularity label of business data.

[0016] This application also provides an electronic device, including:

[0017] A memory, configured to store a computer program;

[0018] A processor, configured to implement the above data popularity prediction method when executing the computer program.

[0019] This application also provides a non-volatile computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the above data popularity prediction method is implemented.

[0020] The present application can first obtain the access log of the business data in a preset time period, and obtain the data features of the business data at each time point from the access log; wherein the data features include data access features, data attribute features, storage device features and storage device performance features, and these data features contain key feature information that can distinguish the heat of the data. Subsequently, the present application can determine the heat label of the business data at each time point according to the data access features to automatically determine the heat of the business data at each time point. Subsequently, the present application can train at least two decision tree models using the data features and heat labels, and use the trained decision tree model to select the target data features for classifying the heat labels, that is, the present application can use the decision tree model to automatically select the target data features that can distinguish the heat of the data in the preset time period, so as to achieve a better data heat prediction effect. Subsequently, the present application can train the time series model using the target data features and heat labels of the business data at each time point, and then obtain the latest access log of the business data, obtain the latest value of the target data feature of the business data from the latest access log, and input the latest value into the trained time series model to obtain the heat label prediction result of the business data. It can be seen that the present application can dynamically select target data features that can distinguish data heat in a preset time period, and then use the target data features to train a time series model, and use the time series model to predict the heat labels of business data, which can solve the technical problems in related technologies that are difficult to predict the heat changes of business data and difficult to adapt to complex and changeable data access methods, thereby achieving better data heat prediction results and adapting to complex and changeable data access modes.

[0021] The present application also provides a data heat prediction device, an electronic device and a computer-readable storage medium, which have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 A flowchart of a data heat prediction method provided in an embodiment of the present application;

[0024] Figure 2 A schematic diagram of another data heat prediction process provided in an embodiment of the present application;

[0025] Figure 3 A flowchart of data collection and preprocessing provided in an embodiment of the present application;

[0026] Figure 4A schematic diagram of storing data migration provided by an embodiment of the present application;

[0027] Figure 5 A structural block diagram of a data heat prediction device provided by an embodiment of the present application;

[0028] Figure 6 A structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0030] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0031] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0032] With the rapid growth of data volume, storage systems are faced with the challenge of balancing high performance and low cost. To this end, a hierarchical structure is usually set up in the storage system according to the performance of the storage medium, and business data is migrated to storage layers with different performances according to the popularity of the data. For example, a hot data layer, a warm data layer, and a cold data layer can be set up in the storage system to store hot data, warm data, and cold data respectively, where hot data, warm data, and cold data are business data with different access popularities. The hot data layer can be composed of high-performance media (such as SSD, Solid State Drive), which is suitable for fast access. The warm data layer can be composed of medium-performance media (such as HDD, Hard Disk Drive), which is suitable for data that is accessed occasionally. The cold data layer can be composed of low-cost media (such as tapes, cold storage), which is suitable for data that is not accessed for a long time. Therefore, how to quickly and flexibly determine the popularity of business data and then migrate the business data to the corresponding storage layer according to the popularity is the key to realizing hierarchical storage. However, in the related art, the detection of the popularity of business data is usually based only on static rules (such as the size of the data, the creation time, or the number of accesses), which not only makes it difficult to predict the change of the popularity of business data, but also difficult to adapt to complex and changeable data access patterns.

[0033] In view of this, aiming at the technical problem of how to flexibly and effectively predict the popularity of business data, this application can provide a data popularity prediction method, which can dynamically select target data features that can distinguish data popularity, and at the same time can use the target data features to train a model, and can use the model to predict the popularity label of business data, so as to achieve a better data popularity prediction effect.

[0034] For ease of understanding, please refer to Figure 1 , Figure 1 which is a flowchart of a data popularity prediction method provided by an embodiment of this application. This method can include:

[0035] S101. Obtain the access log of business data in a preset time period, and obtain the data features of the business data at each time point from the access log; wherein, the data features include data access features, data attribute features, storage device features, and storage device performance features.

[0036] In this step, to predict the future popularity of business data based on its past access situation, the access logs of the business data in a preset time period can be obtained, and the data characteristics of the business data at each time point can be obtained from the access logs to form a time series of data characteristics. Among them, the access log is the log recorded by the storage system when the business data is accessed. The data characteristics can include data access characteristics, data attribute characteristics, storage device characteristics, and storage device performance characteristics. The data access characteristics can reflect the usage and timeliness of the data. For example, it can include access frequency, access time distribution, etc. The data attribute characteristics can reflect the inherent attributes of the data. For example, it can include data size, data type (text, picture, video, etc.), creation time, etc., which can help understand the basic characteristics of the data. The storage device characteristics can reflect the device information of the storage device storing the business data, such as including the type of storage medium (such as SSD, HDD, cache, etc.), cache capacity, and current load. The storage device performance characteristics are used to evaluate the performance of the storage system and can include latency, throughput, and IOPS (input / output operations per second). It should be noted that since the access scenarios and access methods of users for business data are different, the popularity of business data may be affected by factors other than access time and access frequency. That is, the above-mentioned various data characteristics contain key information that can distinguish the popularity of business data. Therefore, in this embodiment, various types of data characteristics related to business data access will be collected as much as possible to dynamically select the target data characteristics suitable for predicting the popularity of business data subsequently.

[0037] Furthermore, to improve the quality of the data characteristics and eliminate the dimensional difference between different data characteristics, in this embodiment, when obtaining the data characteristics, the original data characteristics can be cleaned to remove the outlier values (data significantly different from most of the data) and abnormal duplicate values in the original data characteristics, and the cleaned original data characteristics can be normalized. Among them, the original data characteristics refer to the data characteristics that have not undergone preprocessing.

[0038] Based on this, obtaining the data characteristics of the business data at each time point from the access logs can include:

[0039] Step 11: Obtain the original data characteristics of the business data at each time point from the access logs.

[0040] Step 12: Remove the outlier values and abnormal duplicate values in the original data characteristics, and perform normalization processing on the original data characteristics to obtain the data characteristics.

[0041] It should be noted that considering that the user's access scenarios and access patterns are constantly changing, that is, the data features generated at a later time are closer to the current user access scenarios and access patterns, while the data features generated at an earlier time are less close to the current user access scenarios and access patterns. Therefore, when normalizing data features, different weights can be set for data features at different time points. Specifically, the normalization process in this embodiment can be carried out according to the following formula:

[0042] ;

[0043] where, is the original value of the i-th feature, and are the minimum and maximum values of the i-th feature at time t respectively, is the weight of the i-th feature at time t, satisfying . The preprocessed data features can be expressed as: , where represents the d-th data feature.

[0044] Furthermore, the length of the preset time period in this embodiment is not limited and can be set according to actual application requirements. For example, it can be one day, one week, one month, etc.

[0045] S102. Determine the heat labels of business data at each time point according to the data access characteristics.

[0046] In this step, based on preset rules, the heat labels of business data at each time point can be determined using the data access characteristics to achieve automatic annotation of heat labels. For example, heat labels can be set for business data according to data size, creation time, or access times. The specific heat labels in this embodiment are not limited. For example, they can include hot data, warm data, cold data, etc.

[0047] S103. Train at least two decision tree models using the data features and heat labels, and select target data features for classifying heat labels using the trained decision tree models.

[0048] In this step, the random forest algorithm will be adopted to train at least two decision tree models using the data features and popularity labels of business data, and the trained decision tree models will be used to select the target data features for classifying the popularity labels. The target data features refer to the data features used to predict the popularity of business data. For example, at least two data sets can be randomly selected from the data features and popularity labels of business data, and the corresponding decision tree models can be trained using each data set. Finally, based on the classification results of each decision tree model, the target data features for classifying the popularity labels are selected. In this way, considering that the access scenarios and access patterns of users to business data are always changing, the target data features most suitable for predicting the popularity of business data currently can be selected based on the decision tree models, so as to ensure that the selection of target data features dynamically matches the access scenarios and access patterns of users.

[0049] Specifically, in this embodiment, the importance index values of each data feature can be determined using the trained decision tree model, the data features can be sorted according to the importance index values, and the target data features can be selected from the sorted data features according to the preset selection quantity of the target data features. Among them, the value of the importance index represents the importance degree of the data feature in distinguishing the data popularity.

[0050] Based on this, using the trained decision tree model to select the target data features for classifying the popularity labels may include:

[0051] Step 21: Determine the importance index values of each data feature using the trained decision tree model;

[0052] Step 22: Sort the data features according to the importance index values, and select the target data features from the sorted data features according to the preset selection quantity of the target data features.

[0053] Specifically, the importance index value can be calculated based on the following formula:

[0054] ;

[0055] where T is the number of decision trees, which refers to the number of target data features here. is the reduction in the Gini index of data feature j in the t-th decision tree. The importance index values of each data feature can be expressed as: .

[0056] It should be noted that this embodiment does not limit how to build the decision tree model, and relevant techniques in machine learning can be referred to.

[0057] Of course, the number of target data features may not be fixed. In this case, the features can be sorted according to their importance, and the feature with the highest importance can be selected as the target data feature. A threshold θ can be set to screen the target data features.

[0058] S104. Train a time series model using the target data features and popularity labels of business data at each time point.

[0059] In this step, to predict the future popularity of business data based on its past access situation, and to adapt to the seasonal changes in the popularity of business data, a time series model can be introduced to learn the time series of target data features and popularity labels, so as to use the trained time series model for data popularity prediction. Among them, the time series model is a neural network model that can process time series data, such as an LSTM model (Long Short-Term Memory, long short-term memory artificial neural network).

[0060] It should be noted that this embodiment does not limit how to train the time series model, and relevant techniques in machine learning can be referred to.

[0061] S105. Obtain the latest access log of business data, obtain the latest value of the target data feature of business data from the latest access log, and input the latest value into the trained time series model to obtain the prediction result of the popularity label of business data.

[0062] In this step, this embodiment can use the trained time series model to predict the popularity label of business data. Specifically, the latest access log of business data can be obtained, and the latest value of the target data feature of business data can be obtained from the latest access log. Subsequently, the latest value can be encoded into a feature vector, and the feature vector can be input into the trained time series model, and the hidden state corresponding to the current time point output by the time series model using the feature vector can be received. Furthermore, the softmax function can be used to reduce the dimension of the hidden state to obtain the prediction result of the popularity label of business data.

[0063] Based on this, inputting the latest value into the trained time series model to obtain the prediction result of the popularity label of business data can include:

[0064] Step 31: Encode the latest value into a feature vector, input the feature vector into the trained time series model, and receive the hidden state corresponding to the current time point output by the time series model using the feature vector;

[0065] Step 32: Use the softmax function to reduce the dimension of the hidden state to obtain the prediction result of the popularity label of business data.

[0066] Furthermore, after obtaining the prediction result of the popularity label, the business data can be migrated to the storage layer corresponding to the popularity level of the prediction result of the popularity label, so as to migrate it to the storage layer corresponding to the popularity level in advance before the popularity of the business data changes.

[0067] Based on this, after obtaining the prediction result of the popularity label of business data, it may further include:

[0068] Step 41: According to the prediction result of the popularity label, migrate the business data to the storage layer corresponding to the popularity level corresponding to the prediction result of the popularity label.

[0069] Furthermore, considering that the prediction result of the popularity label of business data may be incorrect, which may lead to incorrect migration of business data. For this reason, after completing the migration of business data, this embodiment can also obtain the data characteristics of the business data after migration, and use the data characteristics after migration to determine the popularity index value of the business data, and this popularity index value represents the actual popularity of the business data. Subsequently, this embodiment can determine whether the popularity index value corresponds to the migrated popularity level. If not, this embodiment needs to collect the data characteristics of the business data after migration, and use this data characteristic to continue training the decision tree model and the time series model to correct the model.

[0070] Based on this, after migrating the business data to the storage layer corresponding to the popularity level corresponding to the prediction result of the popularity label, it may further include:

[0071] Step 51: Obtain the data characteristics of the business data after migration, and use the data characteristics after migration to determine the popularity index value of the business data.

[0072] Step 52: Determine whether the popularity index value corresponds to the popularity level.

[0073] Step 53: If the popularity index value does not correspond to the popularity level, use the data characteristics of the business data after migration to continue training the decision tree model and the time series model.

[0074] It should be noted that this embodiment does not limit how to determine the popularity index value. For example, it can be obtained by weighting the data characteristics of the business data.

[0075] Based on the above embodiments, the present application can first obtain the access logs of business data in a preset time period, and obtain the data characteristics of the business data at each time point from the access logs; wherein, the data characteristics include data access characteristics, data attribute characteristics, storage device characteristics, and storage device performance characteristics, and key characteristic information for distinguishing data heat is implicitly included in these data characteristics. Subsequently, the present application can determine the heat labels of the business data at each time point according to the data access characteristics to automatically determine the heat of the business data at each time point. Subsequently, the present application can train at least two decision tree models using the data characteristics and heat labels, and use the trained decision tree models to select target data characteristics for classifying the heat labels, that is, the present application can automatically select target data characteristics that can distinguish data heat in the preset time period by using decision tree models, so as to achieve a better data heat prediction effect. Subsequently, the present application can train a time series model using the target data characteristics and heat labels of the business data at each time point, and then obtain the latest access logs of the business data, obtain the latest values of the target data characteristics of the business data from the latest access logs, and input the latest values into the trained time series model to obtain the heat label prediction result of the business data. It can be seen that the present application can dynamically select target data characteristics that can distinguish data heat in the preset time period, then can train a time series model using the target data characteristics, and can use the time series model to predict the heat labels of business data, which can solve the technical problems in the related art that it is difficult to predict the heat change of business data and it is difficult to adapt to complex and changeable data access methods, so as to achieve a better data heat prediction result and can adapt to complex and changeable data access patterns.

[0076] Based on the above embodiments, to further improve the expression ability of data characteristics, this embodiment can also fuse data characteristics to obtain at least two comprehensive data characteristics, and then perform data heat prediction based on the comprehensive data characteristics. Based on this, the method can also include:

[0077] S201. Obtain the access logs of business data in a preset time period, and obtain the data characteristics of the business data at each time point from the access logs; wherein, the data characteristics include data access characteristics, data attribute characteristics, storage device characteristics, and storage device performance characteristics.

[0078] S202. Determine the heat labels of the business data at each time point according to the data access characteristics.

[0079] S203. Fuse the data characteristics into at least two comprehensive data characteristics.

[0080] In this step, multiple data characteristics can be fused into at least two comprehensive data characteristics. For example, the following fusion can be performed:

[0081] 1. Specify the combination of access frequency and recent access time to generate a "heat" feature, and the fusion formula is as follows:

[0082] 。

[0083] 2. Specify the combination of "access type" (read / write) and "data size" to generate the "read-write load" feature. The fusion formula is as follows:

[0084] 。

[0085] 3. Specify the combination of "disk where data is located" and "access frequency" to generate the "disk hot spot" feature.

[0086] 4. Specify the combination of "node where data is located" and "network distance" to generate the "access cost" feature. The fusion formula is as follows:

[0087] 。

[0088] 5. Specify the combination of "data type" (text, image, video) and "data size" to generate the "storage requirement" feature.

[0089] 6. Combine "latency" and "throughput" to generate the "performance efficiency" feature. The fusion formula is as follows:

[0090] 。

[0091] It can be seen that in this embodiment, multiple features (such as access frequency and data size) can be combined according to business requirements to generate new comprehensive features (such as the "importance" score of data), thereby further improving the expression ability of data.

[0092] S204. Train at least two decision tree models using the comprehensive data features and heat tags, and use the trained decision tree models to select the target comprehensive data features for classifying the heat tags.

[0093] S205. Train a time series model using the target comprehensive data features and heat tags of the business data at each time point.

[0094] S206. Obtain the latest access log of the business data, obtain the latest value of the target comprehensive data features of the business data from the latest access log, and input the latest value into the trained time series model to obtain the prediction result of the heat tag of the business data.

[0095] Based on the above embodiments, considering that the heat change frequencies of different business data are different, the heat of some business data often changes, and the heat of some business data seldom changes. Therefore, in this embodiment, different decision tree models and time series models can also be set according to different time period lengths of a preset time period, so that the decision tree models and time series models corresponding to each time period length can be used for data heat prediction, thereby improving the flexibility of data heat prediction. Based on this, the method may further include:

[0096] S301. Obtain the access logs of business data in a preset time period, and obtain the data characteristics of the business data at each time point from the access logs; wherein, the data characteristics include data access characteristics, data attribute characteristics, storage device characteristics, and storage device performance characteristics, and the time period length of the preset time period includes at least two types.

[0097] In this embodiment, the time period length of the preset time period may include at least two types. For example, two preset time period lengths of one week and one month can be set. Subsequently, this embodiment can obtain data characteristics based on different preset time period lengths.

[0098] S302. Determine the heat labels of the business data at each time point according to the data access characteristics.

[0099] S303. Train at least two decision tree models corresponding to the time period length of the preset time period by using the data characteristics and heat labels, and use the trained decision tree models to select the target data characteristics for classifying the heat labels.

[0100] In this step, decision tree models corresponding to different time period lengths can be trained by using the data characteristics and heat labels, so as to adaptively select different target data characteristics for data heat prediction in different time period lengths.

[0101] S304. Train a time series model corresponding to the time period length by using the target data characteristics and heat labels of the business data at each time point.

[0102] In this step, time series models corresponding to different time period lengths can be trained by using the target data characteristics and heat labels, so that the time series model can perform data heat prediction based on the target data characteristics in different time period lengths.

[0103] S305. Obtain the latest access log of the business data, obtain the latest value of the target data characteristics of the business data from the latest access log, and input the latest value into the trained time series model corresponding to the duration of the preset time period to obtain the heat label prediction result corresponding to the time period length of the business data.

[0104] Based on the above embodiments, the above data heat prediction method will be introduced based on a specific schematic diagram next.

[0105] Please refer to Figure 2 , Figure 2 , which is a schematic diagram of another data heat prediction process provided by the embodiments of the present application. Through the dynamic hierarchical method implemented by machine learning and real-time data analysis technologies, the present application can monitor the access patterns and heat changes of data in real time, and automatically adjust the storage location of data according to real-time requirements. This method can not only significantly improve the utilization efficiency of storage resources, but also reduce latency, optimize system performance, and reduce the total cost of the storage system. The detailed content is as follows:

[0106] 1. Data collection and preprocessing:

[0107] Please refer to Figure 3 , Figure 3 , which is a flowchart of data collection and preprocessing provided by the embodiments of the present application. Data collection and preprocessing are key steps in implementing dynamic hierarchical storage, aiming to obtain and prepare high-quality data to support subsequent analysis and decision-making. The data acquisition module is responsible for collecting access logs of various types of data, including information such as access time, access frequency, data size, accessing users, data types, etc. These data are classified into the following features:

[0108] Data access features: including access frequency, access time distribution, and data heat level (such as high, medium, low), which are used to reflect the usage and timeliness of data.

[0109] Data attribute features: covering data size, data type (such as text, picture, video, etc.), and data creation time, which help to understand the basic features of data.

[0110] Storage system features: involving storage medium type (such as HDD, SSD, cache, etc.), remaining capacity, and current load, which reflect the operating status of the storage system.

[0111] Performance metrics: including latency, throughput, and IOPS (input / output operations per second), which are used to evaluate the performance of the storage system.

[0112] In the data preprocessing stage, the collected data will go through steps such as cleaning, normalization, and feature combination to form a high-quality data set suitable for input into the machine learning model. Specifically, it includes:

[0113] Feature cleaning: removing invalid or duplicate feature data to ensure the integrity and consistency of the data.

[0114] Feature normalization: normalizing numerical features to eliminate dimensional differences and ensure comparability of different features.

[0115] Feature combination: According to business requirements, multiple features (such as access frequency and data size) are combined to generate new comprehensive features (such as the "importance" score of data), further enhancing the expressiveness of data.

[0116] Through data collection and preprocessing, high-quality feature inputs can be provided for subsequent machine learning models, ensuring the accuracy and effectiveness of the dynamic hierarchical storage strategy.

[0117] 2. Machine learning model construction:

[0118] Random forest can reduce noise, improve the interpretability and generalization ability of the model through feature selection, while LSTM (Long Short-Term Memory network) is good at capturing long-term dependencies in time series data.

[0119] The random forest learning and LSTM algorithms are used serially, combined with the stored data feature information, to construct a more powerful and efficient machine learning model. By using the random forest algorithm to perform importance scoring on the feature data, select features with higher importance, reduce the feature dimension, remove noise features, and output a set of key features; then input the key features selected by the random forest into the LSTM model, and the LSTM algorithm captures the time dependence to predict the target variable (such as the data heat level). Through the construction of the machine learning model, it can efficiently, accurately, and respond in real time to changes in data access patterns.

[0120] 3. Intelligent adjustment of data hierarchical strategy:

[0121] According to the prediction results of the learning model, the hierarchical strategy is intelligently adjusted to divide the data into multiple levels:

[0122] Hot data layer: Stored in high-performance media (such as SSD) for fast access.

[0123] Warm data layer: Stored in medium-performance media (such as HDD), suitable for data with occasional access.

[0124] Cold data layer: Stored in low-cost media (such as tapes, cold storage), suitable for data that is not accessed for a long time.

[0125] Based on the prediction results output by the model in real time as the decision-making basis, data migration is automatically triggered, and the data predicted as "hot" is migrated to the high-performance layer, and the data predicted as "cold" is migrated to the low-cost layer. For example, please refer to Figure 4 , Figure 4A schematic diagram of storage data migration provided by an embodiment of this application. According to the hierarchical metric data output by the LSTM model, the data hierarchical migration process can be started. The data with hierarchical metrics in the range of [0.8, 1] is migrated to the hot data layer, the data in the range of (0.4, 0.8) is migrated to the warm data layer, and the data in the range of [0, 0.4] is migrated to the cold data layer.

[0126] 4. Dynamic Migration and Feedback Optimization Learning Model:

[0127] Data Migration Mechanism: Dynamically adjust the distribution of data among different levels according to the real-time prediction of the model.

[0128] Feedback Optimization: By monitoring the actual access situation of data (performance data corresponding to key features), use the feedback results to optimize the machine learning model and improve the accuracy of prediction.

[0129] Through the solution of this application, the dynamic adjustment of the storage data hierarchical strategy is realized, which can flexibly cope with complex and changeable data access scenarios, thereby significantly improving the utilization rate of storage resources, reducing data access latency, optimizing system performance, and effectively controlling the overall cost of the storage system.

[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0131] The embodiment of this application also provides a data heat prediction device. Please refer to Figure 5 , Figure 5 A structural block diagram of a data heat prediction device provided by an embodiment of this application. This device may include:

[0132] A feature acquisition module 501, configured to acquire the access log of business data in a preset time period, and acquire the data features of business data at each time point from the access log; wherein, the data features include data access features, data attribute features, storage device features, and storage device performance features;

[0133] A labeling module 502, configured to determine the heat label of business data at each time point according to the data access features;

[0134] A data feature selection module 503, configured to train at least two decision tree models using the data features and heat labels, and select target data features for classifying the heat labels using the trained decision tree models;

[0135] A prediction model training module 504, configured to train a time series model using the target data features and heat labels of business data at each time point;

[0136] A prediction module 505, configured to obtain the latest access log of service data, obtain the latest value of the target data feature of the service data from the latest access log, and input the latest value into a trained time series model to obtain a prediction result of the popularity label of the service data.

[0137] Optionally, the apparatus may further include:

[0138] A migration module, configured to migrate the service data to a storage layer corresponding to the popularity level according to the prediction result of the popularity label.

[0139] Optionally, the apparatus may further include:

[0140] A migration evaluation module, configured to obtain the data feature of the service data after migration, and determine the popularity metric value of the service data by using the data feature after migration;

[0141] A migration feedback module, configured to determine whether the popularity metric value corresponds to the popularity level; if the popularity metric value does not correspond to the popularity level, continue to train the decision tree model and the time series model by using the data feature of the service data after migration.

[0142] Optionally, the feature acquisition module 501 may include:

[0143] An acquisition sub-module, configured to obtain the original data features of the service data at each time point from the access log;

[0144] A preprocessing sub-module, configured to remove the outliers and abnormal duplicate values in the original data features, and perform normalization processing on the original data features to obtain data features.

[0145] Optionally, the apparatus may further include:

[0146] A feature fusion module, configured to fuse the data features into at least two comprehensive data features;

[0147] The data feature selection module 503 may be used for:

[0148] Training at least two decision tree models by using the comprehensive data features and the popularity labels, and selecting target comprehensive data features for classifying the popularity labels by using the trained decision tree models;

[0149] The prediction model training module 504 may be used for:

[0150] Training a time series model by using the target comprehensive data features and the popularity labels of the service data at each time point;

[0151] The prediction module 505 may be used for:

[0152] Obtain the latest value of the target comprehensive data feature of the business data from the latest access log, and input the latest value into the trained time series model to obtain the predicted result of the popularity label of the business data.

[0153] Optionally, the data feature selection module 503 may include:

[0154] An index value determination sub-module for determining the importance index value of each data feature by using the trained decision tree model;

[0155] A selection sub-module for sorting the data features according to the importance index value, and selecting the target data features from the sorted data features according to the preset selection quantity of the target data features.

[0156] Optionally, the prediction module 505 may include:

[0157] A model processing sub-module for encoding the latest value into a feature vector, inputting the feature vector into the trained time series model, and receiving the hidden state corresponding to the current time point output by the time series model using the feature vector;

[0158] A dimensionality reduction sub-module for reducing the dimension of the hidden state by using the softmax function to obtain the predicted result of the popularity label of the business data.

[0159] Optionally, the data feature selection module 503 may be used for:

[0160] Training at least two decision tree models corresponding to the time period length of the preset time period by using the data features and popularity labels; the time period length includes at least two types;

[0161] The prediction model training module 504 may be used for:

[0162] Training a time series model corresponding to the time period length by using the target data features and popularity labels of the business data at each time point;

[0163] The prediction module 505 may be used for:

[0164] Inputting the latest value into the trained time series model corresponding to the duration of the preset time period to obtain the predicted result of the popularity label of the business data corresponding to the time period length.

[0165] For the description of the features in the corresponding embodiments of the data popularity prediction device, reference may be made to the relevant descriptions in the corresponding embodiments of the data popularity prediction method, which will not be elaborated here one by one.

[0166] Please refer to Figure 6 , Figure 6The block diagram of a structure of an electronic device provided by an embodiment of the present application. An embodiment of the present application provides an electronic device 10, including a processor 11 and a memory 12. Among them, the memory 12 is used to store a computer program. The processor 11 is used to execute the data heat prediction method provided by the foregoing embodiment when executing the computer program.

[0167] For the specific process of the above data heat prediction method, reference can be made to the corresponding content provided in the foregoing embodiment, and details will not be elaborated here.

[0168] Moreover, as a carrier for resource storage, the memory 12 can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc., and the storage method can be temporary storage or permanent storage.

[0169] In addition, the electronic device 10 further includes a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. Among them, the power supply 13 is used to provide working voltage for each hardware device on the electronic device 10. The communication interface 14 can create a data transmission channel between the electronic device 10 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here. The input / output interface 15 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application requirements, and specific limitations are not imposed here.

[0170] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any of the foregoing embodiments of the data heat prediction method when running.

[0171] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (abbreviated as ROM), random access memories (abbreviated as RAM), mobile hard disks, magnetic disks, optical disks, and other media that can store computer programs.

[0172] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and the computer program implements the steps in any of the foregoing embodiments of the data heat prediction method when executed by a processor.

[0173] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the computer program implements the steps in any of the foregoing embodiments of the data heat prediction method when executed by a processor.

[0174] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0175] The above has introduced in detail a data heat prediction method, device, electronic device, and storage medium provided by this application. Specific examples have been used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for predicting data popularity, characterized in that, Including: Obtain the access logs of business data in a preset time period, and obtain the data characteristics of the business data at each time point from the access logs; wherein, the data characteristics include data access characteristics, data attribute characteristics, storage device characteristics, and storage device performance characteristics; Determine the heat tags of the business data at each time point according to the data access characteristics; Train at least two decision tree models using the data characteristics and the heat tags, and use the trained decision tree models to select the target data characteristics for classifying the heat tags; Train a time series model using the target data characteristics and the heat tags of the business data at each time point; Obtain the latest access log of the business data, obtain the latest value of the target data characteristics of the business data from the latest access log, and input the latest value into the trained time series model to obtain the heat tag prediction result of the business data.

2. The data heat prediction method according to claim 1, wherein After obtaining the heat tag prediction result of the business data, it further includes: According to the heat tag prediction result, migrate the business data to the storage layer corresponding to the heat level corresponding to the heat tag prediction result.

3. The data heat prediction method according to claim 2, wherein After migrating the business data to the storage layer corresponding to the heat level corresponding to the heat tag prediction result, it further includes: Obtain the data characteristics of the business data after migration, and determine the heat index value of the business data using the data characteristics after migration; Judge whether the heat index value corresponds to the heat level; If the heat index value does not correspond to the heat level, continue to train the decision tree model and the time series model using the data characteristics of the business data after migration.

4. The data heat prediction method according to claim 1, wherein Obtaining the data characteristics of the business data at each time point from the access logs includes: Obtain the original data characteristics of the business data at each time point from the access logs; Remove the outlier values and abnormal duplicate values in the original data characteristics, and perform normalization processing on the original data characteristics to obtain the data characteristics.

5. The data heat prediction method according to claim 1, wherein After obtaining the data characteristics of the business data at each time point from the access logs, it further includes: Fuse the data characteristics into at least two types of comprehensive data characteristics; Training at least two decision tree models using the data characteristics and the heat tags, and using the trained decision tree models to select the target data characteristics for classifying the heat tags includes: Train at least two decision tree models using the comprehensive data characteristics and the heat tags, and use the trained decision tree models to select the target comprehensive data characteristics for classifying the heat tags; Training a time series model using the target data characteristics and the heat tags of the business data at each time point includes: Train a time series model using the target comprehensive data characteristics and the heat tags of the business data at each time point; Obtain the latest value of the target data characteristics of the business data from the latest access log, and input the latest value into the trained time series model to obtain the heat tag prediction result of the business data, including: Obtain the latest value of the target comprehensive data feature of the service data from the latest access log, and input the latest value into the trained time series model to obtain the prediction result of the popularity label of the service data.

6. The data heat prediction method according to claim 1, wherein Use the trained decision tree model to select the target data features for classifying the popularity label, including: Use the trained decision tree model to determine the importance index values of each of the data features; Sort the data features according to the importance index values, and select the target data features from the sorted data features according to the preset selection quantity of the target data features.

7. The data heat prediction method according to claim 1, characterized in that Input the latest value into the trained time series model to obtain the prediction result of the popularity label of the service data, including: Encode the latest value into a feature vector, input the feature vector into the trained time series model, and receive the hidden state corresponding to the current time point output by the time series model using the feature vector; Use the softmax function to reduce the dimension of the hidden state to obtain the prediction result of the popularity label of the service data.

8. A data heat prediction device, characterized in that, Include: A feature acquisition module, configured to acquire the access log of the service data in a preset time period, and acquire the data features of the service data at each time point from the access log; wherein, the data features include data access features, data attribute features, storage device features, and storage device performance features; A labeling module, configured to determine the popularity label of the service data at each time point according to the data access features; A data feature selection module, configured to train at least two decision tree models using the data features and the popularity labels, and use the trained decision tree models to select the target data features for classifying the popularity label; A prediction model training module, configured to train a time series model using the target data features and the popularity labels of the service data at each time point; A prediction module, configured to acquire the latest access log of the service data, obtain the latest value of the target data feature of the service data from the latest access log, and input the latest value into the trained time series model to obtain the prediction result of the popularity label of the service data.

9. An electronic device, characterized in that, Include: A memory, configured to store a computer program; A processor, configured to implement the data popularity prediction method according to any one of claims 1 to 7 when executing the computer program.

10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the data popularity prediction method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Data storage method and device, equipment and storage medium

    CN121008752A

  • Distributed storage system and data management method, device and equipment thereof

    CN121277441A