Data hierarchical storage method and device, computer equipment and storage medium
By collecting data access logs and using deep learning models to predict access modes, dynamically adjusting the distribution of data at different storage layers, solving the problems of resource waste and performance bottlenecks in traditional hierarchical storage systems, achieving efficient data management and reducing costs.
Patent Information
- Application Number
- CN202510494666.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-12
AI Technical Summary
Traditional hierarchical storage systems are difficult to efficiently manage dynamic data access patterns, resulting in waste of resources, performance bottlenecks and data inconsistencies, and increase management complexity and costs.
By collecting data access logs, using deep learning models to predict access modes, dynamically adjusting the distribution strategy of data at different storage layers, and realizing data migration and storage.
Improve storage efficiency and access efficiency, reduce resource waste and management costs, and enhance system flexibility and reliability.
Smart Images

Figure CN120469630A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of database technology, and in particular to a data hierarchical storage method, apparatus, computer equipment, and storage medium. Background Art
[0002] Tiered storage technology uses different storage methods to store data on storage devices with different performance characteristics based on metrics such as access frequency, retention time, capacity, and performance. Tiered storage management enables automatic migration of data objects between storage devices. For example, hot data (highly accessed data) is stored on high-performance storage devices such as solid-state drives, while cold data (lower-accessed data) is stored on lower-cost storage devices such as cloud storage or tape libraries. This technology is widely used in industries such as banking, securities, and e-commerce, as well as in fields such as big data analytics and artificial intelligence, to improve data management efficiency, reduce storage costs, and increase the value of data analysis and utilization.
[0003] With the rapid growth of data volumes, traditional tiered storage systems struggle to efficiently manage and quickly respond to dynamic data access patterns. Therefore, a more intelligent approach to optimizing data storage is needed. Current tiered storage systems face the following shortcomings: Poor recognition of hot and cold data labels can lead to incorrect identification of hot and cold files, causing some hot data to be assigned to lower-performing storage devices. This can lead to performance bottlenecks when data is accessed. Due to data redundancy and storage tier division, data updates and synchronization can become slow, impacting overall system performance. Furthermore, some cold data is assigned to higher-performing storage devices, resulting in wasted high-end storage resources, increased enterprise investment costs, and resource waste.
[0004] Since there may be synchronization and update delays between data at different levels, data inconsistency may occur, which will affect the accuracy of data analysis and decision-making. The system needs to frequently synchronize and verify data between multiple storage layers, increasing the complexity and management costs of the system. Summary of the Invention
[0005] The embodiments of the present application provide a data tiered storage method, apparatus, computer device, and storage medium to solve the problem of how to improve the real-time management of data tiered storage to reduce resource waste and data access restrictions.
[0006] In a first aspect, a data tiered storage method includes: Collecting access logs of target data at the current moment, and determining data information of the target data based on the access logs, the data information including access tag information and data tag information; Extracting features from the data information to obtain target access features and target data features; Input the target access feature and the target data feature into a trained deep learning model, and output a predicted access pattern within a preset time period after the current moment; According to the predicted access pattern, the distribution strategy of the target data in at least two storage layers of the database is adjusted to obtain an updated storage strategy. According to the updated storage strategy, the target data is migrated and stored, and the at least two storage layers are disks with different read and write rates.
[0007] In a second aspect, a data tiered storage device includes: A data collection module is used to collect access logs of target data at the current moment, and determine data information of the target data based on the access logs, wherein the data information includes access tag information and data tag information; A feature extraction module is used to extract features from the data information to obtain target access features and target data features; A data prediction module, configured to input the target access features and the target data features into a trained deep learning model and output a predicted access pattern within a preset time period after the current moment; A storage decision module is used to adjust the distribution strategy of the target data in at least two storage layers of the database according to the predicted access pattern to obtain an updated storage strategy, and migrate the target data according to the updated storage strategy, wherein the at least two storage layers are disks with different read and write rates.
[0008] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data hierarchical storage method described in the first aspect when executing the computer program.
[0009] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data hierarchical storage method described in the first aspect is implemented.
[0010] The present application has the following beneficial effects compared to the prior art: the present application collects the access log of the target data at the current moment, determines the data information of the target data based on the access log, performs feature extraction on the data information to obtain target access features and target data features, inputs the target access features and the target data features into a trained deep learning model, outputs the predicted access pattern within a preset time period after the current moment, and adjusts the distribution strategy of the target data in at least two storage layers with different read and write rates of the database based on the predicted access pattern to obtain an updated storage strategy, and migrates the target data for storage based on the updated storage strategy. The model can analyze and predict data access patterns in real time, automatically adjust the distribution of data between different storage layers, and effectively improve storage efficiency and access efficiency compared to existing static storage strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 This is a schematic diagram of an application environment of a data tiered storage method provided in Example 1 of the present application; Figure 2 This is a flowchart of a data hierarchical storage method provided in Example 2 of the present application; Figure 3 This is a data flow diagram of a data hierarchical storage method provided in Example 2 of the present application; Figure 4 This is a flowchart of a data hierarchical storage method provided in Example 3 of the present application; Figure 5 This is a flowchart of a data hierarchical storage method provided in Example 4 of the present application; Figure 6 This is a flowchart of a data hierarchical storage method provided in Example 5 of the present application; Figure 7 This is a flowchart of a data hierarchical storage method provided in Example 6 of the present application; Figure 8 This is a structural diagram of a data tiered storage device provided in Example 7 of the present application; Figure 9 This is a structural diagram of a computer device provided in Example 8 of the present application. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0015] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0016] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0017] In order to illustrate the technical solution of the present application, specific embodiments are provided below.
[0018] The data layered storage method provided in the first embodiment of the present application can be applied as follows: Figure 1 Specifically, the data layer storage method is applied in a database system, which includes Figure 1 The client and server shown communicate over a network to implement operations such as data transmission and data access. The server hosts a database and completes the steps of the aforementioned data tiered storage method. The client, also known as the user end, refers to the program that corresponds to the server and provides local services to clients. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0019] The above-mentioned database systems include but are not limited to databases used for data storage in different industries, such as banking databases, securities databases, and e-commerce databases, to support the needs of corresponding industry services.
[0020] like Figure 2 As shown in the figure, a data layer storage method provided in the second embodiment of the present application is used in Figure 1 Taking the server in FIG. 1 as an example, the data tiered storage method includes the following steps: Step S201 : collecting access logs of target data at the current moment, and determining data information of the target data based on the access logs.
[0021] The current moment is the time at which the data tiered storage method is executed. This method can be triggered by the user or by a set timer. The target data is any data stored in the database. The access records of this data up to the current moment are the access logs, which can reflect the access time and data identifier of the data.
[0022] By performing statistics on the access log, the data access frequency and corresponding timestamp under the data identifier can be determined, and data information such as data type and data size can also be determined. Among them, the data information includes access tag information and data tag information. The access tag information includes but is not limited to access frequency, timestamp, access purpose, etc., and the data tag information includes but is not limited to data type, data size, data source, etc.
[0023] Step S202: extracting features from the data information to obtain target access features and target data features.
[0024] In order to realize data analysis, after obtaining the data information, it is necessary to extract the characteristics of the data, so as to determine the characteristic information in the data that can characterize the essence of the data.
[0025] Since the data information includes access tag information and data tag information, when extracting data features, the two parts of tag information are extracted separately. Finally, the target access features can be obtained corresponding to the access tag information, and the target data features can be obtained corresponding to the data tag information.
[0026] This feature extraction can extract the main features in the information. For example, when extracting information such as access frequency, timestamp, and access purpose, the access frequency value and timestamp value can be extracted as the target access feature. In some cases, the access purpose, etc. may not be written into the target access feature.
[0027] Step S203: Input the target access feature and the target data feature into a trained deep learning model, and output a predicted access pattern within a preset time period after the current moment.
[0028] Among them, the trained deep learning model has the function of predicting access patterns. During training, it uses access features and data features as input and access patterns as output. Therefore, the predicted access pattern can be obtained by inputting the target access features and target data features into the trained deep learning model.
[0029] When training the above-mentioned deep learning model, the preset time is also used as an input parameter for training, and when in use, the preset time is used as an inherent parameter of the trained deep learning model, so that the access pattern directly output is the predicted access pattern within the preset time period after the current moment.
[0030] The deep learning model may be a model including a Transformer, which is a deep learning model based on an attention mechanism, for example, BERT, GPT, RoBERTa, and other models.
[0031] Optionally, the trained deep learning model includes: A first feature encoder, a second feature encoder, a feature fusion device and a classifier; Inputting the target access feature and the target data feature into a trained deep learning model and outputting a predicted access pattern within a preset time period after the current moment includes: Input the target access feature into a first feature encoder, output a target access feature vector, input the target data feature into a second feature encoder, output a target data feature vector; Inputting the target access feature vector and the target data feature vector into the feature fusion device, and outputting a target fusion feature; The target fusion feature is input into a classifier, and a predicted access pattern within a preset time period after the current moment is output.
[0032] Among them, the designed deep learning model includes a first feature encoder, a second feature encoder, a feature fusion device and a classifier. The first feature encoder and the second feature encoder encode two types of feature information respectively, and then fuse the two types of feature information in the feature fusion device, which can cope with data of different dimensions to a certain extent, thereby realizing multi-dimensional fusion.
[0033] The classifier has multi-classification capabilities. Specifically, it can divide the calculated prediction value into different categories through multiple intervals. That is, the number of intervals is related to the number of access patterns. For example, access patterns include hot data patterns and cold data patterns. Accordingly, the classifier can adopt a binary classification method, using 1 to represent hot data patterns and 0 to represent cold data patterns.
[0034] Optionally, the training process of the trained deep learning model includes: Obtain training access features and training data features for N data, as well as calibrated access patterns within a preset time period after corresponding data collection, where N is an integer greater than zero; Input the training access feature into a first feature encoder, output a training access feature vector, input the training data feature into a second feature encoder, output a training data feature vector; Inputting the training access feature vector and the training data feature vector into the feature fusion device, and outputting a training fusion feature; The trained fusion features are input into the classifier, and the classification results are output. According to the classification results and the loss of the calibrated access pattern, the parameters of the first feature encoder, the second feature encoder, the feature fuser and the classifier are adjusted until the iteration conditions are met to obtain a trained deep learning model.
[0035] Among them, for the training of deep learning models, a training set consisting of training access features and training data features is used. Accordingly, a calibrated access pattern within a preset time period needs to be given so that it can be compared with the classification result after classification by the classifier, and then loss functions such as contrast loss and cross entropy loss are used to adjust the parameters in all encoders, fusers, and classifiers to obtain a trained deep learning model after meeting the iteration conditions.
[0036] Iteration conditions include but are not limited to reaching a preset number of iterations, loss meeting preset conditions, etc.
[0037] Step S204: According to the predicted access pattern, the distribution strategy of the target data in at least two storage layers of the database is adjusted to obtain an updated storage strategy, and the target data is migrated and stored according to the updated storage strategy.
[0038] The at least two storage layers are disks with different read / write speeds. For example, the storage layers may be solid-state drives, magnetic disks, optical disks, or tapes, with the read / write speeds of solid-state drives, magnetic disks, optical disks, or tapes decreasing in order. In addition to read / write speeds, other read / write characteristics of the disks may also serve as a basis for updating the storage policy.
[0039] At least two storage layers are set up in the database. Disks with different read and write rates can store data with different access modes. If it is hot data, it needs to be stored on disks with relatively higher read and write rates, which can effectively improve access efficiency. Of course, cold data can also be stored on disks with relatively higher read and write rates, but this will increase the burden on disks with relatively higher read and write rates, and even require the design of more disks with relatively higher read and write rates, resulting in increased costs. Therefore, storing cold data on disks with relatively lower read and write rates can minimize the cost of database deployment without affecting data usage.
[0040] Access patterns can be mapped to storage policies. For example, a cold data pattern might correspond to storage on disks with lower read / write speeds, while a hot data pattern might correspond to storage on disks with higher read / write speeds. After updating a storage policy, data needs to be migrated. If the storage policy remains unchanged, the data's corresponding storage tier can remain unchanged. Of course, data can also be migrated under certain conditions. For example, if a cold data pattern indicates data that has not been accessed for a long time, the data can be migrated to a low-read / write storage area on a disk with a lower read / write speed.
[0041] For example, Figure 3 The figure shows a data flow diagram of a data hierarchical storage method provided in Example 2 of the present application, in which the corresponding data processing functions are modularized, the data acquisition module: uses the log collection module to collect real-time access logs, collects data access frequency, timestamp, data type, data size and other labels, and stores the logs in an efficient database; the feature extraction module: combines the file system metadata to clean and extract data features, and cleans out data labels, including user importance, user behavior, data access frequency, timestamp, data type, data size, historical access trends, data update frequency, etc.; the deep learning model: uses the deep learning model to train data and predict future data access patterns; the policy decision module: formulates a data storage strategy based on the model prediction results, dynamically adjusts the distribution of data between SSDs, high-speed hard drives and ordinary hard drives, and then sends the data storage strategy to the life cycle module, which is responsible for managing the migration of data on different storage devices; the life cycle module: is responsible for migrating data between different storage devices according to the strategy; the real-time optimization module: updates the model and storage strategy in real time based on the monitoring system performance and access status.
[0042] The above-mentioned database can be used in scenarios such as banking systems and securities systems that require the storage of large amounts of data. It is aimed at fields such as big data analysis and artificial intelligence, and can improve data management efficiency, reduce storage costs, and increase the value of data analysis and utilization.
[0043] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0044] The embodiment of the present application collects the access log of the target data at the current moment, determines the data information of the target data based on the access log, performs feature extraction on the data information to obtain target access features and target data features, inputs the target access features and the target data features into a trained deep learning model, outputs the predicted access pattern within a preset time period after the current moment, and adjusts the distribution strategy of the target data in at least two storage layers with different read and write rates of the database based on the predicted access pattern to obtain an updated storage strategy. Based on the updated storage strategy, the target data is migrated and stored. The model can analyze and predict data access patterns in real time, automatically adjust the distribution of data between different storage layers, and effectively improve storage efficiency and access efficiency compared to existing static storage strategies.
[0045] like Figure 4 FIG. 1 is a flow chart of a data hierarchical storage method provided in Embodiment 3 of the present application. In step S202, feature extraction of the data information is performed to obtain target access features and target data features, which may include the following steps: Step S401: clean the data information to obtain cleaned data information.
[0046] Step S402: extracting features from the cleaned data information to obtain target access features and target data features.
[0047] Among them, the collected data information may contain unnecessary information or even erroneous information. Data cleaning can effectively remove unnecessary and erroneous information and obtain cleaned data information.
[0048] Data cleaning is the process of re-examining and verifying data, aiming to discover and correct errors in data files and ensure data consistency and accuracy. The main methods of data cleaning include: Handling missing values: Calculate the missing value ratio for each field and adopt different handling strategies based on the missing value ratio and field importance. Unimportant data or data with a high missing rate can be directly removed, while important data can be supplemented or retrieved from other channels. Handling inconsistent data values: This includes removing illogical characters (such as spaces, numeric symbols, etc.) to ensure the correctness of the data content; Deduplication: Identify and remove duplicate records in a dataset, preserving only unique data records. In some cases, duplicate records may be used to analyze evolution patterns or detect business rule issues; in these cases, duplicate records should not be removed. Handling unreasonable data: discover outliers through binning, clustering, regression, etc., and perform manual processing; Unified data format: When integrating data from multiple sources, ensure consistency in data field formats to avoid errors caused by inconsistent formats; Check data consistency: Ensure that data from different sources are logically consistent and correct inconsistent data; Handling invalid and erroneous values: Identify and correct invalid and erroneous values to ensure data accuracy.
[0049] The embodiments of the present application can improve the quality of data and ensure the accuracy and reliability of data analysis by means of data cleaning.
[0050] like Figure 5 FIG. 4 is a flow chart of a data hierarchical storage method according to a fourth embodiment of the present application. In step S402, feature extraction is performed on the cleaned data information to obtain target access features and target data features. The following steps may be included: Step S501: Obtain target data labels of cleaned data information.
[0051] The target data tags include user importance, user behavior, data access frequency, timestamp, data type, data size, historical access trend and data update frequency.
[0052] Some labels are in the data information, and for the cleaned data information, unnecessary labels or data will be processed, so that the target data label is determined to be a label that can have a feature impact.
[0053] Step S502 , extracting the data value corresponding to each target data tag, and determining the data value corresponding to the data type and the data size as the target data feature.
[0054] Step S503 : determining data values corresponding to the user importance, the user behavior, the data access frequency, the timestamp, the historical access trend, and the data update frequency as target access features.
[0055] Among them, the data value of each target data label is extracted, and the data value corresponding to the data type and data size is used as the target data feature, while the others are used as target access features. The target data feature represents the essence of the data, and the target access feature represents the essence of the data when it is accessed.
[0056] The embodiment of the present application determines target data features and target access features through label recognition and classification, thereby improving the efficiency of data feature extraction.
[0057] like Figure 6 FIG. 2 is a flow chart of a data tiered storage method according to a fifth embodiment of the present application. In step S204, adjusting the distribution strategy of the target data in at least two storage tiers of the database according to the predicted access pattern to obtain an updated storage strategy may include the following steps: Step S601: Obtain the current distribution strategy of the target data in at least two storage layers of the database.
[0058] Step S602: If the predicted access mode is a hot data mode, it is detected whether the current distribution strategy is to store in a storage layer with a higher read and write rate.
[0059] Step S603: If it is detected that the current distribution strategy is not to store in a storage layer with a higher read / write rate, the current distribution strategy is adjusted to store in a storage layer with a higher read / write rate to obtain an updated storage strategy.
[0060] Step S604: If the predicted access mode is a cold data mode, detecting whether the current distribution strategy is to store in a storage layer with a lower read / write rate.
[0061] Step S605: If it is detected that the current distribution strategy is not to store in a storage layer with a lower read / write rate, the current distribution strategy is adjusted to store in a storage layer with a lower read / write rate to obtain an updated storage strategy.
[0062] Among them, it is first necessary to obtain the current distribution strategy of the target data. The current distribution strategy is the strategy updated during historical processing. The current distribution strategy has been executed, and the target data is stored in the corresponding storage layer according to the current distribution strategy.
[0063] This application divides the predicted access pattern into hot data pattern and cold data pattern, so as to effectively classify and detect it with the current distribution strategy, and finally use the new storage strategy to update the current distribution strategy, accurately perform tiered storage, and improve database utilization efficiency.
[0064] like Figure 7 FIG. 2 is a flow chart of a data tiered storage method according to a sixth embodiment of the present application. In step S204, migrating and storing the target data according to the updated storage policy may include the following steps: Step S701: Obtain the read / write rate level of the current storage layer of the target data.
[0065] Step S702, if the updated storage strategy is to store in a storage layer with a lower read / write rate, then when the read / write rate level is not the lowest level, the target data is migrated to a storage layer lower than the read / write rate level for storage; when the read / write rate level is the lowest level, the target data is kept stored in the current storage layer.
[0066] Step S703, if the updated storage strategy is to store in a storage layer with a higher read / write rate, then when the read / write rate level is not the highest level, the target data is migrated to a storage layer higher than the read / write rate level for storage; when the read / write rate level is the highest level, the target data is kept stored in the current storage layer.
[0067] Among them, for migration storage, if the target data has been stored in the storage layer corresponding to the updated storage policy, there is no need to migrate the target data, that is, keep the target data stored in the current storage layer; if the target data is not stored in the storage layer corresponding to the updated storage policy, the target data needs to be migrated.
[0068] When performing migration storage, the database's own data migration function can be used for migration storage, for example, the life cycle module in Example 1, which will not be described in detail in this embodiment of the application.
[0069] The embodiments of the present application accurately predict and optimize. Traditional systems have difficulty accurately predicting the hot and cold changes of data, which may lead to waste of resources. Through deep learning algorithms, it is possible to more accurately predict which data requires fast access, thereby optimizing storage strategies and reducing unnecessary storage and retrieval time, thereby improving system performance. Storage system performance bottlenecks lead to increased data access latency. Through intelligent tiering, high-frequency access data is stored in high-performance storage, reducing access latency and improving overall system performance. Unreasonable data allocation leads to increased storage costs. The embodiments of the present application can effectively identify low-frequency access data and move it to low-cost storage devices, reducing the use of expensive storage and reducing overall costs.
[0070] Through these improvements, the intelligent data tiered storage system based on large models can significantly improve efficiency, reduce costs, and enhance system flexibility and reliability.
[0071] The seventh embodiment of the present application provides a data tiered storage device, which corresponds one-to-one to the data tiered storage method in the above embodiment. Figure 8 As shown, the data hierarchical storage device includes a data acquisition module 81, a feature extraction module 82, a data prediction module 83 and a storage decision module 84. The functional modules are described in detail as follows: The data collection module 81 is used to collect the access log of the target data at the current moment, and determine the data information of the target data according to the access log, wherein the data information includes access tag information and data tag information; A feature extraction module 82 is used to extract features from the data information to obtain target access features and target data features; A data prediction module 83 is configured to input the target access feature and the target data feature into a trained deep learning model and output a predicted access pattern within a preset time period after the current moment; The storage decision module 84 is used to adjust the distribution strategy of the target data in at least two storage layers of the database according to the predicted access pattern to obtain an updated storage strategy, and migrate the target data according to the updated storage strategy, wherein the at least two storage layers are disks with different read and write rates.
[0072] Optionally, the trained deep learning model includes: A first feature encoder, a second feature encoder, a feature fusion device and a classifier; Inputting the target access feature and the target data feature into a trained deep learning model and outputting a predicted access pattern within a preset time period after the current moment includes: Input the target access feature into a first feature encoder, output a target access feature vector, input the target data feature into a second feature encoder, output a target data feature vector; Inputting the target access feature vector and the target data feature vector into the feature fusion device, and outputting a target fusion feature; The target fusion feature is input into a classifier, and a predicted access pattern within a preset time period after the current moment is output.
[0073] Optionally, the training process of the trained deep learning model includes: Obtain training access features and training data features for N data, as well as calibrated access patterns within a preset time period after corresponding data collection, where N is an integer greater than zero; Input the training access feature into a first feature encoder, output a training access feature vector, input the training data feature into a second feature encoder, output a training data feature vector; Inputting the training access feature vector and the training data feature vector into the feature fusion device, and outputting a training fusion feature; The trained fusion features are input into the classifier, and the classification results are output. According to the classification results and the loss of the calibrated access pattern, the parameters of the first feature encoder, the second feature encoder, the feature fuser and the classifier are adjusted until the iteration conditions are met to obtain a trained deep learning model.
[0074] Optionally, the feature extraction module 82 includes: A data cleaning unit, configured to clean the data information to obtain cleaned data information; The feature extraction unit is used to extract features from the cleaned data information to obtain target access features and target data features.
[0075] Optionally, the feature extraction unit includes: A tag acquisition subunit is used to obtain target data tags of the cleaned data information, wherein the target data tags include user importance, user behavior, data access frequency, timestamp, data type, data size, historical access trend and data update frequency; A tag value extraction subunit is used to extract the data value corresponding to each target data tag, and determine the data value corresponding to the data type and the data size as the target data feature; The feature determination subunit is used to determine the data values corresponding to the user importance, the user behavior, the data access frequency, the timestamp, the historical access trend and the data update frequency as target access features.
[0076] Optionally, the storage decision module 84 includes: a current detection and acquisition unit, configured to acquire a current distribution strategy of the target data in at least two storage layers of the database; a first strategy detection unit, configured to detect whether the current distribution strategy is to store data in a storage layer with a higher read / write rate if the predicted access mode is a hot data mode; A first strategy updating unit is configured to adjust the current distribution strategy to store in the storage layer with a higher read / write rate if it is detected that the current distribution strategy is not to store in the storage layer with a higher read / write rate, thereby obtaining an updated storage strategy; a second strategy detection unit, configured to detect, if the predicted access mode is a cold data mode, whether the current distribution strategy is to store data in a storage layer with a lower read / write rate; The second strategy updating unit is configured to adjust the current distribution strategy to be stored in a storage layer with a lower read / write rate if it is detected that the current distribution strategy is not stored in a storage layer with a lower read / write rate, thereby obtaining an updated storage strategy.
[0077] Optionally, the storage decision module 84 includes: A rate level acquisition unit, configured to acquire a read / write rate level of a current storage layer of the target data; a first decision unit configured to, if the updated storage policy is to store the target data in a storage tier with a lower read / write rate, migrate the target data to a storage tier with a lower read / write rate level for storage when the read / write rate level is not the lowest level, and to keep the target data stored in the current storage tier when the read / write rate level is the lowest level; The second decision unit is used to migrate the target data to a storage layer higher than the read-write speed level for storage if the updated storage strategy is to store the data in a storage layer with a higher read-write speed level when the read-write speed level is not the highest level, and to keep the target data stored in the current storage layer when the read-write speed level is the highest level.
[0078] The specific definition of the data tiered storage device can be found in the definition of the data tiered storage method above and will not be repeated here. Each module in the aforementioned data tiered storage device may be implemented in whole or in part via software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in the form of software in a memory in the computer device, so that the processor can call and execute the corresponding operations of each module.
[0079] In the eighth embodiment, the present application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to implement the above-mentioned data hierarchical storage method. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a data hierarchical storage method is implemented.
[0080] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the data hierarchical storage method in the above embodiment is implemented. For example, Figure 2 S201-S204 shown, or Figures 3 to 7 Alternatively, when the processor executes the computer program, the functions of each module / unit in the embodiment of the data hierarchical storage device are realized, for example Figure 8 The functions of the data acquisition module 81, feature extraction module 82, data prediction module 83 and storage decision module 84 are not described here in detail to avoid repetition.
[0081] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the data hierarchical storage method in the above embodiment is implemented, for example Figure 2 S201-S204 shown, or Figures 3 to 7 Alternatively, when the computer program is executed by a processor, the functions of the modules / units in the embodiment of the data hierarchical storage device are realized, for example, Figure 8 The functions of the data acquisition module 81, feature extraction module 82, data prediction module 83 and storage decision module 84 are not described here in detail to avoid repetition.
[0082] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0083] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0084] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A data layered storage method, characterized in that: include: Collecting access logs of target data at the current moment, and determining data information of the target data based on the access logs, the data information including access tag information and data tag information; Extracting features from the data information to obtain target access features and target data features; Input the target access feature and the target data feature into a trained deep learning model, and output a predicted access pattern within a preset time period after the current moment; According to the predicted access pattern, the distribution strategy of the target data in at least two storage layers of the database is adjusted to obtain an updated storage strategy. According to the updated storage strategy, the target data is migrated and stored, and the at least two storage layers are disks with different read and write rates.
2. The data layered storage method according to claim 1, wherein: The trained deep learning model includes: A first feature encoder, a second feature encoder, a feature fusion device and a classifier; Inputting the target access feature and the target data feature into a trained deep learning model and outputting a predicted access pattern within a preset time period after the current moment includes: Input the target access feature into a first feature encoder, output a target access feature vector, input the target data feature into a second feature encoder, output a target data feature vector; Inputting the target access feature vector and the target data feature vector into the feature fusion device, and outputting a target fusion feature; The target fusion feature is input into a classifier, and a predicted access pattern within a preset time period after the current moment is output.
3. The data layered storage method according to claim 2, wherein: The training process of the trained deep learning model includes: Obtain training access features and training data features for N data, as well as calibrated access patterns within a preset time period after corresponding data collection, where N is an integer greater than zero; Input the training access feature into a first feature encoder, output a training access feature vector, input the training data feature into a second feature encoder, output a training data feature vector; Inputting the training access feature vector and the training data feature vector into the feature fusion device, and outputting a training fusion feature; The trained fusion features are input into the classifier, and the classification result is output. According to the classification result and the loss of the calibrated access pattern, the parameters of the first feature encoder, the second feature encoder, the feature fuser and the classifier are adjusted until the iteration conditions are met to obtain a trained deep learning model.
4. The data layered storage method according to claim 1, wherein: The feature extraction of the data information to obtain target access features and target data features includes: Cleaning the data information to obtain cleaned data information; Feature extraction is performed on the cleaned data information to obtain target access features and target data features.
5. The data layered storage method according to claim 4, wherein: The feature extraction of the cleaned data information to obtain target access features and target data features includes: Obtain target data tags for the cleaned data information, wherein the target data tags include user importance, user behavior, data access frequency, timestamp, data type, data size, historical access trend, and data update frequency; Extract the data value corresponding to each target data label, and determine the data value corresponding to the data type and the data size as the target data feature; The data values corresponding to the user importance, the user behavior, the data access frequency, the timestamp, the historical access trend, and the data update frequency are determined as target access features.
6. The data layered storage method according to any one of claims 1 to 5, characterized in that: The step of adjusting the distribution strategy of the target data in at least two storage layers of the database according to the predicted access pattern to obtain an updated storage strategy includes: Obtaining a current distribution strategy of the target data in at least two storage layers of a database; If the predicted access mode is a hot data mode, detecting whether the current distribution strategy is to store data in a storage layer with a higher read / write rate; If it is detected that the current distribution strategy is not to store in the storage layer with a higher read / write rate, the current distribution strategy is adjusted to store in the storage layer with a higher read / write rate to obtain an updated storage strategy; If the predicted access mode is a cold data mode, detecting whether the current distribution strategy is to store data in a storage layer with a lower read / write rate; If it is detected that the current distribution strategy is not for storage in a storage layer with a lower read / write rate, the current distribution strategy is adjusted to be for storage in a storage layer with a lower read / write rate, thereby obtaining an updated storage strategy.
7. The data layered storage method according to claim 6, wherein: The migrating and storing the target data according to the updated storage strategy includes: Obtaining the read and write rate level of the current storage layer of the target data; If the updated storage strategy is to store data in a storage layer with a lower read / write rate, then when the read / write rate level is not the lowest level, the target data is migrated to a storage layer with a lower read / write rate level for storage; when the read / write rate level is the lowest level, the target data is kept stored in the current storage layer; If the updated storage strategy is to store in a storage layer with a higher read / write rate, then when the read / write rate level is not the highest level, the target data will be migrated to a storage layer higher than the read / write rate level for storage; when the read / write rate level is the highest level, the target data will be kept stored in the current storage layer.
8. A data tiered storage device, characterized in that: include: A data collection module is used to collect access logs of target data at the current moment, and determine data information of the target data based on the access logs, wherein the data information includes access tag information and data tag information; A feature extraction module is used to extract features from the data information to obtain target access features and target data features; A data prediction module, configured to input the target access features and the target data features into a trained deep learning model and output a predicted access pattern within a preset time period after the current moment; A storage decision module is used to adjust the distribution strategy of the target data in at least two storage layers of the database according to the predicted access pattern to obtain an updated storage strategy, and migrate the target data according to the updated storage strategy, wherein the at least two storage layers are disks with different read and write rates.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the data hierarchical storage method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data hierarchical storage method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Memory hierarchical optimization method and device, computer equipment and storage medium
CN121166349A
Data storage device supporting large model training and use method
CN121326249A
A data storage device for supporting large model training and a method of use
CN121326249B