Data optimization storage method and system based on artificial intelligence
By extracting features from data and access logs for classification and tiered storage, combined with real-time monitoring and migration optimization, the problem of irrational storage resource allocation in existing technologies is solved, and efficient and continuous optimization of the storage system is achieved.
Patent Information
- Application Number
- CN202510879785.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-30
AI Technical Summary
Existing AI-powered data optimization storage methods cannot accurately reflect the actual access needs of data, resulting in irrational allocation of storage resources, lack of real-time monitoring and dynamic adjustment, and affecting storage system performance.
By extracting features from the decompressed raw data and access logs, performing preliminary classification, formulating a tiered storage strategy, and monitoring the tiered storage status of data in real time, the storage status of data at each level is dynamically monitored, and data migration is performed based on the monitoring results, forming a closed-loop optimization process.
It dynamically adjusts storage resources based on data access patterns and frequency, ensures data is stored on the most appropriate tier, keeps the storage system running efficiently, and continuously monitors and optimizes performance.
Smart Images

Figure CN120723735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of storage optimization technology, and more specifically to an artificial intelligence-based data optimization storage method and system. Background Art
[0002] The rapid development of artificial intelligence technology has brought new opportunities to the field of data storage. By applying artificial intelligence technology to data storage systems, it is possible to achieve intelligent and optimized data storage, improve storage efficiency, and reduce storage costs. The main methods of using artificial intelligence technology to optimize data storage are through data classification and prediction. AI technology is used to predict data access patterns, load data that may be accessed into the cache in advance, reduce data access latency, perform security processing such as encryption and access control on data to protect data security and privacy, and intelligently schedule and optimize storage resources to reduce storage costs.
[0003] However, the above process still has the following disadvantages:
[0004] First, existing AI-based data optimization and storage methods may simply classify data based on data type or size, failing to accurately reflect actual data access needs and dynamically adjust based on those needs. This results in irrational allocation of storage resources and suboptimal storage performance.
[0005] Second, existing AI-based data optimization and storage methods may lack real-time monitoring of tiered storage status data and dynamic data migration, making it difficult to promptly identify and resolve storage issues. This can lead to irrational data distribution across storage tiers, impacting storage system performance.
[0006] Third, existing AI-based data optimization and storage methods may lack a closed-loop optimization mechanism and are unable to continuously monitor and adjust storage system performance, resulting in a gradual decline in storage system performance. Summary of the Invention
[0007] In order to overcome the above-mentioned defects of the prior art, the present invention provides a data optimization storage method and system based on artificial intelligence to solve the problems existing in the above-mentioned background technology.
[0008] The present invention provides the following technical solution: a data optimization storage method based on artificial intelligence, comprising:
[0009] S1: used to collect data to be stored, and clean, compress, and decompress the collected data;
[0010] S2: Extract features from the decompressed raw data and access logs, and perform preliminary data classification based on the extracted features;
[0011] S3: Based on the preliminary classification results and the actual access pattern data of the data, a tiered storage strategy is formulated to store the data on appropriate media;
[0012] S4: tiered storage status based on real-time monitoring data, including data access speed, storage utilization, I / O operation load, performance bottlenecks, and anomalies, and parallel analysis of storage status data at each tier;
[0013] S5: Based on the storage status monitoring results of each layer, dynamically monitor the storage status of data at each layer, determine whether the data stored at each layer meets the current storage tier, and migrate data that does not meet the current storage tier based on the monitoring results;
[0014] S6: Used to monitor data migration results in real time and analyze adjusted system performance to evaluate the migration effect. The evaluation results are then fed back to S2 and S3, forming a closed-loop optimization process for data storage.
[0015] Preferably, the S1 identifies the data sources that need to be collected, including databases, log files, sensors, and external APIs, uses ETI tools to extract data from the data sources, transmits the data to the storage system through a secure channel, and temporarily stores the data in a buffer or a temporary database during the data collection process. The collected data that needs to be stored is then deduplicated, missing value processed, noise and outlier processed, format unified, and verified. The cleaned data is compressed according to a compression algorithm, and the compressed data is transmitted to a central server. The compressed data is decompressed on the central server or cloud platform and restored to its original format.
[0016] Preferably, the S2 collects access logs related to the original data, associates and integrates the decompressed original data with the access logs, and then extracts data access frequency, data access time, data size, data type and business criticality from the integrated data, and preprocesses the extracted feature data. The supervised learning algorithm then automatically classifies and analyzes the preprocessed feature data to divide the original data into different storage categories.
[0017] Preferably, the S3 divides the data into different storage tiers, including a high-performance tier, a standard tier, a low-frequency tier, and an archive tier, by mapping the classification results into business-readable tags and combining the actual access pattern data of the data. Each tier corresponds to different storage media and access delay requirements. According to the automated script, the data is migrated to the classified storage media according to the tiered storage strategy, so as to perform tiered storage of the classified data.
[0018] Preferably, the S4 is used to monitor the access speed, storage utilization, I / O operation load, performance bottlenecks and anomalies of the high-performance layer, standard layer, low-frequency layer and archive layer respectively, and perform real-time analysis on the storage status data of each layer, and calculate the high-performance layer storage status monitoring index, standard layer storage status monitoring index, low-frequency layer storage status monitoring index and archive layer storage status monitoring index respectively, which are used to monitor the data access status of each layer after hierarchical storage and verify whether the classification is accurate.
[0019] Preferably, S5 sets a high-performance status threshold for the high-performance layer, a standard status threshold for the standard layer, and a low-frequency status threshold for the low-frequency layer based on the storage status monitoring results of each layer; the high-performance status threshold, the standard status threshold, and the low-frequency status threshold are set and adjusted based on the analysis of historical access data and storage status monitoring results;
[0020] If the high-performance tier storage status monitoring index is greater than or equal to the high-performance status threshold, the data stored in the high-performance tier meets the current storage tier requirements and no data migration is required. If the high-performance tier storage status monitoring index is less than the high-performance status threshold and the high-performance tier storage status monitoring index is greater than or equal to the standard status threshold, it indicates that the access demand for the current data in the high-performance tier has decreased and AI is required to migrate the data from the high-performance tier to the standard tier.
[0021] If the standard tier storage status monitoring index is less than the standard status threshold, it indicates that the current data storage status does not meet the standard tier requirements, and the data will be migrated from the standard tier to the low-frequency tier. If the standard tier storage status monitoring index is greater than or equal to the high-performance status threshold, the reverse migration mechanism of the standard tier will be triggered, and the data will be migrated from the standard tier to the high-performance tier through AI.
[0022] If the low-frequency tier storage status monitoring index is less than the low-frequency status threshold, it indicates that the current data storage status does not meet the requirements of the low-frequency tier, and the data is migrated from the low-frequency tier to the archive tier. If the low-frequency tier storage status monitoring index is greater than or equal to the standard status threshold, the reverse migration mechanism of the low-frequency tier is triggered, and the data is migrated from the low-frequency tier to the standard tier through AI.
[0023] If the archive layer storage status monitoring index is less than the low-frequency status threshold, it indicates that the data status meets the requirements of the archive layer and no data migration is required. If the archive layer storage status monitoring index is greater than or equal to the low-frequency status threshold, the reverse migration mechanism of the archive layer is triggered, and the data is migrated from the archive layer to the low-frequency layer through AI.
[0024] Preferably, S6 is used to collect system performance indicator data after data migration in real time, analyze the collected system performance indicator data, and calculate the system performance evaluation coefficient. By setting a system performance threshold, the system performance evaluation coefficient is compared with the system performance threshold to determine whether the current migration has achieved the expected effect. When the system performance evaluation coefficient is greater than or equal to the system performance threshold, it means that the system performance after migration is better. When the system performance evaluation coefficient is less than the system performance threshold, it means that the system performance after migration has deteriorated or has not achieved the expected effect. The judgment result at this time is fed back to S2 and S3 to re-execute the data feature extraction and classification of S2, and adjust the tiered storage strategy in S3 to further optimize and monitor the data migration until the system is monitored to be in the best operating state. Stop executing the data migration optimization task, but still need to maintain continuous monitoring of the system, and start the optimization process in time when the system performance deteriorates.
[0025] To achieve the above objectives, the present invention provides the following technical solutions: an artificial intelligence-based data optimization storage system, which implements the artificial intelligence-based data optimization storage method, comprising:
[0026] Data collection module: used to collect data that needs to be stored, and clean, compress and decompress the collected data;
[0027] Data classification module: This module extracts features from the decompressed raw data and access logs and performs preliminary classification of the data based on the extracted features.
[0028] Data tiered storage module: Based on the preliminary classification results and the actual data access pattern data, it formulates a tiered storage strategy and stores the data on the appropriate media;
[0029] Tiered storage status analysis module: This module monitors the tiered storage status of data in real time, including data access speed, storage utilization, I / O operation load, performance bottlenecks, and anomalies, and performs parallel analysis of storage status data at each tier.
[0030] Storage status monitoring module: Based on the storage status monitoring results of each layer, it dynamically monitors the storage status of data at each layer, determines whether the data stored at each layer meets the current storage tier, and migrates data that does not meet the current storage tier based on the monitoring results;
[0031] Data migration assessment module: This module is used to monitor data migration results in real time and analyze the performance of the adjusted system to evaluate the migration effect. The assessment results are then fed back to S2 and S3, forming a closed-loop optimization process for data storage.
[0032] Technical effects and advantages of the present invention:
[0033] (1) By extracting features from the decompressed raw data and access logs and performing preliminary classification of the data based on these features, we can more accurately understand the access patterns and requirements of the data, providing strong support for the subsequent formulation of tiered storage strategies. Based on the preliminary classification results and the actual access pattern data of the data, a tiered storage strategy can be formulated to store the data on appropriate media. Data can be allocated to storage tiers with different performance according to the access frequency and importance of the data, thereby optimizing storage performance.
[0034] (2) Based on the hierarchical storage status of real-time monitoring data and parallel analysis of the storage status data of each layer, it is conducive to timely discovery of problems in the storage system. Based on the storage status monitoring results of each layer, the storage status of data at each layer is dynamically monitored, and data that does not conform to the current storage layer is migrated according to the monitoring results. This can ensure that data is always stored at the most appropriate layer and maintain the efficient operation of the storage system.
[0035] (3) Based on real-time monitoring of data migration results and analysis of adjusted system performance, the migration effect is evaluated and then fed back to the data classification to form a closed-loop optimization process for data storage, which is conducive to continuous monitoring and adjustment of storage system performance to ensure that the storage system is always in the best operating state. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A diagram showing the steps of the method of the present invention.
[0037] Figure 2 This is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0038] The technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. In addition, the forms of the various structures described in the following embodiments are merely examples. The artificial intelligence-based data optimization storage method and system involved in the present invention are not limited to the various structures described in the following embodiments. All other implementations obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0039] like Figure 1 This embodiment provides a data optimization storage method based on artificial intelligence, including:
[0040] S1: Used to collect data that needs to be stored, and clean, compress, and decompress the collected data.
[0041] In this embodiment, the S1 identifies the data sources that need to be collected, including databases, log files, sensors, and external APIs, uses ETI tools to extract data from the data sources, and transmits the data to the storage system through a secure channel. During the data collection process, the data is temporarily stored in a buffer or a temporary database. The collected data to be stored is then deduplicated, missing value processed, noise and outlier processed, format unified, and verified. The cleaned data is compressed according to a compression algorithm, and the compressed data is transmitted to the central server. On the central server or cloud platform, the compressed data is decompressed and restored to its original format.
[0042] S2: Extract features from the decompressed raw data and access logs, and perform preliminary classification of the data based on the extracted features.
[0043] In this embodiment, S2 collects access logs related to the original data, associates and integrates the decompressed original data with the access logs, and then extracts data access frequency, data access time, data size, data type and business criticality from the integrated data, and preprocesses the extracted feature data. The supervised learning algorithm then automatically classifies and analyzes the preprocessed feature data to divide the original data into different storage categories.
[0044] It should be specifically explained that by creating a mapping table, the data identifiers (such as file names, data IDs, etc.) in the access log are associated with the decompressed data files, and the records in the access logs and the corresponding original data file information are integrated into a unified data set. Feature extraction and preprocessing are then performed on the integrated data set. Appropriate supervised learning algorithms, such as decision trees, logistic regression, neural networks, etc., are selected. The preprocessed feature data are divided into training sets, validation sets, and test sets. The selected algorithm is trained on the training data set. The loss function is minimized by optimizing the algorithm parameters (such as learning rate, tree depth), and finally a supervised learning model with predictive capabilities is generated to predict data categories based on the feature data input. In the supervised learning model, the original data is divided into the corresponding category labels and the classification results are output. The category labels include storage priority category, access mode category, business value category, and storage medium recommendation category. Among them, the storage priority category includes "hot data", "warm data", "cold data" or "archive data", the access mode category includes "high frequency access", "low frequency access", "periodic access", and "one-time access", the business value category includes "critical business data", "important business data", "general business data", and "historical data", and the storage medium recommendation category includes "recommended to store in SSD", "recommended to store in HDD", "recommended to store in object storage", and "recommended to delete or archive".
[0045] S3: Based on the preliminary classification results and the actual access pattern data of the data, a tiered storage strategy is formulated to store the data on appropriate media.
[0046] In this embodiment, the S3 maps the classification results into business-readable tags and combines the actual access pattern data of the data to divide the data into different storage tiers, including a high-performance tier, a standard tier, a low-frequency tier, and an archive tier. Each tier corresponds to different storage media and access latency requirements. According to the automated script, the data is migrated to the classified storage media according to the tiered storage strategy, so as to perform tiered storage of the classified data.
[0047] It should be noted that the classification results output by the model, such as storage priority (hot, warm, cold, archive), access mode (high frequency, low frequency, periodic, one-time), etc., are converted into labels that are easy for business personnel to understand. For example, storage priority: hot data "Key business high-frequency data", warm data "Important business occasionally accesses data", cold data "General business low-frequency data", archived data "Historical compliance archive data", access mode: high-frequency access "Multiple visits per day", low frequency visits "Several visits per month", periodic visits "Quarterly / annual cycle visit", one-time visit "Archive after single use"; combine the classification results output by the model with the actual access pattern data of the data itself, and the actual access pattern data of the data itself is obtained by analyzing the access logs, including high-frequency access, low-frequency access, periodic access, one-time access, etc., and divide the data into high-performance tier, standard tier, low-frequency tier and archive tier. Among them, the high-performance tier is for storing hot data and frequently accessed data, using high-performance media SSD, and the access delay is usually in the millisecond level. The standard tier is for storing warm data and occasionally accessed data, using medium-performance media HDD, and the access delay is usually between a few milliseconds and tens of milliseconds. The low-frequency tier is for storing cold data and infrequently accessed data, using object storage or tape library as the medium, and the access delay may be longer, but the cost is lower. The archive tier is for storing archived data and data that is rarely accessed, using tape library or cloud archive storage as the medium, with the longest access delay but the lowest cost. According to the developed automation script, data is automatically migrated to the corresponding storage medium according to the tiered storage strategy.
[0048] S4: Monitors the tiered storage status of data in real time, including data access speed, storage utilization, I / O operation load, performance bottlenecks, and anomalies, and performs parallel analysis of storage status data at each tier.
[0049] In this embodiment, the S4 is used to monitor the access speed, storage utilization, I / O operation load, performance bottlenecks and anomalies of the high-performance layer, standard layer, low-frequency layer and archive layer respectively, and perform real-time analysis on the storage status data of each layer, and calculate the high-performance layer storage status monitoring index, standard layer storage status monitoring index, low-frequency layer storage status monitoring index and archive layer storage status monitoring index respectively, which are used to monitor the data access status of each layer after hierarchical storage and verify whether the classification is accurate.
[0050] It should be noted that the high-performance tier storage status monitoring index is mainly used to monitor the high-speed access and low latency of the high-performance tier. The high-performance tier storage status monitoring index is calculated by analyzing the access speed, storage utilization, and I / O operation load. The specific calculation formula is: ,in, Indicates the actual access speed of the high-performance tier. Indicates the maximum access speed of the high-performance tier. Indicates the actual storage utilization of the high-performance tier. Indicates the maximum storage utilization of the high-performance tier. Indicates the current I / O operation load of the high-performance tier. Indicates the I / O operation load threshold of the high-performance tier. Represents the weight coefficient of the high performance layer;
[0051] The Standard Tier Storage Status Monitoring Index is mainly used to monitor the balance between performance and cost of the Standard Tier. It is calculated by analyzing storage utilization, I / O operation load, and performance bottlenecks. The specific calculation formula is: ,in, Indicates the actual storage utilization of the standard tier. Indicates the maximum storage utilization of the standard tier. Indicates the current I / O operation load of the standard layer. Indicates the maximum I / O operation load of the standard tier. Indicates the performance bottleneck indicator score of the standard layer. Represents the weight coefficient of the standard layer;
[0052] The low-frequency layer storage status monitoring index is mainly used to monitor the infrequently accessed data in the low-frequency layer. The low-frequency layer storage status monitoring index is calculated by analyzing the access speed, storage utilization and abnormal conditions. The specific calculation formula is: ,in, Indicates the actual access speed of the low-frequency layer. Indicates the maximum access speed of the low-frequency layer. Indicates the actual storage utilization of the low-frequency layer. Indicates the maximum storage utilization of the low-frequency layer, represents the abnormality score of the low-frequency layer, Represents the weight coefficient of the low-frequency layer;
[0053] The archive layer storage status monitoring index is mainly used to monitor the long-term storage status of the archive layer. It is calculated by analyzing the access speed, storage utilization and abnormal conditions. The specific calculation formula is: ,in, Indicates the actual access speed of the low-frequency layer. Indicates the maximum access speed of the low-frequency layer. Indicates the actual storage utilization of the low-frequency layer. Indicates the maximum storage utilization of the low-frequency layer, represents the abnormality score of the low-frequency layer, Represents the weight coefficient of the low-frequency layer.
[0054] S5: Based on the storage status monitoring results of each layer, dynamically monitor the storage status of data at each layer, determine whether the data stored at each layer meets the current storage layer, and migrate data that does not meet the current storage layer according to the monitoring results.
[0055] In this embodiment, S5 sets a high-performance status threshold for the high-performance layer, a standard status threshold for the standard layer, and a low-frequency status threshold for the low-frequency layer based on the storage status monitoring results of each layer. The high-performance status threshold, standard status threshold, and low-frequency status threshold are set and adjusted based on the analysis of historical access data and storage status monitoring results.
[0056] If the high-performance tier storage status monitoring index is greater than or equal to the high-performance status threshold, the data stored in the high-performance tier meets the current storage tier requirements and no data migration is required. If the high-performance tier storage status monitoring index is less than the high-performance status threshold and the high-performance tier storage status monitoring index is greater than or equal to the standard status threshold, it indicates that the access demand for the current data in the high-performance tier has decreased and AI is required to migrate the data from the high-performance tier to the standard tier.
[0057] If the standard tier storage status monitoring index is less than the standard status threshold, it indicates that the current data storage status does not meet the standard tier requirements, and the data will be migrated from the standard tier to the low-frequency tier. If the standard tier storage status monitoring index is greater than or equal to the high-performance status threshold, the reverse migration mechanism of the standard tier will be triggered, and the data will be migrated from the standard tier to the high-performance tier through AI.
[0058] If the low-frequency tier storage status monitoring index is less than the low-frequency status threshold, it indicates that the current data storage status does not meet the requirements of the low-frequency tier, and the data is migrated from the low-frequency tier to the archive tier. If the low-frequency tier storage status monitoring index is greater than or equal to the standard status threshold, the reverse migration mechanism of the low-frequency tier is triggered, and the data is migrated from the low-frequency tier to the standard tier through AI.
[0059] If the archive layer storage status monitoring index is less than the low-frequency status threshold, it indicates that the data status meets the requirements of the archive layer and no data migration is required. If the archive layer storage status monitoring index is greater than or equal to the low-frequency status threshold, the reverse migration mechanism of the archive layer is triggered, and the data is migrated from the archive layer to the low-frequency layer through AI.
[0060] S6: Used to monitor data migration results in real time and analyze adjusted system performance to evaluate the migration effect. The evaluation results are then fed back to S2 and S3, forming a closed-loop optimization process for data storage.
[0061] In this embodiment, S6 is used to collect system performance indicator data after data migration in real time, analyze the collected system performance indicator data, and calculate the system performance evaluation coefficient. By setting a system performance threshold, the system performance evaluation coefficient is compared with the system performance threshold to determine whether the current migration has achieved the expected effect. When the system performance evaluation coefficient is greater than or equal to the system performance threshold, it means that the system performance after migration is better. When the system performance evaluation coefficient is less than the system performance threshold, it means that the system performance after migration has deteriorated or has not achieved the expected effect. The judgment result at this time is fed back to S2 and S3 to re-execute the data feature extraction and classification of S2, and adjust the tiered storage strategy in S3 to further optimize and monitor the data migration until the system is monitored to be in the best operating state. The data migration optimization task is stopped, but continuous monitoring of the system is still required to be maintained, and the optimization process is started in time when the system performance deteriorates.
[0062] It should be noted that the system performance evaluation coefficient is obtained by collecting the system performance indicators after migration, such as response time, throughput, resource utilization (CPU, memory, disk I / O, etc.), analyzing and calculating the collected system performance indicators. The specific calculation formula is: ,in, represents the actual value of the current i-th system performance indicator, represents the target value of the i-th system performance indicator, represents the weight coefficient of the i-th system performance indicator, n represents the number of collected system performance indicators; and the system performance threshold is set based on the analysis results of the system's historical performance data.
[0063] like Figure 2The embodiment shown provides an implementation system corresponding to the artificial intelligence-based data optimization storage method, including a data collection module, a data classification module, a data tiered storage module, a tiered storage status analysis module, a storage status monitoring module, and a data migration assessment module. The data collection module is connected to the data classification module, the data classification module is connected to the data tiered storage module, the data tiered storage module is connected to the tiered storage status analysis module, the tiered storage status analysis module is connected to the storage status monitoring module, the storage status monitoring module is connected to the data migration assessment module, and the data migration assessment module is connected to the data classification module and the data tiered storage module.
[0064] The data collection module is used to collect data that needs to be stored, and clean, compress and decompress the collected data;
[0065] The data classification module extracts features from the decompressed raw data and access logs, and performs preliminary classification of the data based on the extracted features;
[0066] The data tiered storage module formulates a tiered storage strategy based on the preliminary classification results and the actual access pattern data of the data, and stores the data on appropriate media;
[0067] The tiered storage status analysis module monitors the tiered storage status of data in real time, including data access speed, storage utilization, I / O operation load, performance bottlenecks and anomalies, and performs parallel analysis on the storage status data of each tier;
[0068] The storage status monitoring module dynamically monitors the storage status of data at each level based on the storage status monitoring results of each level, determines whether the data stored at each level meets the current storage level, and migrates data that does not meet the current storage level based on the monitoring results;
[0069] The data migration evaluation module is used to monitor the data migration results in real time and analyze the adjusted system performance to evaluate the migration effect, and then feed the evaluation results back to S2 and S3 to form a closed-loop optimization process for data storage.
[0070] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0071] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data optimization storage method based on artificial intelligence, characterized in that: include: S1: used to collect data to be stored, and clean, compress, and decompress the collected data; S2: Extract features from the decompressed raw data and access logs, and perform preliminary data classification based on the extracted features; S3: Based on the preliminary classification results and the actual access pattern data of the data, a tiered storage strategy is formulated to store the data on appropriate media; S4: tiered storage status based on real-time monitoring data, including data access speed, storage utilization, I / O operation load, performance bottlenecks, and anomalies, and parallel analysis of storage status data at each tier; S5: Based on the storage status monitoring results of each layer, dynamically monitor the storage status of data at each layer, determine whether the data stored at each layer meets the current storage tier, and migrate data that does not meet the current storage tier based on the monitoring results; S6: Used to monitor data migration results in real time and analyze adjusted system performance to evaluate the migration effect. The evaluation results are then fed back to S2 and S3, forming a closed-loop optimization process for data storage.
2. The data optimization storage method based on artificial intelligence according to claim 1 is characterized in that: The S1 identifies the data sources that need to be collected, including databases, log files, sensors, and external APIs, uses ETI tools to extract data from the data sources, and transmits the data to the storage system through a secure channel. During the data collection process, the data is temporarily stored in a buffer or a temporary database. The collected data that needs to be stored is then deduplicated, missing value processed, noise and outlier processed, format unified, and verified. The cleaned data is compressed according to a compression algorithm, and the compressed data is then transmitted to a central server. The compressed data is decompressed on the central server or cloud platform and restored to its original format.
3. The data optimization storage method based on artificial intelligence according to claim 2 is characterized in that: The S2 collects access logs related to the original data, associates and integrates the decompressed original data with the access logs, and then extracts data access frequency, data access time, data size, data type and business criticality from the integrated data, and preprocesses the extracted feature data. The supervised learning algorithm then automatically classifies and analyzes the preprocessed feature data to divide the original data into different storage categories.
4. The data optimization storage method based on artificial intelligence according to claim 3 is characterized in that: The S3 maps the classification results into business-readable tags and combines the actual access pattern data of the data to divide the data into different storage tiers, including high-performance tier, standard tier, low-frequency tier and archive tier. Each tier corresponds to different storage media and access latency requirements. According to the automated script, the data is migrated to the classified storage media according to the tiered storage strategy, so as to perform tiered storage of the classified data.
5. The data optimization storage method based on artificial intelligence according to claim 4 is characterized in that: The S4 is used to monitor the access speed, storage utilization, I / O operation load, performance bottlenecks and anomalies of the high-performance layer, standard layer, low-frequency layer and archive layer respectively, and perform real-time analysis on the storage status data of each layer, and calculate the high-performance layer storage status monitoring index, standard layer storage status monitoring index, low-frequency layer storage status monitoring index and archive layer storage status monitoring index respectively, which are used to monitor the data access status of each layer after hierarchical storage and verify whether the classification is accurate.
6. The data optimization storage method based on artificial intelligence according to claim 5 is characterized in that: S5 sets a high-performance status threshold for the high-performance layer, a standard status threshold for the standard layer, and a low-frequency status threshold for the low-frequency layer based on the storage status monitoring results of each layer; the high-performance status threshold, the standard status threshold, and the low-frequency status threshold are set and adjusted based on the analysis of historical access data and storage status monitoring results; If the high-performance tier storage status monitoring index is greater than or equal to the high-performance status threshold, the data stored in the high-performance tier meets the current storage tier requirements and no data migration is required. If the high-performance tier storage status monitoring index is less than the high-performance status threshold and the high-performance tier storage status monitoring index is greater than or equal to the standard status threshold, it indicates that the access demand for the current data in the high-performance tier has decreased and AI is required to migrate the data from the high-performance tier to the standard tier. If the standard tier storage status monitoring index is less than the standard status threshold, it indicates that the current data storage status does not meet the standard tier requirements, and the data will be migrated from the standard tier to the low-frequency tier. If the standard tier storage status monitoring index is greater than or equal to the high-performance status threshold, the reverse migration mechanism of the standard tier will be triggered, and the data will be migrated from the standard tier to the high-performance tier through AI. If the low-frequency tier storage status monitoring index is less than the low-frequency status threshold, it indicates that the current data storage status does not meet the requirements of the low-frequency tier, and the data is migrated from the low-frequency tier to the archive tier. If the low-frequency tier storage status monitoring index is greater than or equal to the standard status threshold, the reverse migration mechanism of the low-frequency tier is triggered, and the data is migrated from the low-frequency tier to the standard tier through AI. If the archive layer storage status monitoring index is less than the low-frequency status threshold, it indicates that the data status meets the requirements of the archive layer and no data migration is required. If the archive layer storage status monitoring index is greater than or equal to the low-frequency status threshold, the reverse migration mechanism of the archive layer is triggered, and the data is migrated from the archive layer to the low-frequency layer through AI.
7. The data optimization storage method based on artificial intelligence according to claim 6 is characterized in that: The S6 is used to collect system performance indicator data after data migration in real time, analyze the collected system performance indicator data, and calculate the system performance evaluation coefficient. By setting a system performance threshold, the system performance evaluation coefficient is compared with the system performance threshold to determine whether the current migration has achieved the expected effect. When the system performance evaluation coefficient is greater than or equal to the system performance threshold, it means that the system performance after migration is better. When the system performance evaluation coefficient is less than the system performance threshold, it means that the system performance after migration has deteriorated or has not achieved the expected effect. The judgment result at this time is fed back to S2 and S3 to re-execute the data feature extraction and classification of S2, and adjust the tiered storage strategy in S3 to further optimize and monitor the data migration until the system is monitored to be in the best operating state. The data migration optimization task is stopped, but continuous monitoring of the system is still required. When the system performance deteriorates, the optimization process is started in time.
8. An artificial intelligence-based data optimization storage system, implementing the artificial intelligence-based data optimization storage method according to any one of claims 1 to 7, characterized in that: include: Data collection module: used to collect data that needs to be stored, and clean, compress and decompress the collected data; Data classification module: This module extracts features from the decompressed raw data and access logs and performs preliminary classification of the data based on the extracted features. Data tiered storage module: Based on the preliminary classification results and the actual data access pattern data, it formulates a tiered storage strategy and stores the data on the appropriate media; Tiered storage status analysis module: This module monitors the tiered storage status of data in real time, including data access speed, storage utilization, I / O operation load, performance bottlenecks, and anomalies, and performs parallel analysis of storage status data at each tier. Storage status monitoring module: Based on the storage status monitoring results of each layer, it dynamically monitors the storage status of data at each layer, determines whether the data stored at each layer meets the current storage tier, and migrates data that does not meet the current storage tier based on the monitoring results; Data migration assessment module: This module is used to monitor data migration results in real time and analyze the performance of the adjusted system to evaluate the migration effect. The assessment results are then fed back to S2 and S3, forming a closed-loop optimization process for data storage.
Citation Information
Patent Citations
File data hierarchical storage method and device, medium and electronic equipment
CN118672520A
Data storage method and system based on legal knowledge service platform and storage medium
CN118964496A
File storage management method and system based on computer system resource use
CN119938601A
Cache management optimization method and system
CN119961189A
Artificial intelligence big data processing method and system for smart traffic
CN119964365A
Cited By
Log management system
CN121351425A