Hybrid high-performance storage method and system and medium
By adopting hybrid high-performance storage methods in the storage system, dynamically migrate data and automatically adjust data distribution, the contradiction between storage performance and cost in the existing technology is solved, and the intelligence and flexibility of data management are improved to adapt to the needs of different business scenarios.
Patent Information
- Application Number
- CN202510289622.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing storage solutions are difficult to flexibly switch between block storage and object storage when handling business load fluctuations, resulting in an increase in storage costs and lack of intelligent data management and automatic migration mechanisms, which increases management complexity.
The hybrid high-performance storage method is adopted to determine the data storage location through the metadata server, dynamically migrate data between the block storage cache layer and the object storage layer based on a predetermined strategy, and automatically adjust the data distribution using an intelligent hierarchical management mechanism, and expand or reduce the storage capacity of the block storage cache layer as needed.
It achieves a balance between performance and cost, improves the intelligence and flexibility of data management, adapts to the needs of different business scenarios, and reduces the management complexity of the storage system.
Smart Images

Figure CN120179173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of storage technologies, and in particular, to a hybrid high-performance storage method, system, and medium. Background Art
[0002] With the rapid development of cloud computing technologies, storage systems have become a core component in data management and architecture design. In fields such as big data processing, cloud storage, and enterprise-level data management, how to balance storage costs and performance has become a key issue. Currently, block storage and object storage are two main solutions for cloud storage. Block storage is widely used in scenarios that require quick responses due to its low latency and high-performance characteristics, while object storage has become the preferred choice for long-term storage due to its strong scalability and low cost.
[0003] However, existing storage solutions have certain limitations when dealing with peak and valley business requirements. When the business load fluctuates, it is difficult for the storage system to flexibly switch between block storage and object storage, which may lead to an increase in storage costs. In addition, during the data migration process, there are often high time and cost overheads. When two storage solutions need to be used simultaneously, there is a lack of intelligent data management and automatic migration mechanisms, and users have to manually adjust the storage strategy, increasing the management complexity and failing to fully optimize resources. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a hybrid high-performance storage method, system, and medium, which combines low-cost object storage with high-performance block storage, is conducive to solving the contradiction between storage performance and cost in the prior art, and improves the intelligence and flexibility of data management to meet the requirements of different business scenarios.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions.
[0006] In a first aspect, a hybrid high-performance storage method provided by the present invention adopts the following technical solutions: Receive a data read or write request from a client; Determine the storage location of the requested data through a metadata server; Obtain or store the data from the block storage cache layer or the object storage layer according to the storage location; Dynamically migrate data between the block storage cache layer and the object storage layer based on a predetermined policy, where the block storage cache layer is used to store hot data and the object storage layer is used to store cold data; And, through an intelligent hierarchical management mechanism, automatically adjust the distribution of data between the block storage cache layer and the object storage layer according to preset data level factors.
[0007] Furthermore, in the above-mentioned hybrid high-performance storage method, it further includes: Expanding or reducing the storage capacity of the block storage cache layer as needed to adapt to fluctuations in business load; Only write back the changed data to the block storage cache layer to reduce redundant transmission during the data migration process.
[0008] Furthermore, in the above-mentioned hybrid high-performance storage method, the predetermined policy includes: Set a dynamically adjustable threshold according to the access frequency of data. When the data access frequency exceeds the threshold, migrate the data from the object storage layer to the block storage cache layer; When the data access frequency is lower than the threshold, migrate the data from the block storage cache layer to the object storage layer.
[0009] Furthermore, in the above-mentioned hybrid high-performance storage method, it further includes: Set a time-to-live for the data stored in the block storage cache layer; When the time-to-live of the data expires, migrate the data from the block storage cache layer to the object storage layer.
[0010] Furthermore, in the above-mentioned hybrid high-performance storage method, the intelligent hierarchical management mechanism includes: Classify data according to the importance, life cycle, and access pattern of the data; Based on the classification results, allocate different storage levels and migration strategies for different categories of data.
[0011] Furthermore, in the above-mentioned hybrid high-performance storage method, it further includes: Implement an intelligent prefetching mechanism to predict the data to be accessed according to the historical access pattern of the client and the data usage trend; Pre-load the predicted data from the object storage layer into the block storage cache layer.
[0012] Furthermore, in the above-mentioned hybrid high-performance storage method, the predicting the data to be accessed according to the historical access pattern of the client and the data usage trend includes: Collect the historical data access records of the client; Use machine learning algorithms to analyze the historical data access records to identify the time pattern and correlation of data access; Based on the identified time pattern and correlation, predict the data that may be accessed within a future period of time.
[0013] Furthermore, in the above-mentioned hybrid high-performance storage method, it further includes: Monitor the system load and storage resource usage; Automatically adjust the data distribution ratio between the block storage cache layer and the object storage layer according to the monitoring results.
[0014] Furthermore, in the above hybrid high-performance storage method, it further includes: Perform batch processing on write requests, combining multiple small write operations into larger write operations; Before writing data to the object storage layer, compress and deduplicate the data to reduce storage space occupancy and data transmission volume.
[0015] In a second aspect, a hybrid high-performance storage system provided by the present invention adopts the following technical solutions: A client interface for receiving data read or write requests; A metadata server for determining the storage location of the requested data; A block storage cache layer for storing hot data; An object storage layer for storing cold data; A data migration module for dynamically migrating data between the block storage cache layer and the object storage layer based on a predetermined policy; And an intelligent hierarchical management module for automatically adjusting the distribution of data between the block storage cache layer and the object storage layer according to preset data level factors.
[0016] Furthermore, in the above hybrid high-performance storage system, it further includes: A storage capacity adjustment module for expanding or reducing the storage capacity of the block storage cache layer as needed to adapt to fluctuations in the business load; A data write-back optimization module for only writing back the changed data to the block storage cache layer to reduce redundant transmission during the data migration process.
[0017] A readable storage medium, characterized in that the readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the hybrid high-performance storage method described in any one of claims 1-9 is implemented.
[0018] In a third aspect, a readable storage medium provided by the present invention adopts the following technical solutions: A readable storage medium, the readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the hybrid high-performance storage method described in any one of the above first aspects is implemented.
[0019] In summary, compared with the prior art, the present invention includes at least one of the following beneficial technical effects: The present invention combines the advantages of block storage and object storage to achieve efficient data management and access. The block storage cache layer provides fast data read and write capabilities, suitable for handling hot data; the object storage layer provides a large-capacity and low-cost storage solution, suitable for storing cold data. Through intelligent hierarchical management and dynamic data migration, a balance can be achieved between performance and cost to meet the requirements of different business scenarios. At the same time, this method also improves the flexibility and scalability of the storage system, enabling it to better cope with the challenges brought by data volume growth and access pattern changes. Brief Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of a specific embodiment of a hybrid high-performance storage method of the present invention.
[0022] Figure 2 It is a flowchart of a specific embodiment of a hybrid high-performance storage method of the present invention.
[0023] Figure 3 It is a flowchart of another specific embodiment of a hybrid high-performance storage method of the present invention.
[0024] Figure 4 It is a flowchart of another specific embodiment of a hybrid high-performance storage method of the present invention.
[0025] Figure 5 It is a schematic structural diagram of a specific embodiment of a hybrid high-performance storage system of the present invention. Detailed Description of the Embodiments
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. In addition, it should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.
[0027] It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments of the present application. And in the following embodiments, each embodiment has its own emphasis. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0028] For the method steps described in the embodiments of the present invention, the execution order can be the order described in the specific implementation manner, or can be adjusted according to actual needs on the premise of being able to solve the technical problem. The execution orders of the steps are not listed one by one here.
[0029] Referring Figure 1 , the embodiments of the present invention provide a hybrid high-performance storage method, including: S1, receiving a data read or write request from a client; S2, determining the storage location of the requested data through a metadata server; S3, obtaining or storing the data from the block storage cache layer or the object storage layer according to the storage location; S4, dynamically migrating data between the block storage cache layer and the object storage layer based on a predetermined policy, where the block storage cache layer is used to store hot data and the object storage layer is used to store cold data; S5, automatically adjusting the distribution of data between the block storage cache layer and the object storage layer through an intelligent hierarchical management mechanism according to preset data level factors.
[0030] In some embodiments, the hybrid high-performance storage method adopts a three-layer architecture, including a client, a server (cache layer), and an object storage. The method includes receiving a data read or write request from a client, determining the storage location of the requested data through a metadata server, and obtaining or storing data from the block storage cache layer or the object storage layer according to the storage location.
[0031] After receiving the client request, the metadata server will query the metadata information it maintains to determine the storage location and related attributes of the target data, such as the data server where it is located, the cache status, etc. The metadata server then returns the query result to the client, and the client sends a specific data request to the corresponding data server according to the location information.
[0032] Data is divided into data blocks of a fixed size, and all operations are performed in units of data blocks. During the data storage and acquisition process, the block storage cache layer is used to store hot data, and the object storage layer is used to store cold data. Based on a predetermined policy, data is dynamically migrated between the block storage cache layer and the object storage layer.
[0033] In some embodiments, the predefined policy includes setting a dynamically adjustable threshold according to the access frequency of data. When the data access frequency exceeds the threshold, the data is migrated from the object storage layer to the block storage cache layer; when the data access frequency is lower than the threshold, the data is migrated from the block storage cache layer to the object storage layer. This dynamic migration mechanism ensures that frequently accessed data can be quickly responded to, while reducing storage costs.
[0034] The intelligent hierarchical management mechanism automatically adjusts the distribution of data between the block storage cache layer and the object storage layer according to predefined data level factors. In some embodiments, the intelligent hierarchical management mechanism includes classifying data according to the importance, lifecycle, and access pattern of the data, and assigning different storage levels and migration policies to different categories of data based on the classification results.
[0035] This hybrid high-performance storage method achieves a balance between high performance and low cost by combining the advantages of block storage and object storage. The intelligent hierarchical management and dynamic data migration mechanisms ensure the efficient utilization of storage resources while providing flexible data access performance. In addition, this method can adapt to different business load fluctuations, optimize storage costs while ensuring data access speed, and provide an efficient storage solution for various cloud computing and big data application scenarios.
[0036] Referring to Figure 2 , in some embodiments, the hybrid high-performance storage method includes receiving a data read request from a client. After receiving the data read request, the metadata server is used to determine the storage location of the requested data. The metadata server queries the metadata information it maintains to determine the storage location and related attributes of the target data, such as the data server where it is located, the cache status, etc. The metadata server then returns the query result to the client.
[0037] The client sends a specific data request to the corresponding data server according to the location information. After receiving the file content request from the client, the data server checks whether the file content exists in the local cache. If the file content is cached, the data server directly returns the file content to the client. If the file content is not in the cache, the data server sends a request to the object storage to read the target data and update or create a copy of the file content in the local cache.
[0038] When the data server obtains the target file, it packages the file content and returns it to the client. After the client receives and verifies the returned data, it performs corresponding operations according to business needs, such as reading the file content for processing, presenting it to the user, or further storing it in other systems.
[0039] Referring to Figure 3, in some embodiments, the hybrid high-performance storage method includes receiving a data write request from a client. After receiving the data write request, the metadata server is used to check the data write permission and verify the storage status of the data.
[0040] The metadata server checks, according to the request, whether there is a need to replace existing chunks and allocates storage space for the data. In some embodiments, if the current data needs to be updated in chunks, a suitable storage location is allocated for the data according to the space availability. If the storage space of a data chunk is full, space is reallocated or the storage scheme is adjusted.
[0041] In some embodiments, the data is split into multiple data chunks and stored at a local node. These data chunks are distributed to another adjacent node for backup to prevent data loss caused by a failure of the local node. This backup process ensures the high availability and fault tolerance of the data.
[0042] After the data chunking and backup are completed, the data server updates the distribution information of the file and synchronizes it to the metadata server. The metadata server then maintains the up-to-date storage status to ensure that the data can be located quickly.
[0043] In some embodiments, the metadata server periodically checks in the background the data that has been cached but not yet written back to the object storage. For this data, it is processed according to the configured policy: the data chunks are written back to the object storage to ensure the long-term retention of the data; according to the policy, the data chunks that have been successfully written back are deleted or copies are retained as needed to save storage space.
[0044] In some embodiments, the data is written back to the object storage asynchronously in the background to reduce the impact on performance during the write process. After the write-back, whether to retain copies or delete the data is determined by the storage policy to ensure storage efficiency and cost optimization.
[0045] After the data is successfully written, the data server returns a write completion response to the client. After the client confirms that the data has been successfully written, it continues with subsequent operations.
[0046] In some embodiments, the hybrid high-performance storage method includes expanding or shrinking the storage capacity of the block storage cache layer on demand to adapt to fluctuations in the business load. By dynamically adjusting the storage capacity, this method can provide sufficient cache space during high load periods and release excess resources during low load periods, thereby optimizing resource utilization and reducing costs.
[0047] Specifically, the adjustment of storage capacity can be achieved in various ways. In some embodiments, elastic scaling technology is adopted to automatically increase or decrease storage nodes according to a preset load threshold. When it is detected that the load exceeds the upper threshold, new storage nodes are automatically added to expand the capacity; when the load is lower than the lower threshold, redundant nodes are automatically removed to reduce the capacity.
[0048] In some other embodiments, the capacity is adjusted by dynamically allocating and reclaiming storage space. During high load, additional space is allocated from the reserved storage pool to the block storage cache layer; during low load, unused space is returned to the storage pool. This method can achieve fast capacity adjustment without changing the number of physical nodes.
[0049] To further optimize storage efficiency, the method also includes data write-back optimization. In some embodiments, only the changed data is written back to the block storage cache layer to reduce redundant transmission during the data migration process. This incremental write-back mechanism significantly reduces the data transmission volume and improves the write-back efficiency.
[0050] Specifically, a differential comparison algorithm can be used to identify the changed data blocks. When the data is updated, the new data is compared with the original data, and only the changed part is transmitted. For the unchanged data blocks, the original storage state is retained to avoid unnecessary data transmission.
[0051] In some embodiments, data compression and deduplication technologies can also be combined to further reduce the amount of data to be written back. By compressing the changed data and identifying and removing duplicate data, the data transmission and storage overhead can be significantly reduced.
[0052] In addition, in some embodiments, a batch write-back strategy can be adopted to combine multiple small write operations into larger write operations to reduce the number of I / O operations and improve the write efficiency. By setting appropriate write-back intervals and triggering conditions, the write-back efficiency is maximized while ensuring data consistency.
[0053] By combining the dynamic adjustment of storage capacity and data write-back optimization, this hybrid high-performance storage method can effectively improve the utilization rate of storage resources, reduce operating costs, and adapt to load changes in different business scenarios while ensuring performance.
[0054] In some embodiments, the predefined policy includes setting a dynamically adjustable threshold according to the access frequency of the data. This threshold is used to determine the migration of data between the block storage cache layer and the object storage layer.
[0055] Specifically, when the data access frequency exceeds the set threshold, the data is migrated from the object storage layer to the block storage cache layer; when the data access frequency is lower than the threshold, the data is migrated from the block storage cache layer to the object storage layer. This dynamic migration mechanism ensures that frequently accessed data can be quickly responded to, while reducing storage costs.
[0056] The threshold is set using a dynamic adjustment mechanism to adapt to different data access patterns and storage requirements. In some embodiments, the initial value of the threshold is set based on historical data access statistics and storage capacity. Subsequently, by continuously monitoring the data access pattern and the usage of the storage layer, the threshold is adjusted regularly.
[0057] The threshold adjustment can consider multiple factors, including but not limited to: 1. Determined by the data access frequency distribution, specifically including analyzing the overall data access pattern to determine an appropriate threshold range; 2. Determined by the storage layer capacity utilization rate, specifically including adjusting the threshold according to the current usage of the block storage cache layer and the object storage layer to balance the loads of the two storage layers; 3. Determined by performance metrics, specifically including monitoring data access latency and throughput and adjusting the threshold according to performance goals; 4. Determined according to cost factors, including considering the costs of different storage layers and optimizing the threshold to achieve a balance between performance and cost.
[0058] In some embodiments, a multi-level threshold strategy is adopted. For example, three threshold levels of high, medium, and low are set, corresponding to different data migration priorities. Data above the highest threshold is immediately migrated to the block storage cache layer; data below the lowest threshold is migrated to the object storage layer; data in the middle is determined whether to migrate according to the current storage capacity and load conditions.
[0059] During the threshold adjustment process, to avoid performance overhead caused by frequent migrations, a cooling period mechanism can be introduced. Within a certain period of time after data migration, even if the data access frequency changes, new migration operations are not triggered temporarily to ensure the stability of the storage layer.
[0060] In addition, in some embodiments, the threshold adjustment also considers the data life cycle and business importance. For data with a short life cycle or high business importance, the migration threshold can be appropriately reduced to make it easier to be retained in the block storage cache layer to provide faster access speed.
[0061] Through this dynamic threshold adjustment and data migration strategy, the efficient utilization of storage resources is achieved. While ensuring the quick response of frequently accessed data, the overall storage cost is optimized. This method can adapt to changes in data access patterns under different business scenarios and provide a flexible and efficient storage solution.
[0062] In some embodiments, the hybrid high-performance storage method further includes setting a time-to-live (TTL) for the data stored in the block storage cache layer. The time-to-live refers to the maximum time that data can be retained in the block storage cache layer. By setting the time-to-live, the data in the cache layer can be effectively managed to ensure efficient utilization of storage resources.
[0063] The setting of the time-to-live can be based on multiple factors, including but not limited to the access frequency, importance, data size, and current storage capacity of the data. In some embodiments, a dynamic time-to-live strategy is adopted to automatically adjust the length of the time-to-live according to the real-time monitored data access pattern and storage usage.
[0064] In specific implementation, a timestamp attribute can be added to each data block to record the time when the data enters the block storage cache layer. Regularly check the timestamps of the data blocks, calculate the duration of their storage in the cache layer, and compare it with the preset time-to-live.
[0065] When the time-to-live of the data expires, the data is migrated from the block storage cache layer to the object storage layer. The migration process includes the following steps:
[0066] 1. Regularly scan the data in the block storage cache layer to identify data blocks that have exceeded the time-to-live; 2. Package the identified expired data for migration. In some embodiments, the data can be compressed or encrypted to optimize the transmission efficiency and security;
[0067] 3. Transmit the prepared data packet to the object storage layer. The transmission process is asynchronous to reduce the impact on normal business operations;
[0068] 4. After the object storage layer receives the data, perform integrity verification to ensure lossless data transmission; 5. After the migration is completed, update the metadata information to record the new location and status of the data; 6. Delete the successfully migrated data from the block storage cache layer to free up storage space.
[0069] In some embodiments, to avoid the impact of data migration on performance, a migration window period can be set. Perform data migration operations during off-peak business hours to minimize the interference with normal business.
[0070] In addition, during the data migration process, an incremental migration strategy can also be implemented. For data whose time-to-live is about to expire but is still being frequently accessed, only a part of the data can be migrated, and the most recently accessed part is retained in the block storage cache layer to maintain access performance.
[0071] By implementing a data management and migration strategy based on the time-to-live, the hybrid high-performance storage method can optimize the utilization of storage resources while ensuring data access efficiency, achieving automatic hierarchical storage of hot and cold data. This method not only improves storage efficiency but also reduces management complexity, providing the most suitable storage environment for different types of data.
[0072] In some embodiments, the intelligent hierarchical management mechanism classifies data according to the importance, lifecycle, and access pattern of the data, and assigns different storage levels and migration strategies to different categories of data based on the classification results.
[0073] The data classification process involves evaluations of multiple dimensions. The importance of data can be judged by predefined business rules or machine learning algorithms, taking into account factors such as the impact of the data on business operations and regulatory compliance requirements. The lifecycle assessment is based on the creation time, last access time, and expected retention period of the data. The access pattern analysis is completed by monitoring metrics such as the access frequency and access time distribution of the data.
[0074] Based on these classification results, appropriate storage levels are assigned to different categories of data. For example, data with high importance, short lifecycle, and frequent access may be assigned to the high-performance block storage cache layer to ensure fast access and processing. In contrast, data with low importance, long lifecycle, and low access frequency may be assigned to the lower-cost object storage layer.
[0075] The formulation of the migration strategy is also based on the data classification results. For data with high importance but gradually decreasing access frequency, a progressive migration strategy may be adopted, first migrating part of the data to a lower-level storage and then fully migrating it as the access frequency further decreases. For data with a short lifecycle, a more aggressive migration strategy may be adopted to remove the data from the high-performance storage layer in a timely manner before it expires.
[0076] In some embodiments, it is supported to customize the migration strategy according to the time window and data lifecycle. For example, data migration can be set to be executed during off-peak business hours, or the migration operation can be triggered at a specific time point after the data is created. This flexible strategy configuration can meet the needs of different business scenarios, optimizing the utilization of storage resources while ensuring data access performance.
[0077] The intelligent hierarchical management mechanism dynamically adjusts the classification results and migration strategies by continuously monitoring and analyzing data characteristics. This adaptive mechanism ensures the efficient utilization of storage resources while meeting the storage requirements of different types of data. Through refined data management, a balance between storage performance and cost is achieved, providing a flexible and efficient storage solution for different business scenarios.
[0078] In some embodiments, the hybrid high-performance storage method includes implementing an intelligent prefetching mechanism that predicts the data to be accessed soon based on the client's historical access patterns and data usage trends. This mechanism analyzes the client's data access behavior to identify potential access patterns, and thus preloads the data that may be accessed from the object storage layer to the block storage cache layer in advance.
[0079] The implementation of the intelligent prefetching mechanism involves multiple steps. First, collect and store the client's historical data access records, including information such as the types of data accessed, access times, access frequencies, etc. These records form the basic data set for the prediction model.
[0080] Based on the collected data, use machine learning algorithms to analyze the historical data access records and identify the time patterns and correlations of data access. For example, the algorithm can discover that certain data is frequently accessed within a specific time period, or there is a correlation in the access order between certain data. These patterns and correlations provide important bases for predicting future data access.
[0081] Based on the identified time patterns and correlations, predict the data that may be accessed within a future period of time. The prediction process considers multiple factors, including but not limited to: historical access frequency, recent access time, correlations between data, current time, etc. The prediction results include a list of data that may be accessed and their priorities.
[0082] According to the prediction results, preload data from the object storage layer to the block storage cache layer in advance. The loading process considers multiple factors, such as the predicted access probability, data size, available space in the current cache layer, etc. In some embodiments, a hierarchical loading strategy is adopted, preferentially loading data with a higher access probability while reserving a certain amount of cache space for real-time response to data access requests that are not predicted.
[0083] To optimize the prefetching efficiency, in some embodiments, implement a dynamic adjustment mechanism. This mechanism continuously monitors the actual usage of the prefetched data and evaluates the accuracy of the prediction. According to the evaluation results, dynamically adjust the parameters of the prediction model to improve the prediction accuracy. For example, if it is found that some prefetched data has not been accessed for a long time, the weights of relevant prediction rules may be reduced; conversely, if the prefetched data is frequently accessed, the weights of the corresponding rules may be increased.
[0084] In addition, in some embodiments, the intelligent prefetching mechanism can also consider the overall load situation of the storage system. During periods of low load, the number of prefetched data may be increased; while during periods of high load, the prefetching operations may be reduced to avoid affecting normal data access.
[0085] By implementing an intelligent prefetch mechanism, the hybrid high-performance storage method can load the data that may be accessed in advance from the object storage layer to the block storage cache layer, thereby reducing data access latency and improving overall storage performance. This prediction method based on historical access patterns and data usage trends enables the storage system to manage data more intelligently and provide a faster and more efficient data access experience for clients.
[0086] In some embodiments, the hybrid high-performance storage method includes predicting the data to be accessed according to the historical access patterns and data usage trends of the client. This method analyzes the data access behavior of the client to identify potential access patterns, and thus loads the data that may be accessed in advance from the object storage layer to the block storage cache layer.
[0087] Specifically, based on the collected data (including information such as the type of accessed data, access time, access frequency, etc.), machine learning algorithms are used to analyze the historical data access records to identify the time patterns and correlations of data access. For example, the algorithm identifies that certain data is frequently accessed within a specific time period, or there is an access order correlation between certain data. These patterns and correlations provide important bases for predicting future data access.
[0088] Based on the identified time patterns and correlations, predict the data that may be accessed within a future period of time. The prediction process considers multiple factors, including but not limited to: historical access frequency, recent access time, correlations between data, current time, etc. The prediction results include a list of data that may be accessed and their priorities.
[0089] In some embodiments, the prediction method employs multiple machine learning algorithms, such as time series analysis, association rule mining, clustering analysis, etc. Time series analysis is used to identify the periodic patterns of data access, such as the access rules on a daily, weekly, or monthly basis. Association rule mining helps discover the access associations between different data. For example, data B is usually accessed after accessing data A. Clustering analysis is used to group data with similar access patterns for batch prediction and processing.
[0090] The prediction method can also consider the attributes of the data and metadata information. For example, for structured data, attributes such as the data type, size, creation time, etc. are considered; for unstructured data, information such as the file type, content keywords, etc. may be considered. These additional information helps to improve the accuracy of the prediction. In some embodiments, the prediction method adopts an incremental learning strategy. As new access records are continuously generated, the prediction model is updated regularly to adapt to the changes in the data access pattern. This dynamic adjustment mechanism ensures that the prediction model can reflect the latest data access trends. In some embodiments, the prediction method also considers context information, such as the current time, user identity, device type, etc. These context information helps to improve the precision of the prediction. For example, for different user groups or different types of devices, the prediction method adopts different model parameters. In certain embodiments, the prediction method adopts a multi-model fusion strategy. Multiple prediction models run in parallel, and each model focuses on different data access patterns or time scales. The final prediction result is obtained by synthesizing the outputs of multiple models, improving the robustness and accuracy of the prediction.
[0091] Furthermore, the prediction method also includes an adaptive threshold mechanism. According to the feedback of the prediction accuracy, the confidence threshold of the prediction is dynamically adjusted. When the prediction accuracy is high, the threshold is lowered to prefetch more data; when the accuracy drops, the threshold is raised to reduce unnecessary prefetching.
[0092] Through this prediction method based on historical access patterns and data usage trends, the hybrid high-performance storage method can manage data more intelligently and provide a faster and more efficient data access experience for the client.
[0093] Exemplarily, referring to Figure 4 , based on the prediction method described in the above embodiments, the metadata server predicts the data that the client may need to access through intelligent analysis or preset rules, and decides whether to read the corresponding data from the storage system. When the data server 1 receives a request, it first tries to obtain the data from its own cache. If the data is not found, it requests the corresponding data block from the object storage and stores it in the local cache so that subsequent accesses can directly obtain the data from the cache, improving the access efficiency. Subsequently, the metadata server updates the file layout information to ensure that subsequent data accesses can quickly find the location of the data, reducing the query time and unnecessary storage calls.
[0094] After completing data acquisition and cache update, the metadata server will perform intelligent / rule prediction again based on the access pattern and historical data to analyze which data may become hot data. The data predicted to have a high access frequency will be marked as hot data blocks and transferred from data server 1 to data server 2 to achieve the distribution of hot data. This strategy helps to balance the load, avoid a single data server becoming a performance bottleneck, and also improve the availability of data. After the data is successfully stored in data server 2, the metadata server will update the file layout again and record the new storage location of the data to optimize the access path for subsequent data requests.
[0095] In some embodiments, the hybrid high-performance storage method includes monitoring the load and storage resource usage and automatically adjusting the data distribution ratio between the block storage cache layer and the object storage layer according to the monitoring results.
[0096] Specifically, the monitoring process involves collecting multiple metrics, including but not limited to: CPU usage rate, memory usage rate, storage I / O throughput, network bandwidth utilization, etc. For storage resources, the monitoring content includes the capacity usage, read and write operation frequencies, data access patterns, etc. of the block storage cache layer and the object storage layer. These metrics are obtained through a combination of regular sampling and real-time monitoring to form a comprehensive performance and resource usage profile.
[0097] Based on the monitoring data, the automatic adjustment mechanism evaluates whether the current data distribution is reasonable. In some embodiments, preset thresholds are used to trigger the adjustment. For example, when the usage rate of the block storage cache layer exceeds 80%, or the read operation frequency of the object storage layer increases significantly, the data redistribution process is started.
[0098] The data redistribution process involves multiple steps. In some embodiments, first, the data with a low access frequency in the block storage cache layer is identified and migrated to the object storage layer to release high-performance storage space. At the same time, the data with an increasing access frequency is identified from the object storage layer and migrated to the block storage cache layer to improve the access speed.
[0099] During the adjustment process, multiple factors are considered to determine the optimal data distribution ratio. These factors include but not limited to: the current load level, the expected load change trend, the cost-benefit ratio of different storage layers, changes in the data access pattern, etc. By comprehensively analyzing these factors, the optimal data distribution ratio is dynamically calculated.
[0100] In some embodiments, the automatic adjustment mechanism adopts a progressive strategy to avoid the impact of large-scale data migration on performance. The adjustment process is carried out in multiple small batches, and the system performance is evaluated after each batch is completed. According to the evaluation results, it is decided whether to continue the next batch of adjustments.
[0101] Through this dynamic monitoring and automatic adjustment mechanism, the hybrid high-performance storage method can flexibly adjust the distribution of data among different storage layers according to the actual load and resource usage. This method not only optimizes the utilization of storage resources but also can adapt to the load changes in different business scenarios and provide continuous high-performance storage services.
[0102] In some embodiments, the hybrid high-performance storage method includes batch processing of write requests, combining multiple small write operations into larger write operations. Through the batch processing mechanism, the overhead required for processing each small write operation individually is reduced, improving the overall write efficiency. The batch processing process involves temporarily storing the received write requests in a memory buffer, and when the accumulated data volume reaches a preset threshold or after a predetermined time interval, a batch write operation is triggered. During the batch write process, multiple small write requests are combined into one large write operation, reducing the number of interactions with the storage layer, thereby reducing latency and increasing throughput.
[0103] In some embodiments, the hybrid high-performance storage method compresses and deduplicates data before writing it to the object storage layer to reduce storage space occupancy and data transfer volume. The compression process uses a suitable compression algorithm, such as LZ4 or Snappy, minimizing CPU overhead while ensuring compression efficiency. The deduplication process calculates the hash values of data blocks, identifies and eliminates duplicate data blocks, and only stores unique data blocks. Through compression and deduplication, not only is the storage cost reduced, but also the data transfer time is decreased, improving the overall storage efficiency.
[0104] In some embodiments, the hybrid high-performance storage method adopts an asynchronous background write mechanism to write data to the object storage layer. Asynchronous writing allows the main thread to return immediately after initiating a write request, while the actual write operation is handled by a background thread. This approach reduces the impact of write operations on the performance of the main thread and improves the overall response speed. During the background write process, a queue mechanism is used to manage the data to be written, ensuring that the data is written in order and updating the metadata information after the write is completed. The asynchronous write mechanism also includes failure retry and data consistency checks to ensure the reliability and integrity of the data.
[0105] The hybrid high-performance storage method described in the embodiments of the present invention realizes a high-performance, low-cost, and highly scalable storage solution through a hybrid storage architecture, intelligent cache management, and data tiering optimization. By adopting a combination of block storage and object storage, frequently accessed data is retained in the high-performance cache layer, while less frequently accessed data is stored in the low-cost object storage, effectively balancing storage performance and cost. A cluster-based cache mechanism is introduced to enhance the concurrent access capability of data, reduce the single-node storage bottleneck, and improve the overall throughput of the storage system. In addition, the solution employs intelligent data migration to dynamically adjust the storage location based on the access frequency, optimizing the utilization of storage resources. At the same time, combined with an intelligent prefetching strategy, it reduces data access latency and improves the response speed. Moreover, this solution supports incremental data migration, reducing the cost of data writing and migration, and improving the efficiency and stability of the storage system. In summary, the present invention can effectively improve data access performance, while reducing storage and management costs, and realizing efficient and intelligent data storage management in application scenarios such as cloud storage, big data processing, and high-performance computing.
[0106] The embodiments of the present invention also disclose a hybrid high-performance storage system.
[0107] Referring to Figure 5 , the hybrid high-performance storage system includes a client interface 1, a metadata server 2, a block storage cache layer 3, an object storage layer 4, a data migration module 5, an intelligent tiering management module 6, a storage capacity adjustment module 7, and a data write-back optimization module 8.
[0108] The client interface 1 is used to receive data read or write requests. The client interface 1 can be an API or a graphical user interface, supporting multiple protocols such as HTTP, FTP, etc., to meet the needs of different clients.
[0109] The metadata server 2 is used to determine the storage location of the requested data. The metadata server 2 maintains metadata such as the location information and access permissions of the data, and provides a fast query service. The metadata server 2 adopts a distributed architecture to ensure high availability and scalability.
[0110] The block storage cache layer 3 is used to store hot data. The block storage cache layer 3 adopts high-performance storage devices such as SSDs to provide low-latency, high-throughput data access. The block storage cache layer 3 supports the all-SSD mode, in which all data is stored in SSDs to obtain the best performance.
[0111] The object storage layer 4 is used to store cold data. The object storage layer 4 adopts large-capacity, low-cost storage devices, which are suitable for long-term storage of data with low access frequencies. The object storage layer 4 supports data compression and deduplication to improve storage efficiency.
[0112] The data migration module 5 is used to dynamically migrate data between the block storage cache layer 3 and the object storage layer 4 based on a predetermined policy. The data migration module 5 monitors the data access pattern and determines the data migration timing according to a preset threshold. The data migration process is asynchronous to minimize the impact on system performance.
[0113] The intelligent hierarchical management module 6 is used to automatically adjust the distribution of data between the block storage cache layer 3 and the object storage layer 4 according to preset data level factors. The intelligent hierarchical management module 6 considers factors such as data access frequency, importance, and lifecycle to dynamically optimize the data distribution. The intelligent hierarchical management module 6 supports custom storage policies, allowing users to adjust the data layering rules according to business requirements.
[0114] In some embodiments, the hybrid high-performance storage system adopts a cluster cache mechanism. Compared with the single-machine cache mechanism, the cluster cache mechanism provides higher performance and reliability through the collaborative work of multiple nodes. The cluster cache mechanism supports data sharding and replication to ensure load balancing and fault recovery capabilities.
[0115] In some embodiments, the hybrid high-performance storage system further includes a storage capacity adjustment module 7. The storage capacity adjustment module 7 is used to expand or shrink the storage capacity of the block storage cache layer 3 as needed to adapt to the fluctuations of the business load. The storage capacity adjustment module 7 monitors the system load and resource usage and automatically triggers the capacity adjustment operation.
[0116] In some embodiments, the hybrid high-performance storage system further includes a data write-back optimization module 8. The data write-back optimization module 8 is used to write back only the changed data to the block storage cache layer 3 to reduce the redundant transmission during the data migration process. The data write-back optimization module 8 adopts an incremental update strategy to identify and transmit the changed data blocks to improve the write-back efficiency.
[0117] Through the collaborative work of these components, the hybrid high-performance storage system achieves efficient data management and access, strikes a balance between performance and cost, and meets the requirements of different business scenarios. The flexibility and scalability of the system enable it to adapt to various storage needs and can be applied from small enterprises to large-scale data centers.
[0118] An embodiment of the present invention also discloses a readable storage medium.
[0119] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the hybrid high-performance storage method described in any one of the above embodiments. The computer-readable storage medium may include: any entity or device capable of carrying the computer program, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc. The computer program includes computer program code. The computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), and a software distribution medium, etc.
[0120] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0121] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices.
[0122] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hybrid high-performance storage method, characterized in that: include: Receive data read or write requests from clients; Determining the storage location of the requested data through a metadata server; Obtain or store the data from a block storage cache layer or an object storage layer according to the storage location; Based on a predetermined strategy, dynamically migrate data between the block storage cache layer and the object storage layer, the block storage cache layer is used to store hot data, and the object storage layer is used to store cold data; Furthermore, through an intelligent hierarchical management mechanism, the distribution of data between the block storage cache layer and the object storage layer is automatically adjusted according to preset data level factors.
2. The hybrid high-performance storage method according to claim 1, characterized in that: Also includes: Expand or reduce the storage capacity of the block storage cache layer as needed to adapt to fluctuations in business load; Only the changed data is written back to the block storage cache layer to reduce redundant transmission during data migration.
3. The hybrid high-performance storage method according to claim 1, characterized in that: The predetermined strategy includes: A dynamically adjustable threshold is set according to the access frequency of the data. When the data access frequency exceeds the threshold, the data is migrated from the object storage layer to the block storage cache layer. When the data access frequency is lower than the threshold, the data is migrated from the block storage cache layer to the object storage layer.
4. The hybrid high-performance storage method according to claim 1, characterized in that: Also includes: Set a lifetime for data stored in the block storage cache layer; When the data lifetime expires, the data is migrated from the block storage cache layer to the object storage layer.
5. The hybrid high-performance storage method according to claim 1, characterized in that: The intelligent hierarchical management mechanism includes: Classify data based on its importance, life cycle, and access patterns; Based on the classification results, different storage tiers and migration strategies are assigned to different categories of data.
6. The hybrid high-performance storage method according to claim 1, characterized in that: Also includes: Implement intelligent pre-fetching mechanism to predict the data to be accessed based on the client's historical access patterns and data usage trends; Predicted data is loaded from the object storage layer into the block storage cache layer in advance.
7. The hybrid high-performance storage method according to claim 6, characterized in that: The method of predicting the data to be accessed based on the client's historical access patterns and data usage trends includes: Collect historical data access records of clients; Analyze the historical data access records using a machine learning algorithm to identify temporal patterns and correlations of data access; Based on the identified temporal patterns and associations, predict which data is likely to be accessed in the future.
8. The hybrid high-performance storage method according to claim 1, characterized in that: Also includes: Monitor system load and storage resource usage; Based on the monitoring results, the data distribution ratio between the block storage cache layer and the object storage layer is automatically adjusted.
9. The hybrid high-performance storage method according to claim 1, characterized in that: Also includes: Batch write requests to combine multiple small write operations into larger write operations; Before writing data to the object storage layer, the data is compressed and deduplicated to reduce storage space usage and data transmission volume.
10. A hybrid high-performance storage system, characterized in that: The system comprises: Client interface, used to receive data read or write requests; A metadata server, which is used to determine the storage location of the requested data; Block storage cache layer, used to store hot data; Object storage layer, used to store cold data; A data migration module, used for dynamically migrating data between the block storage cache layer and the object storage layer based on a predetermined strategy; and an intelligent tiered management module, which is used to automatically adjust the distribution of data between the block storage cache layer and the object storage layer according to preset data level factors.
11. The hybrid high-performance storage system according to claim 10, characterized in that: The system further comprises: A storage capacity adjustment module, used to expand or reduce the storage capacity of the block storage cache layer as needed to adapt to fluctuations in business load; The data write-back optimization module is used to write only the changed data back to the block storage cache layer to reduce redundant transmission during data migration.
12. A readable storage medium, characterized in that: The readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the hybrid high-performance storage method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Mixed energy-saving memory system and method of reliability-based high-performance file system
CN106777342A
Consistent hash-based hierarchical mixed storage system and method
CN107844269A
A librados-based distributed NFS system and a construction method thereof
CN109783438A
Unified data access management system and method in multi-storage environment
CN118349169A
Distributed cluster file management method and system
CN118426713A
Cited By
AI-driven adaptive storage layering and cache prefetching system
CN120428926A
Data cache management method and system and electronic equipment
CN121456017A