Multi-level cloud storage cluster construction and data storage method
By building multi-level storage layers and adopting replica strategies, erasure and coding technologies, the problems of performance bottlenecks, high storage costs and insufficient scalability in cloud storage systems are solved, and efficient and reliable data storage and management are achieved.
Patent Information
- Application Number
- CN202510258959.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-17
AI Technical Summary
When handling complex and diverse storage needs, existing cloud storage systems face problems such as performance bottlenecks, high storage costs, insufficient scalability, unbalanced load and insufficient data reliability and consistency management capabilities.
By building a multi-level storage layer, including a cache layer, performance layer and capacity layer, the storage layer of data is dynamically adjusted according to the frequency and importance of data access, and replica strategies and erasure coding technology are used to achieve data redundancy and fault tolerance. At the same time, data distribution is optimized through load balancing and intelligent scheduling mechanisms, and distributed metadata management methods are used to achieve rapid query and consistency maintenance.
It realizes efficient data storage and management, reduces access latency and storage costs, improves system scalability and reliability, and ensures load balancing and data consistency.
Smart Images

Figure CN120162003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to a method for constructing a multi-level cloud storage cluster and storing data. Background Art
[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, the storage and management of massive data have become an urgent problem to be solved. Existing cloud storage systems face the following main challenges when dealing with complex and diverse storage requirements.
[0003] Currently, most cloud storage systems adopt a single storage layer architecture and cannot perform hierarchical storage management according to the access frequency and importance of data. This design results in the mixed storage of frequently accessed data and infrequently accessed data, increasing the random read and write operations of the storage system, thereby limiting the overall performance of the system. Especially in high-concurrency scenarios, it is easy to cause access latency and throughput reduction.
[0004] The load distribution in a distributed storage cluster often becomes uneven due to the dynamic change of the data access pattern or the performance difference of nodes, resulting in some nodes being overloaded and their performance degrading, while some node resources are idle and wasted. In addition, load imbalance may exacerbate the storage hot spot problem of the system, further affecting the stability of the system and the user experience.
[0005] With the continuous growth of data volume, cloud storage systems need to have good scalability. However, existing storage systems face problems such as low metadata management efficiency and high data redistribution cost during expansion. Especially in large-scale distributed storage clusters, when nodes are added or reduced, it is often necessary to rebalance the data distribution, resulting in a short-term decline in system performance. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides a method for constructing a multi-level cloud storage cluster and storing data, which solves the problems of performance bottleneck, high storage cost, insufficient scalability, load imbalance, and insufficient data reliability and consistency management ability in cloud storage systems.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for constructing a multi-level cloud storage cluster and storing data includes the following steps; S1. Construct a multi-level storage layer, including a cache layer, a performance layer, and a capacity layer. Each storage layer has different storage media and is respectively used to store hot data, warm data, and cold data; S2. Determine the data heat by monitoring the data access frequency, access time interval, and data importance; S3. Dynamically adjust the storage level of the data according to the data heat and trigger the hierarchical migration of the data; S4. Implement data redundancy and fault tolerance by adopting a replica strategy and erasure code technology during the storage process; S5. Optimize data distribution through a load balancing and intelligent scheduling mechanism to improve storage performance and system scalability; S6. Use a distributed metadata management method to achieve fast data query and consistency maintenance.
[0008] Preferably, the S1 step specifically includes the following steps; S1.1. Build it through a cache layer using high-speed storage media for storing frequently accessed data; S1.2. Build it through a performance layer using medium-speed storage media for storing warm data; S1.3. Build it through a capacity layer using low-cost storage media for storing infrequently accessed data.
[0009] Preferably, the calculation of data access heat in the S2 step is based on the following formula; , where: represents the access frequency of data ; represents the time interval since the last access; represents the importance weight of data; , , is a weight parameter.
[0010] Preferably, the hierarchical migration mechanism in the S3 step is triggered based on the following conditions; The data heat exceeds the threshold range of the current storage layer; The utilization rate of the storage layer exceeds the set maximum load threshold; The data life cycle expires and needs to be migrated to the low-cost storage layer.
[0011] Preferably, when the data heat exceeds the heat threshold of the current storage layer in the S3 step, data migration is triggered, and the data migration target is determined by the following cost function; , where; is the unit storage cost of the target storage layer k; is the data block size; is the network transmission time; is the computational load of the migration operation; , is a weight parameter.
[0012] Preferably, in the storage process of step S4, hot data adopts a replica redundancy strategy, and the number of replicas R is dynamically adjusted; , where; is the default number of replicas; is an adjustment coefficient; is the data heat; is the heat value of the hottest data.
[0013] Preferably, in step S4, cold data is stored using erasure coding technology, and the erasure coding storage efficiency formula is; , where; k is the number of data shards; m is the number of parity shards.
[0014] Preferably, in step S6, distributed metadata management uses the consistent hashing algorithm for sharded storage, and the sharding rule is; , where: M is a metadata item, representing the data object to be stored; N is the number of metadata storage nodes, representing the number of nodes in the distributed storage system; is the hash function applied to the metadata item M, generating an integer value; represents taking the modulus of the calculation result by N, so as to distribute the hash value to N storage nodes.
[0015] Preferably, in the metadata management of step S6, high-frequency metadata adopts the LRU algorithm for cache optimization, and the cache size is dynamically adjusted according to the data access frequency.
[0016] Preferably, the load balancing mechanism calculates the node load balancing degree based on the following formula; , When the load imbalance degree of node exceeds the preset threshold , data migration between nodes is triggered; where; is the load imbalance degree of node i, used to measure the load status of the node; is the number of requests received by node i; The total number of requests is the total number of requests of all nodes in the entire system; is the storage capacity of node i; The total capacity is the total storage capacity of all nodes in the entire system; is the load imbalance threshold preset for the system, used to determine whether the load needs to be adjusted.
[0017] The present invention provides a method for constructing a multi-level cloud storage cluster and storing data. It has the following beneficial effects: 1. By constructing a multi-level storage layer, the present invention adopts a hierarchical architecture in which the cache layer stores high-frequency data, the performance layer stores medium-frequency data, and the capacity layer stores low-frequency data, enabling high-heat data to be preferentially stored in high-performance media, reducing access latency, and improving the efficiency of data reading and writing. At the same time, combined with the dynamic data migration strategy, it can adjust the storage location in real time according to the data access heat, further optimizing the overall performance of the system.
[0018] 2. By using different storage media for data with different access frequencies, the present invention significantly reduces the storage cost of cold data. High-cost high-speed media are only used to store high-heat data, while cold data is stored in the low-cost capacity layer. In addition, erasure coding technology is used for cold data to further reduce the storage space requirement, greatly saving the storage cost while ensuring data redundancy; 3. By using the distributed consistent hashing algorithm to slice and manage data and metadata, when the number of storage nodes changes, only a small part of the data needs to be redistributed, reducing the overhead of data migration. This distributed management method significantly improves the scalability of the system and is suitable for the dynamic expansion requirements of large-scale cloud storage clusters. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] Please refer to the attached Figure 1 , the embodiments of the present invention provide a method for constructing a multi-level cloud storage cluster and storing data, including the following steps; S1. Build a multi-level storage layer, including a cache layer, a performance layer, and a capacity layer. Each storage layer has different storage media, which are used to store hot data, warm data, and cold data respectively; S2. Determine the data heat by monitoring the data access frequency, access time interval, and data importance; S3. Dynamically adjust the storage level of the data according to the data heat, and trigger the hierarchical migration of the data; S4. During the storage process, adopt the replica strategy and erasure code technology to achieve data redundancy and fault tolerance; S5. Optimize the data distribution through the load balancing and intelligent scheduling mechanism to improve the storage performance and system scalability; S6. Use the distributed metadata management method to achieve fast data query and consistency maintenance.
[0022] Specifically; In this embodiment, by building a multi-level storage layer, including a cache layer, a performance layer, and a capacity layer, the data is reasonably divided and managed in terms of storage levels according to the data access frequency, time interval, and importance. The core goal of this hierarchical storage architecture is to effectively improve the storage system performance while reducing the storage cost. The cache layer is mainly responsible for the fast access of high-frequency data, the performance layer takes into account both the throughput and response speed, and the capacity layer optimizes the resource utilization by storing low-frequency accessed data; In this embodiment, the multi-level storage layer includes a cache layer, a performance layer, and a capacity layer. Each storage layer depends on different types of storage media and management systems to meet the storage requirements of different types of data; Construction of the cache layer.
[0023] Storage medium selection: The cache layer uses high-speed storage media, which have the characteristics of high throughput and low latency.
[0024] The purpose of selecting this storage medium is to meet the fast access requirements of high-frequency data; Storage data range: The cache layer mainly stores data with extremely high access frequencies to ensure low-latency performance of the system in high-concurrency scenarios; The data in this layer is usually short-term active data, such as hot search keywords, real-time transaction information, etc.
[0025] Implementation method: Deploy a distributed cache system as the core management tool for the cache layer.
[0026] Utilize the key-value pair storage mechanism to support efficient data query and management; Data storage rule; When the data heat is higher than the cache layer threshold When there is data, it is preferentially stored in the cache layer; The calculation formula for data heat is as follows: , where: represents the access frequency of the data ; is the time interval between the last access time and the current time; represents the importance weight of the data; , , is the weight parameter, controlling the contribution of each index to the heat.
[0027] The construction of the performance layer; Storage medium selection: The performance layer uses enterprise-level SSDs or high-performance HDDs, taking into account both throughput performance and storage capacity.
[0028] The selection of this storage medium aims to meet the stability requirements of medium-frequency accessed data; Storage data range: The performance layer stores warm data with a moderate access frequency but high throughput requirements, usually the business data of continuous applications; Implementation method: Deploy a distributed file system to provide flexible distributed storage capabilities; The performance layer supports the block storage mode, which is suitable for scenarios that require large-scale read and write throughput; Data storage rules.
[0029] When the data heat is between the heat threshold of the cache layer and the threshold of the performance layer , the data is stored in the performance layer; The construction of the capacity layer; Storage medium selection: The capacity layer uses low-cost and large-capacity media to optimize storage costs; The selection of this storage medium is suitable for low-frequency accessed data; Storage data range: The capacity layer stores cold data, such as archived files, historical log data, etc.
[0030] Implementation method: Deploy an object storage system to achieve efficient management of unstructured data; Combine the multi-region distribution capabilities of cloud storage to achieve off-site backup and high availability of data Data storage rules: When the data heat When the heat is lower than the heat threshold of the performance layer the data is stored in the capacity layer; Hierarchical logic of the multi-level storage layer: In this embodiment, the multi-level storage layer divides data through the following logic: When the data heat the data is stored in the cache layer; When the heat threshold of the cache layer the data is stored in the performance layer; When the data heat the data is stored in the capacity layer; Distributed management and system implementation.
[0031] In this embodiment, the multi-level storage layer is uniformly coordinated and managed through a distributed management system: Data storage mapping; Use the distributed consistent hashing algorithm to determine the storage location of the data: , D is the data identifier, which is used to uniquely identify the specific data object to be stored; is the number of storage nodes, the calculation result of the hash function, indicating that the data identifier D is mapped to a hash value; System operation guarantee: The data storage adopts a multi-copy mode to ensure the reliability of the data in case of high concurrency or node failure; Configure a distributed transaction and real-time monitoring system to ensure data consistency and system stability; Summary: Through the multi-level storage design of the cache layer, performance layer and capacity layer, the present invention realizes the hierarchical management of data with different access characteristics, and optimizes the performance and resource utilization efficiency of the storage system. The cache layer provides extremely high access speed, the performance layer supports high throughput requirements, and the capacity layer significantly reduces the storage cost. This hierarchical storage architecture provides a flexible and efficient solution for complex storage requirements.
[0032] In this embodiment, for step S2, by monitoring the access frequency, access time interval and importance of the data, the access heat of the data is calculated, and based on this, the data is hierarchically divided and storage scheduling management is performed. The calculation model of the data heat is a multi-factor comprehensive evaluation model, which can dynamically reflect the access activity and importance of the data, and provide accurate support for data hierarchical storage and migration; Calculation of data access heat; In this embodiment, the data heat is calculated based on the following three main dimensions: Data access frequency , which is used to measure the access density of data within a given time window; The interval between the last access time of the data and the current time , which is used to reflect the time activity of the data; The importance weight of the data , which is used to quantify the requirements of the business for data priority; The comprehensive calculation formula for data heat is as follows: , where: represents the data 's access frequency, calculated within a certain time window T: , T is usually dynamically configured as 1 hour, 1 day, etc.; is the data 's interval between the last access time and the current time, defined as: ,
[0033] is the current time; is the data 's last access time; is the data 's importance weight, manually assigned according to business logic and data category, or automatically evaluated through a trained model; Higher weights are assigned to core business data; Lower weights are assigned to cold data or archived data; , , are weight parameters that respectively control the contributions of access frequency, time activity, and business importance to heat, and can be dynamically adjusted according to the actual scenario; The implementation process of data heat calculation; Collection of data access logs: Real-time collection of data access records through the log management system; The records include: extraction of data access features such as access timestamp, access count, and data identifier etc.; Statistical calculation of the access frequency of each data based on the access logs ; Calculation from the access timestamp ; Obtained according to data classification or pre-trained model ; Heat calculation and update: For each piece of data Calculate the heat in real time ; Use the time window sliding mechanism to dynamically update and , to maintain the real-time nature of heat calculation; Distribution and hierarchical division of data heat; In this embodiment, according to the data heat value , the data is divided into the corresponding storage layer: Data heat division rule; When the data heat , the data is stored in the cache layer; When the heat threshold of the cache layer , the data is stored in the performance layer; When the data heat , the data is stored in the capacity layer; Adjustment of data heat distribution; The update frequency of data heat is determined by the business load and is generally set to the minute level; For data with abnormal heat fluctuations, special rules can be set in combination with the business scenario to weighted adjust the value.
[0034] Optimization strategy for data heat; Dynamic weight adjustment; By analyzing the influence of different dimensions on the data storage layer, dynamically adjust the weight parameters , , ; , When the access frequency fluctuates greatly, increase the weight; When the real-time requirement of the data is high, increase the weight; When the importance of the data prevails, increase the weight; Time window adjustment: The setting of T directly affects the calculation result; A shorter T is suitable for real-time scenarios; A longer T is suitable for batch data processing scenarios; Automated heat calculation; Use machine learning algorithms to predict future access patterns and dynamically adjust the parameters of the heat calculation model; Formula expansion and refinement:
[0035] Weighted smoothing heat calculation: To avoid drastic fluctuations in heat values, exponential weighted moving average can be introduced to smooth the heat; , is the smoothed heat value of the data ; is the currently calculated heat value of the data ; is the heat value calculated last time of the data ; is the smoothing factor, usually taking values in the range of 0.3 - 0.7; Heat normalization: To facilitate hierarchical partitioning and storage allocation, the heat value is normalized to the range of [0, 1]; , is the normalized heat value of the data ; is the current heat value of the data ; is the minimum heat value of all data in the system; is the maximum heat value of all data in the system; Summary: Through the calculation and real-time update of data heat, the present invention dynamically allocates data between different storage layers, ensuring that high-heat data is preferentially stored and quickly accessed, and low-heat data utilizes low-cost storage resources. This process achieves a balance between storage performance and resource utilization rate, providing an accurate and efficient solution for data management.
[0036] In this embodiment, for step S3, the present invention dynamically adjusts the storage layer of data according to the change of data access heat and triggers the hierarchical migration of data. In a multi-level cloud storage cluster, the heat of data fluctuates with time and business requirements. By designing a dynamic migration mechanism, high-heat data can be migrated to the high-speed storage layer, and low-heat data can be migrated to the low-cost storage layer, thus optimizing storage costs while meeting performance requirements; Dynamic adjustment of storage layer and hierarchical migration; In this embodiment, the core of dynamic adjustment of storage layer and hierarchical migration includes migration trigger, migration target selection, and migration execution; Trigger conditions for data migration; Heat threshold trigger: Data heat When the heat range of the current storage layer is exceeded, migration is triggered. For example: When the data heat is at a certain level, the data is stored in the cache layer; When the heat threshold of the cache layer is reached, the data is stored in the performance layer; When the data heat is at another level, the data is stored in the capacity layer; When the utilization rate of a certain storage layer exceeds the preset maximum threshold migration is triggered to free up storage space; , If is the case, low-heat data is preferentially migrated to the next lower storage layer; Data life cycle trigger; For data with a set life cycle, if its life cycle expires, it needs to be migrated to the capacity layer for long-term storage.
[0037] The expiration trigger condition of data or archived data can be expressed as; , is the current time; is the data creation time; is the maximum life cycle; Selection of data migration target; Migration cost function: The target storage layer for data migration is selected by optimizing the following cost function; , is the unit storage cost of the target storage layer k; is the data block size; is the network transfer time for the data to migrate from the current layer to the target layer, and the formula is: , is the network bandwidth; is the migration computing load; , Weight parameter, used to balance storage cost, network overhead and computing load; Priority selection of the target layer: Preferentially select the target storage layer that can meet the following conditions; , that is, the remaining space in the target storage layer can accommodate the migrated data; Data heat is within the heat threshold range of the target storage layer; Data migration sorting: When the storage layer needs to free up space, sort the data from low to high heat, and preferentially migrate low-heat data; The execution process of data migration; Distributed transaction support: The data migration process involves multiple nodes and data consistency needs to be ensured. This embodiment completes the migration through distributed transaction management, which specifically includes the following steps: Data replication: Copy the data from the current storage layer to the target storage layer; Metadata update: Update the data location pointer in the metadata service; Original data deletion: After confirming that the data storage in the target layer is completed, delete the data in the current layer; Migration mode: Online migration: The data remains available during the migration process, and uninterrupted operation is achieved through parallel writing; Batch migration: Adopt an offline migration method for a large amount of cold data to avoid affecting the real-time performance of the storage layer; Migration optimization strategy: Optimize network transmission: Reduce the amount of transmitted data through compression and incremental migration technologies; Parallel processing: Multiple data migration tasks are executed in parallel among nodes to shorten the migration time; The formula effect and the impact of migration optimization; Minimize migration cost; Data migration cost function is optimized to ensure that the data is preferentially migrated to the most economical and demand-compliant storage layer, reducing storage costs and migration overhead; System load balancing; By migrating to free up the overloaded state of the storage layer (that is ), the overall performance of the system is improved; Access performance optimization; Migrate high-heat data to the cache layer, significantly reducing data access latency; Summary: By dynamically adjusting the storage hierarchy of data and triggering hierarchical migration, the present invention realizes the dynamic allocation and optimization of storage resources. The migration mechanism takes data heat as the core, and through accurate calculation of migration costs and real-time execution of migration, it ensures that high-heat data is stored in a layer with better performance, and low-heat data is stored in a layer with lower costs, providing flexible and efficient dynamic management capabilities for multi-level cloud storage systems.
[0038] In this embodiment, for step S4, the present invention realizes data redundancy and fault tolerance by adopting a replica strategy and erasure coding technology, ensuring data security and availability in a multi-level cloud storage cluster. The data redundancy strategy is dynamically adjusted based on data popularity. Hot data uses multi-replica redundancy to optimize read and write performance, while cold data reduces storage costs through erasure coding technology. Combining with the actual scenario, by designing a reasonable redundancy and fault tolerance mechanism, the fault tolerance ability of the storage system can be effectively improved, and data can be quickly restored when a node fails.
[0039] Data redundancy and fault tolerance mechanism; In this embodiment, the data redundancy and fault tolerance mechanism includes a replica redundancy strategy and erasure coding storage technology, which are respectively applicable to different storage layers and data access scenarios; Replica redundancy strategy; Dynamic adjustment of the number of replicas; The dynamic adjustment formula for the number of replicas R is as follows; , is the default number of replicas; is the adjustment coefficient used to control the influence of popularity on the number of replicas; is the data popularity; is the highest popularity value in the current system; When high-popularity data is stored in the cache layer or performance layer, a multi-replica strategy is adopted to ensure read and write performance under high-concurrency access; Dynamically adjust the number of replicas to reduce the redundant replica quantity of low-popularity data and lower storage overhead; Implementation method: Data is stored in multiple nodes in the form of replicas in a distributed storage system, and each replica is stored on a different storage node; Use a distributed consistency algorithm to ensure data consistency between replicas; Effect of replica redundancy: When a storage node fails, it can quickly switch to other replica nodes to ensure high data availability; The dynamic replica adjustment strategy avoids resource waste caused by excessive replica redundancy; Erasure coding storage technology for cold data; Erasure coding storage efficiency; Cold data storage adopts erasure coding (k, m), and the storage efficiency is: , k is the number of data shards; m is the number of parity shards; is the storage efficiency, representing the proportion of data shards in the total storage capacity; In the event of a node failure, erasure coding can reconstruct the lost data through the remaining shards, and its recovery time is; , is the size of the data to be recovered; n is the number of nodes participating in parallel recovery; is the network bandwidth of a single node; When cold data is stored in the capacity layer, erasure coding technology is adopted to reduce redundant storage overhead, which is particularly suitable for low-frequency access data such as archived files and historical logs; Implementation method: Use a distributed erasure coding algorithm to store data shards and parity shards; During data recovery, combine the distributed computing power to parallelly recover the lost shards and reduce the recovery latency; Fault recovery mechanism; Fault detection: Monitor the status of storage nodes based on the heartbeat protocol, and the interval time is , if a certain node does not respond for more than multiple heartbeat cycles, it is determined as a faulty node; Fault recovery process: Hot data recovery: If a hot data node fails, the system automatically switches to other replica nodes; Trigger a new replica replication operation in the background to restore the number of replicas; Use erasure coding to parallelly reconstruct the lost shards through the remaining k data shards and m parity shards; Assign the recovery task to multiple storage nodes to improve the recovery speed; Recovery performance optimization: Network optimization: Adopt a high-performance transmission protocol to reduce the network latency of recovery data; Computing optimization: Parallelly process erasure coding calculations through a GPU accelerator or FPGA; Dynamic adjustment effect of replica redundancy: Increase the replicas of high-heat data to improve read and write performance; Reduce redundant replicas of low-heat data, release storage resources, and reduce storage costs; Optimization effect of erasure coding technology; In cold data storage, the storage efficiency of erasure coding can significantly reduce storage occupancy, During the data recovery process, the recovery time is shortened through parallelization and distributed recovery technologies , Comprehensiveness of the fault tolerance mechanism; By combining the replica strategy and erasure coding technology, the diversity of hot data and cold data in terms of performance, cost, and fault tolerance requirements can be satisfied simultaneously; Summary: Through replica redundancy and erasure coding technology, the present invention realizes a data redundancy and fault tolerance mechanism for a multi-level storage cluster. Dynamically adjusting the number of replicas improves the flexibility and resource utilization efficiency of the system, while erasure coding technology significantly reduces the storage overhead of cold data. Combined with an efficient fault recovery mechanism, this mechanism can quickly recover data when a node fails, ensuring the high availability and reliability of the storage system.
[0040] In this embodiment, for step S5, the data distribution in the multi-level storage cluster is optimized through a load balancing and intelligent scheduling mechanism to ensure the high performance and scalability of the system. In a multi-level cloud storage environment, the load of storage nodes is often unbalanced due to fluctuations in data access or performance differences among nodes, which may lead to performance bottlenecks or resource waste. The present invention designs a set of dynamic load balancing models and intelligent scheduling algorithms for real-time monitoring of the system's load status and realizing dynamic optimization of the load through data migration or reassignment of tasks; Load balancing and intelligent scheduling mechanism; In this embodiment, the load balancing and intelligent scheduling mechanism is divided into three parts: load status detection, data migration strategy, and intelligent scheduling optimization; Load status detection; Definition of node load; The load of node i is comprehensively calculated from its storage load and access load; , , is a weight parameter used to adjust the influence of storage load and access load on the total node load; and are the used storage and total storage capacity of node i, respectively; and total request count: the request count of node i and the total request count of the entire system; Load balancing degree calculation: The load balancing degree of the system is measured by the load difference among all nodes; , B is the load balancing degree of the system, representing a relative measure of the load difference among all nodes; is the load value of the i-th node; is the maximum value among all node load values; is the minimum value among all node load values; is the sum of all node load values; N is the total number of nodes in the system; is the average value of node loads in the system.
[0041] Load monitoring and triggering conditions: The system periodically calculates the load of each node and the load balancing degree B; If (load balancing degree threshold), trigger data migration or task reallocation operations; When the load of node i , migrate part of the data from this node to a node j with a lower load; The target node j satisfies; , is the size of the migrated data block; is the current load value of the target node j; Migration data priority: Sort according to the access heat of the data from low to high, and preferentially migrate low-heat data; , P(D) is the priority of data D, used to determine the order of data migration; S(D) is the size of data D; H(D) is the access heat of data D; , is the importance weight for balancing the data block size and heat; Migration cost model; The target selection of data migration is constrained by the migration cost function: , is the storage cost of the target node j; is the network transmission time of the migrated data, calculation formula; , is the network bandwidth; Migration implementation: Data migration includes the following steps; Data selection: Select the migration data according to the migration priority P(D); Data copying: Copy the data to the target node; Metadata update: Update the metadata of the data storage location; Original data cleaning: Delete the data copy at the source node; Intelligent scheduling optimization; Task scheduling algorithm: The system dynamically adjusts the task allocation according to the real-time load of the nodes; , is the scheduling priority of node i; is the remaining storage capacity of node i currently; The total storage capacity is the total storage capacity of all nodes in the system; is the current load value of node i; is a small value to prevent the denominator from being zero; Hot data distribution optimization: Hot data is preferentially distributed to nodes with lower load, and the impact of future access hotspots on the system is reduced through the preheating mechanism; Predictive scheduling: Use the time series model to predict the future load changes of the nodes, allocate tasks or migrate data in advance, and avoid uneven load; Formula effect and technical advantages: Optimization effect of load balancing: By real-time monitoring the node load and triggering migration operations, the storage and access pressures within the system are balanced, and the occurrence probability of load hotspots is significantly reduced; Minimize migration costs.
[0042] Migration cost model Optimizes the storage and network overheads during the migration process and ensures efficient data distribution; Performance improvement of intelligent scheduling: Dynamic task scheduling preferentially allocates new tasks to nodes with lower load, improving the system throughput and response speed; Summary: Through the load balancing and intelligent scheduling mechanisms, the present invention effectively optimizes the load distribution of the storage system, avoiding the situations of node overload or resource idleness. Based on the comprehensive solution of real-time monitoring, dynamic migration and predictive scheduling, it provides an efficient and stable operation guarantee for the multi-level cloud storage cluster, and at the same time improves the scalability of the system and the user request processing ability.
[0043] In this embodiment, for step S6, a distributed metadata management method is used to achieve fast data query and consistency maintenance. Since metadata plays an important role in indexing and locating data in a distributed cloud storage system, its performance directly affects data access efficiency and the overall performance of the system. The present invention uses a distributed consistent hashing algorithm for metadata sharding storage, combined with a caching optimization mechanism for high-frequency metadata, effectively improving the access performance of metadata and the ability to maintain distributed consistency; In this embodiment, the distributed metadata management method is divided into three parts: metadata sharding storage, caching optimization of high-frequency metadata, and consistency maintenance; Metadata sharding storage: Metadata is distributed to multiple storage nodes through a consistent hashing algorithm; , is the storage node number of metadata item M; is to calculate the hash value of metadata M; N is the total number of distributed storage nodes; Sharding strategy: The advantage of using a consistent hashing algorithm is that when nodes are added or removed, only some metadata needs to be remapped, reducing the data migration overhead; To avoid uneven distribution of node data, virtual node technology is adopted, and each physical node is divided into several virtual nodes; , is the hash value of data item M mapped to a virtual node; N is the total number of physical nodes; V is the number of virtual nodes corresponding to each physical node; Sharding update and maintenance: When storage nodes are added or removed, recalculate the virtual node range corresponding to the new nodes, and only migrate the metadata within this range; Implementation method: Deploy a distributed metadata management tool to achieve dynamic distribution and highly available management of metadata; Caching optimization of high-frequency metadata; Caching optimization mechanism: For frequently accessed metadata, it is preferentially stored in a cache (such as memory or SSD) to reduce the access overhead to the backend storage.
[0044] The metadata cache is managed using the Least Recently Used (LRU) algorithm, and the cache update rule is: When adding new data, if the cache space is insufficient, the least recently used metadata is removed; After data is accessed, its access record updates the priority in the LRU queue; Cache hit rate calculation: The hit rate formula for metadata cache is; , is the cache hit rate; The number of high-frequency metadata hits is the number of times metadata is directly retrieved from the cache, The total number of metadata requests is the total number of all access requests to metadata in the system; Dynamic adjustment of cache size: Combining system load and access patterns, dynamically adjust the cache capacity, calculation formula; , is the current cache capacity; is the maximum cache capacity; is the current system load; is the adjustment coefficient; Implementation method: Deploy a distributed cache service to support efficient metadata cache access; Distributed consistency maintenance; Consistency protocol: The update and distribution of metadata are implemented using a distributed consistency protocol: The data modification request is first proposed by the master node for modification; The modification proposal takes effect after being approved by more than half of the nodes; After the modification is completed, all replica nodes synchronize the update; Version control: Each metadata item M is attached with a version number , and the version number is used to judge whether there is a conflict during the update operation; , M is the metadata item, representing the specific data object to be operated; is the version number of the metadata item M, used to record the version status of the data; is the old version number of the metadata item M; is the new version number of the metadata item M; If the version numbers are inconsistent, the update request is rejected to avoid data inconsistency caused by latency; Fault handling: When the master node fails, the system selects a new master node through an election mechanism and continues with the distributed management of metadata; Performance optimization: Route read requests to replica nodes preferentially to disperse the access pressure; Concentrate write requests on the master node for processing to ensure consistency; Formula effects and technical advantages; Effect of sharding storage of metadata: The consistent hashing algorithm significantly reduces the migration overhead when nodes are added or removed, and the virtual node technology further balances the data distribution, avoiding single-point overload; Effect of cache optimization: Hit rate of high-frequency metadata Improved, reducing the I / O load on the backend storage system; Dynamically adjusting the cache capacity ensures the stability of the system under high load; Effect of maintaining distributed consistency: The consistency protocol and version control mechanism ensure the consistency of multi-copy metadata, avoiding data conflicts or losses; The fault handling mechanism improves the high availability and fault tolerance of metadata management; Summary; Through the distributed metadata management method, the present invention realizes the efficient storage, fast query, and consistency maintenance of metadata. The distributed consistent hashing algorithm optimizes the efficiency of metadata distribution and update, the cache mechanism for high-frequency metadata significantly reduces the access latency, and the distributed consistency protocol guarantees the reliability and data consistency of the system. This solution provides strong technical support for the metadata management of multi-level cloud storage clusters.
[0045] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-level cloud storage cluster construction and data storage method, characterized in that: The steps include: S1. Build a multi-level storage layer, including a cache layer, a performance layer, and a capacity layer. Each storage layer has different storage media, which are used to store hot data, warm data, and cold data respectively. S2. Determine data popularity by monitoring data access frequency, access time interval, and data importance; S3: Dynamically adjust the data storage level according to data popularity and trigger data tiered migration; S4. During the storage process, replication strategy and erasure coding technology are used to achieve data redundancy and fault tolerance. S5. Optimize data distribution through load balancing and intelligent scheduling mechanisms to improve storage performance and system scalability; S6. Use distributed metadata management methods to achieve fast data query and consistency maintenance.
2. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: The S1 step specifically includes the following steps: S1.1, built using high-speed storage media through a cache layer to store frequently accessed data; S1.2, built using medium-speed storage media through the performance layer to store warm data; S1.
3. The capacity layer is built using low-cost storage media to store infrequently accessed data.
3. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: The calculation of data access heat in step S2 is based on the following formula: ,in: Representation data Frequency of visits; Indicates the last access time interval; Indicates the importance weight of the data; , , is the weight parameter.
4. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: The hierarchical migration mechanism in step S3 is triggered based on the following conditions: The data heat exceeds the threshold range of the current storage layer; The utilization of the storage layer exceeds the set maximum load threshold; When data lifecycle expires, it needs to be migrated to a low-cost storage tier.
5. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: When the data heat in the S3 step exceeds the heat threshold of the current storage layer, data migration is triggered, and the data migration target is determined by the following cost function; , in; is the unit storage cost of the target storage layer k; is the data block size; is the network transmission time; The computational load for the migration operation; , is the weight parameter.
6. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: During the storage process in step S4, the hot data adopts a copy redundancy strategy, and the number of copies R is dynamically adjusted; ,in; is the default number of copies; is the adjustment coefficient; is the data heat; The heat value of the highest heat data.
7. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: In the step S4, cold data is stored using erasure coding technology, and the erasure coding storage efficiency formula is: ,in; k is the number of data shards; m is the number of verification fragments.
8. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: In the step S6, the distributed metadata management adopts the consistent hashing algorithm for shard storage, and the sharding rule is: ,in: M is a metadata item, indicating the data object that needs to be stored; N is the number of metadata storage nodes, which indicates the number of nodes in the distributed storage system; is a hash function applied to the metadata item M, producing an integer value; Express The calculation result is modulo N, so that the hash value is distributed to N storage nodes.
9. A multi-level cloud storage cluster construction and data storage method according to claim 1, characterized in that: In the metadata management in step S6, the high-frequency metadata is cached using the LRU algorithm and the cache size is dynamically adjusted according to the data access frequency.
10. A multi-level cloud storage cluster construction and data storage method according to claim 5, characterized in that: The load balancing mechanism calculates the node load balancing degree based on the following formula; , When the node The load imbalance exceeds the preset threshold When , data migration between nodes is triggered; in; is the load imbalance degree of node i, which is used to measure the load status of the node; is the number of requests received by node i; The total number of requests is the total number of requests from all nodes in the entire system; is the storage capacity of node i; The total capacity is the total storage capacity of all nodes in the entire system; The system preset load imbalance threshold is used to determine whether the load needs to be adjusted.
Citation Information
Patent Citations
Method and system for realizing hierarchical storage and management in cloud storage
CN104462240A
Consistent hash-based hierarchical mixed storage system and method
CN107844269A
Hierarchical storage method based on dynamic threshold adjusting in cloud storage system
CN108810140A
Attenuation type hierarchical storage system and method based on distributed storage system
CN110162273A
Multi-mode low-energy-consumption distributed cloud storage system, electronic equipment and storage medium
CN115963995A
Cited By
Storage node load balancing method and device, equipment and storage medium
CN120390014A
Storage chip-oriented efficient data burning and error correction method and system
CN120472965A
Efficient data programming and error correction method and system for memory chip
CN120472965B
Multi-level data recovery system and method based on redundant storage
CN120578539A
Data storage control method and device, storage medium and electronic equipment
CN120704617A