Data management system and method based on distributed cloud storage

By introducing access log collection, node status monitoring, data thermal evaluation and redundancy strategy decision modules into distributed cloud storage systems, the sharded redundancy strategy is dynamically adjusted, and the balance between resource allocation efficiency and access performance is solved, and efficient dynamic balance between data access and storage costs is achieved.

CN120386630AInactive Publication Date: 2025-07-29YANCHENG CHUANGJIE TECH CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510512123.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In distributed cloud storage systems, it is difficult for the existing technology to balance resource allocation efficiency and access performance, resulting in high-frequency access data caused access delay due to long distances of storage nodes and high replica redundancy. Low-frequency data caused resource waste due to occupancy of high-speed storage media, and lack of comprehensive perception of node load status, network hop count and bandwidth, which can easily lead to performance bottlenecks or resource not being reasonably utilized.

Method used

The access log acquisition module, node status monitoring module, data thermal evaluation module, redundancy strategy decision module and storage execution scheduling module are adopted to dynamically generate sharded redundancy strategies through the three-dimensional thermal evaluation model. A single replica + local cache strategy is adopted for high-frequency shards, and an erasure code storage is adopted for low-frequency shards, which combine load perception and bandwidth-delay joint optimization model to select reading nodes.

Benefits of technology

It realizes a dynamic balance between data access efficiency and storage cost. Through data classification and differentiated storage strategies, access delays are reduced, storage resource waste is reduced, system stability and data aggregation efficiency are improved, and resource optimization is adapted to heterogeneous environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386630A_ABST
    Figure CN120386630A_ABST
Patent Text Reader

Abstract

The invention discloses a data management system and method based on distributed cloud storage, and relates to the technical field of distributed cloud storage, and the system comprises an access log collection module, a node state monitoring module, a data thermal evaluation module, a redundancy strategy decision module and a storage execution scheduling module. A fragmentation access behavior is recorded through an access log acquisition module, a node state monitoring module acquires a node resource state in real time, and a data thermal evaluation module constructs a three-dimensional thermal model based on an access frequency, a network hop count and a storage cost to calculate a thermal value and supports dynamic weight adjustment; the redundancy strategy decision module divides the fragments into high frequency, medium frequency and low frequency according to the thermodynamic value, and single copy + local cache, erasure code storage and dynamic adjustment strategies are adopted respectively; the storage execution scheduling module executes fragmentation operation and optimizes reading node selection, and data access efficiency is improved through load awareness and a bandwidth-delay joint model; according to the invention, dynamic balance between data access efficiency and storage cost is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud storage distribution, and particularly to a data management system and method based on distributed cloud storage. Background Art

[0002] In a distributed cloud storage system, data management faces the challenge of difficult to balance resource allocation efficiency and access performance. Traditional solutions usually adopt a fixed redundancy strategy (such as unified multi-copy storage), without distinguishing data access heat, network locality, and storage cost differences, resulting in access latency for frequently accessed data due to long distances between storage nodes and high copy redundancy, while low-frequency data wastes resources by occupying high-speed storage media. In addition, existing systems lack comprehensive awareness of node load status, network hops, and bandwidth, which easily leads to performance bottlenecks caused by frequent access to high-load nodes, or the unreasonable utilization of high-quality node resources. How to dynamically optimize the storage strategy according to data characteristics and achieve intelligent scheduling of heterogeneous nodes has become a key issue in improving the efficiency of distributed cloud storage systems.

[0003] To solve the above problems, the present invention proposes a data management system and method based on distributed cloud storage. Summary of the Invention

[0004] The purpose of the present invention is to provide a data management system and method based on distributed cloud storage to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A data management system based on distributed cloud storage includes an access log collection module, a node status monitoring module, a data heat evaluation module, a redundancy policy decision module, and a storage execution scheduling module; the access log collection module is responsible for recording the shard access behavior throughout the link to provide raw data for heat evaluation; the node status monitoring module collects the node resource status in real time, obtains the network hops and storage cost between the node and the request end, and provides real-time parameters for heat evaluation and policy execution; the data heat evaluation module is used to connect the access log collection module and the node status monitoring module, and generates a three-dimensional heat evaluation value based on the shard access frequency, network hops, and storage cost; the redundancy policy decision module is connected to the data heat evaluation module, and dynamically generates a shard redundancy policy according to the three-dimensional heat evaluation value. The policies are to adopt a single-copy + local cache policy for high-frequency shards and an erasure code storage policy for low-frequency shards; the storage execution scheduling module performs shard creation, deletion, and migration operations according to the redundancy policy instructions, and selects a shard reading node based on the joint optimization model of node hops and bandwidth.

[0007] The access log collection module includes a log capture unit; the log capture unit is used to collect client request logs and inter-node data transfer logs; the client request logs include user ID, request time, shard ID, and response latency; all node data transfer logs include source / destination node IP, transferred data volume, and time consumption; and the above data is partitioned and stored according to the shard ID, and the logs of the most recent 7 days are retained for heat value calculation.

[0008] The node status monitoring module includes a resource monitoring unit and a load warning unit;

[0009] The resource monitoring unit collects the hardware resource status and network transmission parameters of distributed nodes in real time, provides basic data for calculating the network hop count d(s) and storage cost c(s) for the data heat evaluation module, and provides a basis for judging node load for the storage execution scheduling module; specifically: first, define monitoring indicators, clarify nine core parameters from three dimensions. In terms of computing resources, collect CPU utilization rate, remaining memory space, and disk I / O latency;

[0010] In terms of storage resources, collect shard storage space occupancy, storage medium type, and data redundancy; among them, the storage medium type is marked as high-frequency storage, low-frequency storage, and archival storage, which is used to define the storage cost calculation weight; the data redundancy represents the number of shard replicas currently stored on the node, which is used to determine whether to trigger the replica deletion strategy;

[0011] In terms of network resources, collect the real-time available bandwidth, network hop count d(s) with the request end, and network connection stability; among them, d(s) represents the network hop count between the node and the request end, which is directly obtained by detecting the router hop count in the routing path through the ICMP protocol and is used to evaluate the network locality of data access; the determination method of network connection stability is to count the packet loss rate of the node in the past 5 minutes;

[0012] Secondly is the data collection mechanism, which performs timed polling through the node agent program; at the same time, protocol adaptation is carried out, using the lightweight MQTT protocol for edge nodes to reduce network consumption, and using the gRPC protocol for core nodes to ensure data integrity; and data preprocessing is carried out, filtering abnormal values, and using the moving average method to generate a stable state snapshot; finally is the output interface, providing d(s) and c(s) to the data heat evaluation module, and providing the real-time available bandwidth, CPU / memory utilization rate to the storage execution scheduling module, which is used to judge whether the node is a high-load node; among them, the calculation formula of c(s) is as follows:

[0013] c(s) = storage space × storage medium weight + transmission time consumption;

[0014] The load warning unit identifies abnormal load status by analyzing node resource data in real time and triggers the avoidance strategy of the storage execution scheduling module.

[0015] The load warning unit further includes the following:

[0016] Specifically: First, define the load threshold, set three levels of load, and calculate the load L by comprehensively considering the weighted values of CPU utilization U cpu , memory utilization U mem , and disk I / O utilization U io . The calculation formula is as follows:

[0017] L = ω cpu ·U cpu + ω men ·U men + ω io ·U io ;

[0018] Among them, ω cpu , ω men , ω io are the weight values of U cpu , U mem , U io respectively, and support dynamic adjustment to adapt to the business scenario;

[0019] Among them, in the normal state, when the load ≤ 60% there is no warning and full participation in data reading and writing is allowed; in the warning state, when 60% < load ≤ 80% it is marked as a node to be observed, and its participation in the storage of high-frequency data shards is restricted; in the overload state, when the load > 80% it is marked as a hot node, triggering the node avoidance strategy of the scheduling module. Secondly, there is a dynamic warning logic, which uses continuous state judgment. When the node reports a load entering the warning state or overload state continuously for 3 times, the corresponding level of warning is triggered; and a state transition mechanism is set. When the node load returns to the normal state continuously for 5 times, the warning mark is cleared and full functionality is restored. Then, send a warning message through the distributed event bus, including the node ID, load level, and warning timestamp, for the storage execution scheduling module to capture in real time. Finally, there is coordination with the scheduling module. When the node is in the overload state, the storage execution scheduling module automatically skips this node during the shard reading optimization stage; for nodes in the warning state, in the redundant strategy decision, its role as a cache node for high-frequency shards is restricted, and only low-frequency shards are allowed to be stored.

[0020] The data thermal evaluation module includes a thermal calculation unit and a weight adjustment unit;

[0021] The thermal calculation unit calculates the thermal value of the data shard using a three-dimensional thermal model. The calculation formula is as follows:

[0022]

[0023] Among them, f(s) is the access frequency of the shard in the recent 24 hours, obtained from the access log collection module; d(s) is the number of hops between the node and the requesting network end, provided by the node monitoring module; c(s) represents the storage cost, provided by the node monitoring module; f avg is the average access frequency of the shard updated within a preset time range; c max is the storage cost threshold preset by the system;

[0024] For the heat value evaluation, this formula comprehensively evaluates the access heat, network locality, and storage cost of the above data. The evaluation basis for each factor is as follows:

[0025] Part 1: Compare the recent access frequency f(s) of the shard with f avg the global average access frequency to highlight the access heat of this shard relative to the whole, and use it to adjust the proportion of access heat in the heat value;

[0026] Part 2: Make the network locality have a positive impact on the heat value;

[0027] Part 3: Compare the storage cost c(s) with the preset threshold c max to control the role of storage cost in the heat value;

[0028] At the same time, in order to highlight the impact of the latest access behavior on the heat value, when calculating the access frequency f(s) in the recent 24 hours, the exponential moving average EMA is used for time window processing; EMA can give a higher weight to f(s), so that the heat value can more prominently reflect the change of data access heat. The calculation formula is as follows:

[0029] EMA t =α·f(s) t +(1 - ε)·EMA t-1 ;

[0030] Among them, EMA t is the exponential moving average value at the current time, f(s) t is the access frequency of the shard at the current time, EMA t-1 is the exponential moving average value at the previous moment, and ε is the smoothing coefficient, and its value range is [0, 1];

[0031] The weight adjustment unit provides a parameter configuration interface for the administrator, allowing it to dynamically adjust the weights of α, β, and γ according to different business scenarios.

[0032] The redundancy policy decision module includes a threshold determination unit and a policy generation unit;

[0033] The function of the threshold determination unit is to classify data shards based on the heat value H(s) generated by the data heat evaluation module, providing a basis for subsequent policy generation;

[0034] For threshold presetting and optimization: A high-frequency threshold T is preset in advance H and a low-frequency threshold T L ; These thresholds are not fixed, but the parameters are optimized by fitting historical data; Specifically, collect the heat value data of different shards in the past 7 days, as well as the performance data of these shards during storage and access, and use the least squares method to find the thresholds that can achieve the best balance between storage resource utilization and data access efficiency for the system; Then, according to the optimized thresholds, mark the shards as high-frequency, medium-frequency, and low-frequency; When H(s)>T H , mark it as a high-frequency shard; When H(s)<T L , mark it as a low-frequency shard; When T L ≤H(s)≤T H , mark it as a medium-frequency shard; This classification method can distinguish data with different access hotness, providing a basis for subsequent differential storage policies;

[0035] The policy generation unit formulates redundancy policies for different types of shards according to the classification results of the threshold determination unit for data shards;

[0036] For the high-frequency shard policy: Set the number of high-frequency shard replicas to 1; Within a limited number of hops near the request end, select nodes according to real-time network routing information to create temporary cache shards; The cache period is determined according to the shard heat value, that is The heat value is positively correlated with the cache period; When selecting cache nodes, preferentially select nodes equipped with high-speed storage media;

[0037] For the low-frequency shard policy: Convert the low-frequency shards from multi-copy storage to erasure code storage; Determine the redundancy according to the shard heat value, the lower the heat value, the lower the redundancy; Migrate the low-frequency shards to nodes using low-cost storage media;

[0038] For the medium-frequency shard policy: The medium-frequency shards default to maintaining the current storage policy; The system monitors the resource status in real time. When the storage resources are less than the preset storage threshold, increase the number of shard replicas; When the resources are greater than the preset storage threshold, use the LZ77 compression algorithm to compress and store the shards.

[0039] The storage execution scheduling module includes a shard operation unit and a read optimization unit;

[0040] The sharding operation unit receives redundancy policy instructions from the redundancy policy decision module; when the instructions require high-frequency sharding operations, it calls the distributed storage API to delete remote copies. Before deletion, it checks the copy status. If the copy is being accessed, it waits for the access to end or processes it by means of data migration; for low-frequency sharding, it calls the API to generate erasure code shards. During the generation process, it uses a checksum mechanism to ensure data accuracy; at the same time, it records operation logs, including operation type, shard ID, and operation time information, for backtracking and recovery in case of operation failure; for high-frequency sharding, it creates a temporary cache in nodes within 3 hops of the request end; when writing cached data, it uses a data checksum synchronization mechanism; it sets a TTL expiration mechanism to automatically delete the cache at a cycle; before the cache expires, if the source data is updated, it timely notifies the cache node to update the data through the message queue mechanism;

[0041] The read optimization unit obtains the network hop count d(s), real-time available bandwidth b i and load information of the node from the node status monitoring module; it preferentially selects the node with the smallest network hop count d(s); if the load of this node is higher than a specific threshold, it triggers a bandwidth-latency game model. This model aims to minimize the comprehensive cost of network latency and bandwidth. Suppose there are m optional nodes, and the network latency-related value of node j is l j , which is the conversion value associated with the network hop count d(s), and the real-time available bandwidth is b j , and the objective function is: The constraint condition is where B r represents the bandwidth required for shard aggregation. The meaning of the objective function is to minimize the sum of the ratios of network latency to bandwidth of the selected nodes, and the constraint condition ensures that the total bandwidth of the selected nodes can meet the basic requirements for data reading; x j represents whether to select node j, x j =1 means selection, and x j =0 means non-selection; then it uses the Lagrange multiplier method to solve this model to obtain the optimal node selection combination;

[0042] Then, in accordance with the principle of short network hop count and high available bandwidth, it sets weights for the network hop count and available bandwidth respectively, and dynamically adjusts the weights according to the system network conditions and data reading requirements. The adjustment of the weights is carried out through a rule-based weight adjustment algorithm; it sorts the nodes according to the adjusted weights to generate a read request queue containing node IP, shard offset, and transmission priority, where the transmission priority is determined according to the data heat value and service requirements. It sets priority coefficients for the data heat value and service requirements respectively, and obtains the comprehensive priority by weighted summation of the two.

[0043] A data management method based on distributed cloud storage, comprising the following steps:

[0044] S1. The access log collection module collects client request logs and data transfer logs between nodes, and stores 7-day data partitioned by shard ID; meanwhile, the node status monitoring module collects node hardware resources and network parameters, generates a status snapshot after preprocessing, and the load warning unit analyzes the node load status, and triggers a warning if it is abnormal;

[0045] S2. The data heat evaluation module calculates the data shard heat value based on the access log and node status data using a three-dimensional heat model; the administrator can adjust the calculation weight according to the business scenario through the parameter configuration interface;

[0046] S3. The redundancy strategy decision module classifies data shards into high-frequency, medium-frequency, and low-frequency categories according to the heat value through preset and optimized thresholds; the strategy generation unit formulates redundancy storage strategies for different types of shards: create a temporary cache for high-frequency shards, use erasure code storage for low-frequency shards, and adjust the storage strategy as needed for medium-frequency shards;

[0047] S4. The shard operation unit of the storage execution scheduling module receives the policy instruction, deletes the remote copy of the high-frequency shard, creates a temporary cache and sets an expiration mechanism; generates erasure code shards for low-frequency shards; records logs during the operation process to ensure data accuracy and operation traceability;

[0048] S5. The read optimization unit obtains the node network hop count, available bandwidth, and load information, and preferentially selects the node with the smallest network hop count; if the node load is high, triggers the bandwidth-delay game model to solve the optimal node combination; adjusts the weight according to the network condition and read requirement, and generates a read request queue containing node IP, shard offset, and transmission priority;

[0049] S6. After completing one round of processing, the system determines whether to continue running; if so, returns to the data collection and status analysis step; if it ends, stops running orderly.

[0050] Compared with the prior art, the beneficial effects of the present invention are:

[0051] 1. Data classification and differential storage: Quantify the data heat through a three-dimensional heat evaluation model (access frequency, network hop count, storage cost), adopt the "single copy + local cache" strategy for high-frequency shards to reduce access latency; use erasure code storage for low-frequency shards and migrate them to low-cost media to reduce waste of storage resources, and achieve a dynamic balance between storage cost and access efficiency.

[0052] 2. Load-Aware Intelligent Node Scheduling: Through node status monitoring and load warning mechanisms, high-load nodes are identified in real time and avoidance strategies are triggered (such as automatically skipping overloaded nodes and restricting high-frequency caching for warning nodes). By combining with the bandwidth-latency joint optimization model, read nodes are selected to improve system stability and data aggregation efficiency.

[0053] 3. Dynamic Policy Adaptation and Resource Optimization: Support administrators to adjust the thermal evaluation weight according to business scenarios. Medium-frequency shards can dynamically adjust the number of replicas or compressed storage according to resource status, enhancing the system's adaptability to heterogeneous environments and avoiding resource allocation imbalances caused by fixed policies. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a system architecture diagram of a data management system based on distributed cloud storage according to the present invention;

[0055] Figure 2 It is a redundant policy flow chart of a data management system based on distributed cloud storage according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution,

[0058] A data management system based on distributed cloud storage includes an access log collection module, a node status monitoring module, a data thermal evaluation module, a redundant policy decision module, and a storage execution scheduling module; the access log collection module is responsible for recording shard access behaviors across the entire link to provide raw data for thermal evaluation; the node status monitoring module collects node resource status in real time, obtains the network hop count and storage cost between the node and the request end to provide real-time parameters for thermal evaluation and policy execution; the data thermal evaluation module is used to connect the access log collection module and the node status monitoring module, and generate a three-dimensional thermal evaluation value based on shard access frequency, network hop count, and storage cost; the redundant policy decision module is connected to the data thermal evaluation module, and dynamically generates a shard redundancy policy according to the three-dimensional thermal evaluation value. The policies are respectively a single-copy + local cache policy for high-frequency shards and an erasure code storage policy for low-frequency shards; the storage execution scheduling module executes shard creation, deletion, and migration operations according to the redundant policy instructions, and selects shard read nodes based on the joint optimization model of node hop count and bandwidth.

[0059] The access log collection module includes a log capture unit; the log capture unit is used to collect client request logs and inter-node data transmission logs; the client request logs include user ID, request time, shard ID, and response latency; all node data transmission logs include source / destination node IP, data volume transmitted, and time taken; and the above data is partitioned and stored according to the shard ID, and the logs of the most recent 7 days are retained for heat value calculation.

[0060] The node status monitoring module includes a resource monitoring unit and a load warning unit;

[0061] The resource monitoring unit collects the hardware resource status and network transmission parameters of distributed nodes in real time, provides basic data for calculating the network hop count d(s) and storage cost c(s) for the data heat evaluation module, and provides a basis for judging node load for the storage execution scheduling module; specifically: first, define the monitoring indicators, clarify nine core parameters from three dimensions. In terms of computing resources, collect CPU utilization rate, remaining memory space, and disk I / O latency;

[0062] In terms of storage resources, collect shard storage space occupancy, storage medium type, and data redundancy; among them, the storage medium type is marked as high-frequency storage, low-frequency storage, and archival storage, which is used to define the storage cost calculation weight; the data redundancy represents the number of shard replicas currently stored on the node, which is used to judge whether to trigger the replica deletion strategy;

[0063] [[ID=1,4]]In terms of network resources, collect real-time available bandwidth, network hop count d(s) with the request end, and network connection stability; where d(s) represents the network hop count between the node and the request end, which is directly obtained by detecting the router hop count in the routing path through the ICMP protocol and is used to evaluate the network locality of data access; the determination method of network connection stability is to count the packet loss rate of the node in the past 5 minutes;

[0064] Secondly is the data collection mechanism, which performs timed polling through the node agent program; at the same time, protocol adaptation is performed, using the lightweight MQTT protocol for edge nodes to reduce network consumption and the gRPC protocol for core nodes to ensure data integrity; and data preprocessing is performed, filtering out outliers, and using the moving average method to generate a stable state snapshot; finally is the output interface, providing d(s) and c(s) to the data heat evaluation module, and providing real-time available bandwidth, CPU / memory utilization rate to the storage execution scheduling module, which is used to judge whether the node is a high-load node; among them, the calculation formula of c(s) is as follows:

[0065] c(s) = storage space × storage medium weight + transmission time taken;

[0066] The load warning unit identifies abnormal load status by analyzing node resource data in real time and triggers the avoidance strategy of the storage execution scheduling module.

[0067] The load warning unit further includes the following:

[0068] Specifically: First, define the load threshold, set three levels of load, and calculate the load L by comprehensively calculating the weighted values of CPU utilization U cpu , memory utilization U mem , and disk I / O utilization U io . The calculation formula is as follows:

[0069] L = ω cpu ·U cpu + ω men ·U men + ω io ·U io ;

[0070] Among them, ω cpu , ω men , and ω io are the weight values of U cpu , U mem , and U io respectively, and support dynamic adjustment to adapt to the business scenario;

[0071] Among them, in the normal state, when the load ≤ 60% there is no warning and full participation in data reading and writing is allowed; in the warning state, when 60% < load ≤ 80% it is marked as a node to be observed, and its participation in the storage of high-frequency data shards is restricted; in the overload state, when the load > 80% it is marked as a hot node, triggering the node avoidance strategy of the scheduling module. Secondly, there is the dynamic warning logic, which uses continuous state judgment. When the node reports the load entering the warning state or overload state continuously 3 times, the corresponding level of warning is triggered; and a state transition mechanism is set. When the node load returns to the normal state continuously 5 times, the warning mark is cleared and full functionality is restored. Then, warning messages are sent through the distributed event bus, including the node ID, load level, and warning timestamp, for the storage execution scheduling module to capture in real time. Finally, there is the coordination with the scheduling module. When the node is in the overload state, the storage execution scheduling module automatically skips this node during the shard reading optimization phase; for warning state nodes, their participation as cache nodes for high-frequency shards is restricted in the redundancy strategy decision-making, and only low-frequency shards are allowed to be stored.

[0072] The data heat evaluation module includes a heat calculation unit and a weight adjustment unit;

[0073] The heat calculation unit calculates the heat value of the data shard using a three-dimensional heat model. The calculation formula is as follows:

[0074]

[0075] Among them, f(s) is the access frequency of the shard in the recent 24 hours, obtained from the access log collection module; d(s) is the number of hops between the node and the request network end, provided by the node monitoring module; c(s) represents the storage cost, provided by the node monitoring module; f avg is the average access frequency of the shard updated within the preset time range; c max is the storage cost threshold preset by the system;

[0076] For the heat value evaluation, this formula comprehensively evaluates the access heat, network locality, and storage cost of the above data. The evaluation basis for each factor is as follows:

[0077] Part 1: Compare the recent access frequency f(s) of the shard with f avg the global average access frequency to highlight the access heat of this shard relative to the whole, and use it to adjust the proportion of access heat in the heat value;

[0078] Part 2: Make the network locality have a positive impact on the heat value;

[0079] Part 3: Compare the storage cost c(s) with the preset threshold c max to control the role of storage cost in the heat value;

[0080] At the same time, in order to highlight the impact of the latest access behavior on the heat value, when calculating the access frequency f(s) in the recent 24 hours, the exponential moving average EMA is used for time window processing; EMA can give a higher weight to f(s), so that the heat value can more prominently reflect the change of data access heat. The calculation formula is as follows:

[0081] EMA t = α·f(s) t +(1 - ε)·EMA t-1 ;

[0082] Among them, EMA t is the exponential moving average value at the current time, f(s) t is the access frequency of the shard at the current time, EMA t-1 is the exponential moving average value at the previous moment, and ε is the smoothing coefficient, and its value range is [0, 1];

[0083] The weight adjustment unit provides a parameter configuration interface for the administrator, allowing it to dynamically adjust the weights of α, β, and γ according to different business scenarios.

[0084] The redundancy policy decision module includes a threshold determination unit and a policy generation unit;

[0085] The threshold determination unit classifies data shards based on the heat value H(s) generated by the data heat evaluation module, providing a basis for subsequent policy generation;

[0086] For threshold presetting and optimization: preset the high-frequency threshold T H and the low-frequency threshold T L ; these thresholds are not fixed, but the parameters are optimized by fitting historical data; specifically, collect the heat value data of different shards in the past 7 days, as well as the performance data of these shards during storage and access, and use the least squares method to find the thresholds that can achieve the best balance between storage resource utilization and data access efficiency for the system; then, according to the optimized thresholds, mark the shards as high-frequency, medium-frequency, and low-frequency; when H(s)>T H , mark it as a high-frequency shard; when H(s)<T L , mark it as a low-frequency shard; when T L ≤H(s)≤T H , mark it as a medium-frequency shard; this classification method can distinguish data with different access hotness, providing a basis for subsequent differential storage policies;

[0087] The policy generation unit formulates redundancy policies for different types of shards according to the classification results of the threshold determination unit for data shards;

[0088] For the high-frequency shard policy: set the number of high-frequency shard replicas to 1; within the limited number of hops near the request end, select nodes according to real-time network routing information to create temporary cache shards; the cache period is determined according to the shard heat value, that is the heat value is positively correlated with the cache period; when selecting cache nodes, preferentially select nodes equipped with high-speed storage media;

[0089] For the low-frequency shard policy: convert the low-frequency shards from multi-copy storage to erasure code storage; determine the redundancy according to the shard heat value, and the lower the heat value, the lower the redundancy; migrate the low-frequency shards to nodes using low-cost storage media;

[0090] For the medium-frequency shard policy: the medium-frequency shards default to maintaining the current storage policy; the system monitors the resource status in real time. When the storage resources are less than the preset storage threshold, increase the number of shard replicas; when the resources are greater than the preset storage threshold, use the LZ77 compression algorithm to compress and store the shards.

[0091] The storage execution scheduling module includes a shard operation unit and a read optimization unit;

[0092] The sharding operation unit receives redundancy policy instructions from the redundancy policy decision module; when the instructions require high-frequency sharding operations, it calls the distributed storage API to delete remote copies. Before deletion, it checks the copy status. If the copy is being accessed, it waits for the access to end or processes it by means of data migration; for low-frequency sharding, it calls the API to generate erasure code shards, and during the generation process, it uses the checksum mechanism to ensure data accuracy; at the same time, it records operation logs, including operation type, shard ID, and operation time information, for backtracking and recovery in case of operation failure; for high-frequency sharding, it creates a temporary cache at nodes within 3 hops of the request end; when writing cached data, it adopts a data checksum synchronization mechanism; it sets a TTL expiration mechanism to automatically delete the cache at a cycle; before the cache expires, if the source data is updated, it timely notifies the cache node to update the data through the message queue mechanism;

[0093] The read optimization unit obtains the network hop count d(s), real-time available bandwidth b i and load information of the node from the node status monitoring module; it preferentially selects the node with the smallest network hop count d(s); if the load of this node is higher than a specific threshold, it triggers a bandwidth-latency game model, which aims to minimize the comprehensive cost of network latency and bandwidth. Suppose there are m optional nodes, and the network latency-related value of node j is l j , which is the conversion value associated with the network hop count d(s), and the real-time available bandwidth is b j , and the objective function is: The constraint conditions are where B r represents the bandwidth required for shard aggregation. The meaning of the objective function is to minimize the sum of the ratios of network latency to bandwidth of the selected nodes, and the constraint conditions ensure that the total bandwidth of the selected nodes can meet the basic requirements of data reading; x j represents whether to select node j, x j =1 means selection, and x j =0 means non-selection; then it uses the Lagrange multiplier method to solve this model to obtain the optimal node selection combination;

[0094] Then, in accordance with the principle of short network hop count and high available bandwidth, it sets weights for the network hop count and available bandwidth respectively, and dynamically adjusts the weights according to the system network conditions and data reading requirements. The adjustment of the weights is carried out through a rule-based weight adjustment algorithm; it sorts the nodes according to the adjusted weights to generate a read request queue containing node IP, shard offset, and transmission priority, where the transmission priority is determined according to the data heat value and business requirements. Priority coefficients are set for the data heat value and business requirements respectively, and the two are weighted and summed to obtain the comprehensive priority.

[0095] A data management method based on distributed cloud storage, comprising the following steps:

[0096] S1. The access log collection module collects client request logs and inter-node data transfer logs, and stores 7-day data by shard ID partitioning; meanwhile, the node status monitoring module collects node hardware resources and network parameters, generates a status snapshot through preprocessing, and the load warning unit analyzes the node load status. If it is abnormal, a warning is triggered;

[0097] S2. The data heat evaluation module calculates the data shard heat value using a three-dimensional heat model based on access logs and node status data; administrators can adjust the calculation weights according to the business scenario through the parameter configuration interface;

[0098] S3. The redundancy strategy decision module classifies data shards into high-frequency, medium-frequency, and low-frequency categories according to the heat value through a preset and optimized threshold; the strategy generation unit formulates redundancy storage strategies for different types of shards: create a temporary cache for high-frequency shards, use erasure coding for low-frequency shards, and adjust the storage strategy as needed for medium-frequency shards;

[0099] S4. The sharding operation unit of the storage execution scheduling module receives the policy instruction, deletes the remote copy of the high-frequency shard, creates a temporary cache and sets an expiration mechanism; generates erasure-coded shards for the low-frequency shard; records logs during the operation process to ensure data accuracy and operation traceability;

[0100] S5. The read optimization unit obtains the node network hop count, available bandwidth, and load information, and preferentially selects the node with the smallest network hop count; if the node load is high, it triggers the solution of the optimal node combination by the bandwidth-delay game model; adjusts the weights according to the network condition and read requirements, and generates a read request queue containing node IP, shard offset, and transmission priority;

[0101] S6. After completing one round of processing, the system determines whether to continue running; if so, it returns to the data collection and status analysis steps; if it ends, it stops running orderly.

[0102] Example:

[0103] Suppose the distributed cloud storage system contains 3 storage nodes (Node A, B, and C), and the client initiates a data access request. The access log collection module records the client request logs through the log capture unit, such as user ID "User001", request time "2025-04-17 09:00:00", shard ID "Shard001", and response latency 120ms; at the same time, it records the inter-node transfer logs, such as the data volume of 50MB transferred from Node A to Node B, which takes 80ms, and stores the data of the last 7 days by shard ID partition. In the node status monitoring module, the resource monitoring unit periodically polls and collects data through the node agent program: the CPU utilization rate of Node A is 60%, the remaining memory space is 40GB, and the disk IO latency is 10ms; the storage medium type is high-frequency storage (weight 1.0), the storage space occupied is 20GB, and the data redundancy is 2 (currently storing 2 copies); the network hop count d(s) is 3 (through 3 routers with the request end), the real-time available bandwidth is 100MB / s, and the packet loss rate in the past 5 minutes is 1% (the connection is stable). According to the storage cost formula, it is calculated that c(s)=20×1.0 + 10 = 30. The load warning unit assumes ω cpu 、ω men 、ω io are 0.4, 0.3, and 0.3 respectively, and calculates the load L of Node A = 0.4×60% + 0.3×40% + 0.3×30% = 45%, which is in the normal state;

[0104] The data heat evaluation module calculates the heat value of the shard "Shard001" based on the collected data. The access frequency f(s) in the past 24 hours = 200 times (calculated by EMA, assuming the smoothing coefficient = 0.5, the EMA value of the previous day is 150, and the current EMA value is 0.5×200 + 0.5×150 = 175, the global average access frequency f avg is 100 times, the network hop count d(s) = 3, the storage cost c(s) = 30, and the system preset storage cost threshold c max = 50; according to the three-dimensional heat model formula, it is calculated that H(s) = 3.62. The redundancy strategy decision module presets the high-frequency threshold T H = 3 and the low-frequency threshold T L = 2. Since H(s) = 3.62 > T H , mark "Shard001" as a high-frequency shard.

[0105] The policy generation unit formulates a policy for the high-frequency shard "Shard001": the number of replicas is set to 1, and within 3 hops of the request side, high-speed storage medium nodes are screened (such as node B, with a storage medium weight of 1.0) to create a temporary cache. The cache period is approximately 0.27 days (about 6.6 hours) according to calculations; the sharding operation unit of the storage execution scheduling module receives the instruction, first checks the status of the remote replica (waits or migrates if being accessed), calls the API to delete the redundant replica of node C, creates a cache shard on node B, uses a checksum synchronization mechanism during writing, sets a TTL expiration mechanism (automatically deleted after 6.6 hours), and listens for source data updates through a message queue;

[0106] The read optimization unit obtains the node status: node A has a load of 45% (normal state), node B has a load of 70% (warning state), and node C has a load of 85% (overloaded state). Node A with the smallest network hop count (3 hops) is preferentially selected, but the available bandwidth of node A is 100MB / s, and the available bandwidth of node B is 150MB / s. Since node B is in a warning state, the bandwidth-delay game model is triggered, and the objective function is (l j is a delay-related value, which is positively correlated with the number of hops. Let node A l = 3, node B l = 4, and the constraint condition is that the total bandwidth ≥ the required bandwidth of 50MB / s for shard aggregation. Solving through the Lagrange multiplier method, nodes A (x1 = 1) and B (x2 = 1) are selected, and the total bandwidth of 250MB / s meets the requirements. Sorted by weight (network hop count weight 0.6, bandwidth weight 0.4), a read request queue is generated, and the priority is calculated as 3.62×0.8 + 0.2×1 = 3.096 according to the heat value (3.62) and business requirements (coefficient 0.8). After one round of processing, the system returns to the data collection stage, continuously monitors the logs and node status, and dynamically adjusts the policy to ensure the balance between high-frequency data access efficiency and low-frequency data storage cost.

[0107] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. A data management system based on distributed cloud storage, characterized in that: It includes an access log collection module, a node status monitoring module, a data heat evaluation module, a redundancy strategy decision-making module, and a storage execution scheduling module; the access log collection module is responsible for recording the shard access behavior across the entire link and providing raw data for heat evaluation; the node status monitoring module collects the node resource status in real time, obtains the network hop count and storage cost between the node and the request end, and provides real-time parameters for heat evaluation and policy execution; the data heat evaluation module is used to connect the access log collection module and the node status monitoring module, and generate a three-dimensional heat evaluation value based on the shard access frequency, network hop count, and storage cost; the redundancy strategy decision-making module is connected to the data heat evaluation module, and dynamically generates a shard redundancy strategy according to the three-dimensional heat evaluation value. The strategies are to adopt a single-copy + local cache strategy for high-frequency shards and an erasure code storage strategy for low-frequency shards; The storage execution scheduling module performs shard creation, deletion, and migration operations according to the redundancy strategy instructions, and selects a shard reading node based on the joint optimization model of node hop count and bandwidth.

2. The data management system based on distributed cloud storage according to claim 1, wherein: The access log collection module includes a log capture unit; the log capture unit is used to collect client request logs and inter-node data transmission logs; the client request logs include user ID, request time, shard ID, and response latency; all node data transmission logs include source / destination node IP, transmitted data volume, and elapsed time; and the above data is partitioned and stored according to the shard ID, and the logs of the most recent 7 days are retained for heat value calculation.

3. A data management system based on distributed cloud storage according to claim 1, characterized in that: The node status monitoring module includes a resource monitoring unit and a load warning unit; The resource monitoring unit collects the hardware resource status and network transmission parameters of distributed nodes in real time, provides basic data for calculating the network hop count d(s) and storage cost c(s) for the data heat evaluation module, and provides a basis for judging node load for the storage execution scheduling module; specifically: first, define the monitoring indicators, clarify nine core parameters from three dimensions. In terms of computing resources, collect CPU utilization rate, remaining memory space, and disk I / O latency; In terms of storage resources, collect the occupied space of shard storage, storage medium type, and data redundancy; among them, the storage medium type is marked as high-frequency storage, low-frequency storage, and archival storage, which is used to define the weight of storage cost calculation; the data redundancy indicates the number of shard replicas currently stored on the node, which is used to determine whether to trigger the replica deletion strategy; In terms of network resources, collect the real-time available bandwidth, the network hop count d(s) with the request end, and the network connection stability; among them, d(s) represents the network hop count between the node and the request end, which is directly obtained by detecting the router hop count in the routing path through the ICMP protocol and is used to evaluate the network locality of data access; the determination method of network connection stability is to count the packet loss rate of the node in the past 5 minutes; Secondly, there is the data acquisition mechanism, which performs timed polling through the node agent program; meanwhile, protocol adaptation is carried out. The lightweight MQTT protocol is adopted for edge nodes to reduce network consumption, and the gRPC protocol is adopted for core nodes to ensure data integrity; and data preprocessing is carried out, filtering out outliers, and using the moving average method to generate a stable state snapshot; finally, there is the output interface, which provides d(s) and c(s) to the data thermal evaluation module, and provides the real-time available bandwidth, CPU / memory utilization rate to the storage execution scheduling module, which is used to judge whether the node is a high-load node; among them, the calculation formula of c(s) is as follows: c(s) = storage space × storage medium weight + transmission time; The load warning unit triggers the avoidance strategy of the storage execution scheduling module by analyzing the node resource data in real time and identifying the abnormal load state.

4. A data management system based on distributed cloud storage according to claim 3, characterized in that: The load warning unit further includes the following: Specifically: First, define the load threshold, set three load levels, and use the comprehensive CPU utilization U cpu , memory utilization U mem , Disk IO utilization U io The weighted value of is used to calculate the load L, and the calculation formula is as follows: L = ω cpu ·U cpu + ω men ·U men + ω io ·U io ; Among them, ω cpu , ω men , ω io are the weight values of U cpu , U mem , U io respectively, and support dynamic adjustment to adapt to the business scenario; Among them, in the normal state, when the load ≤ 60% there is no warning and full participation in data reading and writing is allowed; in the warning state, 60% < load ≤ 80% is marked as a node to be observed, and its participation in the storage of high-frequency data shards is restricted; in the overload state, when the load > 80% it is marked as a hot node, triggering the node avoidance strategy of the scheduling module; secondly, there is the dynamic warning logic, which uses continuous state judgment. When the node reports the load and enters the warning state or overload state continuously for 3 times, the corresponding level of warning is triggered; and a state transition mechanism is set. When the node load returns to the normal state continuously for 5 times, the warning mark is cleared and full function participation is restored; then warning messages are sent through the distributed event bus, including the node ID, load level, and warning timestamp, for the storage execution scheduling module to capture in real time; finally, there is the cooperation with the scheduling module. When the node is in the overload state, the storage execution scheduling module automatically skips this node during the shard reading optimization stage; for nodes in the warning state, its role as a cache node for high-frequency shards is restricted in the redundancy strategy decision, and only low-frequency shards are allowed to be stored.

5. A data management system based on distributed cloud storage according to claim 1, characterized in that: The data thermal evaluation module includes a thermal calculation unit and a weight adjustment unit; The thermal calculation unit calculates the thermal value of the data shard using a three-dimensional thermal model, and the calculation formula is as follows: Among them, f(s) is the shard access frequency in the past nearly 24 hours, obtained from the access log collection module; d(s) is the number of network hops between the node and the request, provided by the node monitoring module; c(s) represents the storage cost, provided by the node monitoring module; f avg is the average access frequency of the shards updated within a preset time range; c max is the storage cost threshold preset by the system; For the thermal value evaluation, this formula comprehensively evaluates the three factors of the access heat, network locality, and storage cost of the above data. The evaluation basis for each factor is as follows: Part: Compare the recent access frequency f(s) of the shard with f avg with the global average access frequency to highlight the access popularity of the shard relative to the whole, which is used to adjust the proportion of the access popularity in the heat value; Part: enabling network locality to have a positive impact on the thermal value; Part: Compare the storage cost c(s) with a preset threshold c max to control the role of the storage cost in the thermal value; At the same time, in order to highlight the influence of the latest access behavior on the thermal value, when calculating the access frequency f(s) in the past 24 hours, exponential moving average EMA is used for time window processing; EMA can give f(s) a higher weight, so that the thermal value can more prominently reflect the change of data access heat, and the calculation formula is as follows: EMAt = α·f(s) t + (1 - ε)·EMA t-1 ; Among them, EMA t is the exponentially weighted moving average of the current time, and f(s) t is the shard access frequency at the current time. EMA t-1 is the exponentially weighted moving average of the previous moment, and ε is the smoothing coefficient with a value range of [0, 1]; The weight adjustment unit provides an interface for parameter configuration for the administrator, allowing him to dynamically adjust the weights of α, β, and γ according to different business scenarios.

6. The data management system based on distributed cloud storage according to claim 5, characterized in that: The redundancy strategy decision module includes a threshold determination unit and a strategy generation unit; The role of the threshold determination unit is to classify the data shards according to the thermal value H(s) generated by the data thermal evaluation module, providing a basis for subsequent strategy generation; For threshold presetting and optimization: preset the high-frequency threshold T H and the low-frequency threshold T L ; these thresholds are not fixed, but the parameters are optimized by fitting historical data; specifically, collect the heat value data of different shards in the past 7 days, as well as the performance data of these shards during storage and access, and use the least squares method to find the threshold that can achieve the best balance between storage resource utilization and data access efficiency in the system; then, according to the optimized threshold, mark the shards as high-frequency, medium-frequency, and low-frequency; when H(s)>T H , mark it as a high-frequency shard; when H(s)<T L , mark it as a low-frequency shard; when T L ≤H(s)≤T H , mark it as a medium-frequency shard; this classification method can distinguish data with different access hotness and provide a basis for subsequent differential storage strategies; The policy generation unit formulates redundancy policies for different types of shards according to the classification results of the data shards by the threshold determination unit; For the high-frequency shard policy: set the number of high-frequency shard replicas to 1; within a limited number of hops near the request end, filter nodes according to real-time network routing information to create temporary cached shards; The cache period is determined based on the shard heat value, that is The heat value is positively correlated with the cache period; when selecting a cache node, a node equipped with a high-speed storage medium is preferentially selected; For the low-frequency shard policy: convert the low-frequency shards from multi-replica storage to erasure code storage; determine the redundancy according to the shard heat value, and the lower the heat value, the lower the redundancy; Migrate the low-frequency shards to nodes using low-cost storage media; For the medium-frequency shard policy: the medium-frequency shards default to maintaining the current storage policy; The system monitors the resource status in real time. When the storage resources are less than the preset storage threshold, increase the number of shard replicas; when the resources are greater than the preset storage threshold, compress and store the shards using the LZ77 compression algorithm.

7. A data management system based on distributed cloud storage according to claim 6, characterized in that: The storage execution scheduling module includes a shard operation unit and a read optimization unit; The shard operation unit receives redundancy policy instructions from the redundancy policy decision module; when the instruction requires an operation on high-frequency shards, call the distributed storage API to delete remote replicas. Before deletion, check the replica status. If the replica is being accessed, wait for the access to end or use the data migration method for processing; For low-frequency shards, call the API to generate erasure code shards. During the generation process, use the checksum mechanism to ensure data accuracy; at the same time, record the operation log, including the operation type, shard ID, and operation time information for backtracking and recovery in case of operation failure; for high-frequency shards, create a temporary cache at nodes within 3 hops of the request end; When writing cached data, a data verification and synchronization mechanism is adopted; a TTL expiration mechanism is set to automatically delete the cache at intervals of and update the cache in a timely manner by means of a message queue mechanism if the source data is updated before the cache expires. The read optimization unit obtains the network hop count d(s), the real-time available bandwidth b of the node from the node status monitoring module i and the load information; preferentially select the node with the smallest network hop count d(s); if the load of this node is higher than a specific threshold, trigger the bandwidth-delay game model, which aims to minimize the comprehensive cost of network delay and bandwidth. Suppose there are m alternative nodes, and the network delay related value of node j is l j , which is the conversion value associated with the network hop count d(s), and the real-time available bandwidth is b j , and the objective function is: The constraint conditions are where B r represents the bandwidth required for shard aggregation. The meaning of the objective function is to minimize the sum of the ratios of network delay and bandwidth of the selected nodes, and the constraint conditions ensure that the total bandwidth of the selected nodes can meet the basic requirements of data reading; x j indicates whether to select node j, x j =1 means selection, x j =0 means non-selection; then use the Lagrange multiplier method to solve this model to obtain the optimal node selection combination; Then, according to the principle of short network hops and high available bandwidth, set weights for network hops and available bandwidth respectively, and dynamically adjust the weights according to the system network status and data reading requirements. The adjustment of the weights is carried out through a rule-based weight adjustment algorithm; sort the nodes according to the adjusted weights to generate a read request queue containing node IP, shard offset, and transmission priority, where the transmission priority is determined according to the data heat value and business requirements. Set priority coefficients for the data heat value and business requirements respectively, and sum them weighted to obtain the comprehensive priority.

8. A data management method based on distributed cloud storage, applied to a data management system based on distributed cloud storage as claimed in claims 1-7, characterized in that: It includes the following steps: S1. The access log collection module collects client request logs and inter-node data transfer logs, and stores the 7-day data partitioned by shard ID; at the same time, the node status monitoring module collects node hardware resources and network parameters, generates a status snapshot after preprocessing, and the load warning unit analyzes the node load status. If it is abnormal, trigger a warning; S2. The data heat evaluation module calculates the data shard heat value using a three-dimensional heat model based on the access log and node status data; Administrators can adjust the calculation weights through the parameter configuration interface according to the business scenario; S3. The redundancy policy decision module classifies the data shards into high-frequency, medium-frequency, and low-frequency categories according to the heat value through preset and optimized thresholds; the policy generation unit formulates redundancy storage policies for different types of shards: create temporary caches for high-frequency shards, use erasure code storage for low-frequency shards, and adjust the storage policy as needed for medium-frequency shards; S4. The sharding operation unit of the storage execution scheduling module receives the policy instruction, deletes the remote replicas of the high-frequency shards, creates a temporary cache and sets an expiration mechanism; Generates erasure code shards for low-frequency shards; records logs during the operation process to ensure data accuracy and operation traceability; S5. The read optimization unit obtains the node network hop count, available bandwidth and load information, and preferentially selects the node with the smallest network hop count; if the node load is high, triggers the bandwidth-delay game model to solve the optimal node combination; adjusts the weights according to the network conditions and read requirements, and generates a read request queue containing node IPs, shard offsets and transmission priorities; S6. After completing one round of processing, the system determines whether to continue running; if so, returns to the data collection and status analysis steps; if it ends, stops running orderly.

Citation Information

Cited By

  • Data storage control method and device, storage medium and electronic equipment

    CN120704617A

  • Large model training resource optimization method and system based on Hadoop ecology

    CN121070632A

  • Data storage system for information security

    CN121092072A

  • Erasure code strategy conversion method and device, equipment, medium and product

    CN121116206A

  • Distributed hierarchical data storage system and method

    CN121143721A