Distributed data storage optimization method, device and system based on cloud computing

CN122733809APending Publication Date: 2026-09-11SHANDONG CHANGGUANG INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610901017.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]现有分布式存储过程依赖静态时间阈值或最近最少使用(least recently used,LRU)等规则实施冷热数据判定与缓存准入,仅能捕捉局部的绝对访问频次,无法反映混部掩盖下的周期节拍与突发脉冲特征,导致将周期性批量任务误判为冷数据、突发查询误判为热数据的冷热误判,以及缓存层被大量非持续性热点数据无效占用的缓存污染现象

Benefits of technology

通过在云计算分布式数据存储场景下,对数据分片访问序列与节点混部聚合序列进行混部补偿、节拍与峰形特征提取、融合链路耗时计算预测访问占优度并结合缓存内存约束控制预取,能够在不依赖经验阈值的条件下,有效剥离节点混部干扰、精准区分稳定周期访问与突发脉冲访问,实现预取操作的合理调度,降低冷热数据误判与缓存污染,从而提升热数据缓存层对稳定访问分片的命中率,缩短周期性批量任务的读取时延,并提高缓存资源利用率与跨节点数据搬运效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122733809A_ABST
    Figure CN122733809A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of distributed storage, in particular to a distributed data storage optimization method, device and system based on cloud computing, which solves the technical problem of cold and hot misjudgment in the distributed storage process caused by the dependence on absolute access frequency in the prior art. The method comprises the following steps: acquiring an access count sequence of a data shard in a discrete time window and a mixed department aggregation count sequence of a storage node where the data shard is located in the discrete time window; performing mixed department compensation processing on the access count sequence and the mixed department aggregation count sequence to determine an isolated residual sequence; determining a beat consistency coefficient and a skew penalty amount according to the isolated residual sequence; determining a predicted access dominant degree of the data shard according to the beat consistency coefficient, the skew penalty amount and a link time consumption of completing a pre-fetch operation; and controlling an asynchronous pre-fetch operation on the data shard according to the predicted access dominant degree and a physical memory capacity of a cache node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed storage technology, and more specifically to a distributed data storage optimization method, device, and system based on cloud computing. Background Technology

[0002] With the deep integration of cloud computing and big data, the scale of data is growing exponentially, and structured and unstructured data coexist. Traditional centralized storage architectures are gradually reaching their bottlenecks in terms of scalability, reliability, and high-concurrency access performance, which makes distributed data storage the core infrastructure supporting massive data lakes.

[0003] In typical scenarios such as user behavior data lakes on e-commerce platforms, data sharding is usually distributed across cold and hot tiered distributed cluster nodes based on strategies such as consistent hashing. The system not only needs to handle periodic extraction transformation loading (ETL) tasks that start in the early morning and read large-scale, low-frequency historical wide tables, but also needs to deal with ad-hoc operational queries initiated by multiple business lines during the day, thus forming a complex mixed load of periodic fluctuations and sudden spikes.

[0004] Existing distributed storage processes rely on static time thresholds or least recently used (LRU) rules to determine hot and cold data and grant cache access. They can only capture the absolute access frequency of a local area and cannot reflect the periodic beats and burst pulse characteristics under the cover of mixed distribution. This leads to misjudging periodic batch tasks as cold data and burst queries as hot data, as well as cache pollution caused by a large amount of non-persistent hot data occupying the cache layer. Summary of the Invention

[0005] To address the technical problem of misjudging hot and cold storage in distributed storage processes due to the reliance on absolute access frequency in existing technologies, the present invention aims to provide a distributed data storage optimization method, device, and system based on cloud computing. The specific technical solution adopted is as follows: Firstly, a cloud computing-based distributed data storage optimization method is provided, comprising: obtaining the access count sequence of data shards in a distributed data storage system within a discrete time window, and the mixed-area aggregation count sequence of the storage nodes where the data shards reside within the discrete time window; performing mixed-area compensation processing on the access count sequence and the mixed-area aggregation count sequence to determine an isolation residual sequence used to characterize the relative access intensity change of the data shards after removing the influence of the mixed-area background of the nodes; determining a clock consistency coefficient used to characterize the stability of the access peak time interval of the data shards based on the isolation residual sequence, and determining a skew penalty amount used to characterize the burstiness of the access pulses of the data shards based on the difference between the rise and fall rates of the access peaks in the isolation residual sequence; determining the predicted access dominance of the data shards based on the clock consistency coefficient, the skew penalty amount, and the link time to complete one prefetch operation; and controlling the asynchronous prefetch operation of the data shards based on the predicted access dominance and the physical memory capacity of the cache nodes in the distributed data storage system.

[0006] Based on the above technical solution, in the distributed data storage optimization method based on cloud computing provided by this invention, by performing mixed-location compensation, extracting rhythm and peak features, calculating the access dominance of data shard access sequences and node mixed-location aggregation sequences in the cloud computing distributed data storage scenario, and combining cache memory constraints to control prefetching, it is possible to effectively remove node mixed-location interference, accurately distinguish stable periodic access from burst pulse access, achieve reasonable scheduling of prefetching operations, reduce misjudgment of hot and cold data and cache pollution, thereby improving the hit rate of hot data cache layer for stable access shards, shortening the read latency of periodic batch tasks, and improving cache resource utilization and cross-node data transfer efficiency.

[0007] In conjunction with the first aspect mentioned above, in one possible implementation, the method for performing mixed-part compensation processing on the access count sequence and the mixed-part aggregation count sequence to determine the isolation residual sequence specifically includes: determining a short-scale observation interval based on the average duration of historical traffic peaks in the data shards; determining a long-scale observation interval based on the mixed-part access duration of the storage node where the data shards are located within a business cycle; determining the degree to which the value of the access count sequence deviates from the mean within the short-scale observation interval in the discrete time window, as a local deviation; calculating the average absolute deviation of the mixed-part aggregation count sequence within the long-scale observation interval and the average absolute deviation of the access count sequence within the short-scale observation interval to determine the mixed-part dilution coefficient; calculating the difference values ​​of the access count sequences within the short-scale observation interval, as well as continuous difference values ​​with the same sign, to determine the continuous bias coefficient; and determining the isolation residual based on the local deviation, the mixed-part dilution coefficient, and the continuous bias coefficient, combined with the access count sequence and the mixed-part aggregation count sequence, and forming an isolation residual sequence.

[0008] In conjunction with the first aspect above, in one possible implementation, the method for determining the beat consistency coefficient based on the isolation residual sequence specifically includes: taking the absolute value of the isolation residual sequence to obtain an energy sequence; performing center-weighted smoothing on the energy sequence with a short-scale observation interval as the width to obtain an energy envelope sequence; the weighted smoothing is used to eliminate the time misalignment of adjacent access peaks caused by resource queuing; extracting the window position corresponding to each maximum value in the energy envelope sequence, determining the inter-peak distance corresponding to the window positions of adjacent maximum values, and using the median of the inter-peak distance as the natural beat span of the data shard; determining the median of the absolute value of the difference between the inter-peak distance and the natural beat span as the beat dispersion; and determining the beat consistency coefficient based on the natural beat span and the beat dispersion.

[0009] In conjunction with the first aspect above, in one possible implementation, the method for determining the skew penalty based on the difference in the rising and falling rates of the access peaks in the isolation residual sequence specifically includes: determining the pre-peak rising slope based on the rising amplitude and rising duration between each maximum and the previous adjacent valley in the energy envelope sequence; determining the post-peak falling slope based on the falling amplitude and falling duration between each maximum and the next adjacent valley in the energy envelope sequence; determining the single-peak skew penalty value for each maximum based on the post-peak falling slope and the pre-peak rising slope; and weighting the single-peak skew penalty values ​​using the peak energy of each maximum as the weight to obtain the skew penalty amount.

[0010] In conjunction with the first aspect above, in one possible implementation, the method for determining the predicted access dominance of data shards based on the takt consistency coefficient, skew penalty, and link time to complete one prefetch operation specifically includes: using the natural takt span as the access period of the data shard; determining the access time margin based on the access period and link time, and using the proportion of the access time margin in the access period as the prefetch fulfillment rate; and fusing the takt consistency coefficient, prefetch fulfillment rate, and skew penalty to obtain the predicted access dominance.

[0011] In conjunction with the first aspect mentioned above, in one possible implementation, the method for obtaining link latency specifically includes: obtaining the cross-node network transmission latency between the storage node where the data shard resides and the cache node, the data volume of the data shard, and the average effective bandwidth, average deserialization, and memory write duration of the cache node; determining the data transmission time of the data shard from the storage node where the data shard resides to the cache node based on the data volume of the data shard and the average effective bandwidth of the cache node; and summing the cross-node network transmission latency, data transmission time, and average deserialization and memory write duration to obtain the link latency.

[0012] In conjunction with the first aspect above, in one possible implementation, the method for controlling the asynchronous prefetching operation of data shards based on predicted access dominance and the physical memory capacity of cache nodes in the distributed data storage system specifically includes: sorting data shards within the jurisdiction of the cache nodes in descending order according to predicted access dominance; determining a set of prefetched data shards within the physical memory capacity limit according to the descending order; determining the prefetch sleep waiting time for each data shard in the prefetched data shard set based on the access cycle and link latency; and using the window time corresponding to the maximum value of the current access peak of the current data shard as the timing start point, and issuing an asynchronous prefetch instruction to the cache node after the prefetch sleep waiting time of the current data shard.

[0013] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the access count sequence of data shards within a discrete time window in a distributed data storage system, and the mixed aggregation count sequence of the storage node where the data shards reside within the discrete time window, specifically includes: dividing continuous time into multiple discrete time windows using the minimum input / output scheduling granularity of the storage engine as the width of the discrete time window; filtering internal system traffic to retain business read requests at the object storage engine entry side of the storage node where the data shards reside; generating an event record containing the identifier of the requested data shard and the request arrival time in response to the business read request; counting the number of event records for the data shards within each discrete time window to obtain the access count sequence; and summing the access counts of multiple data shards on the storage node where the data shards reside within the same discrete time window to obtain the mixed aggregation count sequence.

[0014] Secondly, a cloud computing-based distributed data storage optimization device is provided, comprising: a data acquisition unit for acquiring access count sequences of data shards within a discrete time window in a distributed data storage system, and mixed-area aggregation count sequences of storage nodes where the data shards reside within the discrete time window; a mixed-area compensation unit for performing mixed-area compensation processing on the access count sequences and mixed-area aggregation count sequences to determine an isolation residual sequence characterizing the relative access intensity change of the data shards after removing the influence of the mixed-area background of nodes; a skew determination unit for determining a clock consistency coefficient characterizing the stability of the access peak time interval of the data shards based on the isolation residual sequence, and a skew penalty amount characterizing the burstiness of access pulses of the data shards based on the difference between the rise and fall rates of the access peaks in the isolation residual sequence; a predicted access unit for determining the predicted access dominance of the data shards based on the clock consistency coefficient, the skew penalty amount, and the link time for completing one prefetch operation; and a control unit for controlling the asynchronous prefetch operation of the data shards based on the predicted access dominance and the physical memory capacity of the cache nodes in the distributed data storage system.

[0015] Thirdly, a cloud-based distributed data storage optimization system is provided, running in a cloud computing environment, including multiple storage nodes, multiple cache nodes, and a cloud-based distributed data storage optimization device; the distributed data storage optimization device executes any one of the methods in the first aspect; the cache nodes are used to receive and execute asynchronous prefetch instructions issued by the distributed data storage optimization device; the storage nodes are used to store data fragments and provide data when receiving data requests from the cache nodes.

[0016] The present invention has the following beneficial effects: By performing mixed-location compensation, extracting cycle time and peak shape features, calculating access dominance based on data shard access sequences and node mixed-location aggregation sequences in cloud computing distributed data storage scenarios, and combining this with cache memory constraints to control prefetching, we can effectively remove node mixed-location interference, accurately distinguish between stable periodic access and burst pulse access without relying on empirical thresholds. This enables reasonable scheduling of prefetching operations, reduces misjudgment of hot and cold data and cache pollution, thereby improving the hit rate of hot data cache layer for stable access shards, shortening the read latency of periodic batch tasks, and improving cache resource utilization and cross-node data transfer efficiency. Attached Figure Description

[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A system architecture diagram of a cloud computing-based distributed data storage optimization system is provided as an embodiment of the present invention. Figure 2 This is a flowchart illustrating a cloud computing-based distributed data storage optimization method according to an embodiment of the present invention. Detailed Implementation

[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a cloud computing-based distributed data storage optimization method, device, and system proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0021] The following description, in conjunction with the accompanying drawings, details a specific solution for a cloud computing-based distributed data storage optimization method, device, and system provided by this invention.

[0022] Please see Figure 1 This diagram illustrates a system architecture of a cloud-based distributed data storage optimization system according to an embodiment of the present invention. The cloud-based distributed data storage optimization system operates in a cloud computing environment and includes: a cloud-based distributed data storage optimization device 1, multiple storage nodes, and multiple cache nodes. Figure 1 (An example of 3 storage nodes and 2 cache nodes is used for illustration). The cloud-based distributed data storage optimization device 1 serves as the core scheduling and computing hub of the system, coordinating the entire process of data acquisition, feature computation, and prefetching decision-making. Multiple storage nodes are responsible for the persistent storage of data shards and responding to business read requests. Multiple cache nodes are responsible for maintaining the hot data cache space, receiving and executing asynchronous prefetching instructions issued by the cloud-based distributed data storage optimization device 1, and realizing cross-node data retrieval and high-speed caching. The above devices and nodes achieve data interaction and instruction transmission through the cloud computing network, collaboratively completing the optimized scheduling of distributed data storage.

[0023] The cloud-based distributed data storage optimization device 1 includes: a data acquisition unit 11, a mixed-distribution compensation unit 12, a skew determination unit 13, a predictive access unit 14, and a control unit 15.

[0024] Data acquisition unit 11 is the system's data acquisition entry point. Its core function is to acquire the access count sequence of data shards within discrete time windows in the distributed data storage system, as well as the mixed-distribution aggregation count sequence of the storage nodes where the data shards reside within discrete time windows. This unit can be implemented through a control plane server cluster in a cloud computing environment. The cluster has a built-in bypass probe acquisition module, a time-series window partitioning module, a traffic filtering module, and a data synchronization module. The bypass probe acquisition module is deployed on the object storage engine entry side of each storage node, and mirrors and collects read requests before disk read operations are performed. The traffic filtering module removes internal system traffic such as replica repair, background compression, and metadata reconstruction, retaining only valid read requests from cross-business lines. The time-series window partitioning module divides discrete time windows based on the smallest input / output scheduling granularity of the storage engine, and counts the number of data shard accesses and the total number of mixed-distribution accesses within each window. The data synchronization module asynchronously pulls statistical data from each storage node and performs time-series rearrangement, ultimately generating two types of sequences and transmitting them to the mixed-distribution compensation unit 12, providing raw data support for subsequent mixed-distribution interference removal.

[0025] The core function of the mixed-location compensation unit 12 is to perform mixed-location compensation processing on the access count sequence and the mixed-location aggregated count sequence to determine the isolation residual sequence used to characterize the relative access intensity change of data shards after removing the influence of the mixed-location background of nodes. This unit can be implemented through a high-performance computing node in the cloud computing control plane, and has built-in observation interval calculation module, local deviation statistics module, deviation index calculation module, compensation coefficient generation module, and residual generation module. The observation interval calculation module determines the short-scale observation interval based on the average duration of the historical traffic peaks of the data shards, and the long-scale observation interval based on the complete business cycle of the storage node. The local deviation statistics module calculates the degree to which the access count sequence value deviates from the mean of the short-scale observation interval. The deviation index calculation module calculates the average absolute deviation of the sequence in both the long and short-scale intervals. The compensation coefficient generation module calculates the mixed-location dilution coefficient based on the deviation index and determines the continuous bias coefficient by combining the sequence difference value. The residual generation module integrates the local deviation and the compensation coefficient to generate the isolation residual and form a sequence, which is output to the skew determination unit 13 to completely remove the access fluctuation interference caused by the mixed-location of multiple business lines and ensure the accuracy of subsequent feature extraction.

[0026] The core function of the skew determination unit 13 is to determine the beat consistency coefficient based on the isolation residual sequence, and simultaneously determine the skew penalty based on the difference in the rising and falling rates of the access peaks in the isolation residual sequence. This unit can be implemented through a parallel computing module in the cloud computing control plane, and includes an energy sequence processing module, a beat feature extraction module, a peak slope calculation module, and a penalty generation module. The energy sequence processing module takes the absolute value of the isolation residual sequence to generate an energy sequence, and then performs short-scale interval weighted smoothing to generate an energy envelope sequence, eliminating the access peak time misalignment caused by resource queuing. The beat feature extraction module extracts the maximum value position of the energy envelope sequence, calculates the inter-peak time span to determine the natural beat span, and generates the beat consistency coefficient by combining the beat discreteness. The peak slope calculation module matches the preceding and following valley values ​​corresponding to each peak, and calculates the rising slope before the peak and the falling slope after the peak. The penalty generation module calculates the single-peak skew penalty value based on the slope difference, and obtains the skew penalty by weighted average of the peak energy. The two types of feature parameters are output to the prediction access unit 14, respectively characterizing the periodic stability and burstiness of data shard access.

[0027] The core function of the predictive access unit 14 is to determine the predicted access dominance of data shards based on the clock consistency coefficient, skew penalty, and the link time required to complete one prefetch operation. This unit can be implemented through a cloud computing control plane parameter fusion calculation module, which includes a link time statistics module, a prefetch fulfillment calculation module, and a dominance fusion module. The link time statistics module collects cross-node network transmission latency, data shard size, effective bandwidth of cache nodes, and deserialization memory write time, and integrates these to calculate the complete prefetch link time. The prefetch fulfillment calculation module uses the natural clock span as the access period, combines the link time to determine the access time margin ratio, and generates the prefetch fulfillment. The dominance fusion module integrates the clock consistency coefficient, prefetch fulfillment, and skew penalty to calculate the predicted access dominance and outputs it to the control unit 15, providing a unique quantitative weight basis for prefetch scheduling decisions.

[0028] The core function of control unit 15 is to control the asynchronous prefetching operation of data shards based on the predicted access dominance and the physical memory capacity of the cache nodes in the distributed data storage system. This unit can be implemented through the cloud computing control plane scheduling decision module, which includes a sorting and truncation module, a prefetching timing calculation module, and an instruction issuance module. The sorting and truncation module sorts the data shards under the jurisdiction of the cache nodes in descending order of predicted access dominance, and determines the set of prefetched data shards by naturally truncating them based on the physical memory capacity of the cache. The prefetching timing calculation module calculates the prefetching sleep waiting time for each data shard based on the access cycle and link latency. After the current access peak ends and the sleep waiting time is reached, the instruction issuance module issues an asynchronous prefetching instruction to the corresponding cache node, directly driving the prefetching action to be implemented, completing the optimized scheduling closed loop.

[0029] Multiple storage nodes form the foundation for distributed data storage. Their core functions include distributed storage data sharding, responding to business read requests, and cooperating with cache nodes to complete data retrieval operations. Storage nodes can be implemented using distributed storage servers in a cloud computing environment. Each server integrates an object storage engine and a data sharding storage module. The object storage engine is responsible for parsing read / write requests, locating data shards, and completing basic input / output (I / O) interactions. The data sharding storage module adopts a cold / hot tiered architecture, storing low-frequency cold data shards and high-frequency hot data shards separately.

[0030] Multiple cache nodes serve as hot data residing carriers for distributed data storage. Their core functions include maintaining the hot data cache memory space, receiving and executing asynchronous prefetch commands issued by the distributed data storage optimization device 1, and pulling data from storage nodes. Cache nodes can be implemented using cache servers in a cloud computing environment, with each server having a built-in cache management module and command execution module. The cache management module maintains the node's hot cache memory space, records memory usage status, and coordinates data hot / cold switching. The command execution module receives prefetch commands from the control unit 15, performs cold data fragmentation and cross-node transport, deserialization, and memory write operations, and simultaneously feeds back the prefetch execution results to the cloud-based distributed data storage optimization device 1, ensuring the reliability of distributed storage data flow and prefetch execution.

[0031] Please see Figure 2 The diagram illustrates a flowchart of a cloud computing-based distributed data storage optimization method according to an embodiment of the present invention. This cloud computing-based distributed data storage optimization method includes: S1. Obtain the access count sequence of data shards in the distributed data storage system within the discrete time window, and the mixed aggregation count sequence of the storage node where the data shards are located within the discrete time window.

[0032] A distributed data storage system is a cloud computing underlying storage architecture composed of multiple independent storage nodes that collaboratively handle massive data storage and read / write requests. Complete data files are split into independent data units, resulting in data shards, which represent the smallest scheduling granularity in distributed storage. Obtaining the access sequence of a data shard itself is to understand the original access frequency and fluctuation patterns of the target shard. Obtaining the node aggregation sequence is to understand the overall access background fluctuations of the nodes.

[0033] In some implementations, the minimum I / O scheduling granularity of the storage engine is used as the width of the discrete-time window, dividing continuous time into multiple discrete-time windows. The rule for setting the width of the discrete-time window is to directly use the underlying clock interrupt cycle of the storage engine, which also corresponds to the minimum I / O scheduling granularity of the storage engine. This ensures that the time statistical scale is consistent with the execution scale of subsequent prefetching actions, avoiding statistical deviations caused by mismatch between the time window and the storage scheduling unit. For example, when the minimum I / O scheduling granularity of the storage engine is 1 second, the width of the discrete-time window is set to 1 second, dividing the continuous time axis into non-overlapping discrete-time windows in 1-second units.

[0034] At the object storage engine entry point of the storage node where the data shards reside, internal system traffic is filtered to retain only business read requests. A bypass probe module is deployed at the object storage engine entry point of each storage node to collect and filter traffic. This bypass probe module mirrors read requests entering the storage engine before network packets are parsed, data shards are located, and actual disk or object reads begin, thus avoiding performance impact on normal I / O paths. After collecting read requests, internal system traffic such as replica repair, background compression, cold / hot merging, index inspection, and metadata reconstruction is filtered based on the process identifier and service type of the source port, retaining only valid business read requests across business lines. The filtering rule for internal system traffic is a preset service type whitelist, only recognizing read requests marked as belonging to business application processes or business service ports as valid business read requests. For example, read requests whose process identifier belongs to the storage system's background management process set are directly filtered and not included in subsequent statistics.

[0035] In response to a business read request, an event record is generated containing the data shard identifier to be requested and the request arrival time. Each valid business read request corresponds to an event record in the form of an access event tuple, where the data shard identifier comes from the data shard identifier returned by the storage engine object location module, and the request arrival time comes from the storage engine's underlying clock interrupt record.

[0036] The number of event records within each discrete time window of the statistical data shard is counted to obtain an access count sequence. Using discrete time windows as units, the event records corresponding to each valid read request are time-aligned. The discrete time window number to which the request arrival time belongs is associated with the corresponding data shard identifier. The number of valid reads for each data shard within each discrete time window is counted to obtain the data shard window access count. All data shard window access counts are arranged by window number to form the data shard access count sequence. To reduce the data transmission pressure on edge nodes, edge nodes only asynchronously upload the shard identifier, window number, and corresponding window access count to the control plane, without uploading the full raw log, thereby reducing transmission bandwidth consumption.

[0037] The access counts of multiple data shards on the same storage node within the same discrete time window are summed to obtain the mixed-part aggregation count sequence. For any storage node, the access counts of all data shards on that node within the same discrete time window are summed to obtain the node's mixed-part aggregation count. All node mixed-part aggregation counts are arranged according to the window number to form the storage node's mixed-part aggregation count sequence. The mixed-part aggregation count sequence is obtained by aggregating the same-source access event stream within the node and is used to characterize node-level access fluctuations under the background of consistent hashing mixed-part. After receiving the shard identifier, window number, and access count uploaded by the edge node, the control plane performs time rearrangement according to the window number to restore the discrete access frequency sequence of each data shard and the mixed-part aggregation sequence of each node, thereby ensuring the timing accuracy of the sequence data.

[0038] S2. Perform mixed-part compensation processing on the access count sequence and the mixed-part aggregation count sequence to determine the isolation residual sequence used to characterize the relative access intensity change of data fragments after stripping the mixed-part background of nodes.

[0039] Storage nodes typically host data shards from multiple services. The overall access peaks or troughs of a node can mask the true hot / cold characteristics of the target shard (for example, when the entire node is busy, low-frequency shards will appear to be accessed frequently). Mixed-distribution compensation processing addresses this scenario where multiple data shards are stored together on a single storage node, leading to mutual interference in access behavior. It uses algorithms to remove the interference from overall node access fluctuations on the target shard's access characteristics. The core is to eliminate the noise caused by mixed node deployment, isolating the target shard from the complex background fluctuations of the nodes. Generating an isolation residual sequence is to obtain interference-free data that accurately reflects the shard's own access strength changes, avoiding subsequent misjudgments of hot / cold shards based on distorted data, and clearing away interference to accurately extract access characteristics.

[0040] In some implementations, the short-scale observation interval is determined based on the average duration of historical traffic peaks in the data shard. The duration of continuous access to the data shard by recent business read requests, recorded in the control plane task orchestrator, is read. The time span of each consecutive access event is identified and extracted, and the median is taken as the average duration of the historical traffic peaks for that data shard. The short-scale observation interval is calculated by dividing the average duration by the storage engine's minimum input / output scheduling granularity and rounding up to obtain the number of consecutive scheduling windows required to cover the typical macroscopic business bandwidth of the data shard. This short-scale observation interval is a preset dynamic interval that adaptively adjusts according to changes in the actual business access patterns of the data shard. For example, if a data shard is historically frequently read by batch processing tasks, with an average access peak duration of approximately 10 seconds and a scheduling granularity of 1 second, then the short-scale observation interval consists of 10 consecutive discrete time windows. This interval can completely cover and effectively reflect the entire process of local access fluctuations of the data shard in a single business activity.

[0041] The long-scale observation interval is determined based on the duration of mixed access to the storage node where the data shard resides within a business cycle. The natural runtime of recurring scheduled jobs in the control plane task orchestrator is read. This runtime is a complete business runtime determined based on the business scenario; in an e-commerce data lake scenario, it corresponds to a complete natural business day, which can be set to 24 hours. This business cycle is divided by the scheduling granularity to obtain the long-scale observation interval, which covers the full background fluctuations of node-level mixed access within a complete natural business cycle. For example, if the business cycle is 24 hours and the scheduling granularity is 1 second, the long-scale observation interval consists of 86,400 consecutive discrete time windows.

[0042] The degree to which the access count sequence deviates from the mean within the short-scale observation interval in a discrete-time window is determined as the local deviation. First, the window access counts of the data pieces within the short-scale observation interval are averaged to obtain the local mean. This local mean represents the average access level of the data piece within the short-scale observation interval, reflecting the regular access intensity of that piece within a local timeframe. Then, the difference between the current discrete-time window access count and this local mean is calculated to obtain the local deviation for the current discrete-time window. This local deviation characterizes the degree to which the access count of the data piece within the current discrete-time window deviates from its local regular access level. It can be positive or negative; a positive value indicates that the access count is higher than the local mean, and a negative value indicates that it is lower than the local mean. For example, if the short-scale observation interval has 5 windows with access counts of 1, 2, 3, 2, and 1 respectively, and the local mean is 2, then the local deviation is 1 if the current discrete-time window access count is 3; if the current discrete-time window access count is 1, then the local deviation is -1.

[0043] The mean absolute deviation of the aggregated count sequence within a long-scale observation interval and the mean absolute deviation of the access count sequence within a short-scale observation interval are used to determine the dilution factor for the mixed-area phenomenon. First, the local mean absolute deviation of the data pattern is calculated by averaging the absolute differences between the access counts of each window within the short-scale observation interval and the local mean. This deviation reflects the dispersion of local access fluctuations within the data pattern itself; a larger value indicates more severe local access fluctuations. Next, the mean absolute deviation of the node mixed-area phenomenon is calculated by averaging the absolute differences between the node aggregated counts and the node background mean within the long-scale observation interval. The node background mean is the average of the node aggregated counts within the long-scale observation interval. This deviation reflects the dispersion of node-level mixed-area access fluctuations; a larger value indicates more severe overall node access fluctuations. Based on these two mean absolute deviations, the dilution factor for the mixed-area phenomenon is calculated. This factor quantifies the degree to which node-level mixed-area fluctuations average and mask the local fluctuations of the target pattern. Specifically, it is calculated by dividing the node mean absolute deviation for the mixed-area phenomenon by the local mean absolute deviation of the data pattern. Notably, when the local mean absolute deviation of the data pattern is 0, the dilution factor for the mixed-area phenomenon is directly assigned a value of 1.

[0044] The continuous bias coefficient is determined by statistically analyzing the difference values ​​of the access count sequence within a short-scale observation interval and the continuous difference values ​​with the same sign. First, the difference value is calculated for the access counts of adjacent windows within the short-scale observation interval. The difference value is the difference between the access counts of the subsequent window and the access count of the preceding window, reflecting the trend of the access count. Then, the difference values ​​within the interval are traversed, and the segment of continuous differences with the same sign that has the largest sum of absolute values ​​is extracted. The same sign in this segment indicates a continuously rising or falling access trend, and the largest sum of absolute values ​​indicates the strongest trend. The absolute values ​​of this continuous difference are accumulated, and then the absolute values ​​of all differences within the short-scale observation interval are accumulated. The continuous bias coefficient is obtained by dividing the accumulated absolute value of the continuous difference by the accumulated absolute value of all differences. The continuous bias coefficient ranges from [0,1]. The larger the value, the more the local access waveform shows a continuous building or falling trend within that interval, which is more consistent with the gradual queuing and closing process of batch processing tasks in the early morning. If all difference values ​​within a short-scale observation interval are positive, and the continuous difference segment with the largest sum of absolute values ​​covers the entire window, then the continuous bias coefficient approaches 1; if the difference values ​​alternate between positive and negative without a clear continuous trend, then the continuous bias coefficient approaches 0. Specifically, when the cumulative sum of the absolute values ​​of all differences is 0, the continuous bias coefficient is directly assigned the value of 0.

[0045] Based on the local deviation, the dilution factor, and the continuous bias coefficient, the isolation residual is determined by combining the access count sequence and the aggregated count sequence, forming an isolation residual sequence. The isolation residual is calculated as follows: The first term calculates the local deviation divided by the local average absolute deviation of the data shard, standardizing the local deviation to eliminate the influence of the shard's own fluctuation amplitude, obtaining a dimensionless relative fluctuation intensity. Positive and negative values ​​reflect the fluctuation direction. Specifically, when the local average absolute deviation of the data shard is 0, the first term is assigned a value of 0, indicating no fluctuation and therefore no deviation. The second term is the dilution factor, used to reverse the dilution masking effect. The stronger the masking, the greater the compensation amplification, restoring the diluted true access signal. The third term calculates the sum of 1 and the continuous bias coefficient, representing the enhancement of the continuous task trend. The more continuous the trend, the stronger the amplification effect, highlighting the characteristics of early morning batch tasks. Multiplying these three parts together to synthesize the fluctuation direction, dilution compensation, and trend enhancement yields the standardized isolation residual. Then, the standardized isolation residuals of each window are arranged chronologically to form an isolation residual sequence. This sequence has stripped away node dilution interference, strengthened task-oriented access trends, and purely reflects the changes in the true relative access intensity of the shards.

[0046] S3. Determine the beat consistency coefficient, which characterizes the stability of the access peak time interval of data fragmentation, based on the isolation residual sequence, and determine the skew penalty amount, which characterizes the burstiness of access pulses in data fragmentation, based on the difference between the rise and fall rates of the access peaks in the isolation residual sequence.

[0047] The access peak value in the isolation residual sequence is significantly higher than that of adjacent time periods, representing the peak period of concentrated access behavior in the data shards. If only the number of accesses is considered, periodic batch tasks (stable cold data) are easily misclassified as hot data, and sudden queries (brief hot data) are misclassified as cold data. Therefore, by quantifying the stability of the time interval between adjacent access peaks of data shards, a beat consistency coefficient is obtained, which can identify whether shard access has a stable cycle (distinguishing between periodic tasks and random access). Simultaneously, the difference in the rise and fall rates of the access peaks reflects the changing rhythm of access behavior, from which the degree to which data shard access behavior deviates from a stable periodic pattern and exhibits irregular burst characteristics can be analyzed, yielding a skew penalty, which can identify whether access is bursty and impulsive (distinguishing between stable access and brief bursts).

[0048] In some implementations, the absolute value of the isolation residual sequence is taken to obtain the energy sequence. Each element represents the magnitude of the access intensity within the corresponding window, eliminating the influence of positive and negative fluctuations in the isolation residual and unifying it into a non-negative energy representation.

[0049] Using a short-scale observation interval as the width, a center-weighted smoothing process is applied to the energy sequence to eliminate temporal misalignment between adjacent access peaks caused by resource queuing, resulting in an energy envelope sequence. Specifically, the weighted smoothing process can employ a preset symmetrical linearly decreasing weight, with the center of the window as the symmetrical point, where the weight is largest and decreases linearly towards both sides, with a total weight sum of 1. For example, when the short-scale observation interval is 5 windows, the weights can be set to [0.1, 0.2, 0.4, 0.2, 0.1]. The smoothing process merges adjacent sub-peaks caused by resource queuing and node contention into a macroscopic peak, eliminating local temporal misalignment and local jitter caused by resource scheduling misalignment, while maintaining the relative energy strength between peaks.

[0050] The window position corresponding to each maximum in the energy envelope sequence is extracted, and the inter-peak distance corresponding to adjacent maximum window positions is determined. The median of the inter-peak distance is used as the natural beat span of the data shard. First, local maximum positions are extracted from the energy envelope sequence to obtain a set of maximum positions, each position corresponding to a window index of a macroscopic access peak. Then, the difference between two adjacent maximum window positions is calculated and multiplied by the minimum input / output scheduling granularity of the storage engine to obtain the inter-peak distance sequence, which represents the actual time interval between adjacent access peaks. The median of the inter-peak distance sequence is taken to obtain the natural beat span of the data shard. The median is used instead of the mean to avoid interference from extreme values ​​and to represent the typical time interval between adjacent access peaks.

[0051] The median of the absolute value of the difference between the inter-peak distance and the natural beat span is determined as the beat dispersion measure, used to quantify the fluctuation of the time interval between adjacent visiting peaks. The larger the value, the more unstable the beat and the more susceptible it is to sudden interference. The median is used to quantify the overall deviation of the inter-peak interval to avoid the influence of extreme deviation values.

[0052] The beat consistency coefficient is determined based on the natural beat span and beat dispersion, and is used to characterize the stability of the time interval between data fragment access peaks. Specifically, the beat consistency coefficient is calculated by dividing the natural beat span by the sum of the beat dispersion and the natural beat span. The value ranges from [0, 1]. A value closer to 1 indicates a more stable time interval between adjacent access peaks, while a value closer to 0 indicates a more disordered beat. Notably, when the sum of the denominator beat dispersion and the natural beat span is 0, the beat consistency coefficient is directly assigned a value of 0.

[0053] Specifically, when the number of elements in the set of maximum positions is less than 2, the natural beat span and beat consistency coefficient are both directly assigned a value of 0.

[0054] In some implementations, the pre-peak rise slope is determined based on the rise amplitude and rise duration between each maximum and the previous adjacent trough in the energy envelope sequence. For each maximum window position, the nearest local trough window position is searched forward, and the difference between the energy envelope sequences corresponding to the maximum and trough is calculated to obtain the pre-peak rise amplitude. Simultaneously, the difference between the maximum window position and the trough window position is calculated and multiplied by the minimum input / output scheduling granularity of the storage engine to obtain the pre-peak rise duration. The rise amplitude is divided by the rise duration to calculate the pre-peak rise slope, quantifying the rise speed during the access peak formation phase. A gentle rise slope corresponds to the gradual establishment process of periodic task-type access, while a steep rise slope corresponds to the rapid initiation process of burst access.

[0055] The post-peak descent slope is determined by the magnitude and duration of the descent between each maximum and the next adjacent trough in the energy envelope sequence. For each maximum window position, the nearest local trough window position is searched backward, and the difference between the energy envelope sequences corresponding to the maximum and trough is calculated to obtain the post-peak descent magnitude. Simultaneously, the difference between the trough window position and the maximum window position is calculated and multiplied by the minimum input / output scheduling granularity of the storage engine to obtain the post-peak descent duration. The descent slope is calculated by dividing the descent magnitude by the descent duration, quantifying the descent rate during the peak decay phase. A gentle descent slope corresponds to the gradual fading process of periodic task-type access, while a steep descent slope corresponds to the rapid decay process of burst access.

[0056] Based on the post-peak descent slope and the pre-peak ascending slope, a single-peak skew penalty value is determined for each maximum value to characterize the waveform skewness of a single access peak. Specifically, the difference between the pre-peak ascending slope and the post-peak descent slope is calculated, divided by the sum of the pre-peak ascending slope and the post-peak descent slope, and the maximum value between this result and 0 is taken as the single-peak skew penalty value, ranging from [0, 1]. The penalty value is greater than 0 only when the pre-peak ascending slope is greater than the post-peak descent slope. The larger the pre-peak ascending slope is relative to the post-peak descent slope, the closer the penalty value is to 1. Taking the maximum value between 0 and 0 ensures that only asymmetric access peaks where the rise is faster than the fall are penalized, distinguishing between the asymmetric sharp waveform of sudden pulse access and the relatively symmetrical smooth waveform of periodic task access. Specifically, when the sum of the post-peak descent slope and the pre-peak ascending slope is 0, the single-peak skew penalty value is assigned to 0.

[0057] Using the peak energy of each maximum (i.e., the corresponding energy envelope sequence value) as a weight, a weighted average of the single-peak skew penalty values ​​is performed to obtain the skew penalty amount. Specifically, each single-peak skew penalty value is multiplied by the corresponding peak energy, summed, and then divided by the sum of all peak energies to quantify the degree of skewness of the overall access waveform of the data slice. A larger value indicates a greater bias towards burst pulse access, while a smaller value indicates a greater bias towards stable task access, ensuring that peaks with higher energy have a greater impact on the overall skewness.

[0058] S4. Determine the predictive access dominance of data shards based on the cycle consistency coefficient, skew penalty, and link time required to complete one prefetch operation.

[0059] Prefetching refers to the scheduling action by which a cache node proactively pulls data shards from the storage node to its local cache before the data shards are actually accessed. The total time consumed from issuing the prefetch instruction to successfully writing the data shard into the cache node's memory is denoted as the link latency. If the latency is too long, by the time the data is pulled into the cache, the access peak has passed, rendering prefetching meaningless and wasting cache resources. Therefore, by integrating the tick consistency coefficient (stability), skew penalty (burstiness), and link latency (timeliness), a comprehensive index is used to evaluate the value of shards from multiple dimensions, quantifying their priority prefetching value and obtaining the predicted access dominance.

[0060] In some implementations, the cross-node network transmission latency between the storage node and the cache node where the data shard resides, the data size of the data shard, and the average effective bandwidth, average deserialization, and memory write time of the cache node are obtained. Specifically, the cross-node network transmission latency is read from the average value recorded by the network link status monitoring module within the most recent statistical period, with the statistical period optionally being 5 minutes, reflecting the average network latency between the storage node and the cache node. The data size of the data shard is read from the target data shard logical size recorded in the object metadata management module, which is a fixed value in bytes or MB. The average effective bandwidth of the cache node is read from the average effective throughput of the cache node I / O statistician within the current statistical period, with the statistical period optionally being 1 minute, reflecting the actual data transmission capacity of the cache node. The average deserialization and memory write time is read from the average processing time recorded in the cache node object parsing monitoring log, with the statistical period optionally being 1 hour, reflecting the average time spent by the cache node processing data and completing memory writes.

[0061] Based on the data size of the data shards and the average effective bandwidth of the cache nodes, the data transfer time from the storage node where the data shard resides to the cache node is determined. Specifically, this is calculated by dividing the data size of the data shard by the average effective bandwidth of the cache node. Notably, when the average effective bandwidth of the cache node is 0, it indicates that the cache node currently has no effective transmission capacity, and the data transfer time is directly assigned the preset maximum value of 3600 seconds (i.e., 1 hour) to ensure that the link latency is sufficiently large and to avoid calculation errors.

[0062] The link latency is obtained by summing the cross-node network transmission latency, data transmission time, and average deserialization and memory write time, which represents the total time cost of the prefetch operation from initiation to data readiness.

[0063] In some implementations, the natural beat span is used as the access period for data shards. This span is the median of the inter-peak distance sequence, which characterizes the typical time interval between adjacent access peaks of a data shard.

[0064] The access time margin is determined based on the access period and link latency. The proportion of the access time margin to the access period is used as the prefetch fulfillment rate. Specifically, the access time margin is calculated by subtracting the link latency from the access period. It represents the remaining safe time margin after the prefetch operation is completed. A positive value indicates that there is remaining time, while a negative value indicates that the link latency exceeds the access period and prefetching cannot be completed. The prefetch fulfillment rate is calculated by dividing the access time margin by the access period. This normalizes the time margin, and the result is the maximum value between the margin and zero, ensuring that the result is non-negative. A prefetch fulfillment rate of 0 indicates that the prefetch operation cannot be completed within the access period and has no fulfillment value. Notably, when the access period is 0, the prefetch fulfillment rate is directly assigned a value of 0.

[0065] The predicted access dominance is obtained by fusing the clock consistency coefficient, prefetch fulfillment rate, and skew penalty. The fusion method is multiplication. The first term is the clock consistency coefficient, which characterizes the stability of access; a larger value indicates greater stability. The second term is the prefetch fulfillment rate, which characterizes the feasibility of prefetching; a larger value indicates easier completion. The third term is 1 minus the skew penalty, which characterizes the non-burst nature of the access waveform; a larger value indicates a smoother waveform. Multiplying these three terms yields a comprehensive prefetch priority index. If any dimension is zero, the overall dominance is zero, ensuring that only shards that simultaneously meet all three conditions have prefetch priority. The predicted access dominance ranges from [0, 1]. A value closer to 1 indicates that the shard simultaneously possesses a stable access clock, fulfillable prefetch feasibility, and a smooth task-type waveform, resulting in a higher prefetch priority; a value of 0 indicates that the shard does not meet the core conditions and is not included in the prefetch set.

[0066] S5. Based on the predicted access dominance and the physical memory capacity of the cache nodes in the distributed data storage system, control the asynchronous prefetching operation of data shards.

[0067] The physical memory of cache nodes is limited, making it impossible to cache all data shards. Blind prefetching can lead to low-value data occupying the cache (cache pollution), which in turn reduces the hit rate of hot data. Prioritizing cache resources for high-value shards based on predicted access dominance, combined with physical memory capacity constraints, avoids cache overflow and maximizes memory utilization. Asynchronous prefetching allows for return without waiting for prefetch instructions to complete, without blocking other system scheduling tasks, ensuring the core system scheduling process and not affecting normal business access. The ultimate goal is to maximize the cache hit rate of hot data, shorten the read latency of periodic tasks, completely solve the cache pollution problem, and improve the overall operating efficiency of the distributed storage system with limited cache resources.

[0068] In some implementations, data shards within the scope of a cache node are sorted in descending order based on predicted access dominance.

[0069] The prefetch data shard set is determined in descending order, within the physical memory capacity limit. The physical memory capacity of the cache node is a preset system parameter, representing the maximum available hot cache memory for the node, defined by the system configuration file. The logical size of each data shard is accumulated sequentially in descending order. When the accumulated total memory usage first exceeds the available physical memory capacity of the cache node, subsequent shards are no longer included, and all shards before the boundary are defined as the prefetch data shard set. If the logical size of a data shard is greater than the remaining available memory, that shard is not included in the set to avoid a single shard occupying all remaining memory, preventing other high-value shards from being cached. This memory constraint-based truncation ensures that the total memory usage of the prefetch set does not exceed the node's hardware limit, preventing cache overflow, while maximizing the cache hit rate of high-value shards and reducing cache pollution.

[0070] Based on the access cycle and link latency, the prefetch sleep wait time for each data shard in the prefetch data shard set is determined. For each maximum value, the time difference between its corresponding window position and the first valley window position after the peak is recorded as the single-peak post-peak duration of that peak, and the median of the single-peak post-peak durations of all peaks is taken as the post-peak duration of that data shard. The specific calculation rule for the prefetch sleep wait time is as follows: subtract the link latency from the access cycle, then subtract the post-peak duration, and compare the result with 0, taking the larger one. When the calculation result is greater than 0, the sleep wait time is the calculated result, which means that starting from the peak moment, the prefetched data will be ready just before the next access peak is initiated; when the calculation result is less than or equal to 0, it means that even if prefetching is done immediately, it cannot be completed within the access cycle, the sleep wait time is 0, and the prefetch command is issued immediately. Precisely control the timing of prefetch instructions to avoid premature occupancy of cache resources, while ensuring that data is written to memory before access peaks arrive, maximizing the caching benefits of prefetching.

[0071] At the window corresponding to the maximum value of the current access peak of the current data shard, as the starting point for timing, an asynchronous prefetch command is issued to the cache node after the prefetch sleep wait time of the current data shard. Specifically, the access peak status of each data shard is tracked in real time by continuously receiving window access counts uploaded by edge nodes, and the end time of the current access peak is identified. After the access peak ends, the prefetch sleep wait time is started. After the timer expires, an asynchronous prefetch command is issued to the cache node, which includes the identifier of the target data shard and the address of the source storage node. After receiving the command, the cache node performs cross-node transport, deserialization, and memory write operations for the data shard, migrating the cold data shard from the cold data baseline layer to the hot data cache layer. For shards not included in the prefetch set, no prefetch operation is performed to avoid ineffective occupation of cache resources.

[0072] Based on the above technical solution, in the scenario of distributed data storage in cloud computing, by performing mixed-location compensation, extracting beat and peak features, calculating the access dominance of data shard access sequences and node mixed-location aggregation sequences, and combining cache memory constraints to control prefetching, it is possible to effectively remove node mixed-location interference, accurately distinguish between stable periodic access and burst pulse access without relying on empirical thresholds, realize reasonable scheduling of prefetching operations, reduce misjudgment of hot and cold data and cache pollution, thereby improving the hit rate of hot data cache layer for stable access shards, shortening the read latency of periodic batch tasks, and improving cache resource utilization and cross-node data transfer efficiency.

[0073] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0074] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0075] In this embodiment of the invention, the functional units of the cloud computing-based distributed data storage optimization device can be divided according to the above method example. For example, each function can be divided into its own functional units, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0076] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings and the disclosure, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In this invention, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several of the functions listed in this invention.

[0077] Although the invention has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the scope of the invention. Accordingly, this specification and drawings are merely illustrative of the invention and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its scope. Thus, if such modifications and modifications of the invention fall within the scope of the invention and its equivalents, the invention is also intended to include such modifications and modifications.

Claims

1. A cloud computing-based distributed data storage optimization method, characterized in that, include: Obtain the access count sequence of data shards in a distributed data storage system within a discrete time window, and the mixed aggregation count sequence of the storage node where the data shards are located within the discrete time window; Mixed-area compensation processing is performed on the access count sequence and the mixed-area aggregation count sequence to determine the isolation residual sequence used to characterize the relative access intensity change of data fragments after stripping the mixed-area background of nodes; The clock consistency coefficient, used to characterize the stability of the access peak time interval of the data fragment, is determined based on the isolation residual sequence, and the skew penalty amount, used to characterize the burstiness of the access pulse of the data fragment, is determined based on the difference between the rise and fall rates of the access peaks in the isolation residual sequence. The predictive access dominance of the data shard is determined based on the cycle consistency coefficient, the skew penalty, and the link time to complete one prefetch operation. The asynchronous prefetching operation on the data shards is controlled based on the predicted access dominance and the physical memory capacity of the cache nodes in the distributed data storage system.

2. The cloud computing based distributed data storage optimization method of claim 1, wherein, Mixed-part compensation processing is performed on the access count sequence and the mixed-part aggregation count sequence to determine the isolation residual sequence, including: The short-scale observation interval is determined based on the average duration of the historical traffic peaks of the data shards, and the long-scale observation interval is determined based on the mixed access duration of the storage node where the data shards are located within a business cycle. The degree to which the numerical value of the access count sequence deviates from the mean within the short-scale observation interval in the discrete time window is determined as the local deviation. The average absolute deviation of the mixed part aggregation count sequence within the long-scale observation interval and the average absolute deviation of the access count sequence within the short-scale observation interval are statistically compared to determine the mixed part dilution coefficient. The difference values ​​of the access count sequence within the short-scale observation interval and the continuous difference values ​​with the same sign are statistically analyzed to determine the continuous bias coefficient; Based on the local deviation, the dilution factor of the mixed part, and the continuous bias factor, the isolation residual is determined by combining the access count sequence and the aggregation count sequence of the mixed part, and the isolation residual sequence is formed.

3. The cloud computing based distributed data storage optimization method of claim 2, wherein, Determining the beat consistency coefficient based on the isolated residual sequence includes: The absolute value of the isolated residual sequence is taken to obtain the energy sequence; Using the short-scale observation interval as the width, the energy sequence is subjected to center-weighted smoothing to obtain an energy envelope sequence; the weighted smoothing is used to eliminate the time misalignment of adjacent access peaks caused by resource queuing; Extract the window position corresponding to each maximum value in the energy envelope sequence, determine the inter-peak distance corresponding to adjacent maximum window positions, and use the median of the inter-peak distance as the natural beat span of the data segment; The median of the absolute value of the difference between the interpeak distance and the natural beat span is determined as the beat dispersion measure. The beat consistency coefficient is determined based on the natural beat span and the beat dispersion.

4. The cloud computing based distributed data storage optimization method of claim 3, wherein, The skew penalty is determined based on the difference between the rise and fall rates of the access peak in the isolated residual sequence, including: The upward slope before the peak is determined based on the rise amplitude and rise duration between each maximum value and the previous adjacent trough value in the energy envelope sequence. The post-peak descent slope is determined based on the magnitude and duration of the descent between each maximum and the next adjacent trough in the energy envelope sequence. Based on the post-peak descent slope and the pre-peak ascending slope, determine the single-peak skew penalty value for each maximum value; The skew penalty is obtained by weighting the peak energy of each maximum value and averaging the single-peak skew penalty values.

5. The distributed data storage optimization method based on cloud computing according to claim 3, characterized in that, The predicted access dominance of the data shard is determined based on the cycle time consistency coefficient, skew penalty, and link time to complete one prefetch operation, including: The natural beat span is used as the access period for the data fragment; The access time margin is determined based on the access period and the link consumption time, and the proportion of the access time margin in the access period is used as the pre-fetch fulfillment rate. The predicted access dominance is obtained by fusing the beat consistency coefficient, the prefetch fulfillment rate, and the skew penalty.

6. The distributed data storage optimization method based on cloud computing according to claim 5, characterized in that, Obtaining the link time includes: The cross-node network transmission latency between the storage node where the data shard is located and the cache node, the data volume of the data shard, and the average effective bandwidth, average deserialization and memory write time of the cache node are obtained. Based on the data volume of the data fragment and the average effective bandwidth of the cache node, the data transmission time of the data fragment from the storage node where the data fragment is located to the cache node is determined; The link latency is obtained by summing the cross-node network transmission latency, the data transmission time, and the average deserialization and memory write time.

7. The distributed data storage optimization method based on cloud computing according to claim 5, characterized in that, Based on the predicted access dominance and the physical memory capacity of the cache nodes in the distributed data storage system, the asynchronous prefetching operation on the data shards is controlled, including: The data shards within the jurisdiction of the cache node are sorted in descending order according to the predicted access dominance. In accordance with the descending order, within the limit of the physical memory capacity, determine the set of prefetch data fragments; Based on the access cycle and the link consumption time, determine the prefetch sleep wait time for each data shard in the prefetch data shard set; At the window time corresponding to the maximum value of the current access peak of the current data shard, as the starting point of the timing, after the prefetch sleep waiting time of the current data shard, an asynchronous prefetch instruction is sent to the cache node.

8. The distributed data storage optimization method based on cloud computing according to claim 1, characterized in that, Obtaining the access count sequence of a data shard in a distributed data storage system within a discrete time window, and the mixed aggregation count sequence of the storage node where the data shard is located within the discrete time window, including: The continuous time is divided into multiple discrete time windows, with the smallest input / output scheduling granularity of the storage engine as the width of the discrete time window. On the object storage engine entry side of the storage node where the data shard is located, the internal traffic of the filtering system retains business read requests; In response to a business read request, an event log is generated containing the identifier of the requested data fragment and the time the request arrived; The number of event records in each discrete time window of the data shard is counted to obtain the access count sequence; The access counts of multiple data shards on the storage node where the data shard is located within the same discrete time window are summed to obtain the mixed-part aggregated count sequence.

9. A distributed data storage optimization device based on cloud computing, characterized in that, include: The data acquisition unit is used to acquire the access count sequence of data shards in the distributed data storage system within a discrete time window, and the mixed aggregation count sequence of the storage node where the data shards are located within the discrete time window. The mixed-part compensation unit is used to perform mixed-part compensation processing on the access count sequence and the mixed-part aggregation count sequence to determine the isolation residual sequence used to characterize the relative access intensity change of the data fragment after stripping the mixed-part background of the node. The skew determination unit is used to determine, based on the isolation residual sequence, a beat consistency coefficient characterizing the stability of the access peak time interval of the data fragment, and a skew penalty amount characterizing the burstiness of the access pulse of the data fragment based on the difference between the rise and fall rates of the access peaks in the isolation residual sequence. The predictive access unit is used to determine the predictive access dominance of the data fragment based on the clock consistency coefficient, the skew penalty, and the link time to complete one prefetch operation. The control unit is used to control the asynchronous prefetching operation of the data shards based on the predicted access dominance and the physical memory capacity of the cache nodes in the distributed data storage system.

10. A distributed data storage optimization system based on cloud computing, running in a cloud computing environment, characterized in that, It includes multiple storage nodes, multiple cache nodes, and cloud-based distributed data storage optimization equipment; The distributed data storage optimization device performs the method as described in any one of claims 1 to 8; The cache node is used to receive and execute the asynchronous prefetch instruction issued by the distributed data storage optimization device; The storage node is used to store data shards and provide data when it receives a data request from the cache node.