Distributed cache optimization method and device, electronic equipment and storage medium

By introducing hot data prediction and reinforcement learning models into the distributed cache system, and dynamically adjusting the cache strategy, the problems of low cache hit rate and unbalanced resource allocation are solved, and more efficient resource utilization and system performance improvement are achieved.

CN120336007AActive Publication Date: 2025-07-18北京领雁科技股份有限公司

Patent Information

Application Number
CN202510408706.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The problems of low cache hit rate, lack of dynamic adjustment capabilities and uneven resource allocation in traditional distributed cache systems have led to system performance degradation and resource waste.

Method used

By introducing hotspot data prediction models and reinforcement learning models, future hotspot data are predicted and loaded into the cache in advance, and dynamically adjust the cache strategy based on real-time load information and performance indicators to achieve balanced allocation and optimal management of resources.

Benefits of technology

Improve cache hit rate, reduce miss rate, ensure optimal cache management under different load and request modes, and avoid waste of resources between nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336007A_ABST
    Figure CN120336007A_ABST
Patent Text Reader

Abstract

The invention provides a distributed cache optimization method and device, electronic equipment and a storage medium, and the method comprises the steps: inputting historical request data of a plurality of hotspot data into a hotspot data prediction model, and carrying out the data preprocessing, feature extraction and future access frequency prediction processing of the historical request data, predicting the access frequency of each piece of hotspot data in a future time period; determining data information of the target hotspot data of which the access frequency is greater than a preset access frequency, and performing node caching on the data information of the target hotspot data based on the real-time load resource use information of each cache node; and dynamically adjusting the caching strategy of each caching node based on a reinforcement learning model and the multiple pieces of performance index data of each caching node. The optimal cache management is ensured to be realized under different loads and request modes, the balanced allocation of resources is realized, and the resource waste among the nodes is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed caching technology, and in particular to an optimization method, device, electronic device and storage medium for distributed caching. Background Art

[0002] In a distributed caching system, the effective utilization of cache resources and the improvement of cache hit rate are key factors in improving system performance. Traditional distributed caching systems adopt static cache replacement strategies (such as LRU, FIFO), and usually encounter the following problems: low cache hit rate: due to the static nature of the cache strategy, it cannot effectively cope with changes in different loads and request patterns, resulting in a low cache hit rate and affecting system performance. Lack of dynamic adjustment ability: existing cache systems cannot dynamically adjust the cache strategy in real time according to factors such as changes in user requests and node resources, resulting in waste of cache resources. Unbalanced resource allocation: in a distributed environment, the loads of multiple cache nodes vary greatly, and traditional methods cannot coordinate the cache content between nodes, resulting in unbalanced allocation of cache resources. Therefore, how to optimize distributed caching has become a technical problem that cannot be underestimated. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide an optimization method, device, electronic device and storage medium for distributed caching, which can predict future hot data in advance, load hot data into the cache in advance, reduce cache miss rate, introduce a reinforcement learning model to continuously adjust the cache strategy according to real-time feedback, ensure optimal cache management under different loads and request patterns, achieve balanced allocation of resources, and avoid waste of resources between nodes.

[0004] The embodiment of this application provides an optimization method for distributed caching, and the optimization method includes:

[0005] Input the historical request data of multiple hot data into the hot data prediction model, perform data preprocessing, feature extraction and future access frequency prediction processing on the historical request data, and predict the access frequency of each hot data in the future time period; wherein, the hot data prediction model is obtained by iteratively training a long short-term memory network model;

[0006] Determine the data information of the target hot data whose access frequency is greater than the preset access frequency, and perform node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node;

[0007] Dynamically adjust the cache strategy of each cache node based on the reinforcement learning model and multiple performance index data of each cache node.

[0008] In a possible implementation, for each of the hot data, inputting the historical request data of multiple hot data into a hot data prediction model, performing data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predicting the access frequency of each hot data in a future time period, including:

[0009] Extracting time features, access frequency features, and data ID embedding vectors from the historical request data of the hot data based on the hot data prediction model, and constructing time series for the time features, access frequency features, and data ID embedding vectors according to a preset multiple of time windows, and determining the time series features under each time window;

[0010] Performing data cleaning processing on the time series features under each time window, and determining the target time series features under each time window;

[0011] Performing short-term feature capture on the target time series features under each time window based on the first long short-term memory network layer of the hot data prediction model, and outputting a short-term feature capture sequence for each time window;

[0012] Inputting the short-term feature capture sequence for each time window into a second long short-term memory network layer for long-term trend feature capture, and determining a long-term feature capture sequence for each time window;

[0013] Inputting the long-term feature capture sequence for each time window into the fully connected network layer of the hot data prediction model for processing, and outputting the access frequency of the hot data in a future time period.

[0014] In a possible implementation, the hot data prediction model is determined through the following steps:

[0015] Inputting the historical request data of multiple sample hot data into the long short-term memory network model, processing the historical request data of each sample hot data, and determining the predicted access frequency of each sample hot data in a future time period;

[0016] Based on the actual access frequency and the predicted access frequency of each sample hot data in a future time period, determining the root mean square error value and the mean absolute error;

[0017] If either the root mean square error value or the mean absolute error is greater than a preset threshold, optimizing the network parameters of the long short-term memory network model, and continuing to train the optimized long short-term memory network model until both the root mean square error value and the mean absolute error are less than or equal to the preset threshold, and then determining the hot data prediction model.

[0018] In a possible implementation, caching the data information of target hot data based on the real-time load resource usage information of each cache node includes:

[0019] Detecting whether the cache node with the lowest real-time load resource usage has enough cache space to accommodate the data information of the target hot data;

[0020] If so, loading and caching the data information of the target hot data in the future time period into the cache node with the lowest real-time load resource usage;

[0021] If not, deleting the low-frequency data in the cache node with the lowest real-time load resource usage to complete the expansion of the cache node, and caching the data information of the target hot data in the future time period into the expanded cache node.

[0022] In a possible implementation, dynamically adjusting the cache policy of each cache node based on the reinforcement learning model and multiple performance index data of each cache node includes:

[0023] Performing feature extraction on the multiple performance index data of each cache node to determine the state feature vector of each cache node; wherein, the performance index data includes cache hit rate, average request delay, node load, and cache space utilization rate;

[0024] Inputting the state feature vector of each cache node into the reinforcement learning model, and performing reinforcement optimization on the cache policy adjustment actions of the state feature vectors of each cache node based on the greedy algorithm and a preset reward mechanism;

[0025] Dynamically adjusting the corresponding cache node based on the cache policy adjustment actions of each cache node; wherein, the cache policy adjustment actions include cache content adjustment actions and load balancing.

[0026] In a possible implementation, for the load balancing adjustment action, dynamically adjusting the corresponding cache node based on the cache policy adjustment actions of each cache node; wherein, the cache policy adjustment actions include cache content adjustment actions or load balancing adjustment actions, includes:

[0027] Based on the distributed cache cluster and the consistent hashing algorithm, migrating the cached data of the hot data in the high-load cache node to the low-load cache node.

[0028] In a possible implementation, after determining the data information of the target hot data whose access frequency is greater than the preset access frequency and caching the data information of the target hot data based on the real-time load resource usage information of each cache node, the optimization method further includes:

[0029] After receiving an update request for the cached data sent by the client, the cache node forwards the update request to the leader cache node in the distributed cache cluster;

[0030] The update information of the cached data is written to the log of the leader cache node. After the log synchronization is completed, the leader cache node broadcasts the cache update operation to all cache nodes through the Kafka message queue.

[0031] An embodiment of the present application also provides an optimization device for a distributed cache. The optimization device includes:

[0032] An access frequency prediction module, configured to input the historical request data of multiple hot data into a hot data prediction model, perform data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predict the access frequency of each hot data in a future time period; wherein, the hot data prediction model is obtained by iteratively training a long short-term memory network model;

[0033] A cache allocation module, configured to determine the data information of the target hot data whose access frequency is greater than the preset access frequency, and perform node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node;

[0034] A dynamic optimization module, configured to dynamically adjust the cache policy of each cache node based on a reinforcement learning model and multiple performance index data of each cache node.

[0035] An embodiment of the present application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the optimization method for the distributed cache as described above are executed.

[0036] An embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the optimization method for the distributed cache as described above are executed.

[0037] An optimization method, device, electronic device and storage medium for distributed caching provided by an embodiment of the present application input historical request data of multiple hot data into a hot data prediction model, perform data preprocessing, feature extraction and future access frequency prediction processing on the historical request data, and predict the access frequency of each hot data in a future time period; wherein, the hot data prediction model is obtained by iteratively training a long short-term memory network model; determine the data information of target hot data whose access frequency is greater than a preset access frequency, and perform node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node; dynamically adjust the caching policy of each cache node based on a reinforcement learning model and multiple performance index data of each cache node. Predict future hot data in advance and load hot data into the cache in advance to reduce cache miss rates. Introduce a reinforcement learning model to continuously adjust the caching policy according to real-time feedback to ensure optimal cache management under different loads and request patterns, achieve balanced allocation of resources, and avoid resource waste between nodes.

[0038] To make the above objects, features and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 A flowchart of an optimization method for distributed caching provided by an embodiment of the present application;

[0041] Figure 2 A schematic structural diagram of an optimization device for distributed caching provided by an embodiment of the present application;

[0042] Figure 3 A schematic structural diagram of an optimization device for distributed caching provided by an embodiment of the present application;

[0043] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without creative efforts falls within the scope of protection of the present application.

[0045] First, the applicable application scenarios of the present application are introduced. The present application can be applied to the field of distributed cache technology.

[0046] Based on this, the embodiments of the present application provide an optimization method for distributed cache, which predicts future hot data in advance and loads the hot data into the cache in advance to reduce the cache miss rate. A reinforcement learning model is introduced to continuously adjust the cache policy according to real-time feedback to ensure optimal cache management under different loads and request patterns, achieve balanced allocation of resources, and avoid resource waste between nodes.

[0047] Please refer to Figure 1 , Figure 1 which is a flowchart of an optimization method for distributed cache provided by the embodiments of the present application. As Figure 1 shown in

[0048] S101: Input the historical request data of multiple hot data into the hot data prediction model, perform data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predict the access frequency of each hot data in the future time period; wherein, the hot data prediction model is obtained by iteratively training a long short-term memory network model.

[0049] In this step, the historical request data of multiple hot data is input into the hot data prediction model, data preprocessing, feature extraction, and future access frequency prediction processing are performed on the historical request data, and the access frequency of each hot data in the future time period is predicted.

[0050] Among them, the historical request data includes: Timestamp: Records the time of each cache request. Data ID: Identifies the data being accessed. Request type: Such as read or write. Access frequency: The number of accesses within a specified time window.

[0051] In a possible implementation, for each of the hot data, inputting the historical request data of multiple hot data into the hot data prediction model, performing data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predicting the access frequency of each hot data in the future time period includes:

[0052] A: Extract time features, access frequency features, and data ID embedding vectors from the historical request data of the hot data based on the hot data prediction model, and construct time series according to a preset multiple time windows for the time features, access frequency features, and data ID embedding vectors to determine the time series features under each time window.

[0053] Here, convert the timestamp into features that are easy to understand, such as hours, minutes, whether it is a peak period (from 9 am to 6 pm on weekdays), whether it is a holiday, etc. Example: Convert the timestamp 2025-01-14 14:30:00 to [14, 30, 1, 0], where 14 is the hour, 30 is the minute, 1 is the weekday peak period, and 0 is a non-holiday. Calculate the number of accesses of each data ID within a specific time period (such as 1 hour, 1 day), and then normalize the frequency to [0, 1] to make all feature values within a unified range to avoid the influence of the magnitude difference between features on the analysis result. Example: Suppose a certain product has been accessed 100 times in the past 1 hour. If the set maximum access frequency is 200, then the normalized frequency = 100 / 200 = 0.5. Split the access frequency into time series features according to a fixed time window.

[0054] For example, construct time series feature inputs with a fixed time window (such as the past 1 hour, divided into 12 5-minute windows). Example: [Window 1: 50 times, Window 2: 60 times,..., Window 12: 100 times].

[0055] Here, the data ID embedding vector provides a vector representation for each data ID through an embedding technique, and inputs the data ID embedding vector together with the time features and access frequency features into the model to improve the prediction accuracy.

[0056] B: Perform data cleaning processing on the time series features under each time window to determine the target time series features under each time window.

[0057] Here, requests with significantly abnormal removal frequencies in each time window of the time series features (such as sudden spikes or abnormally low frequencies caused by crawlers or incorrect operations) are removed, and interpolation is used to fill in the data for missing time periods (if there is no data for certain time periods, interpolation is used to complete the missing values to ensure data continuity), and the target time series features in each time window are determined.

[0058] C: The first long short-term memory network layer of the hot data prediction model captures short-term features of the target time series features in each time window and outputs a short-term feature capture sequence for each time window.

[0059] Here, according to the first long short-term memory network layer of the hot data prediction model, short-term features of the target time series features in each time window are captured, and a short-term feature capture sequence for each time window is output.

[0060] D: The short-term feature capture sequence for each time window is input into the second long short-term memory network layer for long-term trend feature capture, and a long-term feature capture sequence for each time window is determined.

[0061] Here, the short-term feature capture sequence for each time window is input into the second long short-term memory network layer for long-term trend feature capture, and a long-term feature capture sequence for each time window is determined.

[0062] Among them, the two long short-term memory network layers respectively extract short-term and long-term dependencies in the time series.

[0063] E: The long-term feature capture sequence for each time window is input into the fully connected network layer of the hot data prediction model for processing, and the access frequency of the hot data in the future time period is output.

[0064] Here, after the processing of the long-term feature capture sequence for each time window by the fully connected layer, the access frequency of each hot data ID in the future time period is generated.

[0065] In a specific embodiment, based on LSTM (Long Short-Term Memory Network), it is used to predict the future access frequency of each data ID. The core goal is to capture the pattern of access frequency through historical access data (time series) so as to predict future access volume. Input features: including historical access frequency, time features (such as hour, week, holiday), and embedding vector of data ID. LSTM layer: Two layers of LSTM are adopted. The first layer captures short-term patterns, and the second layer captures long-term trends. Fully connected layer: Converts the features output by LSTM into predicted values of future access frequency. Output feature: After the processing of the fully connected layer, the future access frequency prediction of each data ID is generated.

[0066] In a possible implementation, the hot data prediction model is determined through the following steps:

[0067] (1): Input the historical request data of multiple sample hot data into the long short-term memory network model, process the historical request data of each sample hot data, and determine the predicted access frequency of each sample hot data in the future time period.

[0068] Here, input the historical request data of multiple sample hot data into the long short-term memory network model, process the historical request data of each sample hot data, and determine the predicted access frequency of each sample hot data in the future time period.

[0069] Among them, the determination process of the predicted access frequency of the sample hot data is consistent with the determination process of the access frequency of the above hot data in the future time period, and this part will not be elaborated here.

[0070] Here, when the amount of historical access data is small, the model may be underfitted (i.e., unable to capture effective patterns). Data augmentation expands the sample space by adding reasonable perturbations (noise) or simulating data. Adding random noise: Add small random variations to the original data to generate "augmented samples". Example: The original access frequencies are [100, 120, 150], add 5% random noise: [105, 126, 157] to enhance the model's robustness to data fluctuations. Simulating the access trend of data ID: Create virtual samples similar to the characteristics of hot data on the existing data. Example: Assume that the access frequency of hot data shows a trend of "decreasing after a peak" [200, 250, 300, 270, 240], and the simulated data ID follows a similar trend but with a slightly adjusted scale [150, 190, 220, 210, 180] to enable the model to adapt to more potential access patterns.

[0071] (2): Based on the actual access frequency and the predicted access frequency of each sample hot data in the future time period, determine the root mean square error value and the mean absolute error.

[0072] Among them, the root mean square error value RMSE and the mean absolute error MAE are determined through the following formula:

[0073]

[0074] Among them, is the predicted access frequency of the i-th sample hot data in the future time period, y i is the actual access frequency of the i-th sample hot data in the future time period, and N is the total number of samples.

[0075] (3): If either the root mean square error value or the mean absolute error is greater than a preset threshold, optimize the network parameters of the long short-term memory network model, and continue to train the optimized long short-term memory network model until both the root mean square error value and the mean absolute error are less than or equal to the preset threshold, and then stop training to determine the hot data prediction model.

[0076] Here, if either the root mean square error value or the mean absolute error is greater than a preset threshold, optimize the network parameters of the long short-term memory network model, and continue to train the optimized long short-term memory network model until both the root mean square error value and the mean absolute error are less than or equal to the preset threshold, and then stop training to determine the hot data prediction model.

[0077] S102: Determine the data information of the target hot data whose access frequency is greater than the preset access frequency, and perform node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node.

[0078] In this step, determine the data information of the target hot data whose access frequency is greater than the preset access frequency, and perform node caching on the data information of the target hot data according to the real-time load resource usage information of each cache node.

[0079] In a possible implementation manner, the performing node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node includes:

[0080] a: Detect whether the cache node with the lowest real-time load resource usage has enough cache space to accommodate the data information of the target hot data.

[0081] Here, detect whether the cache node with the lowest real-time load resource usage has enough cache space to accommodate the data information of the target hot data.

[0082] b: If so, load and cache the data information of the target hot data in the future time period into the cache node with the lowest real-time load resource usage.

[0083] Here, if so, load and cache the data information of the target hot data in the future time period into the cache node with the lowest real-time load resource usage.

[0084] c: If not, delete the low-frequency data in the cache node with the lowest real-time load resource usage to complete the expansion of the cache node, and cache the data information of the target hot data in the future time period into the expanded cache node.

[0085] Here, if not, delete the low-frequency data in the cache node with the lowest real-time load resource usage to complete the expansion of the cache node, and cache the data information of the target hot data in the future time period into the expanded cache node.

[0086] In a specific embodiment, the system, according to the prediction results of the LSTM model (such as [A101: 300 times, B202: 200 times]), pre-loads this data from the hard disk or distributed storage into the cache in advance. Example: Data IDA101 is predicted to be accessed 300 times and is pre-loaded into the cache node in advance. Data IDB202 is predicted to be accessed 200 times and is also loaded into the cache. When the cache space is insufficient, the LRU (Least Recently Used) policy is adopted to eliminate low-frequency data to make room for hot data. For example, the system checks whether there is enough space in the current cache to accommodate the predicted hot data. Example: The remaining cache space can only store 2MB of data, while the hot data A101 and B202 require a total of 3MB. According to the LRU policy, the data that has been accessed least recently is deleted from the cache. Example: Suppose the access frequencies of C303 and D404 are relatively low, where the access frequency of C303 is 10 times and the access frequency of D404 is 8 times. By judging, the data with the least access in the cache is preferentially eliminated. After deleting the low-frequency data, A101 and B202 are loaded into the cache to complete the pre-loading of hot data. Dynamically adjust the cache capacity and data distribution of each node to optimize the cache hit rate and resource utilization.

[0087] S103: Dynamically adjust the cache policy of each cache node based on the reinforcement learning model and multiple performance metric data of each cache node.

[0088] In this step, the cache policy of each cache node is dynamically adjusted according to the reinforcement learning model and multiple performance metric data of each cache node.

[0089] Among them, each cache node is embedded with a monitoring module, which collects the cache hit rate, average request latency, node load, and cache space utilization rate through the statistical commands provided by Redis. The node load includes CPU utilization, memory utilization, and disk I / O. The monitoring module encapsulates the collected data into the JSON format and transmits it to the central control system through Kafka. The transmitted data will be written into the time series database (InfluxDB) for subsequent query and analysis.

[0090] In this application, performance thresholds can also be configured: the cache hit rate is lower than 80%, the average latency is higher than 100 ms, and the node CPU utilization rate exceeds 90%. Once the set threshold is exceeded, the central control system will trigger an alarm mechanism and notify the developers via email or instant messaging tools (such as Slack). Grafana can also be used to build the following dashboards: Cache Hit Rate Trend Chart: showing the change of the cache hit rate of each node over time; Request Latency Distribution Chart: displaying the changes in the average latency and 95% response time; System Load Heat Map: showing the distribution of node loads in the cluster; Cache Space Utilization Analysis: helping to determine whether it is necessary to expand the cache capacity.

[0091] In one possible implementation, dynamically adjusting the cache policy of each cache node based on the reinforcement learning model and multiple performance metric data of each cache node includes:

[0092] I: Extract features from the multiple performance metric data of each cache node to determine the state feature vector of each cache node; wherein, the performance metric data includes cache hit rate, average request latency, node load, and cache space utilization rate.

[0093] Here, extract features from the multiple performance metric data of each cache node to determine the state feature vector of each cache node. For example, the state feature vector S t [H t 、L t 、U t 、O t , where H t is the cache hit rate, L t is the average request latency, U t is the node load, and O t is the cache space utilization rate.

[0094] Wherein, the performance metric data includes cache hit rate, average request latency, node load, and cache space utilization rate.

[0095] Here, set thresholds and an alarm mechanism, define load warning thresholds, for example: CPU usage > 80%, memory utilization > 70%. Network bandwidth usage > 85%. When the load exceeds the threshold, the system will trigger an alarm and mark the node as a high-load node.

[0096] II: Input the state feature vector of each cache node into the reinforcement learning model, and perform reinforcement optimization on the cache policy adjustment actions of the state feature vector of each cache node based on the greedy algorithm and a preset reward mechanism.

[0097] Here, the state feature vectors of each cache node are input into the reinforcement learning model, and the caching policy adjustment actions of the state feature vectors of each cache node are strengthened and optimized according to the greedy algorithm and the preset reward mechanism.

[0098] Among them, the specific operations of the caching policy include cache content adjustment and load balancing operations. Cache content adjustment includes: 1) Replacement policy: Select which cached data should be eliminated (such as based on the LRU policy or preferentially eliminating non-hot data for hot data); 2) Cache preloading: Dynamically load high-priority data according to predicted hot data. Load balancing operations include data migration, migrating hot data from high-load nodes to low-load nodes, and completing data migration by calling the interface through gRPC communication.

[0099] Here, the greedy algorithm is mainly used for the balance between exploration and exploitation. In reinforcement learning, the adjustment actions of the caching policy are selected according to the current state features. The core of the greedy algorithm is to select the action that is most likely to bring the optimal result according to the current state. For example, select the action with the largest current Q value with a probability of 1 - ∈ (that is, adopt the known optimal policy), and randomly select an action with a probability of ∈ (that is, explore new possible policies). This method can help the system optimize based on the existing policy and also explore new possible policies. The reward mechanism provides a feedback signal for reinforcement learning and is used to evaluate the effect of caching policy adjustment. By setting appropriate reward and punishment rules, the model can be guided to adjust towards the optimization goal (such as increasing the cache hit rate, reducing the request latency, etc.). For example: Cache hit rate reward: If the cache hit rate increases, a positive reward is given; if the cache hit rate decreases, there is no reward or a negative punishment is given. Latency penalty: If the latency increases, a negative punishment is given according to the increased latency time. Load balancing reward: If the load tends to be balanced, a positive reward is given; if a certain node has too high a load, a penalty is imposed.

[0100] Here, the reward mechanism design includes: cache hit rate reward R h is:

[0101] R h = ΔH t × 10

[0102] Example: The hit rate increases from 80% to 85%, and the reward is (85 - 80) × 10 = 50.

[0103] Among them, the average request latency penalty R l is

[0104] R l = -ΔL t × 5

[0105] Example: The latency increases by 10 ms, and the penalty is -10 × 5 = -50.

[0106] Here, the negative node load balancing reward r u gives a reward when the node load is close to equilibrium:

[0107] R u = -|U max - U min | × 2

[0108] where the formula for the total reward R is:

[0109] R = R h + R l + R u

[0110] III: Dynamically adjust the corresponding cache nodes based on the cache policy adjustment actions of each of the said cache nodes; wherein, the cache policy adjustment actions include cache content adjustment actions or load balancing adjustment actions.

[0111] Here, the corresponding cache nodes are dynamically adjusted according to the cache policy adjustment actions of each cache node.

[0112] Among them, cache content adjustment includes replacement policies: selecting which cached data should be eliminated (such as based on the LRU policy or preferentially eliminating non-hot data for hot data). Cache preloading: Dynamically load high-priority data according to predicted hot data. Load balancing adjustment includes load mitigation: migrating hot data from high-load nodes to low-load nodes, and the implementation method is to call the interface through gRPC communication to complete data migration. Action examples: replacing low-priority data, releasing cache space, or migrating hot data from node A to node B with lower load.

[0113] For example, the current situation: Node A: Hit rate = 55%, Load = 85%. Node B: Hit rate = 85%, Load = 40%. The adjustment plan is: Node B increases its cache capacity to store more hot data. Migrate the hot data with ID = 456 from node A to node B to increase the hit rate of node A. The result after adjustment: The hit rate of node A is increased to 70%, and the overall performance is significantly optimized.

[0114] In a specific embodiment, the specific steps for the reinforcement learning model to adopt DQN are as follows: Initialize the Q network: Use a two-layer fully connected neural network, input the state vector St, output the Q values of each action, and define and train the network parameters through PyTorch. Policy selection: The ε-greedy algorithm is used for the balance between exploration and exploitation: randomly select an action with probability ε, and select the action corresponding to the maximum Q value with probability 1 - ε. After each action is executed, adjust the policy according to the reward. Experience replay: Store the historical states, actions, and rewards in the replay pool, and draw samples in batches to train the network to improve the training efficiency.

[0115] In a possible implementation, for the load balancing adjustment action, the cache nodes corresponding thereto are dynamically adjusted based on the cache policy adjustment actions of each of the cache nodes; wherein, the cache policy adjustment action includes a cache content adjustment action or a load balancing adjustment action, including:

[0116] Based on the distributed cache cluster and the consistent hashing algorithm, the cached data of the hot data in the cache nodes with high load is migrated to the cache nodes with low load.

[0117] Here, RedisCluste and Consistent Hashing are used to locate the target node through consistent hashing, and the hot data is copied from the high-load node to the low-load node, ensuring that during the migration process, the least amount of data needs to be redistributed, reducing the performance impact.

[0118] In a possible implementation, after determining the data information of the target hot data whose access frequency is greater than the preset access frequency and caching the data information of the target hot data based on the real-time load resource usage information of each of the cache nodes, the optimization method further includes:

[0119] (1): After the cache node receives an update request for the cached data sent by the client, it forwards the update request to the leader cache node in the distributed cache cluster.

[0120] Here, when each cache node starts up, it registers its own information (IP, port, resource status, etc.) with the cluster management module. After the node completes registration, it joins the distributed cache cluster. The Raft consensus protocol is adopted, and a leader node is elected in the cluster to be responsible for log management, cache update, and global coordination. Other cache (Follower) nodes listen to the status changes of the Leader and perform synchronization operations. The client initiates an update request for the cached data (such as writing or deleting), and the update request reaches the Follower node, and the Follower forwards the request to the Leader node.

[0121] (2): The update information of the cached data is synchronously written to the log of the leader cache node. After the log synchronization is completed, the leader cache node broadcasts the cache update operation to all cache nodes through the Kafka message queue.

[0122] Here, the Leader node writes the update operation to the log (including key, value, timestamp, etc.) and broadcasts it to all Follower nodes. When the majority of nodes (more than half) confirm receiving the log entry, the Leader applies the operation to the cache and notifies the client that the operation is successful.

[0123] Among them, after the Raft log synchronization is completed, the Leader node broadcasts the cache update operation to all nodes through the Kafka message queue. Each node subscribes to a topic (such as Cache_Update) and listens for update events. Kafka persists the message log to ensure that data loss caused by network latency or node offline can be re-consumed after the node recovers.

[0124] In this application, conflict detection and version control can also be implemented: Conflict resolution based on timestamps: The update operation is attached with a timestamp, and the node selects the data with a newer timestamp during synchronization. Conflict detection based on version numbers: The version number is incremented for each update operation, and the node compares the version numbers to determine the latest data version. Conflict handling strategy: If the timestamps and version numbers are the same, the system defaults the data of the Leader to the final version. Example: Nodes A and B update the cache key = 789 at the same time. The timestamp of node A is 2025-01-14T15:05:00Z. The timestamp of node B is 2025-01-14T15:04:59Z. Then the system selects the update of node A and broadcasts and synchronizes it to all nodes.

[0125] In a specific embodiment, this application can also implement periodic consistency checks and snapshot synchronization. Periodic status verification: The system periodically calculates the cache hash values of all nodes and compares whether there are data differences. When a difference is detected, log replay or snapshot synchronization is triggered. Snapshot synchronization optimization: Raft periodically generates snapshot files to replace the long log chain and reduce the recovery time. The snapshot contains the complete state information of the cache, and each node can directly load the snapshot for recovery. Example: The cache hash value of node A is hash_A = abc123, and that of node B is hash_B = xyz789. The system detects that node B is missing logs. Node A sends the snapshot file to node B, and node B loads the snapshot to complete the consistency repair. Fault tolerance and exception recovery: During network partitioning, only the Leader within the partition is allowed to perform write operations, and other partitions suspend writing and wait for the network to recover to repair the data through log synchronization. Node failure recovery: After the failed node recovers, it uses the Kafka message queue to re-consume the update operations. If the node has been offline for a long time, it directly recovers the data through the Raft snapshot. Example: Node B recovers 10 minutes after a failure: Consume 50 lost update messages from Kafka. If there are too many missing logs, request the latest snapshot file from the Leader node A.

[0126] In this application, performance data such as cache hit rate, request latency, and node load are collected through a monitoring module, and the data is transmitted in real time to a central control system via Kafka, and historical data is stored in combination with a time series database (InfluxDB). A deep learning model (such as LSTM or Transformer) is used to predict the hot data access pattern, dynamically allocate cache resources, and give priority to processing hot data to improve system performance. By dynamically expanding and shrinking the node cache capacity, combined with the consistent hashing algorithm and RedisCluster, the balanced migration of data is realized to ensure that the cache pressure of high-load nodes is effectively dispersed. Performance thresholds such as cache hit rate and latency are configured. When the metrics are abnormal, an alarm is triggered and relevant personnel are informed through a notification mechanism for timely intervention. Grafana is used to build a visual dashboard for performance metrics, supporting multi-dimensional data display such as cache hit rate trends, latency distribution, and node load, and generating a system optimization report at the same time.

[0127] An optimization method for a distributed cache provided by an embodiment of this application, the optimization method includes: inputting historical request data of multiple hot data into a hot data prediction model, performing data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predicting the access frequency of each hot data in a future time period; wherein, the hot data prediction model is obtained by iteratively training a long short-term memory network model; determining the data information of target hot data with an access frequency greater than a preset access frequency, and performing node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node; dynamically adjusting the cache policy of each cache node based on a reinforcement learning model and multiple performance index data of each cache node. Predict future hot data in advance and load hot data into the cache in advance to reduce the cache miss rate. Introduce a reinforcement learning model to continuously adjust the cache policy according to real-time feedback to ensure optimal cache management under different loads and request patterns, achieve balanced allocation of resources, and avoid resource waste between nodes.

[0128] Please refer to Figure 2 、 Figure 3 , Figure 2 which is one of the structural schematic diagrams of an optimization device for a distributed cache provided by an embodiment of this application; Figure 3 which is the second structural schematic diagram of an optimization device for a distributed cache provided by an embodiment of this application. As Figure 2 shown in

[0129] An access frequency prediction module 210, which is used to input the historical request data of multiple hotspot data into a hotspot data prediction model, perform data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predict the access frequency of each hotspot data in a future time period; wherein, the hotspot data prediction model is obtained by iteratively training a long short-term memory network model;

[0130] A cache allocation module 220, which is used to determine the data information of target hotspot data whose access frequency is greater than a preset access frequency, and perform node caching on the data information of the target hotspot data based on the real-time load resource usage information of each cache node;

[0131] A dynamic optimization module 230, which is used to dynamically adjust the cache policy of each cache node based on a reinforcement learning model and multiple performance index data of each cache node.

[0132] Furthermore, when the access frequency prediction module 210 is used to input the historical request data of multiple hotspot data into the hotspot data prediction model for each hotspot data, perform data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predict the access frequency of each hotspot data in a future time period, the access frequency prediction module 210 specifically is used for:

[0133] Extract time features, access frequency features, and data ID embedding vectors from the historical request data of the hotspot data based on the hotspot data prediction model, and construct time series for the time features, access frequency features, and data ID embedding vectors according to a preset plurality of time windows to determine the time series features under each time window;

[0134] Perform data cleaning processing on the time series features under each time window to determine the target time series features under each time window;

[0135] Perform short-term feature capture on the target time series features under each time window based on the first long short-term memory network layer of the hotspot data prediction model, and output the short-term feature capture sequences of each time window;

[0136] Input the short-term feature capture sequences of each time window into a second long short-term memory network layer for long-term trend feature capture to determine the long-term feature capture sequences of each time window;

[0137] Input the long-term feature capture sequences of each time window into the fully connected network layer of the hotspot data prediction model for processing, and output the access frequency of the hotspot data in a future time period.

[0138] Further, as Figure 3 shown, the optimization device for distributed cache further includes a model training module 240, and the model training module 240 determines the hot data prediction model through the following steps:

[0139] Input the historical request data of multiple sample hot data into the long short-term memory network model, process the historical request data of each sample hot data, and determine the predicted access frequency of each sample hot data in the future time period;

[0140] Based on the actual access frequency and the predicted access frequency of each sample hot data in the future time period, determine the root mean square error value and the mean absolute error;

[0141] If either the root mean square error value or the mean absolute error is greater than a preset threshold, optimize the network parameters of the long short-term memory network model, continue to train the optimized long short-term memory network model until both the root mean square error value and the mean absolute error are less than or equal to the preset threshold, and then stop training to determine the hot data prediction model.

[0142] Further, when the cache allocation module 220 is used to perform node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node, the cache allocation module 220 specifically is used for:

[0143] Detect whether the cache node with the lowest real-time load resource usage has enough cache space to accommodate the data information of the target hot data;

[0144] If so, load and cache the data information of the target hot data in the future time period into the cache node with the lowest real-time load resource usage;

[0145] If not, delete the low-frequency data in the cache node with the lowest real-time load resource usage to complete the expansion of the cache node, and cache the data information of the target hot data in the future time period into the expanded cache node.

[0146] Further, when the dynamic optimization module 230 is used to dynamically adjust the cache policy of each cache node based on the reinforcement learning model and the multiple performance index data of each cache node, the dynamic optimization module 230 specifically is used for:

[0147] Extract features from the multiple performance index data of each cache node to determine the state feature vector of each cache node; wherein, the performance index data includes cache hit rate, average request latency, node load, and cache space utilization rate;

[0148] Input the state feature vector of each of the cache nodes into a reinforcement learning model, and perform reinforcement optimization on the cache policy adjustment actions of the state feature vectors of each of the cache nodes based on the greedy algorithm and a preset reward mechanism;

[0149] Dynamically adjust the corresponding cache nodes based on the cache policy adjustment actions of each of the cache nodes; wherein, the cache policy adjustment actions include cache content adjustment actions and load balancing.

[0150] Further, when the dynamic optimization module 230 is used for the load balancing adjustment action, and dynamically adjust the corresponding cache nodes based on the cache policy adjustment actions of each of the cache nodes; wherein, the cache policy adjustment actions include cache content adjustment actions or load balancing adjustment actions, the dynamic optimization module 230 is specifically configured to:

[0151] Migrate the cached data of the hot data in the high-load cache nodes to the low-load cache nodes based on the distributed cache cluster and the consistent hashing algorithm.

[0152] Further, as Figure 3 shown, the optimization device of the distributed cache further includes a node synchronization module 250, and the node synchronization module 250 is used for:

[0153] After the cache node receives an update request sent by the client for the cached data, forward the update request to the leader cache node in the distributed cache cluster;

[0154] Synchronously write the update information of the cached data to the log of the leader cache node, and after the log synchronization is completed, the leader cache node broadcasts the cache update operation to all cache nodes through the Kafka message queue.

[0155] An optimization device for distributed caching provided by an embodiment of the present application, the optimization device includes: an access frequency prediction module, configured to input historical request data of multiple hot data into a hot data prediction model, perform data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predict the access frequency of each hot data in a future time period; wherein, the hot data prediction model is obtained by iteratively training a long short-term memory network model; a cache allocation module, configured to determine data information of target hot data whose access frequency is greater than a preset access frequency, and perform node caching on the data information of the target hot data based on real-time load resource usage information of each cache node; a dynamic optimization module, configured to dynamically adjust the cache policy of each cache node based on a reinforcement learning model and multiple performance index data of each cache node. Predict future hot data in advance and load the hot data into the cache in advance to reduce the cache miss rate. Introduce a reinforcement learning model to continuously adjust the cache policy according to real-time feedback to ensure optimal cache management under different loads and request patterns, achieve balanced allocation of resources, and avoid resource waste between nodes.

[0156] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown in

[0157] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 runs, the processor 410 communicates with the memory 420 through the bus 430. When the machine-readable instructions are executed by the processor 410, they can execute the steps of the optimization method for distributed caching in the method embodiment as shown above Figure 1 . The specific implementation manner can refer to the method embodiment and will not be elaborated here.

[0158] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it can execute the steps of the optimization method for distributed caching in the method embodiment as shown above Figure 1 and Figure 2 . The specific implementation manner can refer to the method embodiment and will not be elaborated here.

[0159] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0160] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0161] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0163] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0164] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An optimization method for distributed caching, characterized in that, The optimization method includes the following: Input the historical request data of multiple hot data into the hot data prediction model, perform data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predict the access frequency of each hot data in the future time period; wherein, the hot data prediction model is obtained by iteratively training the long short-term memory network model; Determine the data information of the target hot data whose access frequency is greater than the preset access frequency, and perform node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node; Dynamically adjust the cache policy of each cache node based on the reinforcement learning model and the multiple performance index data of each cache node.

2. The optimization method according to claim 1, wherein For each hot data, the step of inputting the historical request data of multiple hot data into the hot data prediction model, performing data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predicting the access frequency of each hot data in the future time period includes: Extract the time feature, access frequency feature, and data ID embedding vector from the historical request data of the hot data based on the hot data prediction model, and construct a time series for the time feature, access frequency feature, and data ID embedding vector according to a preset multiple of time windows, and determine the time series feature under each time window; Perform data cleaning processing on the time series feature under each time window to determine the target time series feature under each time window; Perform short-term feature capture on the target time series feature under each time window based on the first long short-term memory network layer of the hot data prediction model, and output the short-term feature capture sequence of each time window; Input the short-term feature capture sequence of each time window into the second long short-term memory network layer for long-term trend feature capture, and determine the long-term feature capture sequence of each time window; Input the long-term feature capture sequence of each time window into the fully connected network layer of the hot data prediction model for processing, and output the access frequency of the hot data in the future time period.

3. The optimization method according to claim 1, wherein Determine the hot data prediction model through the following steps: Input the historical request data of multiple sample hot data into the long short-term memory network model, process the historical request data of each sample hot data, and determine the predicted access frequency of each sample hot data in the future time period; Based on the actual access frequency and predicted access frequency of each sample hot data in the future time period, determine the root mean square error value and the mean absolute error; If either the root mean square error value or the mean absolute error is greater than the preset threshold, optimize the network parameters of the long short-term memory network model, continue to train the optimized long short-term memory network model until both the root mean square error value and the mean absolute error are less than or equal to the preset threshold, and then determine the hot data prediction model.

4. The optimization method according to claim 1, wherein Node caching of the data information of the target hot data based on the real-time load resource usage information of each cache node includes: Detecting whether the cache node with the lowest real-time load resource usage has enough cache space to accommodate the data information of the target hot data; If so, loading and caching the data information of the target hot data in the future time period into the cache node with the lowest real-time load resource usage; If not, deleting the low-frequency data in the cache node with the lowest real-time load resource usage to complete the expansion of the cache node, and caching the data information of the target hot data in the future time period into the expanded cache node.

5. The optimization method according to claim 1, wherein Dynamically adjusting the cache policy of each cache node based on the reinforcement learning model and multiple performance index data of each cache node, including: Performing feature extraction on the multiple performance index data of each cache node to determine the state feature vector of each cache node; wherein, the performance index data includes cache hit rate, average request delay, node load, and cache space utilization rate; Inputting the state feature vector of each cache node into the reinforcement learning model, and performing reinforcement optimization on the cache policy adjustment action of the state feature vector of each cache node based on the greedy algorithm and the preset reward mechanism; Dynamically adjusting the corresponding cache node based on the cache policy adjustment action of each cache node; wherein, the cache policy adjustment action includes cache content adjustment and load balancing.

6. The optimization method according to claim 5, wherein Regarding the load balancing adjustment action, dynamically adjusting the corresponding cache node based on the cache policy adjustment action of each cache node; wherein, the cache policy adjustment action includes a cache content adjustment action or a load balancing adjustment action, including: Based on the distributed cache cluster and the consistent hashing algorithm, migrating the cached data of the hot data in the high-load cache node to the low-load cache node.

7. The optimization method according to claim 1, characterized in that After determining the data information of the target hot data whose access frequency is greater than the preset access frequency and performing node caching of the data information of the target hot data based on the real-time load resource usage information of each cache node, the optimization method further includes: After the cache node receives an update request for the cached data sent by the client, forwarding the update request to the leader cache node in the distributed cache cluster; Writing the update information of the cached data into the log of the leader cache node, and after the log synchronization is completed, the leader cache node broadcasts the cache update operation to all cache nodes through the Kafka message queue.

8. An optimization device for distributed caching, characterized in that, The optimization device includes: An access frequency prediction module, configured to input the historical request data of multiple hot data into the hot data prediction model, perform data preprocessing, feature extraction, and future access frequency prediction processing on the historical request data, and predict the access frequency of each hot data in the future time period; wherein, the hot data prediction model is obtained by iteratively training the long short-term memory network model. A cache allocation module, configured to determine data information of target hot data whose access frequency is greater than a preset access frequency, and perform node caching on the data information of the target hot data based on the real-time load resource usage information of each cache node; A dynamic optimization module, configured to dynamically adjust the cache policy of each cache node based on a reinforcement learning model and multiple performance index data of each cache node.

9. An electronic device, characterized in that, Comprising: A processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are run by the processor, the steps of the optimization method for distributed caching according to any one of claims 1 to 7 are executed.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, the steps of the optimization method for distributed caching according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Intelligent hotspot data prediction and caching method

    CN112637273A

  • Edge cache deployment strategy based on wireless edge network

    CN118555615A

  • Database caching strategy adjusting method, device and equipment

    CN118796896A

  • Database cache management method and system based on deep learning

    CN119441300A

Cited By

  • Edge cloud collaborative distribution method and device for model resources, electronic equipment and medium

    CN121531032A