Edge cache replacement method based on deep reinforcement learning
By adopting the edge cache replacement method based on deep reinforcement learning in mobile edge computing, the problems of limited storage space of edge nodes and low efficiency of cache replacement strategies are solved, and low-latency and high-efficiency content cache and replacement are achieved.
Patent Information
- Application Number
- PCT/CN2024/079954
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-03-04
- Publication Date
- 2025-05-08
AI Technical Summary
In mobile edge computing, edge node storage space is limited, and existing cache replacement policies are difficult to effectively identify and cache popular content that users are interested in, resulting in network link congestion and high request latency.
The edge cache replacement method based on deep reinforcement learning is adopted. By building a system network model, the cache status of the MEC server is obtained, the total delay of the user's content is calculated, the optimization objective function is established, the Markov decision-making process is constructed, and the A3C algorithm is used for dynamic content replacement.
Taking into account the dynamic changes in timing delay and popularity, the average delay of user-acquisition content is minimized, the cache hit rate is improved, and the average delay of obtaining content is reduced.
Smart Images

Figure CN2024079954_08052025_PF_FP_ABST
Abstract
Description
An edge cache replacement method based on deep reinforcement learning Technical Field
[0001] The present invention belongs to the field of mobile communication technology, and in particular relates to an edge cache replacement method based on deep reinforcement learning. Background Art
[0002] With the continuous development of wireless communication technology, users have increasingly higher requirements for content latency. Obtaining content from remote data centers may not meet users' low-latency needs. Mobile Edge Computing (MEC) technology has emerged as the times require. Mobile Edge Computing provides computing and caching services from the edge of the network, which can reduce network latency, improve throughput and user experience. Therefore, the study of edge caching has become one of the hot research topics in the field of wireless communications, and solving the problem of limited storage space at edge nodes is key. The caching solution needs to identify and cache popular content that most users are interested in. Based on the popularity of the content, popular content is cached on the edge cloud, which can reduce network link congestion and request latency, thereby improving user experience. Content caching technology is not only an important research topic in mobile edge computing, but also an important means to meet users' low-latency needs.
[0003] In MEC edge caching, cache placement strategies often execute over long intervals, necessitating the integration of cache replacement strategies. In mobile communication networks, user location changes can lead to dynamic changes in popularity, making the development of efficient cache replacement methods crucial. Designing edge cache replacement methods using deep reinforcement learning can predict content popularity and cache popular content to reduce network link congestion and request latency while protecting client data privacy. Therefore, in-depth research into scenarios where popularity changes dynamically is crucial for determining cache replacement strategies.
[0004] Summary of the Invention
[0005] To solve the above problems in the prior art, the present invention proposes an edge cache replacement method based on deep reinforcement learning, which includes:
[0006] S1. Build a system network model;
[0007] S2. Obtain the cache status of the MEC server in the system network model, and calculate the total delay for a single user to obtain the cached content based on the cache status of the MEC server;
[0008] S3. Calculate the total latency for all users to obtain content based on the total latency of cached content for a single user.
[0009] S4. Constructing an optimization objective function based on the total delay of users obtaining content;
[0010] S5. Construct Markov decision making based on the objective function;
[0011] S6. According to Markov decision making, a dynamic content replacement algorithm based on deep reinforcement learning is used to replace the cache of the MEC server.
[0012] Beneficial effects of the present invention:
[0013] Taking into account the dynamic changes in timing delay and popularity, the present invention uses a deep reinforcement learning training model to achieve the minimum average delay in user content acquisition by caching cached content and requested content. The present invention establishes a user delay model based on the state of the MEC content cache, and combines transmission delay and computation delay to obtain the total delay in user content acquisition. The problem is then modeled using the Multidimensional Decision Processing (MDP) to define a four-tuple. The system state refers to cached content and requested content, the action function refers to cache replacement, and the reward function refers to minimizing the average delay in content acquisition. Finally, the A3C algorithm is used to solve the problem, maximizing the system's cumulative reward. This can further improve the cache hit rate of content and reduce the average delay in content acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] FIG1 is a structural diagram of a system network model of the present invention;
[0015] FIG2 is a network structure diagram of the A3C algorithm of the present invention;
[0016] FIG3 is a comparative simulation diagram of the change of average delay of various algorithms under different amounts of content of the present invention;
[0017] FIG4 is a comparative simulation diagram of the changes in cache hit rate of various algorithms under different amounts of content according to the present invention;
[0018] FIG5 is an overall flow chart of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0020] A method for replacing edge cache based on deep reinforcement learning, as shown in FIG5 , includes:
[0021] S1. Build a system network model;
[0022] S2. Obtain the cache status of the MEC server in the system network model, and calculate the total delay for a single user to obtain the cached content based on the cache status of the MEC server;
[0023] S3. Calculate the total latency for all users to obtain content based on the total latency of cached content for a single user.
[0024] S4. Constructing an optimization objective function based on the total delay of users obtaining content;
[0025] S5. Construct Markov decision making based on the objective function;
[0026] S6. According to Markov decision making, a dynamic content replacement algorithm based on deep reinforcement learning is used to replace the cache of the MEC server.
[0027] As shown in Figure 1, the system network model includes: core network, base station, edge server, remote server and user. Constructing the system network model includes: according to the characteristics of MEC, obtaining the cache capacity of MEC as Cap MEC and the cache status of MEC content st f , define the set of all users as User = {1,2,...,Q}, the set of all content files as File = {1,2,...,F}, and the corresponding file size set as Size = {Z1,Z2,...,Z F The user sends a request to the base station, and the base station checks whether the edge server has cached the requested content. If so, the requested content is returned to the user; otherwise, the base station sends the content request to the core network.
[0028] Calculating the total latency for a single user to obtain cached content involves defining the transmission latency and computational latency for the user to obtain content based on the cache status of the MEC content. Then, the total latency for a single user to obtain cached content is calculated based on the transmission latency and computational latency.
[0029] (1) Calculate the transmission delay for user q to obtain content f:
[0030] When st f =1, the transmission delay of the content received by the user is:
[0031] Among them, R base,q Represents the transmission rate of the wireless link.
[0032] in, Represents the transmit power of the base station; represents the channel gain; represents the noise power spectral density.
[0033] When st f = 0, the transmission delay for user q to obtain content f is:
[0034] Among them, r f Represents the backhaul link rate of content f.
[0035] The transmission delay time for user q to obtain content f trans,q,f for:
[0036] (2) Calculate the computational delay for user q to obtain content f:
[0037] When st f =1, the computational delay for user q to obtain content f is:
[0038] Among them, Z f Represents the size of the content f; Represents the number of cycles executed per second by the CPU at the base station.
[0039] When st f = 0, the calculation delay for user q to obtain content f is:
[0040] in, Represents the number of cycles executed per second by the cloud CPU.
[0041] The computational delay timecomputing,q,f for user q to obtain content f is: timecomputing,q,f=st f ·time1_computing,q,f+(1-st f )·time2_computing,q,f
[0042] (3) Combining the transmission delay and computation delay, we can get the total delay time for user q to obtain content f q,f For: time q,f =time trans,q,f +timecomputing,q,f
[0043] Calculate the total delay for all users to obtain content. In a certain time interval τ, the content request set sent by user q is Request q The time delay for a single user to obtain a single content can be obtained as:
[0044] Construct an optimization objective function to minimize the average latency of obtaining content. The optimization problem can be constructed as follows:
[0045] Among them, Q, They represent the total number of users and the number of content requests per user within a time interval τ, respectively; the constraint ψ(1) means that the total volume of content cached by MEC does not exceed the storage capacity; the constraint ψ(2) refers to the cache status of the content. F represents the total number of users who obtain content f, and Z f Indicates the size of the content f obtained by the user, st f Indicates the cache status of the content, Cap MEC Indicates the cache capacity of MEC, and File indicates the collection of all content files.
[0046] Establishing Markov Decision Process: Model the dynamic content replacement problem as a Markov decision process. MDP is usually achieved by It is defined by factors such as:
[0047] Step 1: System state space.
[0048] In each time slot, the base station provides information about the cache placement status and content request status, which is used as the state space of the system. The expression is as follows:
[0049] in, Refers to the base station A collection of content that is always cached.
[0050] in, Refers to the base station The collection of content requests received at that moment.
[0051] Step 2: The action space of the system.
[0052] When the base station receives a user's request, it needs to determine whether to cache and which content to replace. Therefore, the system's action space is expressed as follows:
[0053] Among them, when When , it means that the current request content will not be cached. When , it represents the current content to be cached and replaces the first content.
[0054] Step 3: The state transition probability matrix of the system.
[0055] Based on the probability of the current state and action reaching the next state, the state transition probability matrix of the system is expressed as follows:
[0056] Step 4: The reward function of the system.
[0057] The immediate reward obtained after performing an action can be expressed as follows:
[0058] in, It is a positive integer greater than the average content acquisition delay.
[0059] Step 5: State-value function and state-action value function evaluation strategy.
[0060] The purpose of MDP is to find the best strategy for the decision maker to maximize the cumulative reward. The strategy is expressed as Refers to the status Take action Therefore, the base station is The process at all times is:
[0061] (1) At the starting point of the moment, the base station can observe the current status
[0062] (2) The base station performs actions based on strategy π
[0063] (3) The system is based on and Get cumulative rewards The status changes to
[0064] (4) The system sends rewards to the base station, and enters moment, and then repeat the process;
[0065] Accumulated Rewards The discount factor η can be expressed as The greater the value of η, the more emphasis is placed on cumulative rewards. Finally, by finding the optimal cache strategy π * , so that the expected cumulative reward is maximized, that is:
[0066] In MDP, policies are generally evaluated by state-value functions and state-action-value functions. The state-value function is defined as:
[0067] The dynamic content replacement algorithm based on deep reinforcement learning includes: Since the action space for establishing the Markov decision process is discrete, the cache replacement strategy is implemented through the A3C algorithm in the algorithm of the present invention. The network structure of the A3C algorithm is shown in Figure 2. The Actor-Critic method is a combination of the above two methods. Unlike the general Actor-Critic method, the A3C algorithm replaces the Q value with the advantage function in the calculation of the policy gradient. The formula of the advantage function is as follows:
[0068] Then, an unbiased estimate of the advantage function can be used to obtain the loss function of the Actor network, which is expressed as follows:
[0069] The gradient update of θ is:
[0070] Where λ represents the learning rate; Represents the advantage function
[0071] A3C puts the entropy term of the policy π into the Actor network with a coefficient of β. The gradient update of θ becomes:
[0072] in, Represents the entropy value of each time slot π.
[0073] The Critic network evaluates the behavior of the Actor network by adjusting the Q value. Since the advantage function in the Actor network is only the behavior value function Therefore, only fitting This chapter evaluates the Critic network through the mean square error loss function, namely:
[0074] The update method is as follows:
[0075] in, Refers to the learning rate of the Critic network.
[0076] This example is conducted on a real dataset, the Movielens 1M dataset, which records ratings of 3,900 movies by 6,040 users.
[0077] Parameter settings of A3C network. Among them, the Actor network learning rate λ and the Critic network learning rate The number of neurons in the hidden layer of the Actor network and the number of neurons in the hidden layer of the Critic network are 256 and 256 respectively. The discount factor η is 0.99 and the entropy factor β is 0.1.
[0078] Experimental simulation parameter settings. Among them, the noise power spectrum density -174dBm / Hz, MEC cache capacity Cap MEC 100MB, the base station bandwidth Band , CPU cycle frequency and transmit power They are set to 10MHz, 16GHz and 20W respectively, the CPU cycle frequency of the core network is set to 64GHz, and the channel gain Set to 30.6+36.7log 10 d MEC,u , the transmission rate between the core network and the base station r f Set to 10MB / s.
[0079] The experiment compared this method with several traditional prediction methods. The prediction methods used as controls include the LRU algorithm, the LFU algorithm, and the FIFO algorithm:
[0080] (1) LRU algorithm: When the MEC cache capacity is full, the content with the longest time interval from the last request will be replaced.
[0081] (2) LFU algorithm: When the MEC cache capacity is full, the content with the least user request frequency will be replaced.
[0082] (3) FIFO algorithm: When the MEC cache capacity is full, the content with the longest time will be replaced.
[0083] Figure 3 shows the changing trend of average latency for different amounts of content. It can be seen from Figure 3 that the average latency of the four algorithms increases as the number of files continues to increase. This is because the cache capacity of the edge server is limited. As the number of files increases, the probability that the file requested by the user is not in the MEC server increases, and it takes more time to obtain the content from the core network. From the overall trend, it can be seen that the algorithm in this chapter is better than the other three algorithms. When the number of files is the same, the average transmission delay of the algorithm in this chapter is the lowest, followed by the LRU algorithm, FIFO algorithm, and LFU algorithm. When the number of content is 100, the average transmission delay of the algorithm in this chapter is 0.03s, 0.04s, and 0.05s lower than the PSO algorithm, MPC algorithm, and RC algorithm, respectively. Therefore, the algorithm in this chapter performs better than the other three algorithms in terms of latency.
[0084] Figure 4 shows the change in cache hit rate for different content counts. As can be seen from Figure 4, the cache hit rate of all four algorithms decreases as the content count increases. This is because the edge server's cache capacity is limited. As the content count increases, the probability that the user's requested file is not available on the MEC server increases, requiring more time to retrieve the content from the core network. The overall trend shows that the algorithm in this chapter outperforms the other three algorithms. When the number of files is the same, the algorithm in this chapter has the highest cache hit rate, followed by the LRU algorithm, the FIFO algorithm, and the LFU algorithm. When the number of content is 100, the cache hit rate of the algorithm in this chapter is 6%, 20%, and 24% higher than the LRU algorithm, the FIFO algorithm, and the LFU algorithm, respectively. Therefore, the algorithm in this chapter achieves a better optimization effect on cache hit rate than the other three algorithms.
[0085] In summary, by adjusting various parameters, we demonstrate that our proposed dynamic content replacement algorithm based on deep reinforcement learning can achieve a near-global optimal solution and achieve the best performance in various scenarios. Compared to the other three algorithms, our proposed algorithm achieves superior optimization results in terms of average latency and cache hit rate. This allows the system to maximize rewards and find the most appropriate cache replacement strategy, effectively improving cache efficiency.
[0086] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An edge cache replacement method based on deep reinforcement learning, characterized in that: include: S1, build system network model; S2. Obtain the cache status of the MEC server in the system network model, and calculate the total delay for a single user to obtain the cache content based on the cache status of the MEC server; S3. Calculate the total delay for all users to obtain content based on the total delay of cached content of a single user; S4. Constructing an optimization objective function based on the total delay of users obtaining content; S5. Construct Markov decision according to the objective function; S6. According to Markov decision making, a dynamic content replacement algorithm based on deep reinforcement learning is used to replace the cache of the MEC server.
2. The edge cache replacement method based on deep reinforcement learning according to claim 1, characterized in that: The system network model includes: core network, base station, edge server, remote server and user; the user sends a request to the base station, and the base station checks whether the edge server caches the requested content. If so, the requested content is returned to the user; otherwise, the base station sends a content request to the core network.
3. The edge cache replacement method based on deep reinforcement learning according to claim 1, characterized in that: Calculating the total delay for a single user to obtain cached content includes: calculating the transmission delay for the user to obtain the content, and calculating the calculation delay for the user to obtain the content; and obtaining the total delay for the user to obtain the content according to the transmission delay for the user to obtain the content and the calculation delay for the user to obtain the content.
4. The edge cache replacement method based on deep reinforcement learning according to claim 1, characterized in that: The total delay for all users to obtain content is calculated as: Among them, q represents the user, Q represents the total number of users, f represents the content to be obtained, and time q,f represents the total delay for user q to obtain content f.
5. The edge cache replacement method based on deep reinforcement learning according to claim 1, characterized in that: The optimization objective function is: Where Q represents the total number of users, represents the number of content requests by a single user within a time interval τ; the constraint ψ(1) means that the total volume of content cached by MEC does not exceed the storage capacity; the constraint ψ(2) refers to the cache status of the content; F represents the total number of users obtaining content f, Z f Indicates the size of the content f obtained by the user, st f Indicates the cache status of the content, Cap MEC Represents the cache capacity of MEC, and File represents the collection of all content files.
6. The edge cache replacement method based on deep reinforcement learning according to claim 1, characterized in that: Constructing Markov decision includes: defining the system state space, the system action space and the system reward function; constructing the state-action value function evaluation strategy according to the system state space, the system action space and the system reward function.
7. The edge cache replacement method based on deep reinforcement learning according to claim 6, characterized in that: Defining the system state space, system action space, and system reward function includes: Define the system state space: In each time slot, the base station provides information about the cache placement state and content request state, and uses the cache placement state and content request state as the state space, expressed as: in, For base stations The collection of content placed in the cache at all times, For base stations The set of content requests received at any moment, For base stations The state space at the moment; Define the system action space: Among them, a 0 Express action; when When , it means that the current request content is not cached; when When , it represents the current content to be cached and replaces the first Contents; The total number of contents that are cached for the current content and replaced in the cache space; Set up the system reward function: in, is a positive integer greater than the average content acquisition delay. For the current state, is the current action, time all is the total delay for all users to obtain content, Q represents the total number of users, Represents the number of content requests by a single user within a time interval τ.
8. The edge cache replacement method based on deep reinforcement learning according to claim 6, characterized in that: The strategy for constructing a state-action value function evaluation includes: Step 1: The time is the starting point, and the base station observes the current state Step 2: The base station performs actions based on strategy π Step 3: System based and Get cumulative rewards Status updated to Step 4: The system sends rewards to the base station and enters time, and repeat steps 1 to 3; Step 5: Count the system's cumulative rewards and obtain the best cache strategy π based on the cumulative rewards * ; Step 6: Evaluate the policy through the state-value function and state-action value function.
9. The edge cache replacement method based on deep reinforcement learning according to claim 1, characterized in that: The dynamic content replacement algorithm based on deep reinforcement learning is used to replace the cache of the MEC server. The dynamic content replacement algorithm based on deep reinforcement learning consists of an environment, multiple agents, and a global network, where the global network includes the system state, the Actor network, and the Critic network. The agent and the global network have the same network structure. Specifically, it includes: S61. Define advantage function; in, is the state-action value function, is the state value function; S62. Calculate the policy gradient based on the advantage function Among them, θ represents the neural network parameters, According to the status Take Action The probability of is the action space, and E is the overall mean; S63. Perform an unbiased estimate of the advantage function based on the policy gradient to obtain the loss function of the Actor network. S64. Update the θ gradient according to the loss function of the Actor network. The update formula is: Where λ represents the learning rate; represents the advantage function; S65. Optimize the gradient parameter θ according to the strategy π. The expression of the optimized parameter θ is: in, represents the entropy value of each time slot π, and β is the coefficient; S64, the Critic network evaluates the behavior of the Actor network by adjusting the Q value; the mean square error loss function is calculated according to the optimized parameter θ, and the Critic network is evaluated by the mean square error loss function, which is: in, To obtain cumulative rewards, η is the discount factor, for The state value function at the moment; S65. According to the mean square error loss function, The gradient is updated in the following way: in, Refers to the learning rate of the Critic network; S66. Replace the cache of the MEC server according to the optimized dynamic content replacement algorithm.
Citation Information
Patent Citations
Panoramic video edge collaborative cache replacement method based on deep reinforcement learning
CN113282786A
Dynamic grouping Internet of Vehicles caching method based on reinforcement learning in MEC environment
CN114374741A
Intelligent data caching method based on cooperative sensing
CN114786200A
Resource collaborative scheduling method and system of edge cloud Internet of Vehicles
CN114885303A
Dynamic latency-responsive cache management
US20220224776A1
Cited By
Cloud side information collaboration and content delivery method for air-based high-mobility network
CN120416918A
Entropy driving step length self-adaption-based diffusion reinforcement learning channel access method
CN121463143A
An entropy-driven step length self-adaptive diffusion reinforcement learning channel access method
CN121463143B