Edge Caching Policy Learning Method in Communication Scenarios

The DQN algorithm optimizes edge caching in large-scale wireless communication systems by learning caching strategies, reducing error rates and maximizing spectral efficiency through a deep neural network training approach.

CN115567402BActive Publication Date: 2025-07-15GUANGZHOU UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211194168.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-07-15
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

In large-scale communication systems, the complexity of traditional optimization algorithms has increased exponentially, making it difficult to calculate the optimal cache strategy, and the existing technology cannot effectively solve the problem of wireless edge cache.

Method used

Deep reinforcement learning (DQN) algorithm is adopted to establish wireless communication system models and cache models, use Monte Carlo training samples to optimize cache strategies, learn the optimal cache decision network, and optimize edge cache strategies based on channel information and user needs.

Benefits of technology

In the case of unknown file popularity, the average spectrum efficiency of signal transmission is improved and the average bit error rate is reduced, and the communication performance is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115567402B_ABST
    Figure CN115567402B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of wireless communication and mobile communication, and discloses an edge caching policy learning method in a communication scenario, including the following steps: The first step: Establish a wireless communication system model; The second step: Establish a caching model; The third step: Obtain feedback according to the environment and optimize the policy to obtain an algorithm optimized for the target, including state, action, and feedback; The fourth step: The state, action, and value function are obtained and updated through Monte Carlo training samples, and finally an optimal caching decision network is learned. This edge caching policy learning method in the communication scenario explores how to use the DQN algorithm to make files with high file popularity pass through channels with high channel spectral efficiency and low channel error rate respectively under the condition that the file popularity is unknown. Thereby minimizing the average bit error rate of signal transmission and maximizing the average spectral efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of wireless communication and mobile communication, and specifically to a method for learning an edge caching strategy in a communication scenario. Background Art

[0002] In modern social life, there is an increasing demand for data, including short video, intelligent transportation, intelligent medical application data, and so on. Users have higher and higher requirements in terms of both data transmission rate and communication quality. For a communication network, the transmission link between a remote service center and a proximal base station, or between base stations, is usually referred to as a backhaul. When users frequently obtain data through a remote service center, it will bring a huge backhaul load, while edge caching technology can reduce the backhaul load and improve communication performance. The specific process is to deploy a storage device near the user. When the user initiates a file request, first look for the required file in the storage device of the proximal base station. If it is stored, the file is directly transmitted to the user without having to obtain data from the remote end. However, the capacity of the proximal storage device is limited, so we need to design a reasonable caching strategy to optimize the overall performance.

[0003] Some traditional algorithms have been proposed to solve the problem of wireless edge caching. However, when the system scale is large, the complexity of traditional optimization algorithms grows exponentially, which makes it difficult to obtain calculation results through traditional algorithms. The problem of not being able to obtain results due to high complexity in the case of a large system scale is well solved by the way of neural network learning. For this reason, we propose a method for learning an edge caching strategy in a communication scenario. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for learning an edge caching strategy in a communication scenario, and solves the above problems.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present invention provides the following technical solution: A method for learning an edge caching strategy in a communication scenario, including the following steps:

[0008] The first step: Establish a wireless communication system model;

[0009] The second step: Establish a caching model;

[0010] The third step: Obtain feedback according to the environment and optimize the strategy to obtain an algorithm for target optimization, including state, action, and feedback;

[0011] The fourth step: The state, action, and value function are obtained and updated through Monte Carlo training samples, and finally an optimal caching decision network is learned.

[0012] Preferably, the wireless communication system model in the first step includes the following:

[0013] The cache-aided microcellular network transmission model contains M base stations and N users, denoted as M = {1, 2,..., M} and N = {1, 2,..., N} respectively. The distance from the m-th base station to the n-th user is denoted as d m,n , and the signal transmission power from the edge base station to the user terminal is denoted as p t , and the signal reception power at the user is p r , and the received signal:

[0014]

[0015] where Re(z) represents the real part of the complex number z, u(t) is the equivalent low-pass signal of the transmitted signal, f c is the carrier frequency, is the product of the transmitting and receiving antennas in the line-of-sight radiation direction, e -j2πd / λ is the phase shift generated by the signal propagation distance, n r ~CN(0, p n ) is additive white Gaussian noise with a mean of 0 and a variance of p n . For multi-node communication, the spectral efficiency between the m-th base station and the n-th user is given by the following formula:

[0016] r m,n = log2(1 + γ 2 m,n );

[0017] where γ 2 m,n represents the signal-to-noise ratio of the signal transmission from the proximal edge base station m to the user n, and the bit error rate is:

[0018]

[0019] where b1 and b2 are two constants determined by the specific signal form.

[0020] Preferably, the cache model in the second step is:

[0021] Assume that each user requests a different content set, and the content set requested by the n-th user is denoted as F n , and assume that the storage capacity of the edge base station is C;

[0022] Assume that the content popularity in the t time slot follows the Zipf distribution, where the popularity of the c-th content requested by the n-th user is denoted as c u , written as:

[0023]

[0024] where λ n > 0 is the Zipf factor, representing the popularity of the content in the content set for user n;

[0025] Define a set C m , which represents the cached content at the m-th base station. If c n ∈C m , it means that the c-th content requested by the n-th user is cached at the m-th base station. The channel condition is represented by the signal-to-noise ratio of the receiving edge base station, where a higher signal-to-noise ratio may indicate less fading, and the cache optimization problem is obtained:

[0026]

[0027]

[0028] where I m,n,c represents the cache decision variable. When I m,n,c = 1, it means that c n ∈C m , and γ 2 0 represents the transmission signal-to-noise ratio of the uncached file.

[0029] Preferably, the

[0030] 1) State:

[0031] The state mainly describes the channel information and represents the channel information with the signal-to-noise ratio, that is, γ 2 m,n , m ∈ M, n ∈ N. The system state space is represented by the set S ∈ R M×N as:

[0032]

[0033] 2) Action:

[0034] The action space of DQN at time slot t is represented as:

[0035]

[0036] 3) Reward:

[0037] The spectrum efficiency and bit error rate obtained between the m-th base station and the n-th user are respectively used to characterize the reward corresponding to the executed action. Specifically, the total reward obtained by executing action A (t) at time slot t is defined as:

[0038]

[0039] Preferably, in the fourth step, Monte Carlo learning uses a policy for sampling to obtain a sampling trajectory:

[0040] <S (0) ,A (0) ,R (1) ,...,S (T-1) ,A (T-1) ,R (T) >.

[0041] Preferably, the sampling trajectory includes a vector composed of the user's requests for files, and the other is a vector γ composed of the signal-to-noise ratio of the channel 2 i,j , i ∈ M, j ∈ N. The product of the two parts in the sampling trajectory represents the reward, and the larger its value, the better the performance. For the state-action pair (S, A), the value function of its optimal policy at the current time slot t can be written as:

[0042]

[0043] where represents the reward for executing the optimal policy at the current time slot t, and the resulting optimal caching policy is:

[0044] π * (i) = argmaxR π(A) (i);

[0045] And further use to update the value function for the next time slot:

[0046]

[0047] A coefficient α t can be used to replace it. It is usually a very small positive value. The update of the state value function can be obtained by incremental summation:

[0048]

[0049] where S is the current state and A is the action;

[0050] Using the DQN learning method, it can be used to represent the characteristics of the interactions between various states and actions:

[0051]

[0052] Q π t (S′, A′) ≈ Q π t (S′, A′; θ′);

[0053] Then, according to the update of the state value function, the value of the target network is set to:

[0054] y′ = R A S→S′ + γQ π t (S′, A′; θ′);

[0055] To ensure obtaining the optimal network output at time slot t+1, the deep neural network should be trained with the target value function as the supervised label. Therefore, the loss function of the network can be defined as:

[0056] Loss(θ) = E[y′ - Q π t (S′, A′; θ′)] 2 ;

[0057] The goal of DQN is to minimize the difference between the two. Then, gradient descent is performed to update the network parameters:

[0058]

[0059] Finally, an optimal caching decision network is learned.

[0060] (III) Beneficial Effects

[0061] Compared with the prior art, the present invention provides an edge caching policy learning method in a communication scenario, having the following beneficial effects:

[0062] 1. For the edge caching policy learning method in this communication scenario, considering the case where the file popularity is unknown, it discusses how to use the DQN algorithm to make files with high file popularity pass through channels with high channel spectral efficiency and low channel error rate. Thereby minimizing the average bit error rate of signal transmission and maximizing the average spectral efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic diagram of a wireless communication system model;

[0064] Figure 2 It is a schematic diagram of the average bit error rate based on the DQN algorithm;

[0065] Figure 3 It is a schematic diagram of the average spectral efficiency obtained based on the DQN algorithm;

[0066] Figure 4 It is a schematic diagram of the average bit error rate;

[0067] Figure 5 It is a schematic diagram of the average spectral efficiency. DETAILED IMPLEMENTATION MANNER

[0068] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] Please refer to Figures 1-5 ,

[0070] Wireless communication system model:

[0071] We assume that the cache-assisted microcellular network transmission model near the user contains M base stations and N users, which are respectively represented as M = {1, 2,..., M} and N = {1, 2,..., N}. The distance from the m-th base station to the n-th user is denoted as d m,n , and the signal transmission power from the edge base station to the user end is denoted as p t , and the signal reception power at the user is p r . We use the free space path loss model for analysis, thus generating the received signal:

[0072]

[0073] where Re(z) represents the real part of the complex number z, u(t) is the equivalent low-pass signal of the transmitted signal, f c is the carrier frequency, is the product of the transmitting and receiving antennas in the line-of-sight radiation direction. e -j2πd / λ is the phase shift generated by the signal propagation distance, n r ~CN(0, p n ) represents the additive white Gaussian noise with a mean of 0 and a variance of p n at the receiving end. In addition, for multi-node communication, the spectral efficiency between the m-th base station and the n-th user is given by the following formula:

[0074] r m,n = log2(1 + γ 2 m,n ), (2)

[0075] where γ 2 m,n represents the signal-to-noise ratio of the signal transmission from the proximal edge base station m to the user n. The bit error rate can be approximated by the Q function (reference: F. Zhou, M. Du, Y. Wang, and G. Luo), and is expressed as:

[0076]

[0077] where b1 and b2 are two constants determined by the specific signal form.

[0078] Caching model:

[0079] Assume that each user requests a different content set, and the content set requested by the nth user is denoted as F n . Since in actual engineering practice, the number of files cached by the edge base station is much smaller than the number of user requests. Assume that the storage capacity of the edge base station is C, that is, MC << ∑ n∈N |F n |, where << means much less than, and |·| represents the cardinality of the content set involved. Since only part of the content can be cached in the SBS, without loss of generality, assume that the content popularity in the t time slot follows the Zipf distribution. The popularity (denoted by c u ) of the cth content requested by the nth user can be written as:

[0080]

[0081] where λ n > 0 is the Zipf factor, indicating the popularity of the content in the content set for user n. However, in this work, the Zipf factor is unknown for each user n. Consider such a situation that for a user n, the user's demands for a specified number of files are different. The different demands of a single user for multiple files can be characterized by the above-mentioned Zipf distribution. If only to meet the communication needs of the user, the files with greater corresponding demands of the user can be cached in the edge base station. However, what we should consider more is the overall performance. Assume that another user has a higher demand for another file, then it should be cached more, which will have a greater improvement in the overall performance. And what we do is exactly to consider the demands of several users for several files as a whole. We always have to find the ones with greater demands in the whole for caching, so as to make the overall performance optimal. The above analysis is considered only from the degree of users' different demands for files without considering the channel. In the actual environment, we also have to consider different channel environments. Imagine if the files with high demands can be transmitted through channels with good channel conditions every time for the user, this will further reduce the communication loss. Specifically, we define a set C m , which represents the cached content at the mth base station. If c n ∈C m, it means that the c-th content requested by the n-th user is cached at the m-th base station. On the other hand, the channel condition is represented by the signal-to-noise ratio of the receiving-edge base station, where a higher signal-to-noise ratio may indicate less fading, and vice versa. Therefore, we can obtain the cache optimization problem:

[0082]

[0083]

[0084] where I m,n,c represents the cache decision variable. When I m,n,c = 1, it means that c n ∈C m . γ 2 0 represents the transmission signal-to-noise ratio of the uncached file.

[0085] The optimization algorithm used in this paper:

[0086] In this paper, the DQN algorithm in reinforcement learning is used to optimize the cache decision. Reinforcement learning is an algorithm that obtains feedback based on the environment and optimizes the strategy to achieve the target optimization. It includes state, action, and feedback reward.

[0087] Combined with the objective we want to optimize, the descriptions of state, action, and feedback are as follows:

[0088] 1) State:

[0089] The state mainly describes the information of the system environment. Here, the environment is the channel information, and the signal-to-noise ratio is used to represent the channel information, that is, γ 2 m,n , m ∈ M, n ∈ N. Therefore, the system state space can be represented by the set S ∈ R M×N as:

[0090]

[0091] This set contains multiple states. We make decisions by continuously observing the rewards obtained from transitioning from one state to another. The state transition probability here is unknown. We can directly learn the cache policy using machine learning methods without knowing the state transition probability.

[0092] 2) Action:

[0093] To improve the efficiency of cache decision-making and avoid cache conflicts, we have formulated the following two constraints. First, each SBS has a fixed capacity and can only cache a limited number of contents. Second, for each content, only one SBS can cache the content requested by the user. In the method we proposed, the action introduced in formula (5) can indicate whether a specific content is cached on a certain SBS. Therefore, the action space of the proposed DQN at time slot t can be expressed as:

[0094]

[0095] 3) Reward:

[0096] Appropriately setting the reward can enable the DQN to better find actions to approach the target. Here, the reward information can be obtained according to formula (5). In the DQN algorithm we proposed, when considering the design performance, the spectral efficiency and bit error rate obtained from formulas (2) and (3) are respectively used to characterize the reward corresponding to the executed action. Specifically, the total reward obtained by executing action A at time slot t (t) is defined as:

[0097]

[0098] where g is the function corresponding to (2) or (3) when discussing the corresponding performance. g(γ 2 m,n ) is essentially related to the SNR at the signal reception and is mainly determined by it.

[0099] Introduction and optimization process description of the algorithm

[0100] DQN is an improved Q-learning algorithm. For the problem we studied, since the transition probability of the state space is unknown, we adopt a model-free reinforcement learning algorithm. Among them, the state-action-value function is obtained and updated through Monte Carlo training samples. Specifically, Monte Carlo learning is to use a certain policy for sampling to obtain sampling trajectories:

[0101] <S (0) ,A (0) ,R (1) ,...,S (T-1) ,A (T-1) ,R (T) > (9)

[0102] After multi-step sampling, the average value of the cumulative rewards of the state-action pairs of multiple sampling values is taken to obtain the estimated state-action value function. Therefore, it is necessary to design a sampling trajectory to obtain the optimal or sub-optimal policy for estimating the state-action value function that may achieve the maximum expected reward.

[0103] Since a well-designed sampling trajectory can significantly improve the algorithm efficiency and thus reduce the convergence time. The sampling trajectory in our work consists of two parts. One is a vector composed of the request probabilities of files by users (assuming it follows the Zipf distribution, but the Zipf parameter is unknown), and the other is a vector γ composed of the signal-to-noise ratios of channels 2 i,j , i ∈ M, j ∈ N. The product of the two parts in the sampling trajectory represents the reward, and the larger its value, the better the performance. For the state-action pair (S, A), the value function of the optimal policy at the current time slot t can be written as:

[0104]

[0105] where represents the reward for executing the optimal policy at the current time slot t, and the resulting optimal caching policy is

[0106] π * (i) = argmaxR π(A) (i) (11)

[0107] And further utilize to update the value function of the next time slot

[0108]

[0109] More generally, can be replaced by a coefficient α t , which is usually a very small positive value and can be considered as the learning rate.

[0110] After the above analysis, the update of the state value function can be obtained through incremental summation:

[0111]

[0112] where S is the current state, A is the action, and the value function is updated every time a step is executed. The above is the value function estimation in the finite state space. However, the state space of the problem we consider is continuous and has infinitely many states. Therefore, we adopt a deep neural network to approximate the value function, that is, adopt the DQN learning method, which can be used to represent the characteristics of the interaction between various states and actions. That is:

[0113]

[0114] Q π t (S′, A′) ≈ Q π t (S′, A′; θ′) (15)

[0115] Then, according to Equation (13), let the value of the target network be:

[0116] y′ = R A S→S′ + γQ π t (S′, A′; θ′) (16)

[0117] To ensure obtaining the optimal network output (i.e., policy π) at time slot t + 1, the deep neural network should be trained with the target value function as the supervision label. Therefore, the loss function of the network can be defined as:

[0118] Loss(θ) = E[y′ - Q π t (S′, A′; θ′)] 2 (17)

[0119] The goal of DQN is to minimize the difference between the two. Then, gradient descent is performed to update the network parameters:

[0120]

[0121] Finally, an optimal caching decision network is learned.

[0122] Experimental results and data

[0123] Training and test samples are selected from a sample pool consisting of randomly generated 2×10 4 large-scale channel signal-to-noise ratios. To evaluate the performance of the proposed algorithm, we compare the DQN algorithm with the random algorithm. Specifically, we compare the average spectral efficiency and the average bit error rate obtained by the two algorithms within a certain training interval.

[0124] As Figure 2 shown, the average bit error rate based on the DQN algorithm gradually decreases as the number of epochs increases, and its value tends to be stable after about 10 4 epochs of training, where its value is approximately 2.2×10 -3 . However, the average bit error rate obtained by the random algorithm remains stable, and its value is approximately 4×10 -3 . The average bit error rate performance obtained by the DQN algorithm within the same training interval is generally significantly better than that of the random algorithm.

[0125] As Figure 3As shown, the average spectral efficiency obtained by the proposed DQN algorithm increases with the increase of epochs. After about 10,000 epochs of training, the value of the average spectral efficiency tends to be stable. The value of the average spectral efficiency obtained by the random algorithm is about 2.6 bits per second per hertz (bps / Hz), which is less than that of the proposed DQN-based algorithm, and its value is about 3.4 bits per second per hertz (bps / Hz).

[0126] In addition, we change the transmission power value p of the signal t , and compare the performance of the two algorithms under different transmission power values. Since both the average bit error rate and the average spectral efficiency obtained by the DQN algorithm tend to be stable after about 10,000 epochs of training, we choose the results after training 10,000 epochs to observe the performance of the DQN algorithm. It can be seen in Figure 4 that when we increase the transmission power of the signal, the average bit error rate becomes smaller, and the proposed DQN-based algorithm always gives better results than the random algorithm. At a larger signal transmission power, the advantages of the DQN algorithm are more clearly demonstrated. Figure 5 It shows that the average spectral efficiency increases with the increase of the signal transmission power. The average spectral efficiency obtained by the DQN algorithm is significantly greater than that of the random algorithm, and the greater the signal transmission power, the more obvious this difference is.

[0127] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. Edge caching policy learning method in a communication scenario, characterized in that It includes the following steps: The first step: Establish a wireless communication system model; The second step: Establish a caching model; The third step: Obtain feedback according to the environment and optimize the strategy to obtain an algorithm for target optimization, including state, action, and feedback; The fourth step: The state, action, and value function are obtained and updated through Monte Carlo training samples, and finally an optimal caching decision network is learned; The caching model in the second step is: Assume that each user requests a different content set, and the content set requested by the nth user is denoted as F n , and assume that the storage capacity of the edge base station is C; Assume that the content popularity in time slot t follows the Zipf distribution, where the popularity of the c-th content requested by the n-th user is denoted by c and written as: n denoted by where λ n > 0 is the Zipf factor, representing the popularity of the content in the content set for user n; Define a set C m , which represents the cached content at the m-th base station. If c n ∈ C m , it means that the c-th content requested by the n-th user is cached at the m-th base station. The channel condition is represented by the signal-to-noise ratio of the receiving edge base station, where a higher signal-to-noise ratio may indicate less fading, and the cache optimization problem is obtained: Among which I m,n,c represents the caching decision variable. When I m,n,c = 1, it means that c n ∈C m , and γ 2 0 represents the signal-to-noise ratio of the transmitted uncached file at the receiver.

2. The method for learning edge caching policy in the communication scenario according to claim 1, wherein: The wireless communication system model in the first step includes the following: The cache-aided microcellular network transmission model consists of M base stations and N users, denoted as M = {1, 2, ..., M} and N = {1, 2, ..., N} respectively. The distance from the m-th base station to the n-th user is denoted as d m,n , and the signal transmission power from the edge base station to the user terminal is denoted as p t , and the signal reception power at the user is p r , received signal: where Re(z) represents the real part of the complex number z, u(t) is the equivalent low-pass signal of the transmitted signal, and f c is the carrier frequency, is the product of the transmitting and receiving antennas in the line-of-sight radiation direction, and e -j2πd / λ is the phase shift caused by the signal propagation distance, and n r ~CN(0, p n ) represents the additive white Gaussian noise with a mean of 0 and a variance of p n at the receiving end. For multi-node communication, the spectral efficiency between the m-th base station and the n-th user is given by the following equation: r m,n = log2(1 + γ 2 m,n ); wherein γ 2 m,n represents the signal-to-noise ratio at the receiving end for the signal transmission from the m-th base station to the n-th user, and the bit error rate is: Where b1 and b2 are two constants determined by the specific signal form.

3. The method for learning an edge caching policy in the communication scenario according to claim 1, wherein Define the key parameters of reinforcement learning as follows: 1) State: The state mainly describes the channel information, and the signal-to-noise ratio is used to characterize the channel information, i.e., γ 2 m,n , m ∈ M, n ∈ N, and the set S ∈ R M×N is used to represent the system state space as: 2) Action: The action space of DQN at time slot t is expressed as: 3) Reward: The spectrum efficiency and bit error rate obtained between the m-th base station and the n-th user are respectively used to characterize the reward corresponding to the executed action; Specifically, action A is executed at time slot t (t) The total reward obtained is defined as:

Citation Information

Patent Citations

  • Internet of vehicles edge caching method based on multi-agent deep reinforcement learning

    CN113094982A