A UAV-assisted mobile edge network cache optimization method

Through real-time popularity prediction and deep reinforcement learning, optimized cache content deployment, combined with the hidden Markov model to schedule drones, the problems of low cache hit rate and high transmission delay in the collaborative cache mechanism of drone-assisted ground base stations are solved, and efficient cache content transmission is achieved.

CN116054905BActive Publication Date: 2025-08-26SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211149783.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-08-26
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

In the co-caching mechanism of drone-assisted ground base stations, there are problems such as low cache hit rate, high content acquisition latency, insufficient network throughput and low energy efficiency.

Method used

A real-time popularity prediction algorithm based on user historical requests is adopted, combined with deep reinforcement learning and hidden Markov model, optimized cache content deployment and drone scheduling to improve cache hit rate and reduce transmission delay.

Benefits of technology

It improves the cache hit rate, reduces the transmission delay of cached content, and improves the energy utilization rate and network throughput of the drone.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116054905B_ABST
    Figure CN116054905B_ABST
Patent Text Reader

Abstract

The present invention discloses the field of wireless communications technology, and more particularly, relates to a method for optimizing a mobile edge network cache using drones. The method comprises the following steps: S1: Based on historical user requests, an online, real-time popularity prediction algorithm is used to screen the most popular files and determine their popularity; S2: Based on the file popularity calculated in S1, a deep reinforcement learning method is used to deploy cached content; and S3: Based on step S2, a hidden Markov model is used to schedule drones during file transmission. This method can improve the network cache hit rate and drone energy utilization, while reducing cached content transmission latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a method for optimizing cache in a mobile edge network assisted by a drone. Background Art

[0002] In recent years, with the large-scale commercialization of 5G, both industry and academia have been researching and exploring 6G, the next-generation mobile communications technology. 6G will fully utilize low-, medium-, and high-frequency spectrum resources to achieve seamless global coverage across air, space, and ground. Furthermore, 6G will meet the demand for secure and reliable wireless connectivity anytime, anywhere. To cope with the massive data traffic, Mobile Edge Computing (MEC) has been introduced as a promising model. MEC leverages computing resources and caching capabilities close to the terminal to perform latency-sensitive and compute-intensive tasks, providing high-quality service to mobile users.

[0003] The explosive growth of smart IoT devices has greatly facilitated our daily lives. IoT applications require a vast amount of information input and rely heavily on high-speed, low-latency data transmission. However, backhaul links are often bandwidth-constrained, struggling to meet transmission requirements. This shortcoming poses a significant challenge to mobile networks. Edge caching, a key component of MEC, offers a promising approach to alleviating backhaul loads and addressing these challenges. Caching leverages cache resources at edge nodes and provides access to popular content close to users. Frequently requested content can be cached at the edge of the network during off-peak hours. This content is then distributed to requesters via high-speed, low-cost edge servers, rather than repeatedly transmitted across the backhaul network. However, when mobile users leave the cellular coverage radius, the requested content may not be successfully delivered to them. Alternatively, when users enter a new cellular network, the requested content may not be cached, resulting in additional latency and bandwidth consumption. Drones can be rapidly deployed to complement traditional cellular networks, providing air-to-ground connectivity and high-speed transmission. Their high mobility can be leveraged to extend their coverage area, enabling the delivery of cached content through drone dispatch. The main evaluation metrics for drone-assisted air-ground collaborative network caching mechanisms include: content cache hit rate, content retrieval latency, network throughput, energy efficiency, and user experience quality. The cache hit rate is one of the most important metrics for evaluating drone-assisted air-ground collaborative network caching mechanisms. It can be expressed as the total number of user request hits divided by the total number of requests. A higher cache hit rate indicates a greater probability that a user will be able to directly retrieve a file from the cache node's cache space. The cache hit rate can, to a certain extent, indicate the rationality of the caching strategy design. Content retrieval latency is another crucial metric in caching services. It indicates the total time cost from a user generating a data request to retrieving the complete data. Content retrieval latency is primarily composed of request transmission latency and content transmission latency. Network throughput refers to the maximum data transmission rate in the network. The faster the data transmission rate, the faster the user can retrieve a large number of files from the cache space. Furthermore, network throughput is related to the communication link between the user and the cache server. Summary of the Invention

[0004] In view of this, the present invention discloses a UAV-assisted mobile edge network cache optimization method to solve the optimization problem of UAV-assisted ground base station collaborative cache mechanism.

[0005] The technical solution provided by the present invention is specifically a drone-assisted mobile edge network cache optimization method, comprising the following steps:

[0006] S1: Based on the user's historical request situation, an online real-time popularity prediction algorithm is used to filter out the most popular files and obtain the file popularity;

[0007] S2: Based on the file popularity obtained in S1, the cache content is deployed using deep reinforcement learning methods;

[0008] S3: Based on step S2, the hidden Markov model is used to realize the scheduling of drones during the file transmission process.

[0009] Furthermore, the online real-time popularity prediction algorithm described in S1 specifically includes the following steps:

[0010] S11: Collect the user's historical request information and determine the initial popularity of the file content I t , the initial popularity I t is a random variable greater than zero, which determines the number of requests for content in a time period t. t The distribution of F t (x(d)) indicates that x(f) is the feature vector of file f, which includes the feature type and name;

[0011] S12: Based on the popularity conversion function, the file popularity conversion function is used to describe the change in file popularity during the observation period. The popularity conversion function h(t) describes the change in popularity of a content during period T. h(t) is defined as Formula 1:

[0012]

[0013] in, Indicates that the number of times the user requested content is k;

[0014] S13: Predict the user's short-term user preference p within the observation period t,u,f ; Take the feature value of the file as the input of the model, and regard the user's preference for a certain file as a binary classification problem. After a certain period of learning, update the user's learned parameters and predict the user's short-term user preference p within the observation period. t,u,f ;

[0015] S14: Calculate the loss function value and update the parameters;

[0016] S15: Calculate long-term user popularity.

[0017] Furthermore, the S15 is specifically as follows: firstly, the number of long-term requests of the user X is calculated. k , which represents the number of requests for content k within the observation period T; X k It can be expressed as:

[0018]

[0019] Will Substitute the above formula and we can get X k =I k *H(Tt k +1), for long-term content popularity, use Φ(n) to represent the cumulative number of requests for the most popular content, and calculate Φ(n).

[0020] Furthermore, step S2 adopts an intelligent cache file placement algorithm based on a deep reinforcement learning method, specifically comprising the following steps:

[0021] S21: Determine the representation of states, actions, and rewards in the algorithm;

[0022] S22: Determine the reward function;

[0023] S23: training algorithm and updating network parameters;

[0024] Furthermore, the S21 is specifically that in the intelligent content placement algorithm under study, the cache setting of each content in the drone or base station is the state in the problem. Therefore, in the time period t, the state is defined as s t , according to the above analysis, s t It is a matrix with N (N nodes, including all drones and base stations) rows and F (F files) columns, which represents the cache relationship between nodes and content.

[0025] Furthermore, the goal of the content placement algorithm is to maximize the hit rate of cached content, so the average hit rate of cached content is used as the reward function. The reward function can be expressed by formula (3):

[0026] r t =(S i ,A j )=α*L t -β*(1-L t ) (3)

[0027] Among them, α and β are two constant coefficients of the equilibrium reward, L t is the average probability that each file is cached at time slot t.

[0028] Furthermore, S23 specifically includes selecting an action, that is, the cache node caches a certain file. There are two ways to select an action: one is to randomly select an action according to a probability ε, and the other is to select the action that can maximize the Q value of the approximate function based on the approximate function. This process is iterated continuously until the algorithm ends, and the network parameters are continuously updated during this process.

[0029] Furthermore, S3 uses the hidden Markov model to implement the drone scheduling algorithm, which specifically includes the following steps:

[0030] S31: Determine the initial parameters of the drone scheduling algorithm based on the hidden Markov model;

[0031] S32: Update parameters based on observation results and formula model;

[0032] S33: estimating the model parameters of the hidden Markov model according to the optimization objective function;

[0033] S34: repeatedly update the parameters of S32 to make them converge to the parameters of S33;

[0034] Furthermore, S31 is specifically that the hidden Markov model consists of states, state transition probability matrices and observation states, and is represented by S=(s1, s2, ...s N ), A=[a ij ] |S|*|S| ,O=(v1,v2,...v M ) represent the state, state transition probability matrix and observation state in the hidden Markov model respectively; S=(s1,s2,...s N ) is a finite state set, representing the state of the drone in the drone scheduling problem, O=(v1,v2,...v M ) is a finite set of observations; C miss Represents a set of files requested by the user but cannot be transmitted to the user, R miss is a set of satisfied user requests, and we can get the set K = C miss ∩R miss ; A=[a ij ] |S|*|S| , a ij =p(X t+1 =s j |X t =s i ) is the state transition probability; B=[b j (k)],b ij =p(O t =v j |X t =s i ) is the output probability; π=[π1,π2,...,π N ],π i =P[q1=s i ] is the initial state distribution; the symbol λ is used to describe the parameters of the hidden Markov model, and the hidden Markov model can be represented by the five-tuple λ = [S, O, π, A, B]; for the drone scheduling algorithm, the specific meaning of the above state set and probability matrix in this problem is that S is the state of the base station responsible for the drone, O is the value of the observed state of each drone, A is the state transition probability; B is v iBy possible base station s i The probability of observation completed; Since S and O are two independent and identically distributed variables, the following equation is derived:

[0035]

[0036] Among them, O t+1 It represents the state of the drone at time t+1 in the observation matrix of the hidden Markov model, X t+1 It represents the state of the UAV at time t+1.

[0037] Let p{X t+1 =q j ,O t+1 =v t+1}=α t+1 (q j ,v t+1 ), then formula (4) can be written as:

[0038]

[0039] According to the definitions of the transfer matrix and the observation probability matrix, the above equation can be written as:

[0040]

[0041] Among them, q j and v t+1 are the specific values ​​of the observation matrix and the UAV state at time t+1 respectively;

[0042] The present invention provides a drone-assisted mobile edge network cache optimization method, which can improve the hit rate of network cache, reduce the transmission delay of cache content, and improve the energy utilization rate of drones.

[0043] Specifically, this method proposes an online, real-time popularity prediction algorithm based on historical user requests. This algorithm accurately and in real time analyzes the popularity of files on the network and selects the most popular files. Furthermore, with the goal of maximizing the file cache hit rate, this method proposes an intelligent cache file placement algorithm based on deep reinforcement learning. This algorithm optimizes the energy efficiency of drones when caching files and transferring cached files to users. It utilizes a drone scheduling algorithm based on a hidden Markov model to improve the energy efficiency of drones providing file transfer services and reduce file transfer latency. Evaluation metrics include file cache hit rate, file transfer latency, and drone energy consumption and energy efficiency.

[0044] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 A diagram illustrating an application scenario of a drone-assisted mobile edge network cache optimization technology provided by an embodiment of the present invention;

[0048] Figure 2 This is a schematic diagram of the file cache placement algorithm based on deep reinforcement learning according to the disclosed embodiment of the present invention;

[0049] Figure 3 This is a simulation experiment effect diagram of the local caching probability of different content on different nodes according to the present invention;

[0050] Figure 4 This is an experimental effect diagram of the cache hit rate changing with the cache capacity of the drones when the number of the drones is 6;

[0051] Figure 5 This is an experimental effect diagram of the cache hit rate changing with the base station cache capacity when the number of base stations is 3 according to the present invention;

[0052] Figure 6 This is a graph showing the experimental results of the average delay of a drone transmitting 100MB of content under different cache capacities according to the present invention;

[0053] Figure 7 This is a graph showing the experimental results of the average delay of a base station transmitting 100MB of content under different cache capacities according to the present invention;

[0054] Figure 8 This is a diagram showing the energy efficiency of a drone under different transmission content numbers according to the present invention;

[0055] Figure 9 This is an experimental effect diagram of the total energy consumption of the drone under different numbers of transmission contents described in the present invention; DETAILED DESCRIPTION

[0056] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of systems consistent with certain aspects of the present invention, as detailed in the appended claims.

[0057] The main evaluation metrics for drone-assisted air-ground collaborative network caching mechanisms include: content cache hit rate, content retrieval latency, network throughput, energy efficiency, and user experience quality. The cache hit rate is one of the most important metrics for evaluating drone-assisted air-ground collaborative network caching mechanisms. It can be expressed as the total number of user request hits divided by the total number of requests. A higher cache hit rate indicates a greater probability that a user will be able to directly retrieve a file from the cache node's cache space. This, to a certain extent, can indicate the rationality of the caching strategy design. Content retrieval latency is another crucial metric in caching services. It indicates the total time cost from a user generating a data request to retrieving the complete data. Content retrieval latency is primarily composed of request transmission latency and content transmission latency. Network throughput refers to the maximum data transmission rate in the network. The faster the data transmission rate, the faster the user can retrieve a large number of files from the cache space. Furthermore, network throughput is related to the communication link between the user and the cache server.

[0058] In order to improve the hit rate of network cache, reduce the transmission delay of cached content, and improve the energy utilization of drones, this embodiment provides a drone-assisted mobile edge network cache optimization method, including the following steps:

[0059] S1: To improve the utilization of limited cache space, this method uses a file popularity prediction algorithm based on user historical requests. This algorithm can analyze the popularity of cached file contents. In addition, the algorithm also strives to ensure the long-term file popularity characteristics.

[0060] The specific implementation method is: the first step is to use a file popularity prediction algorithm based on user historical requests. This step uses the storage information of the ground base station to access the user's historical file access records and determine the initial popularity of the file.

[0061] The second step is to use the popularity conversion function to describe the changes in file popularity during the observation period.

[0062] The third step is to store the features of each file through a vector and judge whether the user is interested in a certain file. If p t,u,f If it is greater than 0.5, it means that user u is interested in file f and willt,u,f Store them in a new collection, evaluate the model parameters of the popularity prediction algorithm based on the defined loss function, and update the parameters;

[0063] The fourth step is to calculate the long-term popularity of the file based on the long-term popularity calculation model and save it in a collection, in preparation for the next step of screening out the most popular files to be cached on the drone node or base station node.

[0064] S2: An intelligent cache placement algorithm based on deep reinforcement learning. It mainly improves the reward function during the file cache placement process. Compared with more traditional cache placement algorithms, it can significantly improve the cache hit rate and reduce the transmission latency of cached content.

[0065] This paper improves the cache placement algorithm by mainly introducing a deep reinforcement learning algorithm to implement file cache placement. Figure 2 This is an overview of the file cache placement algorithm based on deep reinforcement learning. DQN integrates deep learning and reinforcement learning to improve the efficiency and accuracy of training. DQN consists of two parts: the current network and the target network. The current network and the target network are two convolutional neural networks (CNNs). The former is used to estimate the true Q value, and the latter is used to calculate the target Q value. However, considering the huge state space and action space in this problem, storing them in matrices will lead to dimensionality disaster. At the same time, excessive computational overhead will produce long delays, thus affecting the performance of the algorithm. DQN performs better than Q-learning in the face of multiple state and action spaces. Because the state and action can be used as inputs of CNN, the Q value can be generated by the neural network. It replaces the formula based on value iteration with the use of neural networks to calculate the Q value. It is worth noting that the experience replay mechanism is introduced in DQN. The experience replay mechanism plays a vital role in reproducing the correlation between experiences during training. The parameters of the current network and the target network are represented by Θ respectively. C and Θ T Indicates that these two neural networks with different parameters play different roles in the execution of the DQN algorithm. C The current network is used to predict the current Q value, and after a period of time, Θ C Deploy to the target network instead of Θ T To achieve parameter update. The Stochastic Gradient Descent (SGD) method can be used to update the parameters. After a certain number of iterations, the parameters Θ of the target network T It can be represented by Θ C Alternative.

[0066] A set S is used to represent the state of the system. From the perspective of the particularity of the problem and the consistency with the above problem formula, in the intelligent content placement algorithm, the cache of each content cached by the drone or base station is set to the state in the problem. Therefore, in time period t, the state is defined as s t , according to the above analysis, s t It is a matrix with N (N nodes, including all drones and base stations) rows and F (F files) columns, which represents the cache relationship between nodes and content. Action means that the agent decides to change the current state to another new state. The decision that causes the state change is called an action. If a good decision is made, the agent will get a high reward, otherwise, it will not get any reward or penalty. The action in time slot t is represented by a t , which is a matrix consisting of N (N nodes) rows and F (F files) columns. The elements in this matrix are binary variables, for example, The value of is equal to 1, which means that node i will cache content j. If the value of is equal to 0, it means that node i does not cache content j. The goal of the content placement algorithm is to maximize the hit rate of cached content. Therefore, the average hit rate of cached content is used as the reward function. The immediate reward for taking an action in the current state is calculated based on the reward function.

[0067] S3: Using the hidden Markov model to schedule drones during file transmission mainly improves the latency of file cache transmission and the energy utilization of drones.

[0068] The content placement algorithm cannot guarantee that the placed content can be delivered to the user directly from the drone or base station. In other words, the content cached on the drone may require the drone to move to a certain location to achieve content delivery, so from the perspective of improving user QoE, the cache hit rate still has room for improvement. It is an effective method to use the outstanding flexibility and maneuverability of drones to transmit requested content that is unreachable due to the long distance, thereby improving the cache content hit rate. This section will discuss how to effectively schedule drones to further improve the successful transmission rate of files with high energy efficiency. Generally speaking, the energy consumption of drones mainly consists of two parts: mechanical energy consumption for drone flight or hovering and communication energy consumption. The hovering energy consumption and moving energy consumption of drones are represented by P and P, respectively. hover and P flightRepresents. The main goal of the drone scheduling algorithm is to optimize the scheduling management of drones. When drones provide content transmission to users, communication energy consumption is much lower than mechanical energy consumption. After the drone scheduling is completed, if there is no need to redeploy, the drone will temporarily stay at its latest position. This part will be discussed in detail in the next section when studying the structure of the drone state matrix. Scheduling drones to deliver missed content is a supplementary process of the entire cache placement scheme. Its purpose is to ensure that the cached content can be transmitted to users as much as possible. With the completion of drone scheduling, the cache scheme will enter the next round of cache content placement. In the drone scheduling algorithm, the position coordinates of drones and users can be collected through ground base stations. The energy consumption of the drone mentioned above can be expressed by formula (6):

[0069]

[0070] use Represents the departure and destination of the scheduled UAV. Therefore, the flight time of the schedule can be calculated using formula (7):

[0071]

[0072] The hidden Markov model consists of three main parts: state, state transition probability matrix and observation state. The following will discuss the specific meaning of the parameters of the hidden Markov model and give the modeling process of the drone scheduling algorithm studied in this chapter. N ), A=[a ij ] |S|*|S| ,O=(v1,v2,...v M ) represent the state, state transition probability matrix and observation state in the hidden Markov model respectively. S=(s1,s2,...s N ) is a finite state set, representing the state of the drone in the drone scheduling problem, O=(v1,v2,...v M ) is a finite set of observations. C miss Represents a set of files requested by the user but cannot be transmitted to the user, R miss is a set of satisfied user requests, and we can get the set K = C miss ∩R miss A=[a ij ] |S|*|S| , a ij =p(X t+1 =s j |X t =s i ) is the state transition probability. j (k)],b ij =p(Ot =v j |X t =s i ) is the output probability. π=[π1,π2,...,π N ],π i =P[q1=s i ] is the initial state distribution. The symbol λ is used to describe the parameters of the hidden Markov model, so the hidden Markov model can be represented by the five-tuple λ = [S, O, π, A, B]. For the drone scheduling algorithm, the specific meaning of the above state set and probability matrix in this problem is that S is the state of the base station responsible for the drone, O is the value of the observed state of each drone, and A is the state transition probability. B is v i By possible base station s i Since S and O are two independent and identically distributed variables, the following equation can be derived:

[0073]

[0074] Let p{X t+1 =q j ,O t+1 =v t+1}=α t+1 (s j ,v t+1 ), then formula (8) can be written as:

[0075]

[0076] According to the definitions of the transfer matrix and the observation probability matrix, the above equation can be written as:

[0077]

[0078] The drone scheduling management architecture of the drone-assisted network consists of two parts: drones and their related base stations. In order to concisely represent the proposed drone scheduling algorithm, it is assumed that the drone's related base stations are unique, N and M represent the number of base stations and drones respectively, and the vector is the observation sequence at time t, element is a binary variable that indicates whether a base station is associated with a drone that can provide the missed content to the user, W t Is a matrix with N rows and M columns, used to represent the state of the drone. can be collected by a central controller, which is determined by the number of unsatisfied user requests. For example, This means that the drone under the control of base station N can provide the content requests that are not satisfied from the user. t(i, j) is the probability of the random process, indicating that the state at time t is S i In the case of , the state at time t+1 is S j The probability, P t (i, j) can be expressed as follows using formula (11):

[0079] P t (i,j)=P(O,O t =S i ,O t+1 =S j |λ)=(a t (i)a ij b j (O t+1 )β t+1 (j)) / P(O|λ) (11)

[0080] Among them, a t (i) = P(O1O2...O t ,S i =q t |λ) and β t (i) = P(O t+1 O t+2 ...O T ,S i =q T |λ) represent the forward and backward probabilities respectively. Therefore, the state at time t is S i The probability can be expressed as:

[0081]

[0082] Based on the above discussion, the parameters in the hidden Markov model can be approximated as follows:

[0083]

[0084]

[0085] π i =P t (i) (15)

[0086] The main process of the drone scheduling algorithm based on the hidden Markov model is as follows:

[0087] (1) Determine the initial parameters of the HUS; (2) Update the parameters based on the observation results and the above formula; (3) Estimate the parameters of the hidden Markov model based on the optimization objective function; (4) Iteratively update the parameters of (2) to make them converge to the parameters of (3); (5) Output the sequence of UAV scheduling.

[0088] To test the efficiency of the present invention, a set of comparative experiments were designed. The experimental comparison showed that the present invention has a good effect in improving cache hit rate and reducing file transfer delay. The UAV-assisted ground base station caching strategy was simulated in an area of ​​5km*5km. The distribution of users in the network follows HPPP, with a density of λ u =10 -4 (per square meter). The flight speed of the drone is set between 15m / s-30m / s. The maximum storage capacity of the drone and the base station is set to C u =12GB and C b = 2TB. Mobile networks offer numerous services that utilize base station storage capacity, but user content delivery is only one of these services. Therefore, despite a base station's maximum storage capacity of 2TB, only 12GB of storage capacity is available for caching content delivered to users. According to 3GPP TR 36.777, the uplink and downlink data rates for drones are set at 50Mbps, and the data rates for base station and drone control commands are set at 80kbps each.

[0089] The caching strategy was evaluated through simulation experiments within 30,000 time slots, with each time slot set to 10ms, assuming that the network environment remains unchanged during the small time interval. To verify the performance of the proposed Intelligent Content Placement (ICP) strategy, five benchmarks were selected for comparison with ICP. One of the comparative experiments used an improved particle swarm optimization algorithm (PSOA) to collaboratively solve the edge cache content placement problem. The content popularity analysis based on Zipf distribution is defined as CPA1, and the content popularity prediction algorithm proposed in this implementation plan is defined as CPA2. To evaluate the proposed intelligent cache placement algorithm ICP, five comparative experiments were conducted: CPA1 combined with PSOA (CPA1+PSOA), CPA1 combined with DQN (CPA1+DQN), and CPA2 combined with PSOA (CPA2+PSOA). In addition, to more comprehensively demonstrate the performance of the proposed ICP, comparative experiments were also conducted using a greedy cache content placement scheme and a local optimal scheme.

[0090] Content Hit Ratio (CHR), Energy Efficiency (EE), and Average Content Delivery Latency (ACDL) are three metrics used to evaluate the performance of ICP. The specifications of these metrics are as follows: CHR is the cache hit ratio, which represents the ratio between the number of user requests for content and the number of user requests that can be satisfied. ACDL represents the average time it takes to complete the content delivery of a set of content requests. The total content delivery latency consists of three parts: transmission latency, flight time of the scheduled drone, and queue time for requests to be responded to. These three delays are discussed in detail in the content delivery latency section. EE represents energy efficiency. Energy consumption is a decisive factor limiting the life cycle of a drone. Therefore, more attention is paid to the energy consumption of the drone during hovering and flying. EE is the total file size that the drone can successfully transmit under a fixed energy consumption.

[0091] Figure 3 The following simulation results show the local caching probability of different content on different nodes. The caching probabilities of 20 pieces of content are partially displayed to illustrate the process of content selection and where they should be cached. For example, file 8 has the highest caching probability on BS3, and file 5 has the second highest caching probability on BS3. Files 5 and 8 will then be cached on BS3. Therefore, each node will cache files in descending order of their caching probability until the cache space is full.

[0092] Through analysis and experiments, we found that when the number of base stations is 3 and the number of drones is 6, the change in CHR is most obvious as the available cache space of drones and base stations changes. Therefore, the number of base stations and drones is set to 3 and 6 respectively.

[0093] Figure 4 and Figure 5 The simulation results of the other five comparative experiments and the ICP algorithm proposed in this chapter are shown respectively. Figure 4 and Figure 5 In Figure 1, we can observe how the cache hit rate of both the drone and base station varies with cache size. CHR increases not only with increasing drone cache capacity, but also with increasing cache capacity. This is because user requests far outnumber the cached content. Furthermore, as cache capacity increases, more content is selected for caching. The more cached content, the more user requests can be responded to by the drone or base station. We can observe that the proposed ICP algorithm outperforms the other five comparative experiments in terms of cache hit rate, evaluating both the drone and base station caches.

[0094] like Figure 4 As shown in the figure, when comparing the differences between the proposed ICP and the greedy method, it is found that when the cache capacity is 6GB, ICP improves the CHR by at least 9%. At the same time, ICP has different degrees of improvement on the cache hit rate compared with other comparative experiments. Figure 5 The proposed ICP significantly improves cache hit rate. When the BS cache capacity is 2GB, the maximum difference between ICP and the greedy condition reaches 47.23%. This phenomenon is due to two reasons: first, the greedy strategy is short-sighted, while the ICP strategy used in this implementation considers long-term benefits. Furthermore, ICP utilizes the DQN algorithm and a content popularity prediction algorithm designed based on historical user content requests.

[0095] By comparing CPA2+DQN and CPA1+DQN, we can find that they are both based on the DQN algorithm, and CP2+DQN outperforms CPA1+DQN in CHR. This comparison result also shows that the proposed content popularity prediction algorithm has a good effect on the placement of cached content. It should be noted that Figure 4 The gap between these six solutions in the drone cache is less than Figure 5 .

[0096] Although the proposed content popularity prediction algorithm was applied in the CPA2+PSOA experiment, its performance was still lower than that of ICP. This shows that the DQN algorithm is better than the particle swarm algorithm in the cache content placement problem. At the same time, it was found that PSOA may lead to unstable experimental results. In the comparative experiments of CPA1+PSOA and CPA2+PSOA, CHR does not always increase with the expansion of cache capacity. Through analysis, it can be found that Figure 4 The CHR of CPA2+PSO shows a downward trend. Figure 5 The CHR of CPA1+PSOA and CPA2+PSOA also shows a downward trend, which is because PSOA can easily get suboptimal solutions. Figure 4 and Figure 5 When the UAV capacity reaches 12GB, the CHRs of the four caching schemes are all above 75%. However, when the BS capacity is 12GB, only the proposed ICP and CPA1+DQN exceed 75% in CHR. This is because the proposed base station-UAV dual-layer caching architecture adopts a base station-UAV dual-layer caching architecture. When the user cannot obtain the target content from the UAV cache, as long as the base station has cached the requested content, it will provide the content to the user. In other words, the BS cache acts as a backup.

[0097] According to the definition of ACDL, ACDL consists of three parts: transmission delay, flight delay of scheduled drones, and queuing delay of requests to be responded to. Figure 6 and 7 The trend of ACDL as the cache capacity of the UAV and base station varies from 2 to 12 GB is shown. Figure 6 and Figure 7 It shows that, with the increase of cache capacity, the ACDL of all six schemes will decrease. Figure 6 and Figure 7 The results in

[15] also show that the performance of the proposed ICP in ACDL is better than that of the other five groups of comparative experiments. Figure 6 The ACDL of the drone shown in the figure, when the cache capacity is 12GB, the ACDL obtained by ICP is 47.1% lower than the greedy strategy, 43.5% lower than the local optimal strategy, 33.64% lower than CPA1+PSOA, 31.8% lower than CPA2+PSOA, and 17.1% lower than CPA1+DQN. Similarly, in Figure 7 In the ACDL of the base station shown, when the cache space is 12GB, the ACDL of ICP is 27.67%, 22.79%, 26.24%, 9%, and 0.9% lower than the above strategies, respectively. This is because in the ICP strategy, all drones and base stations place cached content in a collaborative caching manner, and ICP can accurately predict the popularity of content. In addition, ICP considers the topology of the entire network and user requests based on the overall network situation, deploying user-favorite content on cache nodes as close to the user as possible.

[0098] In contrast, cooperative caching is not considered in the greedy optimal strategy and the local optimal strategy. For the PSOA solution, although it cannot achieve performance similar to CAP1+DQN and ICP, its performance is better than the greedy and local optimal strategies. Figure 7 It shows that when the buffer capacity is the same, the ACDL of the BS is lower than that of the UAV. This is because the transmission power of BSs is greater than that of the UAV.

[0099] In order to verify the energy consumption efficiency of the UAV scheduling algorithm proposed in this invention during the execution process, the following four UAV scheduling solutions were selected as comparative experiments in the simulation experiment. These four solutions are: Stationary UAV (SUAV): the UAV remains stationary at a fixed hovering point; Random Move UAV (RUAV): the UAV moves randomly at a random speed in the network; Maximum Speed ​​UAV (MSUAV): the UAV flies at its maximum flight speed; Stable Trajectory UAV (STUAV): the UAV moves within a predetermined trajectory. Figure 8 and 9 The energy efficiency and total energy consumption of the proposed drone scheduling algorithm and four sets of comparative experiments are shown.

[0100] from Figure 8 and 9 As can be seen in the figure, the EE of HUS, RUAV, and STUAV is significantly higher than that of SUAV and MSUAV. This phenomenon is due to the SUAV strategy being unable to deliver content to users outside the drone's transmission range. In the MSUAV solution, the drone maintains high-speed and rapid movement, increasing energy consumption. Furthermore, due to the drone's high-speed movement, most of its content transmission processes are terminated before successful delivery. Generally speaking, energy efficiency increases with the amount of content delivered, as the drone has more opportunities to provide content to users with the same energy consumption.

[0101] The figure shows that HUS achieves the best performance, with STUAV being the closest to HUS. However, the gap between them widens as the amount of content delivered increases. This is because STUAV restricts the drone's flight trajectory, allowing it to deliver requested content to users as long as the drone has cached it. Compared to the proposed approach, this increases the drone's energy consumption, resulting in STUAV's lower energy efficiency. RUAV, on the other hand, moves randomly at a random speed. Content is delivered to requested users as long as they are within its coverage area. Figure 9 is the total energy consumption of each method in a time interval, from Figure 9 It can be seen that as the amount of delivery content increases, the total energy consumption of the drone under each strategy will also increase.

[0102] The energy efficiency and total energy consumption of the five drone scheduling schemes under different numbers of drones are as follows: Figure 8 and 9As shown above, based on the discussion of drone energy consumption, energy consumption is closely related to flight distance. Since hovering energy consumption is much lower than flying energy consumption, the SUAV has the lowest total energy consumption. The MSUAV solution has the highest energy consumption, and the proposed solution has the second highest energy consumption among the five solutions. However, due to the proposed solution's successful delivery of more content, its energy efficiency is the best of all solutions.

[0103] Except for the SUAV and MSUAV solutions, the EE of the other three methods increases with the increase of the number of drones. The reason why the EE of SUAV and MSUAV always remains at a low level is related to the emergence of Figure 8 The reason for the results in is the same. It is worth noting that Figure 8 The total energy consumption is generally lower than Figure 9 This is because when more drones are deployed, when a certain number of drones cooperate with each other, the transmission time of the content will be reduced and the corresponding energy consumption will also be reduced.

[0104] In summary, in this embodiment, the content popularity analysis and prediction algorithm is first used to screen out the content that is more popular with users in a certain time period, and the changes in the popularity of the content in the next time period are analyzed based on the user's historical requests. This method is used to store more popular content in a limited cache space to improve the utilization efficiency of the cache space. Secondly, in order to improve the file cache hit rate and reduce the file transmission delay, and improve the user experience, the cached content should be placed close to the requested user. For this purpose, the present invention adopts a file placement algorithm based on the deep reinforcement learning algorithm DQN. Finally, a drone scheduling algorithm based on the HMM hidden Markov model is proposed. The advantage of the HMM-based drone scheduling algorithm is that it takes into account the topology of the entire network and schedules the drone as a whole in an energy-efficient manner.

[0105] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.

Claims

1. A UAV-assisted mobile edge network cache optimization method, characterized in that: The steps include: S1: Based on the user's historical request situation, an online real-time popularity prediction algorithm is used to filter out the most popular files and obtain the file popularity; S2: Based on the file popularity obtained in S1, the cache content is deployed using deep reinforcement learning methods; S3: Based on step S2, the hidden Markov model is used to realize the scheduling of drones during the file transmission process; S3 uses the hidden Markov model to implement the drone scheduling algorithm, which includes the following steps: S31: Determine the initial parameters of the drone scheduling algorithm based on the hidden Markov model; S32: Update parameters based on observation results and formula model; S33: estimating the model parameters of the hidden Markov model according to the optimization objective function; S34: repeatedly update the parameters of S32 to make them converge to the parameters of S33; Specifically, the hidden Markov model consists of states, state transition probability matrices, and observation states, and is represented by S=(s1, s2, ...s N ), A=[a ij ] |S|*|S| ,O=(v1,v2,...v M ) represent the state, state transition probability matrix and observation state in the hidden Markov model respectively; S=(s1,s2,...s N ) is a finite state set, representing the state of the drone in the drone scheduling problem, O=(v1,v2,...v M ) is a finite set of observations; C miss Represents a set of files requested by the user but cannot be transmitted to the user, R miss is a set of satisfied user requests, and we can get the set K = C miss ∩R miss ; A=[a ij ] |S|*|S| , a ij =p(X t+1 =s j |X t =s i ) is the state transition probability; B = [b j (k)],b ij =p(O t =v j |X t =s i ) is the output probability; π=[π1,π2,...,π N ],π i =P[q1=s i ] is the initial state distribution; the symbol λ is used to describe the parameters of the hidden Markov model, and the hidden Markov model can be represented by the five-tuple λ = [S, O, π, A, B]; for the drone scheduling algorithm, the specific meaning of the above state set and probability matrix in this problem is that S is the state of the base station responsible for the drone, O is the value of the observed state of each drone, A is the state transition probability; B is v i By possible base station s i The probability of observation completed; Since S and O are two independent and identically distributed variables, the following equation is derived: Among them, O t+1 It represents the state of the drone at time t+1 in the observation matrix of the hidden Markov model, X t+1 It represents the state of the UAV at time t+1; Let p{X t+1 =q j ,O t+1 =v t+1 }=α t+1 (q j ,v t+1 ), then formula (4) can be written as: Among them, q j and v t+1 are the specific values ​​of the observation matrix and the UAV state at time t+1 respectively; According to the definitions of the transfer matrix and the observation probability matrix, the above equation can be written as:

2. The method for optimizing cache of a mobile edge network assisted by a drone according to claim 1, wherein: The online real-time popularity prediction algorithm described in S1 specifically includes the following steps: S11: Collect the user's historical request information and determine the initial popularity of the file content I t , the initial popularity I t is a random variable greater than zero, which determines the number of requests for content in a time period t. t The distribution of F t (x(d)) indicates that x(f) is the feature vector of file f, which includes the feature type and name; S12: Based on the popularity conversion function, the file popularity conversion function is used to describe the change of file popularity during the observation period; the popularity conversion function h(t) describes the change of popularity of a content during period T, and h(t) is defined as formula (1): in, Indicates that the number of times the user requested content is k; S13: Predict the user's short-term user preference p within the observation period t,u,f ; Take the feature value of the file as the input of the model, and regard the user's preference for a certain file as a binary classification problem. After a certain period of learning, update the user's learned parameters and predict the user's short-term user preference p within the observation period. t,u,f ; S14: Calculate the loss function value and update the parameters; S15: Calculate long-term user popularity.

3. The method for optimizing cache in a mobile edge network assisted by a drone according to claim 2, wherein: The specific step of S15 is as follows: first calculate the number of long-term requests of the user X k , which represents the number of requests for content k within the observation period T; X k It can be expressed as: Will Substitute the above formula and we can get X k =I k *H(Tt k +1), for long-term content popularity, use Φ(n) to represent the cumulative number of requests for the most popular content, and calculate Φ(n).

4. The method for optimizing cache in a mobile edge network assisted by a drone according to claim 1, wherein: Step S2 uses an intelligent cache file placement algorithm based on a deep reinforcement learning method, specifically including the following steps: S21: Determine the representation of states, actions, and rewards in the algorithm; S22: Determine the reward function; S23: Train the algorithm and update the network parameters.

5. The UAV-assisted mobile edge network cache optimization technology according to claim 4, characterized in that: Specifically, in the intelligent content placement algorithm under study, the cache of each content in the drone or base station is set to the state in the problem; therefore, in time period t, the state is defined as s t , according to the above analysis, s t It is a matrix with N rows and F columns, which represents the cache relationship between nodes and content, where N represents N nodes, including all drones and base stations, and F represents F files.

6. The UAV-assisted mobile edge network cache optimization technology according to claim 4, characterized in that: Specifically, the goal of the content placement algorithm is to maximize the hit rate of cached content. Therefore, the average hit rate of cached content is used as the reward function. The reward function can be expressed by formula (3): r t =(S i ,A j )=α*L t -β*(1-L t ) (3) Among them, α and β are two constant coefficients of the equilibrium reward, L t is the average probability that each file is cached at time slot t.

7. The UAV-assisted mobile edge network cache optimization technology according to claim 4, characterized in that: Specifically, S23 selects an action, that is, the cache node caches a certain file. There are two ways to select an action: one is to randomly select an action according to a probability ε, and the other is to select the action that can maximize the Q value of the approximate function based on the approximate function. This process is continuously iterated until the algorithm ends, and the network parameters are continuously updated during this process.