Unmanned aerial vehicle cache placement and location deployment method based on high-altitude platform assisted scenario
By optimizing the cache placement and location deployment of UAVs through the HAPS-assisted HSAC algorithm, and utilizing the SAC algorithm with hybrid action space, the problems of insufficient energy and coverage of UAVs in the air access network are solved, and more efficient user content acquisition is achieved.
Patent Information
- Application Number
- CN202310010368.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-04
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-01-04
AI Technical Summary
How can we design a cache placement scheme and location deployment for UAVs with limited energy, storage, and coverage to optimize the distance between UAVs and users and reduce content acquisition latency, thereby meeting the communication needs of ground users?
We adopt the HSAC algorithm in HAPS-assisted scenarios and combine it with the DQN concept. We optimize the buffer placement and location deployment of UAVs through the SAC algorithm in a hybrid action space. We use Actor and Critic networks for action evaluation and training to optimize the location and buffer content of UAVs to reduce transmission latency and distance.
It effectively reduces the average transmission latency for users to obtain content, improves the communication efficiency of users within the UAV coverage area, and solves the problems of insufficient energy and coverage of UAVs in the air access network.
Smart Images

Figure CN116094573B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication technology, and particularly relates to a UAV cache placement and position deployment method based on a HSAC algorithm in a HAPS assisted scenario. BACKGROUND
[0002] With the continuous development of mobile communication technology, the ground network cannot meet the various needs of emerging industries and users, and the future network will expand into an integrated space-ground network architecture. Therefore, the aerial access network composed of high, medium and low dimensional satellites or high and low altitude aircraft platforms is one of the important components of the future integrated space-ground network.
[0003] As a component of the aerial access network, the deployment of HAPS is relatively static, has a wide coverage and a longer endurance, and although the UAV is slightly insufficient in energy and coverage, it is flexible in deployment and has a shorter transmission delay. With the upper HAPS as the center control of the aerial base station, the lower UAV is deployed at the network edge to provide communication services for users, which can alleviate the traffic load pressure of the future ground network.
[0004] In recent years, content caching at the network edge has been proposed and used in content-centric cellular networks, and the UAV is located near the user and has high flexibility. The application of edge caching technology in UAV communication can further improve the performance of the aerial access network such as delay. In addition, the high mobility of the UAV enables it to provide services according to the location and needs of the ground users.
[0005] However, the energy, storage and coverage range of the UAV are limited, therefore, how to design the cache placement scheme of the UAV and how to control the movement of the UAV are important problems faced in the research and application process of the aerial access network. SUMMARY
[0006] The HAPS-assisted scenario based on the HSAC algorithm is used to design and implement a UAV cache placement and location deployment method. The HAPS acts as an air base station to assist a single UAV in content caching and location movement, thereby providing content services for a cluster of ground mobile users. The scheme mainly includes the following steps: (1) establishing a UAV communication system model based on the HAPS-assisted scenario, which includes HAPS, UAV and users; (2) constructing a user movement model based on the Gaussian Markov movement model, and constructing a user request content model based on the Zipf distribution; (3) proposing a UAV cache placement and location deployment scheme based on the HSAC algorithm to optimize the distance between the UAV and the user and the average content acquisition delay of the user.
[0007] The UAV communication in the HAPS-assisted scenario is shown in Figure 1 . It is assumed that in a certain area, the ground network traffic load is too large, and the ground base station cannot meet the communication needs of all users. The air access network composed of HAPS-UAV has the characteristics of large capacity, wide coverage, low delay and high reliability, and can provide timely and effective content services for ground users in this case.
[0008] It is assumed that there is a UAV and K users in the coverage range of HAPS, and the user set is denoted as Due to the overload of the nearby ground base station, the UAV needs to move to provide services for the users. The HAPS is deployed quasi-statically in the air and serves as an air base station to provide communication services for the UAV. The UAV is equipped with a cache unit to provide content services for users at the network edge. There are F contents in the network, and the content set is denoted as
[0009] The deployment position of the HAPS is (0, 0, H), and the position coordinates of the UAV and the user k, can be expressed as w t =(x t ,y t ,z t ), whether the user k requests the content f, is denoted as When the user k requests the content f, then otherwise whether the UAV caches the content f is denoted as When the UAV caches the content f, then otherwise
[0010] In the above scenario, a channel model based on the propagation of line of sight (LOS) probability is established. The path loss of LOS transmission and NLOS transmission is respectively expressed as:
[0011]
[0012]
[0013] where d is the distance between the two ends of the transmission, f G is the carrier frequency, c is the speed of light, η LOS ,η NLOS is the path loss factor. The probabilities of LOS transmission and NLOS transmission are:
[0014] Pr(l LOS ) = (1 + Xexp(-Y[δ - X])) -1 (3)
[0015] Pr(l NLOS ) = 1 - Pr(l LOS )(4)
[0016] where X, Y are constant parameters, which depend on the geographical environment; δ is the elevation angle between the transmitting end and the receiving end. Considering the LOS transmission probability between the UAV and user k, the average path loss is expressed as Between the HAPS and the UAV, since only the LOS transmission is considered, the average path loss is expressed as as follows:
[0017]
[0018]
[0019] where d is the distance between the UAV and user k at time t, and d t is the distance between the HAPS and the UAV, expressed as:
[0020]
[0021]
[0022] Thus, the transmission signal-to-noise ratios between the UAV and user k, and between the HAPS and the UAV, are as follows:
[0023]
[0024] where P k , P are the powers of the UAV and HAPS downlink transmissions, respectively, and σ 2 is the power of the Gaussian noise. According to the Shannon formula, with B k , B representing the UAV downlink transmission bandwidth and the HAPS downlink transmission bandwidth, respectively, the UAV downlink transmission rate and the HAPS downlink transmission rate can be obtained as:
[0025]
[0026] Therefore, the average transmission delay between the UAV and the user is expressed as:
[0027]
[0028] where s is the file size.
[0029] In order to make the UAV better cover more users and avoid the user moving out of the coverage range, the distance between the UAV and the user is as part of the optimization. Then, the system optimization goal takes the weighted sum of the average transmission delay of the user obtaining the content and the average distance between the UAV and the user as:
[0030]
[0031] where λ1+λ2=1. Then the optimization problem can be expressed as follows:
[0032]
[0033] where d max represents the coverage range of the UAV.
[0034] Since the user has mobility, the Gaussian-Markov model is used to describe the mobility of the user. The movement of the user k at time t refers to the speed value, direction value and a random variable at time t-1, and has:
[0035]
[0036]
[0037] The above two expressions represent the movement speed and direction of user k at time t. Wherein 0≤ρ0≤1 is a parameter for adjusting randomness; is the average value of the user speed and direction when t→∞; is a random variable conforming to Gaussian distribution. Therefore, the user position coordinates at the tth time can be expressed as
[0038]
[0039] In addition, the Zipf distribution is used to describe the popularity of the content and the probability of the user requesting the content, that is, the probability of user k requesting content f is as follows:
[0040]
[0041] where, is the rank of content f at time t, and η is the Zipf factor, the smaller η is, the higher the probability of content being requested is. Since the preferences of each user are not exactly the same, if user k requests content f at time slot t, set the state transition probability matrix, which represents the probability of user k requesting content f' at time slot t+1.
[0042] For the above optimization problem, a Markov decision process is constructed. The state space includes the position w t of the UAV the position of the user the cached content of the UAV the requested content of the user the moving speed of the user and the moving direction of the user are represented as:
[0043]
[0044] The action space includes the cached content U t of the UAV at time t t the moving heading angle θ the moving pitch angle t and the moving speed v are represented as:
[0045]
[0046] Since the hybrid action space contains discrete actions and continuous actions , the action space
[0047] The reward is the optimization objective r t , as shown in equation (12).
[0048] To solve the above problem, the SAC algorithm is improved based on the DQN idea to realize the SAC algorithm under the hybrid action space. The continuous action parameters θ are obtained by the Actor network π t (·|s θ is the parameter of the Actor network. Then the Critic network outputs the discrete action and the evaluation of the Actor network output, and φ j is the parameter of the Critic network. In order to eliminate the overestimation in the process of policy improvement, the double Q learning technique is used to train the Actor network in the SAC algorithm, that is, two Critic networks j=1,2 are used to evaluate the Actor network, and the minimum value of the two is selected as the final Q value. Finally, the discrete action j=1,2, the continuous action hybrid action
[0049] The loss function of the critic network is:
[0050]
[0051] where y(r t ,s t+1 ,d t ) is the target value of the critic network, and is expressed as:
[0052]
[0053] where, represents the action value obtained according to the current policy, α is the temperature control coefficient.
[0054] The loss function of the actor network is expressed as:
[0055]
[0056] where, In addition, let represent the expected minimum entropy value, and the loss function of the temperature control coefficient α is expressed as:
[0057]
[0058] The above (20), (22), and (23) update the network parameters in the gradient descent update manner.
[0059] In combination with the HAPS-assisted UAV communication scenario in the application, the flow of the HSAC algorithm is shown in Table 1.
[0060] Table 1: HSAC algorithm based on cache and deployment optimization
[0061]
[0062] BRIEF DESCRIPTION OF DRAWINGS
[0063] Figure 1 UAV communication in a HAPS-assisted scenario
[0064] Figure 2 HSAC algorithm framework
[0065] Figure 3 Iterative convergence curve DETAILED DESCRIPTION
[0066] The application designs a simulation system in python3.6, designs a circular area with HAPS as the center and a radius of 10km, deploys UAV in the coverage of HAPS, and there are 5 users near the UAV, the users have mobility, the UAV provides content service for the users, if the content required by the user is not cached by the UAV, the UAV requests the HAPS to resend to the user.
[0067] The improved SAC algorithm is used to solve the optimization problem of position deployment and cache placement scheme of the UAV in the system, and the average transmission delay of the user to obtain the content and the transmission distance between the UAV and the user are reduced. The specific steps are as follows:
[0068] Step 1: initialize the parameters in each network and the target network.
[0069] Step 2: initialize the positions of HAPS, UAV and user, and observe the initial state of the current environment, as shown in formula (18).
[0070] Step 3: obtain the action of the UAV, as shown in formula (19). First, output the continuous action of the UAV corresponding to all contents from the actor network Second, the Q value output by the critic network is to obtain the action of the cached content; finally, according to the discrete action, select the corresponding continuous action
[0071] Step 4: the UAV executes the action selected in step 3, calculates the reward according to formula (12), and then updates the positions of the UAV and the user, the cached content of the UAV, the requested content of the user, the moving speed and the moving direction, to obtain the state s t+1 .
[0072] Step 5: store (s t , r t ,s t+1 ) in the experience replay area .
[0073] Step 6: in each update step, extract a group of data from , and update the network parameters according to formulas (20)-(23).
[0074] Step 7: repeat steps (3), (4), (5) and (6) until the optimal reward is obtained, and the best cache placement strategy and position deployment scheme of the UAV are obtained.
[0075] The simulation parameters are shown in Table 2.
[0076] Table 2 Simulation parameter settings
[0077]
Claims
1. A method for cache placement and location deployment of unmanned aerial vehicles (UAVs) based on high-altitude platform (HAP) assisted scenarios, characterized in that, Comprise: Step 1: Establish a UAV communication system model based on HAPS assisted scenario, the model includes HAPS, UAV and user; Step 2: Based on the Gaussian Markov movement model, the user's movement model is constructed, and the user request content model is constructed based on Zipf distribution; Step 3: A UAV cache placement and location deployment scheme based on HSAC algorithm is proposed to optimize the distance between UAV and user and the average acquisition content delay of user; Considering the mobility of users, Gaussian-Markov model is used as the user's movement model, the movement at time t is referred to the speed value, direction value at time t-1 and a random variable, which has: The above two equations represent the moving speed and direction of user k at time t; where 0≤ρ0≤1 is a parameter to adjust randomness; is the average value of user speed and direction when t→∞; is a random variable following Gaussian distribution; thus, the user position coordinates at the tth time are expressed as Zipf distribution is used to divide the content popularity and user request probability, which has: The probability of user k requesting content f is: where, is the rank of content f at time t, and η is the Zipf factor, the smaller the value of η, the higher the probability of content being requested; since the preferences of each user are not exactly the same, if user k requests content f at time slot t, the state transition probability matrix is set, which represents the probability of user k requesting content f' at time slot t+1; F represents a total of F contents in the network.
2. The method of claim 1, wherein, A UAV communication system model based on HAPS assisted scenario is established, which has: Suppose there is a UAV and K users under the coverage of HAPS, and the user set is denoted as Due to the overload of nearby ground base stations, the UAV needs to move to provide services for users; HAPS is quasi-statically deployed in the air and provides communication services for the UAV as an air base station; The UAV is equipped with a cache unit to provide content services for users at the edge of the network; there are F contents in the network, and the content set is denoted as The deployment location of HAPS is (0, 0, H), the location coordinates of UAV and user k, can be expressed as w t = (x t , y t , z t ), whether user k requests content f, denoted as when user k requests content f, then otherwise whether UAV caches content f is denoted as when UAV caches content f, then otherwise 3. The method of claim 2, wherein, A channel model based on line of sight (LOS) probability propagation is established, which has: The path loss of LOS transmission and NLOS transmission is represented as: where d is the distance between the two ends of the transmission and f G is the carrier frequency, c is the speed of light, and η LOS ,η NLOS is the path loss factor; the probabilities of LOS and NLOS transmissions are: Pr(l LOS ) = (1 + X exp(-Y [δ - X])) -1 (3) Pr(l NLOS ) = 1 - Pr(l LOS ) (4) where X, Y are constant parameters depending on the geographical environment; δ is the elevation angle between the transmitting and receiving ends; considering the LOS probability transmission between the UAV and the user k, the average path loss is expressed as Between the HAPS and the UAV, since only the LOS transmission is considered, the average path loss is expressed as as follows: wherein the distance between the UAV and the user k at time t and the distance d between the HAPS and the UAV t is represented as:
4. The method of claim 3, wherein, The signal-to-noise ratio and transmission rate of UAV and HAPS downlink transmission process are calculated, which has: The transmission signal-to-noise ratio between UAV and user k, and between HAPS and UAV is as follows: wherein P k , P are the power of the UAV and HAPS downlink transmission respectively, σ 2 is the power of Gaussian noise; according to Shannon formula, B k , B represent the UAV downlink transmission bandwidth and the HAPS downlink transmission bandwidth respectively, and the UAV downlink transmission rate and the HAPS downlink transmission rate are obtained as:
5. The method of claim 4, wherein, Optimizing the average transmission delay of the acquired content of the user and the distance between the UAV and the user, wherein, in order for the UAV to better cover more users, avoid the user moving out of its coverage range, the distance between the UAV and the user As part of the optimization, there is: The average transmission delay between UAV and user is represented as: Where s is the file size; Then, the system optimization objective is the weighted sum of the average transmission delay of user to acquire content and the average distance between UAV and user, which is: Where λ1+λ2=1; Then the optimization problem is expressed as follows: wherein d max represents the coverage of the UAV.
6. The method of claim 1, wherein, A Markov decision process is constructed, which has: The state space includes the position w of the UAV t , the position of the user The cached content of the UAV The user requested content The moving speed of the user And the moving direction Is expressed as: The action space includes the cache content U of the UAV at time t t , the moving heading angle θ t , the pitch angle , and the speed v t , and is expressed as: Since the hybrid action space contains discrete actions and continuous actions Thus, the action space The reward is the optimization objective r t As shown in equation (12).
Citation Information
Patent Citations
Three-dimensional deployment and power distribution joint optimization method for flight base station of unmanned aerial vehicle
CN113206701A
Content distribution method in multi-unmanned aerial vehicle assisted edge computing system
CN115550937A