A joint optimization method for UAV base station location and sub-band allocation

By adopting a joint optimization method of deep reinforcement learning and graph shading algorithms in the drone base station network communication system, the performance limitation and scenario limitation problems in the optimization of base station location and subband allocation are solved, and more efficient network service performance and communication quality are achieved.

CN119545369BActive Publication Date: 2025-05-06THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510089672.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

In the existing UAV base station network communication system, the optimization method for base station location and subband allocation has problems such as high number of iterations, limited performance, many scenario limiting factors, and difficult user precise location to obtain, resulting in affecting the communication effect.

Method used

The joint optimization method of drone base station position adjustment based on deep reinforcement learning and subband allocation based on graph coloring is adopted. By constructing the Markov decision-making process, the agent is trained to optimize the base station position, and subband allocation is combined with the graph coloring algorithm to optimize the network service performance.

Benefits of technology

It realizes that in dynamic and flexible drone network scenarios, supports user movement and handover, fully considers user transmission rate, drone energy consumption and user coverage, and improves network service performance and communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119545369B_ABST
    Figure CN119545369B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of wireless communication technology, and in particular, relates to a method for jointly optimizing the location of a drone base station and sub-band allocation, the steps of which are: first, determining the adjustment items to be executed in this time step; for sub-band allocation: initializing the sub-band allocation parameters and constructing an interference relationship diagram, cyclically adjusting the frequency band selection probability of the drone base station until convergence or reaching a preset number of cycles, and obtaining an optimized allocation scheme. For base station position adjustment: evenly dividing the scene into multiple square grids, and constructing a Markov decision process for the location adjustment of the drone base station; obtaining the state of each drone base station and inputting it into its own intelligent body to obtain the location adjustment result. Compared with existing methods, the method involved in the present invention can jointly optimize the location adjustment and sub-band allocation, can fully consider multiple factors such as transmission rate, user coverage, and base station energy consumption, is suitable for high-dynamic scenarios, and does not rely on precise user coordinates, and can effectively improve network service performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a method for jointly optimizing the location of an unmanned aerial vehicle base station and sub-band allocation. Background Art

[0002] At present, people's demand for extensive communication coverage and high-quality communication services is increasing. Mobile base stations represented by drone base stations have become one of the important candidate solutions for future network coverage and capacity enhancement due to their high dynamics, greater line-of-sight probability and flexible deployment characteristics. In particular, for sudden and temporary application scenarios, such as hot spots and emergency disaster relief, the deployment of drone base stations can make full use of their flexibility advantages to make up for the lack of ground network coverage and achieve fast and efficient networking. In this case, dynamically adjusting the location of base stations according to the communication needs of users is the key to providing effective coverage and high-quality services. In addition, the line-of-sight (LoS) link characteristics of drone base stations make the transmission loss difference between adjacent base station nodes small, resulting in multiple strong interference sources, and the interference intensity changes dynamically with the location of the mobile base station, which seriously limits the network capacity. Therefore, it is also necessary to dynamically allocate and adjust the sub-bands used by base stations in the network to achieve better network service performance.

[0003] The communication scenario of drone base station networking has attracted extensive attention from researchers, and there are many related studies focusing on the adjustment of location and sub-band allocation. In existing studies, for the problem of drone base station location adjustment, researchers mainly use location adjustment schemes based on clustering algorithms or heuristic algorithms. Among them, the clustering algorithm uses the coordinates of ground mobile users to divide users into multiple clusters and obtain the coordinates of the cluster center of each user cluster. Let the drone serve one of the user clusters, and let the drone go to the cluster center to achieve location adjustment or further use the heuristic algorithm to achieve accurate optimization of the base station location. The drone base station location adjustment scheme based on heuristic algorithms mainly uses particle swarm or genetic algorithm to solve the optimal location through multiple iterative searches. In addition, in recent years, reinforcement learning algorithms have been widely used in the field of communication research. Some researchers have designed an intelligent drone base station location adjustment scheme based on reinforcement learning, using intelligent agents to make decisions on the location of drone base stations, which also achieves improvements in network performance.

[0004] On the other hand, in order to reduce the interference caused by multiple drones reusing the same frequency, allocating different frequency channels to different drone base stations has also attracted the attention of some researchers. This problem is mainly solved by establishing a mathematical model based on graph coloring, adjusting the selection probability through multiple iterations, and finally finding a sub-band allocation combination that meets the convergence conditions.

[0005] First, the UAV base station location adjustment scheme based on clustering scheme only relies on the user location for decision-making, and it is difficult to consider other factors that affect the operation of the UAV base station and the performance of network services, such as UAV base station energy consumption, user coverage, user fairness, etc.; at the same time, although the cluster center point can shorten the user's connection distance and achieve a theoretically better position, depending on the different user distribution, the performance at this position often has a lot of room for improvement.

[0006] Secondly, other base station location adjustments based on traditional solutions are mostly based on heuristic algorithms, which often require multiple iterative searches during the convergence process. At the same time, their optimization performance is often affected by factors such as the search range or the number of iterations, and there are problems such as low decision-making performance and high computational complexity. Considering the highly dynamic characteristics of drone networks, using this solution is not conducive to real-time decision-making in the scenario of drone base station networking and communication.

[0007] In addition, although intelligent position adjustment based on reinforcement learning schemes solves the above problems, existing research often has certain limitations in decision-making in terms of initial deployment, user movement, user association and other scenario factors, and cannot meet the application requirements of UAV networks with strong flexibility. For example, some schemes require optimization decisions in scenarios where initial deployment has been completed well, require user locations to be fixed, or require UAV base stations to serve users within a specified range without supporting user switching. At the same time, most schemes need to rely on accurate user location coordinates and other information as algorithm input to generate the state information required for intelligent agent decision-making. However, the user's precise location coordinates are often difficult to obtain in actual situations, and there is a problem of low acquisition accuracy, which affects decision-making performance.

[0008] Finally, although the existing research on sub-band allocation schemes is relatively sufficient, the existing research only considers the sub-band allocation optimization problem of UAV base stations at fixed locations or preset fixed trajectories. Few studies have considered the location of UAV base stations and sub-band allocation problems together, resulting in limited optimization performance. In order to achieve further improvement in network performance, the solution to this joint problem needs further research. Summary of the invention

[0009] The present invention provides a method for jointly optimizing the location of a UAV base station and the allocation of sub-bands, so as to solve the technical problems that the method for jointly optimizing the location of a UAV base station and the allocation of sub-bands in a UAV base station networking communication system in the related art has the following problems: high number of iterations, limited performance, many scene restriction factors, and difficulty in obtaining the precise location of users, which affect the communication effect.

[0010] The technical solution adopted by the present invention is:

[0011] A method for jointly optimizing the location of a drone base station and sub-band allocation comprises the following steps:

[0012] Step 1: According to the position adjustment count value at the current time step, determine whether the number of UAV base station position adjustment optimizations that have been performed after the last sub-band allocation optimization has reached the set number of times. If the number of times has reached the set number, execute step 2 once and restart the position adjustment count. Otherwise, execute step 3.

[0013] Step 2, initialize the optimization parameters of the sub-band allocation of the UAV base station and construct an interference relationship diagram; then cyclically adjust the frequency band selection probability of the UAV base station until convergence or reaching a preset number of cycles to obtain an optimized sub-band allocation scheme; adjust the sub-band allocation of the UAV base station based on the optimized sub-band allocation scheme;

[0014] Step 3, according to the set grid size, the scene is evenly divided into multiple square grids, where the global grid size is larger than the local grid size; and the state space, action space and reward function in the Markov decision process of the UAV base station position adjustment are constructed; the reward function is designed based on the comprehensive factors of user transmission rate, user coverage and UAV base station energy consumption, and the optimization cumulative reward is used as the optimization goal; then the local state and global state of each UAV base station in the scene are obtained to generate the complete state information of each UAV; each UAV inputs its complete state information into its own intelligent body to obtain the position adjustment result; the position of the UAV base station is adjusted based on the position adjustment result.

[0015] Furthermore, after step 1, the following steps are also included:

[0016] Update the scene information, including calculating whether a line-of-sight link is formed; then calculate the channel gain and signal-to-interference-and-noise ratio (SINR) between the drone base station and the user; calculate the user switching based on the Max-SINR criterion, that is, let the user switch to the drone base station with the largest SINR, and calculate the user transmission rate.

[0017] Furthermore, step 2 specifically includes the following steps:

[0018] Step 201, initialize the optimization parameters of the sub-band allocation of the drone base station, calculate the interference intensity and construct an interference relationship diagram according to the channel state between the user and the drone base station;

[0019] Step 202, cyclically executing sub-band selection, state checking and selection probability updating until a convergence condition is met or a preset number of cycles is reached, and an optimized sub-band allocation scheme is obtained;

[0020] Step 203, compare the optimized sub-band allocation scheme with the sub-band allocation scheme before optimization, and select a scheme with a higher total user transmission rate to adjust the sub-band allocation of the drone base station.

[0021] Furthermore, the state space, action space and reward function in the Markov decision process of the drone base station position adjustment in step 3 are specifically:

[0022] The state space includes global state and local state. The global state at the moment is expressed as:

[0023]

[0024] In the formula, , and Respectively represent the number of drone base stations, the number of users, and the number of users in the uncovered state in each global grid; represents the average user transmission rate per unit bandwidth in each global grid; and They represent the maximum user coverage duration and the average user coverage duration in each global grid respectively; Indicates the number of users served by each drone base station;

[0025] The local state at the moment is expressed as:

[0026] In the formula, Represents the drone base station in each global grid The number of service users; , , , and Respectively represent the number of users in each local grid, the number of drone base stations, the number of served users, the average number of users served by other drones, and the number of users in the uncovered state; is the average user transmission rate per unit bandwidth in each local grid; For each user in the local grid The average achievable transmission rate per unit bandwidth; and They represent the maximum user coverage duration and the average user coverage duration in each local grid respectively;

[0027] Time step The state space at time It consists of the local state and global state of each drone base station, expressed as:

[0028] s ( t ) = [ o a ( t ), o p , 1 ( t ), o p , 2 ( t ),..., o p , M ( t ) ]

[0029] In the formula, M is the number of drone base stations;

[0030] Action Space :Assume that the base station moves a fixed distance and only makes decisions on the moving direction, then represents a discrete sequence of directions in which the base station can move, where is the number of movable directions, The action of making a decision is recorded as ;

[0031] Reward Function : is the time step , In Status Take Action Then transfer to The rewards received, The reward function is expressed as:

[0032]

[0033] In the formula, Reward weight for local transmission rate; for The transmission rate of the service users; is the local energy consumption penalty weight; Base stations for all drones Energy consumption; is the global non-coverage penalty weight; is the total duration of all users in the uncovered state; reward weight for global transfer rate; is the time step The service user transmission rate of all drone base stations.

[0034] Furthermore, a Markov decision process for adjusting the location of the UAV base station is constructed, and its optimization objective design includes the following:

[0035] The optimization goal of the drone base station position adjustment is to optimize the cumulative reward, that is, the cumulative reward of all drone base stations at each time step in an adjustment cycle, expressed as:

[0036]

[0037] in, is the time step Rewards for adjusting the positions of all drone base stations; and They are the cumulative weight of non-coverage penalty and the cumulative weight of reward; is the number of position adjustments inserted between two sub-band allocation adjustments; The second position adjustment is one adjustment round, then Indicates the number of adjustment rounds that have been completed since the starting time step; Base station for drones The mobile energy consumption;

[0038] The optimization target of the UAV base station position adjustment is designed as the weighted sum of all user transmission rates, user uncovered time, and UAV base station energy consumption in multiple position adjustments between two sub-band allocations, where:

[0039] User transmission rate: the communication transmission rate of the drone base station serving ground users;

[0040] User uncovered time: When a user receives drone base station service, if the signal-to-noise ratio between the user and the base station is less than the set threshold, the user is judged to be in an uncovered state in the current time step, and the current user's uncovered time increases by one time step interval; if the SINR between the user and the base station is greater than or equal to the set threshold, the current user's uncovered time count is cleared;

[0041] Drone base station energy consumption: energy consumption generated by the hovering and movement of the drone base station.

[0042] Furthermore, in step 3, the local state and global state of each drone base station in the scene are obtained to generate complete state information of each drone, including:

[0043] The updated scene information is mapped to each global grid and local grid according to the location of the user and the drone base station, and the quantity statistics and average value calculation are performed to generate the local state and global state of the drone base station; then the global state is combined with the local state of the current drone base station to obtain the complete state of each drone base station.

[0044] The advantages of the present invention compared with the prior art are:

[0045] 1) Aiming at the highly dynamic and flexible drone base station networking communication scenario, the present invention constructs a drone base station position adjustment and sub-band optimization problem model while supporting user movement and switching, considers the multi-dimensional network service utility including user transmission rate, drone base station energy consumption, and user coverage as optimization targets, and designs a joint optimization scheme of position adjustment based on deep reinforcement learning and sub-band allocation based on graph coloring. Compared with the existing solutions of optimizing position alone, optimizing sub-band allocation alone, or optimizing sub-band allocation in a preset trajectory, the proposed scheme can make better use of the flexibility of drone networking and effectively achieve further improvement of network service performance.

[0046] 2) In view of the high number of iterations and limited performance of traditional location adjustment schemes, the scene restrictions of intelligent solutions, and the difficulty in obtaining the precise location of users, this invention proposes an intelligent location adjustment scheme based on deep reinforcement learning, designs a corresponding distributed Markov decision process, and trains deep reinforcement learning agents to optimize the base station location. This scheme can support users to move and switch during network operation; at the same time, in the reward design, it can fully consider the multi-factor network service utility indicators of user transmission rate, drone energy consumption and user coverage to avoid the problem of limited optimization performance; in addition, in the state design, the gridded global and local observation scene information is considered as the input for the intelligent agent's decision-making, and accurate user coordinate information is not required, allowing a small error in the user's location information without affecting the decision-making effect, which has better feasibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is an overall flow chart of the method for jointly optimizing the location of the UAV base station and the sub-band provided by the present invention.

[0048] Figure 2 It is a single sub-band allocation iteration process of a UAV base station of the UAV base station position and sub-band joint optimization method provided by the present invention.

[0049] Figure 3 It is a general execution flow chart of sub-band allocation of the method for joint optimization of UAV base station location and sub-band provided by the present invention.

[0050] Figure 4 It is an intelligent base station position adjustment algorithm process based on deep reinforcement learning for the drone base station position and sub-band joint optimization method provided by the present invention.

[0051] Figure 5 It is a network service performance curve diagram of the UAV base station location and sub-band joint optimization method provided by the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention.

[0053] A joint optimization method for UAV base station location and sub-band allocation, such as Figure 1 As shown, including but not limited to the following steps:

[0054] Step 1, decision item judgment for the current time step. According to the current position adjustment count, determine whether the number of drone base station position adjustment optimizations that have been executed after the last sub-band allocation optimization has reached the set number. After reaching the set number, perform the sub-band optimization allocation in step 2 and restart the position adjustment count. By judging the decision items, multiple drone base station position adjustment optimizations are inserted into a longer sub-band allocation optimization interval. The sub-band allocation optimization cycle is made longer to reduce the strong interference of the drone base station to other users and avoid the overhead caused by frequent reallocation; the drone base station position adjustment optimization cycle is shorter to dynamically adapt to user movement and topology changes.

[0055] The UAV base station networking communication scenarios considered include Base Stations and Ground users , users and base stations can be Move within the 3D cube area.

[0056] Assume the maximum running time of the network is , with time intervals Divide the running time into time steps, that is, the present invention uses The scene is observed and adjusted for time step intervals. t ∈ [ 1 , t max ] , the coordinates of all users and base station locations are expressed as q u ( t ) = [ q u , 1 ( t ), q u , 2 ( t ),..., q u , K ( t )] and q v ( t ) = [ q v , 1 ( t ), q v , 2 ( t ),..., q v , M ( t )] ,in, q u , k ( t ) = [ x u , k ( t ), y u , k ( t ), z u , k ( t )] , q v , m ( t ) = [ x v , m ( t ), y v , m ( t ), z v , m ( t )] Respectively and The location coordinates of .

[0057] The base stations in the scenario are networked on the same frequency, so all base stations can work on the same frequency band. Sub-band , the bandwidth of each sub-band is Each base station uses a sub-band to provide services to users. Vector f ( t ) = [ f 1 ( t ), f 2 ( t ), … . f M ( t )] Indicates that all base stations are at time step Sub-band allocation scheme.

[0058] After obtaining the adjustment items to be performed in the current time step, the scene information update operation is performed. Specifically, calculate whether a LoS link is formed; further calculate the channel gain and SINR between the drone base station and the user; calculate the user switching according to the Max-SINR criterion, that is, let the user switch to the drone base station with the largest SINR; calculate the user transmission rate.

[0059] When the drone in the scene acts as an aerial base station to provide services to ground users, there is a high probability that a LoS link will be formed, which will affect the channel quality between the service users and the interference intensity between non-service users. With drone base station The LoS link probability is expressed as:

[0060] P km LoS ( θ km ( t )) = 1 1 + a exp( − b [ θ km ( t ) − a ]) (1)

[0061] in, , is a parameter affected by environmental factors; for = The inclination angle between the ground user and the drone base station at the moment. Therefore, the LoS link probability is a function of the inclination angle between the user and the drone. At the same time, the link between the user and the base station has different channel gains depending on whether it is line-of-sight or non-line-of-sight. Then the user With drone base station The channel gains of the LoS link and the non-LoS link between them are expressed as:

[0062] (2)

[0063] in, express Moment User With drone base station The three-dimensional distance of is the system carrier frequency; is the speed of light; and They represent the additional factors of LoS link and non-LoS link respectively, which are affected by environmental factors. Therefore, the channel gain between the user and the base station is expressed as H ( t ) = [ h km ( t )] K × M ,in Indicates that at time step , and The channel gain between or Since the interval between two adjacent moments is The relative position relationship between the ground user and the UAV and the environmental factors change less, so it can be considered that the line-of-sight or non-line-of-sight nature of the link is constant over a certain period of time. Keep unchanged, while maintaining Then with probability Re-determine the new link properties.

[0064] In downlink transmission, the relationship between the user and the base station is expressed by C ( t ) = [ c km ( t )] K × M express, Indicates that at time step , Access ,otherwise Among them, there are .

[0065] When the relative position between the user and the base station is constantly changing, the user will make a handover decision at each time step and choose to switch to the base station with the largest SINR. yes time Access The SINR is expressed as:

[0066] (3)

[0067] in Transmit power for the drone base station; is the noise power; for time Access The interference signal power received at time is expressed as:

[0068] (4)

[0069] Among them, when hour ,otherwise .

[0070] The minimum distance between the user and the access base station should be SINR can maintain normal communication connection. hour, Access ,like , then it is regarded as In the uncovered state, no service is available, the transmission rate is 0, and no bandwidth resources. Otherwise, Consumable bandwidth resources for data transmission. At time step The duration of time when it is in the uncovered state is ,whenever To regain service, .

[0071] If the user service adopts the full-buffer service model, then Access The theoretical transmission rate that can be obtained is estimated using Shannon's formula, expressed as:

[0072] (5)

[0073] in, Represents the time step Inside The number of users served.

[0074] Step 2, perform the sub-band allocation optimization of the UAV base station. In this embodiment, according to the result of the adjustment item judgment in step 1, if the sub-band allocation optimization needs to be performed at the current time step, then steps 201 to 203 are performed, otherwise the UAV base station position adjustment of step 3 is performed.

[0075] In this embodiment, the interference graph is assumed to be , used to represent the interference relationship between base stations, where the drone base station As the vertices of the graph. is the interference threshold. When the following conditions are met, at the vertex An edge is constructed between the two base stations, which is considered to be an interference relationship between the two base stations.

[0076] (6)

[0077] Based on the dynamic graph, the coloring problem can be formulated as follows:

[0078] (7)

[0079] in, Indicates the frequency allocation plan Under the conditions of drone base station The frequency of use of all vertices connected by edges is Are different, otherwise As can be seen from formula (7), the graph coloring-based base station sub-band allocation problem aims to find a sub-band allocation scheme so that two UAV base stations that are susceptible to interference from adjacent UAV base stations do not use the same sub-band as much as possible, thereby reducing interference to service users.

[0080] According to the above sub-band allocation problem model, the overall process of the solution is as follows: Figure 2As shown. and Respectively represent the current cycle number and the maximum cycle number. The algorithm execution flow in one cycle is as follows Figure 3 As shown. Among them, Indicates base station The current loop number of the execution algorithm; Reset cycle for fixed state; is the convergence state detection period; Adjusting parameters for the probability of random selection of sub-bands; for The specific sub-band allocation steps are as follows:

[0081] Step 201, initialize the optimization parameters of the sub-band allocation of the drone base station and construct an interference relationship diagram. and Calculate the current channel state information . Based on the relationship between the user and the drone base station , get the drone base station The user set of the service is According to formula (6), the current channel state is initialized and frequency band allocation scheme Interference graph of For each drone base station, initialize , .make P G , m = [ p G , 1 , p G , 2 ,..., p  G , N f ] Indicates drone base station The probability of selecting a certain sub-band.

[0082] Step 202, each drone base station independently runs the sub-band allocation algorithm. First, if it reaches the fixed state reset period, clear The fixed state of ; Then, if If it is not in a fixed state, then according to the current According to probability Select a sub-band After receiving and satisfactory status detection: if The sub-bands used by all drone base stations constituting the edge are are different, called Reach a satisfactory state; if all drone base stations are in a satisfactory state, the frequency allocation scheme converges; if converged, the loop ends; otherwise, When it is not in a fixed state, check whether it is in a satisfactory state. If the current sub-band is satisfied, The corresponding probability , and further set is a fixed state, allowing it to maintain its own sub-band selection within a certain number of cycles; otherwise, Will adjust ,reduce The probability of being selected. Before the end of the loop, update the number of loops ;

[0083] When all drone base stations complete the sub-band selection in turn, the sub-band allocation scheme changes from Adjust to , complete a cycle, update When the maximum number of cyclic adjustments is reached or the algorithm converges, the sub-band allocation cycle of the current time step ends and the sub-band allocation scheme is obtained. In one cycle, all drone base stations in the scene execute the allocation algorithm in sequence. If the result converges, the cycle is stopped to obtain the allocation solution. , otherwise continue looping until convergence or the number of loops is exhausted.

[0084] Step 203, using formulas (2)-(5) of the scene information update in step 1 to simulate the calculation usage scheme The total user transmission rate at that time. If it is greater than the use plan The transmission rate at that time is determined As the final solution, , otherwise let . The UAV base station is instructed to change to a new sub-band according to the sub-band allocation scheme.

[0085] Step 3: Execute the adjustment of the drone base station position. In this embodiment, the drone base station position adjustment solution is based on The time steps are regarded as an episode, and the optimization goal is to maximize the total reward within an episode.

[0086] In order to facilitate the agent to obtain information in the environment, The scenes are divided into a square global grid of equal size and local grids of the same size, where The drone base station can observe the information in all global grids and a small number of local grids near itself. The grids are used as the side length, and a square area is defined as the local observation range of the UAV base station. Then the number of local grids that each UAV base station can observe is Based on the above grid division scheme, the Markov decision process for adjusting the location of the UAV base station is further constructed as follows.

[0087] State Space and observation space : At time step , drone base station in the scene Global status can be obtained and the local state within its own local observation range , get the status . The global state is expressed as:

[0088] (8)

[0089] in, , and Respectively represent the number of drone base stations, the number of users, and the number of users in the uncovered state in each global grid; is the average user transmission rate per unit bandwidth in each global grid; and They represent the maximum user coverage duration and the average user coverage duration in each global grid respectively; Indicates the number of users served by each drone base station.

[0090] The observed local state is expressed as:

[0091] (9)

[0092] in, Indicates that each global grid The number of service users; , , , and Respectively represent the number of users in each local grid, the number of drone base stations, the number of served users, the average number of users served by other drones, and the number of users in the uncovered state; is the average user transmission rate per unit bandwidth in each local grid; For each user in the local grid The average achievable transmission rate per unit bandwidth; and They represent the maximum user coverage duration and the average user coverage duration in each local grid, respectively.

[0093] So the time step The state space at the moment can be composed of the local state and global state of each drone base station, expressed as:

[0094] s ( t ) = [ o a ( t ), o p , 1 ( t ), o p , 2 ( t ),..., o p , M ( t ) ] (10)

[0095] Action Space :Assume that the base station moves a fixed distance and only makes decisions on the moving direction, then is a discrete sequence of directions in which the base station can move, where is the number of movable directions. The action of making a decision is recorded as .

[0096] Reward Function : is the time step , In Status Take Action Then transfer to The reward obtained. To ensure that the time step defined in 1.8 The utility function of is consistent with The reward function can be expressed as:

[0097] (11)

[0098] in, Reward weight for local transmission rate; for The transmission rate of the service users; is the local energy consumption penalty weight; is the global non-coverage penalty weight; reward weight for global transfer rate; , is the service user transmission rate of all drone base stations; is the total duration of all users in the uncovered state; is the energy consumption of the drone base station when it moves. Specifically, when the drone base station serves ground users, communicating, hovering, or moving will all consume energy. Considering that the energy consumption generated by communication is relatively small, the present invention ignores the energy consumption of communication and only considers the energy consumption generated by the hovering and movement of the drone base station, which is expressed as:

[0099] (12)

[0100] in, express At time step speed, , are blade profile power and derived power, respectively, is the tip speed of the UAV rotor blade, is the average rotor induced speed in the hovering state, is the fuselage drag ratio, is the air density, is the rotor solidity, is the rotor disk area.

[0101] From the perspective of MDP optimization cumulative reward, the cumulative reward finally optimized by the MDP is the cumulative reward of all drone base stations at each time step in an adjustment cycle, which is used as the optimization target of drone base station position adjustment and is expressed as:

[0102] (13)

[0103] in, is the time step Rewards for adjusting the positions of all drone base stations; and are the cumulative weight of non-coverage penalty and the cumulative weight of reward respectively; let the first sub-band allocation adjustment and subsequent The second position adjustment is one adjustment round, then Indicates the number of adjustment rounds that have been completed since the starting time step.

[0104] According to the above Markov decision process, the decision-making scheme for adjusting the location of the drone base station based on deep reinforcement learning is as follows:

[0105] Step 301: Each UAV base station obtains the local state and global state of each UAV base station in the scene according to the set observation range and the local state definition and global state definition in formulas (8) and (9), and combines the global state with the local state to obtain the complete state information of each UAV base station. o m ( t ) = [ o a ( t ), o p , m ( t )] .

[0106] Step 302, each drone base station is based on At the same time, it makes its own position adjustment decision. Input the state information of each drone base station in the scene into its own agent to obtain the Q value of each action. This value represents the value of each action in the current state. Select the action with the largest Q value and obtain the output result of the position adjustment. .

[0107] Step 303: After obtaining the optimized position adjustment output, the position of each drone base station is adjusted based on the optimized position adjustment result. All adjustment operations of the current time step are completed and transferred to the next time step.

[0108] In this embodiment, before executing steps 301 to 303 to optimize the position adjustment of the drone base station, it is necessary to first perform agent training to achieve the best position adjustment effect. After multiple trainings, it is observed that the optimization performance of the agent reward function is improved and the training is stopped after it remains stable. Then, the trained agent can be used to make position adjustment decisions. The intelligent base station position algorithm process and network training method based on deep reinforcement learning are as follows: Figure 4 As shown. Among them, is the random exploration probability; is the current value network parameter; is the target network parameter; is the target network update interval. The drone base station position adjustment agent training process is based on steps 301 to 303, and its specific execution steps are as follows:

[0109] Before executing step 301, the centralized agent sends the latest agent network parameters to all drone base stations, so that the local agent of the drone base station can use the latest network parameters to make decisions.

[0110] After step 301, it is also necessary to simultaneously obtain the status information observed by each drone base station. After that, perform the following operations: If the current time step is not the first time step of the episode, then further obtain the reward corresponding to the adjustment operation of the previous time step . The state, action, reward and current state of the previous time step Combine to get a complete experience sample . Upload the experience samples of all drone base stations to the centralized agent and store them in the experience replay.

[0111] In step 302, the action mode is selected to be modified to be based on probability Randomly select actions with probability based on And the agent outputs the result and selects the action.

[0112] After step 302, it is also necessary to and Temporarily store it to get the complete experience sample for the next time step.

[0113] After executing step 303, when all drone base stations have completed the position adjustment decision, the centralized agent samples experience samples from the experience playback and performs network training, updates the current value network parameters, and Each decision updates the target value network parameters once.

[0114] Through the embodiment provided by the present invention, the decision-making project is judged at each running time step to obtain the adjustment project that needs to be performed at the current time step. If it is necessary to optimize the sub-band allocation of the drone base station, the optimization parameters of the sub-band allocation of the drone base station are initialized and the interference relationship diagram is constructed; the sub-band selection, state check and selection probability update are executed cyclically until convergence or the preset number of cycles is reached; the sub-band allocation of the drone base station is adjusted based on the optimized sub-band selection; if the sub-band allocation optimization is not required or the sub-band allocation optimization has been completed, the drone position adjustment optimization is performed. The Markov decision process for the position adjustment of the drone base station is constructed; the local state and global state of each drone base station in the scene are obtained, and the complete state information of each drone base station is generated; the dual-depth Q network intelligent agent is trained to make decisions, and the state information of each drone base station in the scene is input into its own intelligent agent to obtain the output result of the position adjustment; the position of the drone base station is adjusted based on the optimized position adjustment result; the joint optimization method of the drone base station position and sub-band allocation in the related technology is solved, which has the technical problems of high number of iterations, limited performance, many scene restriction factors, and difficult acquisition of the user's precise position, which affects the communication effect. The present invention implements a joint optimization method of position adjustment based on deep reinforcement learning and sub-band allocation based on graph coloring to achieve better communication system service performance.

[0115] In order to verify the performance of the method proposed in the present invention, a relevant simulation environment was built to simulate the movement of UAV base stations and users, and to perform simulated calculations of wireless channel quality and simulated operation of system service processes.

[0116] The following describes the joint optimization method of the drone base station location and sub-band allocation according to an embodiment of the present invention in conjunction with optional examples.

[0117] (1) Base station location adjustment agent training method

[0118] 600 scenarios were deployed as training sets, 50 of which were selected for each round of training, and 12 rounds of training were conducted continuously. The scenarios in each round were run for 5 episodes in a row, and the sub-band allocation adjustment was interspersed before each episode. In order to speed up convergence, in 4 of the training rounds, the greedy algorithm was used to directly obtain the action with the maximum reward, and the experience sample data of the greedy algorithm was used for agent training. In addition, 5 additional scenarios were deployed as test sets for simulation testing.

[0119] (2) Initial location planning of drone base stations

[0120] In the simulation scenario, the drone base station location, user location, and sub-band allocation are all randomly generated.

[0121] (3) User mobility model

[0122] The maximum moving speeds of users and base stations are and At each time step, the user in the scene moves along a random initial direction, and each interval Random new movement direction.

[0123] (4) Network service performance calculation

[0124] As users move, the positional relationship between the base station and the user changes. In particular, as the line-of-sight and non-line-of-sight properties of the link change, the channel quality between the base station and the service user, and the interference relationship between the base station and the non-service user will change, affecting the service performance of the system. , continuously adjust the location coordinates of the base station , the sub-band allocation vector of the base station f ( t ) = [ f 1 ( t ), f 2 ( t ), … . f M ( t )] , to achieve improved service performance.

[0125] In summary, the present invention defines a network service performance utility function based on user transmission rate and comprehensively considering the energy consumption of the drone base station and user coverage as a measurement indicator of network service performance, which is expressed as:

[0126] (14)

[0127] (5) Simulation parameters and results:

[0128] Table 1 Scenario simulation parameters

[0129]

[0130] Figure 5 It is a network service performance curve diagram of the joint optimization method of the UAV base station location and sub-band allocation provided by the present invention. Figure 5 The network service performance results of the three algorithms in the current test scenario are statistically analyzed, namely, cluster position adjustment and fixed sub-band allocation (labeled as "cluster position adjustment"), cluster position adjustment and graph coloring sub-band allocation (labeled as "cluster position + sub-band adjustment"), and the base station position adjustment based on deep reinforcement learning and the joint optimization scheme of graph coloring sub-band allocation proposed in the present invention (labeled as "intelligent position + sub-band adjustment"). It can be seen that the design scheme of the present invention can improve the network service performance and provide good communication services.

[0131] It can be understood that the present invention proposes a joint optimization method for adjusting the position of a drone base station and sub-band allocation in a drone base station communication network, which realizes the dynamic adjustment of the position and sub-band of the drone base station and improves the quality of network service; as time goes by, the position of the user changes, and at the same time, the channel gain between the user and the base station will also change accordingly. At the same time, with the multiple position adjustment operations of the drone base station, the interference relationship between the drone base station and the user is also changing. Compared with the base station position and sub-band allocation optimization method used in this article, other methods have limited improvements in network service performance. Compared with other methods, the method used in the present invention can obtain better network service performance in a changing environment and improve the communication quality of the user.

Claims

1. A joint optimization method for UAV base station location and sub-band allocation, characterized in that: The following steps are involved: Step 1: According to the position adjustment count value at the current time step, determine whether the number of UAV base station position adjustment optimizations that have been performed after the last sub-band allocation optimization has reached the set number of times. If the number of times has reached the set number, execute step 2 once and restart the position adjustment count. Otherwise, execute step 3. Step 2, initialize the optimization parameters of the sub-band allocation of the UAV base station and construct the interference relationship diagram; Then, the frequency band selection probability of the UAV base station is adjusted cyclically until convergence or the preset number of cycles is reached, and an optimized sub-band allocation scheme is obtained; Adjust the sub-band allocation of the UAV base station based on the optimized sub-band allocation scheme; Step 3: According to the set grid size, the scene is evenly divided into multiple square grids, where the global grid size is larger than the local grid size; and the state space, action space and reward function in the Markov decision process of the drone base station position adjustment are constructed; the reward function is designed based on the comprehensive factors of user transmission rate, user coverage and drone base station energy consumption, and the optimization cumulative reward is taken as the optimization goal; Then, the local state and global state of each UAV base station in the scene are obtained to generate the complete state information of each UAV; each UAV inputs its complete state information into its own intelligent body to obtain the position adjustment result; and the position of the UAV base station is adjusted based on the position adjustment result.

2. The method for joint optimization of UAV base station location and sub-band allocation according to claim 1, characterized in that: After step 1, also include: Update the scene information, including calculating whether a line-of-sight link is formed; then calculate the channel gain and signal-to-interference-and-noise ratio (SINR) between the drone base station and the user; calculate the user switching based on the Max-SINR criterion, that is, let the user switch to the drone base station with the largest SINR, and calculate the user transmission rate.

3. The method for joint optimization of UAV base station location and sub-band allocation according to claim 2, characterized in that: Step 2 specifically includes the following steps: Step 201, initialize the optimization parameters of the sub-band allocation of the drone base station, calculate the interference intensity and construct an interference relationship diagram according to the channel state between the user and the drone base station; Step 202, cyclically executing sub-band selection, state checking and selection probability updating until a convergence condition is met or a preset number of cycles is reached, and an optimized sub-band allocation scheme is obtained; Step 203, compare the optimized sub-band allocation scheme with the sub-band allocation scheme before optimization, and select a scheme with a higher total user transmission rate to adjust the sub-band allocation of the drone base station.

4. The method for joint optimization of UAV base station location and sub-band allocation according to claim 1, characterized in that: The state space, action space and reward function of the Markov decision process for adjusting the position of the drone base station in step 3 are as follows: The state space includes global state and local state. The global state at the moment is expressed as: ; In the formula, , and Respectively represent the number of drone base stations, the number of users, and the number of users in the uncovered state in each global grid; represents the average user transmission rate per unit bandwidth in each global grid; and They represent the maximum user coverage duration and the average user coverage duration in each global grid respectively; Indicates the number of users served by each drone base station; The local state at the moment is expressed as: In the formula, Represents the drone base station in each global grid The number of service users; , , , and Respectively represent the number of users in each local grid, the number of drone base stations, the number of served users, the average number of users served by other drones, and the number of users in the uncovered state; is the average user transmission rate per unit bandwidth in each local grid; For each user in the local grid The average achievable transmission rate per unit bandwidth; and They represent the maximum user coverage duration and the average user coverage duration in each local grid respectively; Time step The state space at time It consists of the local state and global state of each drone base station, expressed as: ; In the formula, M is the number of drone base stations; Action Space :Assume that the base station moves a fixed distance and only makes decisions on the moving direction, then represents a discrete sequence of directions in which the base station can move, where is the number of movable directions, The action of making a decision is recorded as ; Reward Function : is the time step , In Status Take Action Then transfer to The rewards received, The reward function is expressed as: ; In the formula, Reward weight for local transmission rate; for The transmission rate of the service users; is the local energy consumption penalty weight; Base stations for all drones Energy consumption; is the global non-coverage penalty weight; is the total duration of all users in the uncovered state; reward weight for global transfer rate; is the time step The service user transmission rate of all drone base stations.

5. The method for joint optimization of drone base station location and sub-band allocation according to claim 4, characterized in that: The Markov decision process for adjusting the location of the UAV base station is constructed, and its optimization objective design includes the following: The optimization goal of the drone base station position adjustment is to optimize the cumulative reward, that is, the cumulative reward of all drone base stations at each time step within an adjustment cycle, expressed as: ; in, is the time step Rewards for adjusting the positions of all drone base stations; and They are the cumulative weight of non-coverage penalty and the cumulative weight of reward; is the number of position adjustments inserted between two sub-band allocation adjustments; The second position adjustment is one adjustment round, then Indicates the number of adjustment rounds that have been completed since the starting time step; Base station for drones The mobile energy consumption; The optimization target of the UAV base station position adjustment is designed as the weighted sum of all user transmission rates, user uncovered time, and UAV base station energy consumption in multiple position adjustments between two sub-band allocations, where: User transmission rate: the communication transmission rate of the drone base station serving ground users; User uncovered time: When a user receives drone base station service, if the signal-to-noise ratio between the user and the base station is less than the set threshold, the user is judged to be in an uncovered state in the current time step, and the current user's uncovered time increases by one time step interval; if the SINR between the user and the base station is greater than or equal to the set threshold, the current user's uncovered time count is cleared; Drone base station energy consumption: energy consumption generated by the hovering and movement of the drone base station.

6. The method for joint optimization of UAV base station location and sub-band allocation according to claim 1, characterized in that: In step 3, the local and global states of each drone base station in the scene are obtained to generate complete state information for each drone, including: The updated scene information is mapped to each global grid and local grid according to the location of the user and the drone base station, and the quantity statistics and average value calculation are performed to generate the local state and global state of the drone base station; then the global state is combined with the local state of the current drone base station to obtain the complete state of each drone base station.

Citation Information

Patent Citations

  • Cluster communication frequency decision-making method based on a hierarchical matching game

    CN112020021A

  • Unmanned aerial vehicle and ground network spectrum transaction implementation method based on block chain

    CN115866768A