Multi-scenario-based Air Base Station Deployment Method and System

Through the deployment allocation model of reinforcement learning, the location of the air base station is optimized, and the on-demand deployment of drone aerial base stations in different scenarios is solved, efficient communication capacity and coverage supplementation is achieved, and communication needs in hot spots and emergency areas are met.

CN116095696BActive Publication Date: 2025-07-11BEIJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211659446.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-22
Publication Date
2025-07-11
Estimated Expiration
2042-12-22

AI Technical Summary

Technical Problem

In the prior art, the deployment strategy of drone aerial base stations is difficult to achieve high dynamic deployment on demand in different scenarios, and the calculation load and communication load are high, which cannot effectively meet the communication needs of hot spot areas and the coverage of emergency areas.

Method used

Adopt the deployment allocation model based on reinforcement learning, obtain the scene observation status through loops, select the air base station to perform the allocation action, optimize its location to enhance system capacity and coverage, and update the model parameters through the reward mechanism to optimize the deployment strategy.

Benefits of technology

It realizes efficient deployment of air base stations in different scenarios, reduces computing and communication load, improves communication capacity and coverage, and meets the various needs of the ground communication network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116095696B_ABST
    Figure CN116095696B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-scenario-based aerial base station deployment method and system. The method includes: cyclically obtaining the scenario observation status; according to the currently obtained scenario observation status, using a deployment allocation model to select an aerial base station from an intelligent group to perform the allocated action; wherein, the deployment allocation model is updated after evaluating its previous execution action based on the reward obtained by using the selected aerial base station to perform the allocated action when the number of cycles of obtaining the scenario observation status meets a preset number of cycles. The present invention can, according to different scenarios, determine to deploy aerial base stations in hot spots to enhance communication capacity, or deploy multiple aerial base stations in emergency areas for coverage supplementation to restore normal communication, so as to optimize multiple communication metrics; and can, according to different scenarios, automatically adjust the positions of aerial base stations to jointly optimize each communication objective and meet different requirements of the ground communication network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 5G communication technology, and in particular to a method and system for deploying aerial base stations based on multiple scenarios. Background Art

[0002] In recent years, with the continuous deepening of research on drone technology and wireless communication technology, drones have also played an important role in the world's economic growth. Due to the advantages of drones such as miniaturization, low cost, high mobility, easy deployment, ready-to-use and excellent line-of-sight links, the mobile aerial base station deployment (Drone Base Station, DBS) obtained by adding base station functions to drones has greater significance for improving the performance of communication networks. Deploying aerial base stations on demand and ensuring the service quality of users are key technologies in aerial wireless communication networks. The core of aerial base station deployment is aerial location planning. A good location planning solution can improve the service quality of users and reduce interference between aerial base station deployments. At the same time, it can also remove redundant base stations in aerial base stations, reduce operating costs, and reduce energy consumption.

[0003] When the power communication network is in operation, there will be local hot spots, that is, some temporary areas with high user density, and some emergency areas, that is, areas where the original power communication network is paralyzed. Users in hot spots have high capacity demands. Once they exceed the service capacity of the base stations deployed on the ground, communication will be interrupted; in emergency areas, ground base stations are often unable to provide normal communication services due to machine damage or power system failure. In this case, drones equipped with micro communication equipment as air base stations to expand the current network capacity to meet the surging business data demand has become a new solution.

[0004] However, no matter what the application scenario is, the number of drones is always limited. In this case, how to achieve on-demand and highly dynamic deployment of aerial base stations and minimize the computational and communication loads during the deployment strategy optimization process is of great significance to the practical application of aerial networks. Summary of the invention

[0005] The present invention provides an aerial base station deployment method and system based on multiple scenarios, which are used to solve the problem of how to reasonably deploy aerial base stations in the prior art. It can comprehensively optimize multiple communication indicators according to the characteristics of the scenarios, and use reinforcement learning methods to optimize the deployment positions of each drone aerial base station.

[0006] The present invention provides a multi-scenario-based aerial base station deployment method, including: cyclically obtaining the scenario observation status; according to the currently obtained scenario observation status, using a deployment allocation model to select an aerial base station from a smart group to perform the allocated action; wherein, the deployment allocation model is updated after evaluating the previous execution action based on the number of cycles of obtaining the scenario observation status meeting a preset number of cycles and the reward obtained by using the selected aerial base station to perform the allocated action.

[0007] According to the multi-scenario-based aerial base station deployment method provided by the present invention, the step of using a deployment allocation model to select an aerial base station from a smart group to perform the allocated action according to the currently obtained scenario observation status includes: determining a task to be executed according to the currently obtained scenario observation status; the task to be executed includes enhancing the system capacity within a target area to a preset capacity and covering a specified area within the target area; according to the currently obtained scenario observation status, using the deployment allocation model to select an aerial base station from the smart group and allocate corresponding actions, and controlling the selected aerial base station to perform the corresponding allocated actions; obtaining the next scenario observation status and the current reward obtained after the currently selected aerial base station performs the current allocated action, and according to the next obtained scenario observation status, using the deployment allocation model to re-select an aerial base station from the smart group to perform the next allocated action until the task to be executed is completed; wherein, based on the number of cycles of obtaining the scenario observation status meeting a preset number of cycles, using the current reward to evaluate the corresponding previous execution action to update the deployment allocation model.

[0008] According to the multi-scenario-based aerial base station deployment method provided by the present invention, the step of updating the deployment allocation model by evaluating the corresponding previous execution action based on the number of cycles of obtaining the scenario observation status meeting a preset number of cycles and using the current reward includes: randomly selecting the allocated action performed by the aerial base station, the reward obtained by performing the allocated action, and the scenario observation status before and after performing the allocated action based on the number of cycles of obtaining the scenario observation status meeting a preset number of cycles; obtaining a loss function according to the allocated action, the reward obtained by performing the allocated action, and the scenario observation status before and after performing the corresponding allocated action, and updating the weight parameters of the action network of the deployment allocation model based on minimizing the loss function; and updating the weight parameters of the evaluation network of the deployment allocation model according to the allocated action, the reward obtained by performing the allocated action, and the scenario observation status before and after performing the corresponding allocated action, and combining the sampling policy gradient.

[0009] A method for deploying an aerial base station based on multiple scenarios provided by the present invention, the obtaining of the current reward after the selected aerial base station executes the current allocation action includes: based on the currently selected aerial base station executing the allocation action, obtaining the corresponding coverage rate, system capacity, and energy consumption; based on the coverage rate, the system capacity, and the current energy consumption, obtaining the corresponding reward.

[0010] A method for deploying an aerial base station based on multiple scenarios provided by the present invention, the obtaining of the coverage rate based on the currently selected aerial base station executing the allocation action includes: based on the allocation action executed by the currently selected aerial base station, determining the current scenario observation state corresponding to the currently selected aerial base station; based on the current scenario observation state, obtaining the current number of ground users; based on the allocation action executed by the currently selected aerial base station, obtaining the number of users covered by the currently selected aerial base station; based on the current number of ground users and the number of users covered by the currently selected aerial base station, obtaining the coverage rate corresponding to the currently selected aerial base station.

[0011] A method for deploying an aerial base station based on multiple scenarios provided by the present invention, the obtaining of the system capacity based on the currently selected aerial base station executing the allocation action includes: according to the currently selected aerial base station, obtaining the corresponding transmit power and thermal noise power; according to the transmit power, the thermal noise power, and the total propagation loss corresponding to the currently selected aerial base station obtained in advance, obtaining the signal-to-noise ratio of the wireless signal received by the user; according to the signal-to-noise ratio, obtaining the user data transmission rate corresponding to the currently selected aerial base station; according to the user data transmission rate, obtaining the system capacity corresponding to the currently selected aerial base station.

[0012] A method for deploying an aerial base station based on multiple scenarios provided by the present invention, the obtaining of the energy consumption based on the currently selected aerial base station executing the allocation action includes: based on the currently selected aerial base station executing the allocation action, obtaining the equipment movement energy consumption; based on the equipment movement energy consumption and the signal transmission energy consumption obtained in advance, obtaining the energy consumption.

[0013] The present invention also provides a system for deploying an aerial base station based on multiple scenarios, including: a state acquisition module that circularly acquires the scenario observation state; a base station deployment module that, according to the currently acquired scenario observation state, uses a deployment allocation model to select an aerial base station from a smart group to execute the allocation action; wherein, the deployment allocation model is updated after evaluating and updating the previous execution action based on the number of circular acquisitions of the scenario observation state meeting a preset number of circular acquisitions and the reward obtained by using the selected aerial base station to execute the allocation action.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned multi-scenario-based aerial base station deployment methods are implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned multi-scenario-based aerial base station deployment methods are implemented.

[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned multi-scenario-based aerial base station deployment methods are implemented.

[0017] The multi-scenario-based aerial base station deployment method and system provided by the present invention obtain the scene observation status by looping, so as to determine to deploy aerial base stations in hot spots to enhance communication capacity according to different scenarios, or deploy multiple aerial base stations in emergency areas for coverage supplementation to restore normal communication, so as to optimize multiple communication metrics; by using the deployment allocation model, select aerial base stations and allocate actions to control the movement of aerial base stations based on the allocated actions, so that the deployed aerial base stations reach the optimal positions to meet the different needs of the ground communication network. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is one of the flow diagrams of the multi-scenario-based aerial base station deployment method provided by the present invention;

[0020] Figure 2 is another flow diagram of the multi-scenario-based aerial base station deployment method provided by the present invention;

[0021] Figure 3 is the structural diagram of the multi-scenario-based aerial base station deployment system provided by the present invention;

[0022] Figure 4 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Figure 1 The flowchart of a method for deploying aerial base stations based on multiple scenarios according to the present invention is shown. The method includes:

[0025] S11. Continuously obtain the scenario observation status;

[0026] S12. According to the currently obtained scenario observation status, use the deployment allocation model to select an aerial base station from the intelligent group to perform the allocated action; wherein, the deployment allocation model is updated after evaluating the previous execution action based on the reward obtained by using the selected aerial base station to perform the allocated action when the number of cycles of obtaining the scenario observation status meets the preset number of cycles.

[0027] It should be noted that S1N in this specification does not represent the sequence of the method for deploying aerial base stations based on multiple scenarios. The following specifically combines Figure 2 to describe the method for deploying aerial base stations based on multiple scenarios of the present invention.

[0028] Step S11. Continuously obtain the scenario observation status.

[0029] It should be noted that the scenario observation status includes the three-dimensional position information of all aerial base stations in the target area and the two-dimensional position information of all users in the target area. The scenario observation status can be represented as s n (t) = {b n (t), g n (t)}, where s n (t) represents the scenario observation status obtained at the t-th moment, b n (t) represents the three-dimensional positions of all aerial base stations, and g n (t)} represents the two-dimensional position information of all users.

[0030] Step S12. According to the currently obtained scenario observation status, use the deployment allocation model to select an aerial base station from the intelligent group to perform the allocated action; wherein, the deployment allocation model is updated after evaluating the previous execution action based on the reward obtained by using the selected aerial base station to perform the allocated action when the number of cycles of obtaining the scenario observation status meets the preset number of cycles.

[0031] In this embodiment, with reference to Figure 2, according to the currently acquired scene observation state, using the deployment allocation model, select the aerial base station from the intelligent group to perform the allocated actions, including:

[0032] SA1, determine the tasks to be performed according to the currently acquired scene observation state; the tasks to be performed include enhancing the system capacity in the target area to a preset capacity and covering a designated area in the target area.

[0033] It should be noted that by determining the tasks to be performed in the current target area in advance based on the currently acquired scene observation status, the aerial base station can be controlled to move according to the tasks to be performed, so that the aerial base station can be deployed in hot spots to enhance the communication capacity, or multiple aerial base stations can be deployed in emergency areas to supplement coverage to restore normal communication, thereby deploying the aerial base stations to the optimal position to meet the different needs of the ground communication network.

[0034] SA2, based on the currently acquired scene observation status, uses the deployment allocation model to select an aerial base station from the intelligent group and assign corresponding actions, and controls the selected aerial base station to perform the corresponding assigned actions.

[0035] In this embodiment, the actions are represented as: Among them, a n (t) represents the assigned action, v(t) represents the moving speed, Indicates the yaw angle, and index indicates the aerial base station selected to perform the action.

[0036] SA3, obtains the next scene observation state and the current reward obtained after the currently selected aerial base station executes the current assigned action, and based on the next acquired scene observation state, uses the deployment allocation model to reselect the aerial base station from the intelligent group to execute the next assigned action until the task to be executed is completed; wherein, based on the number of cycles for obtaining the scene observation state meeting the preset number of cycles, the current reward is used to evaluate the corresponding previous execution action to update the deployment allocation model.

[0037] Specifically, the current reward obtained after the selected aerial base station performs the current allocation action is obtained, including: performing the allocation action based on the currently selected aerial base station to obtain the corresponding coverage, system capacity and energy consumption; based on the coverage, system capacity and energy consumption, to obtain the corresponding reward.

[0038] First, perform an allocation action based on the currently selected aerial base station to obtain the coverage rate, including: determining the current scene observation state corresponding to the currently selected aerial base station based on the allocation action performed on the currently selected aerial base station; obtaining the current number of ground users based on the current scene observation state; obtaining the number of users covered by the currently selected aerial base station based on the allocation action performed by the currently selected aerial base station; and obtaining the coverage rate corresponding to the currently selected aerial base station based on the current number of ground users and the number of users covered by the currently selected aerial base station.

[0039] In this embodiment, the coverage rate is expressed as: where α t represents the coverage rate at time t, N represents the current number of ground users, and N' represents the number of users covered by the currently selected aerial base station.

[0040] It should be added that when the wireless signal power received by the user terminal is too low, the quality of the communication service will be severely affected. Only when the power P r received by the user is greater than P rmin will the user be considered covered by the drone aerial base station.

[0041] Assume that the transmission power allocated by the drone aerial base station to each user is fixed. Then only when the straight-line distance between the user terminal and the drone aerial base station is less than a certain threshold d min will the path loss of the wireless signal be less than a certain threshold PL min and the power received by the user can exceed P rmin . The user is covered by the aerial base station and can obtain a wireless communication service with better quality of service. It should be noted that after each movement of the aerial base station, it is necessary to recalculate the distance between it and each ground user to determine the number of users it covers.

[0042] In addition, perform an allocation action based on the currently selected aerial base station to obtain the system capacity, including: obtaining the corresponding transmission power and thermal noise power according to the currently selected aerial base station; obtaining the signal-to-noise ratio of the wireless signal received by the user according to the transmission power, thermal noise power, and the total propagation loss corresponding to the currently selected aerial base station obtained in advance; obtaining the user data transmission rate corresponding to the currently selected aerial base station according to the signal-to-noise ratio; and obtaining the system capacity corresponding to the currently selected aerial base station according to the user data transmission rate.

[0043] It should be noted that the system capacity is expressed as:

[0044]

[0045] where C represents the system capacity and c k represents the user data transmission rate.

[0046] Furthermore, c k = Blog2(1 + SNR k ), where B represents the system spectral bandwidth, and SNR k is the signal-to-noise ratio of the wireless signal received by the user.

[0047] Furthermore, where P k represents the transmit power, P N represents the thermal noise power, and PL represents the total propagation loss.

[0048] It should be noted that the UAV aerial base station provides communication services for ground user terminals during high-altitude flight. The communication link can be regarded as the air-to-ground channel between the UAV and the ground terminal. Among them: The wireless signal emitted by the UAV aerial base station propagates in free space at high altitude, which belongs to line-of-sight propagation, and the corresponding loss is the path loss; When the wireless signal approaches the ground, it needs to penetrate the buildings in the city, and the wireless signal attenuates again. This part is non-line-of-sight propagation, and the corresponding loss is the additional loss. Therefore, obtaining the total propagation loss corresponding to the currently selected aerial base station includes: obtaining the line-of-sight propagation loss based on the straight-line distance, the speed of light, the carrier frequency, and the path loss between the selected aerial base station and the ground user; obtaining the non-line-of-sight propagation loss based on the straight-line distance, the speed of light, the carrier frequency, and the additional loss between the selected aerial base station and the ground user; and obtaining the total propagation loss according to the line-of-sight propagation loss and the non-line-of-sight propagation loss.

[0049] In this embodiment, the line-of-sight propagation loss is expressed as:

[0050]

[0051] where PL LoS represents the line-of-sight propagation loss, f c represents the carrier frequency, d represents the straight-line distance between the selected aerial base station and the ground user, c represents the speed of light, and η LoS represents the path loss. It should be noted that η LoS depends on environmental factors.

[0052] The non-line-of-sight propagation loss is expressed as:

[0053]

[0054] where PL NLoS represents the non-line-of-sight propagation loss, fc c represents the carrier frequency, d represents the straight-line distance between the selected aerial base station and the ground user, c represents the speed of light, and η NLoS represents the additional loss. It should be noted that η NLoS depends on environmental factors.

[0055] The total propagation loss is expressed as:

[0056] PL = P LoS × PL LoS + P NLoS × PL NLoS

[0057] Among them, PL represents the total propagation loss, and P LoS represents the probability of the wireless signal performing line-of-sight propagation, and PL LoS represents the line-of-sight propagation loss, and P NLoS represents the probability of the wireless signal performing non-line-of-sight propagation, and PL NLoS represents the non-line-of-sight propagation loss.

[0058] The probability of the wireless signal performing line-of-sight propagation is expressed as:

[0059]

[0060] Among them, P LoS represents the probability of the wireless signal performing line-of-sight propagation, a and b represent environmental parameters, and θ represents the elevation angle between the ground user and the aerial base station.

[0061] The probability P of the wireless signal performing non-line-of-sight propagation NLoS , is expressed as:

[0062] P NLoS = 1 - P LoS

[0063] In addition, based on the currently selected aerial base station, an allocation action is performed to obtain the energy consumption, including: based on the currently selected aerial base station, an allocation action is performed to obtain the device movement energy consumption; based on the device movement energy consumption and the pre-obtained signal transmission energy consumption, the energy consumption is obtained.

[0064] It should be noted that the energy consumption is expressed as:

[0065] P = P t + P e

[0066] Among them, P represents the energy consumption, P t represents the signal transmission energy consumption of the aerial base station communication module, and P e represents the device movement energy consumption.

[0067] Furthermore, the device movement energy consumption is expressed as:

[0068]

[0069] Among them, P e represents the device movement energy consumption, (u x , u y) represents the position coordinates to which the selected aerial base station moves after performing the assigned action, (u x0 , u y0 ) represents the position coordinates of the selected aerial base station before performing the assigned action, e1 represents the energy consumption of the UAV aerial base station moving horizontally by one unit, δ represents a parameter, and the specific value is determined by the UAV. δe1 represents the energy consumption of the UAV aerial base station moving vertically by one unit, u z represents the distance that the aerial base station moves in the vertical direction.

[0070] Secondly, based on the coverage rate, system capacity, and current energy consumption, corresponding rewards are obtained, including: based on the coverage rate, obtaining the coverage rate reward of the selected aerial base station for performing the current assigned action and the previous action; based on the system capacity, obtaining the system capacity reward of the selected aerial base station for performing the current assigned action and the previous action; based on the energy consumption, obtaining the energy consumption reward; based on the coverage rate reward, system capacity reward, and energy consumption reward, obtaining the corresponding reward.

[0071] It should be added that the reward is expressed as:

[0072] r(t) = r c0 (t) * w co + r ca (t) * w ca + r ec (t) * w ec

[0073] Among them, r(t) represents the reward, r co (t) represents the coverage rate reward, r ca (t) represents the system capacity reward, r ec (t) represents the energy consumption reward, w co , w ca and w ec respectively represent the weight parameters of the corresponding rewards, indicating the preference for the goal, and can be dynamically adjusted according to the user distribution.

[0074] Further,

[0075] r co (t) = Δα t

[0076] r ca (t) = ΔC

[0077] r ec (t) = -P t - P e

[0078] It should be noted that Δα tIt can be determined according to the difference in coverage rate between the current allocation action and the previous action performed by the selected aerial base station. ΔC can be determined according to the difference in system capacity between the current allocation action and the previous action performed by the selected aerial base station. r ec (t) can be determined according to the signal transmission energy consumption P in the energy consumption t and the device movement energy consumption P e .

[0079] In an alternative embodiment, the deployment allocation model may adopt the MADDPG algorithm. The specific process of the algorithm includes:

[0080] Initialize the agent environment and create and initialize the agent;

[0081] Initialize Memory D with a capacity of N;

[0082] Initialize the current action network μ of the agent and randomly generate weights θ;

[0083] Initialize the current critic network Q of the agent and randomly generate weights ω;

[0084] Initialize the target action network μ' of the agent with weights θ' = θ;

[0085] Initialize the target critic network Q' of the agent with weights ω' = ω;

[0086] Initialize the distribution state of the ground user terminals;

[0087] Loop through episode = 1, 2, …, M:

[0088] Initialize the position state s of the aerial base station;

[0089] Initialize a random action noise noise to achieve action exploration;

[0090] Loop through step = 1, 2, …, T:

[0091] Initialize the scenario observation state o according to the position information of the aerial base station and the users;

[0092] Select an action a for the agent based on the current action network μ i = μ θi (o i ) + noise;

[0093] Execute the action a i and restrict the drone from flying out of the map range;

[0094] Receive a reward r from the environment and transfer the state to o new ;

[0095] Obtain the respective target weights ω according to the current distribution of ground users r , calculate the total reward value r = ω r r.

[0096] In an alternative embodiment, based on the number of cycles for obtaining the scene observation state conforming to a preset number of cycles, use the current reward to evaluate the corresponding previous execution action to update the deployment allocation model, including:

[0097] SB1, based on the number of cycles for obtaining the scene observation state conforming to a preset number of cycles, randomly select the allocation action performed by the aerial base station, the reward obtained from performing the allocation action, and the scene observation states before and after performing the allocation action, denoted as (o, a, r, o new ), where o represents the scene observation state before performing the corresponding allocation action, o new represents the scene observation state after performing the corresponding allocation action, a represents the corresponding allocated action, and r represents the reward obtained from performing the corresponding allocated action.

[0098] It should be noted that each time a scene observation state is obtained, after the corresponding selected aerial base station performs an action once, (o, a, r, o new ) will be obtained. Store each obtained (o, a, r, o new ) in the data sample D, and randomly select a data sample from D

[0099] SB2, based on the allocated action, the reward obtained from performing the allocated action, and the scene observation states before and after performing the corresponding allocated action, obtain the loss function, and based on minimizing the loss function, update the weight parameters of the action network of the deployment allocation model.

[0100] It should be noted that the loss function is expressed as:

[0101]

[0102]

[0103] where J(ω i ) represents the loss function corresponding to the selected j-th data sample, represents the j-th data sample, y j represents the expected reward, represents the state-action value function corresponding to the current action network, γ represents the decay factor, represents the state-action value function corresponding to the target action network, represents the action corresponding to the num-th aerial base station at the j-th observation, a′ numDenote the action obtained by the target action network for the num-th aerial base station, a′ k Denote the action obtained by the target action network at observation k Denote the target action network at observation k, ω i Denote the weight parameters of the deployment allocation model action network

[0104] SB3, and the reward execution obtained by executing the allocated action and the scene observation states before and after executing the corresponding allocated action, and update the weight parameters of the deployment allocation model evaluation network in combination with the sampling policy gradient

[0105]

[0106] Among them, Denote the gradient of the loss function Denote the gradient of the weight parameters Denote the current action network Denote the observation i in the j-th sample data Denote the gradient of the action allocated to the selected aerial base station Denote the state-action value function at observation i, a i Denote the action allocated to the selected aerial base station, μ i Denote the current action network at observation i, θ i Denote the weight parameters of the deployment allocation model evaluation network

[0107] In summary, in the embodiments of the present invention, the scene observation state is cyclically obtained to determine, according to different scenes, to deploy aerial base stations in hot spots to enhance communication capacity, or deploy multiple aerial base stations in emergency areas for coverage supplement to restore normal communication, so as to optimize multiple communication metrics; by using the deployment allocation model, select aerial base stations and allocate actions to control the movement of aerial base stations based on the allocated actions, so that the deployed aerial base stations reach the optimal positions to meet the different needs of the ground communication network

[0108] Next, a multi-scenario-based aerial base station deployment system provided by the present invention will be described. The multi-scenario-based aerial base station deployment system described below can be mutually corresponding and referred to the multi-scenario-based aerial base station deployment method described above

[0109] Figure 3 The structural schematic diagram of a multi-scenario-based aerial base station deployment system of the present invention is shown. The system includes

[0110] A state acquisition module 31 that cyclically acquires the scene observation state

[0111] The base station deployment module 32 selects an aerial base station from the intelligent group to perform the allocation action according to the currently obtained scene observation state by using the deployment allocation model. The deployment allocation model is updated after evaluating the previous execution action based on the reward obtained by using the selected aerial base station to perform the allocation action when the number of cycles of obtaining the scene observation state meets the preset number of cycles.

[0112] In this embodiment, the base station deployment module 32 includes: a task determination unit that determines a task to be executed according to the currently obtained scene observation state. The task to be executed includes enhancing the system capacity in the target area to a preset capacity and covering a specified area in the target area; a base station deployment unit that selects an aerial base station from the intelligent group and allocates corresponding actions according to the currently obtained scene observation state by using the deployment allocation model, and controls the selected aerial base station to execute the corresponding allocation action; a loop update unit that obtains the next scene observation state and the current reward obtained after the currently selected aerial base station executes the current allocation action, and selects an aerial base station from the intelligent group again to execute the next allocation action according to the next obtained scene observation state by using the deployment allocation model until the task to be executed is completed. Based on the number of cycles of obtaining the scene observation state meeting the preset number of cycles, the current reward is used to evaluate the corresponding previous execution action to update the deployment allocation model.

[0113] Specifically, the loop update unit includes: a data acquisition subunit that obtains the corresponding coverage rate, system capacity, and energy consumption based on the currently selected aerial base station executing the allocation action; a reward acquisition subunit that obtains the corresponding reward based on the coverage rate, system capacity, and energy consumption.

[0114] Specifically, the data acquisition subunit includes: a state acquisition grandson unit that determines the current scene observation state corresponding to the currently selected aerial base station based on the allocation action executed by the currently selected aerial base station; a user number acquisition grandson unit that obtains the current number of ground users based on the current scene observation state; a covered user number acquisition grandson unit that obtains the number of users covered by the currently selected aerial base station based on the allocation action executed by the currently selected aerial base station; a coverage rate acquisition grandson unit that obtains the coverage rate of the currently selected aerial base station based on the current number of ground users and the number of users covered by the currently selected aerial base station.

[0115] In an alternative embodiment, the data acquisition subunit further includes: a power acquisition sub-subunit that obtains the corresponding transmission power and thermal noise power according to the currently selected air base station; a signal-to-noise ratio acquisition sub-subunit that obtains the signal-to-noise ratio of the wireless signal received by the user according to the transmission power, the thermal noise power, and the total propagation loss of the currently selected air base station obtained in advance; a transmission rate acquisition sub-subunit that obtains the user data transmission rate corresponding to the currently selected air base station according to the signal-to-noise ratio; and a capacity acquisition sub-subunit that obtains the system capacity corresponding to the currently selected air base station according to the user data transmission rate.

[0116] Further, the loop update unit further includes: a first loss acquisition sub-unit that obtains the line-of-sight propagation loss based on the straight-line distance between the selected air base station and the ground user, the speed of light, the carrier frequency, and the path loss; a second loss acquisition sub-unit that obtains the non-line-of-sight propagation loss based on the straight-line distance between the selected air base station and the ground user, the speed of light, the carrier frequency, and the additional loss; and a total loss acquisition sub-unit that obtains the total propagation loss according to the line-of-sight propagation loss and the non-line-of-sight propagation loss.

[0117] In an alternative embodiment, the data acquisition subunit further includes: a first energy acquisition sub-subunit that obtains the device movement energy consumption based on the allocation action performed by the currently selected air base station; and a second energy consumption acquisition sub-subunit that obtains the energy consumption based on the device movement energy consumption and the signal transmission energy consumption obtained in advance.

[0118] In addition, the reward acquisition subunit includes: a first reward acquisition sub-subunit that obtains the coverage reward for the currently selected air base station performing the current allocation action and the previous action based on the coverage rate; a second reward acquisition sub-subunit that obtains the system capacity reward for the currently selected air base station performing the current allocation action and the previous action based on the system capacity; a third reward acquisition sub-subunit that obtains the energy consumption reward based on the energy consumption; and obtains the corresponding reward based on the coverage reward, the system capacity reward, and the energy consumption reward.

[0119] In an alternative embodiment, the loop update unit further includes: a sample selection sub-unit that randomly selects the allocation action performed by the air base station, the reward obtained by performing the allocation action, and the scenario observation states before and after performing the allocation action based on the number of loops for obtaining the scenario observation state meeting the preset number of loops; a first parameter update sub-unit that obtains the loss function according to the allocated action, the reward obtained by performing the allocated action, and the scenario observation states before and after performing the corresponding allocation action, and updates the weight parameters of the action network of the deployment allocation model based on minimizing the loss function; and a second parameter optimization sub-unit that updates the weight parameters of the evaluation network of the deployment allocation model according to the allocated action, the reward obtained by performing the allocated action, and the scenario observation states before and after performing the corresponding allocation action, and in combination with the sampling policy gradient.

[0120] In summary, in the embodiments of the present invention, the state acquisition module cyclically acquires the scene observation state, so as to determine to deploy an aerial base station in a hot spot area to enhance the communication capacity according to different scenes, or deploy multiple aerial base stations in an emergency area for coverage supplement to restore normal communication, so as to optimize multiple communication metrics; the base station deployment module uses the deployment allocation model to select an aerial base station and allocate actions, so as to control the aerial base station to move based on the allocated actions, so that the deployed aerial base station reaches the optimal position to meet the different needs of the ground communication network.

[0121] Figure 4 An entity structure diagram of an electronic device is exemplified, as Figure 4 shown, the electronic device may include: a processor 41, a communication interface 42, a memory 43, and a communication bus 44. Among them, the processor 41, the communication interface 42, and the memory 43 communicate with each other through the communication bus 44. The processor 41 can call the logical instructions in the memory 43 to execute the method for deploying an aerial base station based on multiple scenarios, and the method includes: cyclically acquiring the scene observation state; according to the currently acquired scene observation state, using the deployment allocation model, selecting an aerial base station from the intelligent group to execute the allocated actions; wherein, the deployment allocation model is updated after evaluating the previous execution action based on the reward obtained by using the selected aerial base station to execute the allocated actions when the number of cycles of acquiring the scene observation state meets the preset number of cycles.

[0122] In addition, when the logical instructions in the above-mentioned memory 43 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0123] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-scenario-based aerial base station deployment method provided by the above-mentioned various methods. The method includes: cyclically obtaining the scenario observation state; according to the currently obtained scenario observation state, using the deployment allocation model, selecting an aerial base station from the intelligent group to perform the allocated action; wherein, the deployment allocation model is updated after evaluating its previous execution action based on the number of cycles of obtaining the scenario observation state meeting the preset number of cycles and the reward obtained by using the selected aerial base station to perform the allocated action.

[0124] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the multi-scenario-based aerial base station deployment method provided by the above-mentioned various methods. The method includes: cyclically obtaining the scenario observation state; according to the currently obtained scenario observation state, using the deployment allocation model, selecting an aerial base station from the intelligent group to perform the allocated action; wherein, the deployment allocation model is updated after evaluating its previous execution action based on the number of cycles of obtaining the scenario observation state meeting the preset number of cycles and the reward obtained by using the selected aerial base station to perform the allocated action.

[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for deploying an aerial base station based on multiple scenarios, characterized in that Including: Repeatedly obtain the scene observation state; According to the currently obtained scene observation state, use the deployment allocation model to select an aerial base station from the intelligent group to perform the allocation action; wherein, the deployment allocation model is updated after evaluating its previous execution action based on the number of cycles of obtaining the scene observation state meeting the preset number of cycles and the reward obtained by using the selected aerial base station to perform the allocation action; The step of, according to the currently obtained scene observation state, using the deployment allocation model to select an aerial base station from the intelligent group to perform the allocation action includes: obtaining the current reward obtained after the selected aerial base station performs the current allocation action; The step of obtaining the current reward obtained after the selected aerial base station performs the current allocation action includes: Based on the selected aerial base station currently performing the allocation action, obtain the corresponding coverage rate, system capacity, and energy consumption; Based on the coverage rate, the system capacity, and the current energy consumption, obtain the corresponding reward; The step of, based on the selected aerial base station currently performing the allocation action, obtaining the system capacity includes: According to the currently selected aerial base station, obtain the corresponding transmit power and thermal noise power; According to the transmit power, the thermal noise power, and the total propagation loss corresponding to the currently selected aerial base station obtained in advance, obtain the signal-to-noise ratio of the wireless signal received by the user; According to the signal-to-noise ratio, obtain the user data transmission rate corresponding to the currently selected aerial base station; According to the user data transmission rate, obtain the system capacity corresponding to the currently selected aerial base station.

2. The multi-scenario-based aerial base station deployment method according to claim 1, wherein The step of, according to the currently obtained scene observation state, using the deployment allocation model to select an aerial base station from the intelligent group to perform the allocation action includes: According to the currently obtained scene observation state, determine the task to be executed; the task to be executed includes enhancing the system capacity in the target area to the preset capacity and covering the specified area in the target area; According to the currently obtained scene observation state, use the deployment allocation model to select an aerial base station from the intelligent group and allocate corresponding actions, and control the selected aerial base station to perform the corresponding allocation actions; Obtain the next scene observation state, and according to the next obtained scene observation state, use the deployment allocation model to re-select an aerial base station from the intelligent group to perform the next allocation action until the task to be executed is completed; wherein, based on the number of cycles of obtaining the scene observation state meeting the preset number of cycles, use the current reward to evaluate the corresponding previous execution action to update the deployment allocation model.

3. The method for deploying an aerial base station based on multiple scenarios according to claim 2, wherein The step of, based on the number of cycles of obtaining the scene observation state meeting the preset number of cycles, using the current reward to evaluate the corresponding previous execution action to update the deployment allocation model includes: Based on the number of cycles of obtaining the scene observation state meeting the preset number of cycles, randomly select the allocation action performed by the aerial base station, the reward obtained by performing the allocation action, and the scene observation states before and after performing the allocation action; Obtain a loss function based on the assigned action, the reward execution obtained from executing the assigned action, and the scenario observation states before and after executing the corresponding assigned action, and update the weight parameters of the action network of the deployment assignment model based on minimizing the loss function; And update the weight parameters of the evaluation network of the deployment assignment model according to the assigned action, the reward execution obtained from executing the assigned action, and the scenario observation states before and after executing the corresponding assigned action, and in combination with the sampling policy gradient.

4. The method for deploying an aerial base station based on multiple scenarios according to claim 1, wherein The obtaining of the coverage rate based on the assigned action executed by the currently selected aerial base station includes: Determine the current scenario observation state corresponding to the currently selected aerial base station based on the assigned action executed by the currently selected aerial base station; Obtain the current number of ground users based on the current scenario observation state; Obtain the number of users covered by the currently selected aerial base station based on the assigned action executed by the currently selected aerial base station; Obtain the coverage rate corresponding to the currently selected aerial base station based on the current number of ground users and the number of users covered by the currently selected aerial base station.

5. The method for deploying an aerial base station based on multiple scenarios according to claim 1, wherein The obtaining of the energy consumption based on the assigned action executed by the currently selected aerial base station includes: Obtain the equipment movement energy consumption based on the assigned action executed by the currently selected aerial base station; Obtain the energy consumption based on the equipment movement energy consumption and the pre-obtained signal transmission energy consumption.

6. An air base station deployment system based on multiple scenarios, characterized in that It includes: A state acquisition module that circularly acquires scenario observation states; A base station deployment module that, according to the currently acquired scenario observation state, uses the deployment assignment model to select an aerial base station from the intelligent group to execute the assigned action; wherein, the deployment assignment model is updated after evaluating its previous execution action based on the number of cycles of acquiring scenario observation states meeting a preset number of cycles and the reward obtained from using the selected aerial base station to execute the assigned action; The base station deployment module further includes: a cyclic update unit that acquires the current reward obtained after the selected aerial base station executes the current assigned action; The cyclic update unit includes: A data acquisition sub-unit that obtains the corresponding coverage rate, system capacity, and energy consumption based on the assigned action executed by the currently selected aerial base station; A reward acquisition sub-unit that obtains the corresponding reward based on the coverage rate, the system capacity, and the current energy consumption; The data acquisition sub-unit includes: A power acquisition grandchild unit that obtains the corresponding transmission power and thermal noise power according to the currently selected aerial base station; A signal-to-noise ratio acquisition grandchild unit that obtains the signal-to-noise ratio of the wireless signal received by the user according to the transmission power, the thermal noise power, and the total propagation loss corresponding to the currently selected aerial base station pre-acquired; A transmission rate acquisition grandchild unit that obtains the user data transmission rate corresponding to the currently selected aerial base station according to the signal-to-noise ratio; A capacity acquisition grandchild unit that obtains the system capacity corresponding to the currently selected aerial base station according to the user data transmission rate.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multi-scenario based aerial base station deployment method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-scenario based aerial base station deployment method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for determining position of air base station and electronic equipment

    CN112512115A

  • A communication network hotspot area capacity enhancement-oriented multi-air base station deployment method

    CN113873434A

  • Unmanned aerial vehicle autonomous deployment method based on energy consumption optimization in irrigation area scene

    CN115119174A