A scheduling control method and device for a wireless communication system

Through the MADDPG algorithm, the trajectory planning and access control of UAV are optimized, and the data scheduling and energy transmission problems between UAV and ground users are solved, and efficient energy management and stable data transmission of wireless communication systems are realized.

CN116471694BActive Publication Date: 2025-08-01CHINA THREE GORGES CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211393207.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-08-01
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

In the prior art, UAV trajectory planning and access control policies have not been effectively optimized, resulting in the data scheduling and energy transmission between UAV and ground users being disturbed by the environment, making it difficult to maintain stability and efficiency in a dynamic environment.

Method used

Multi-agent deep reinforcement learning (MADDPG) algorithm is adopted, combined with Markov decision-making process (MDP), and the trajectory planning and access control strategy of UAV are optimized, and energy efficiency is maximized by jointly optimizing the UAV's flight trajectory, transmission mode and transmission strategy of ground users.

Benefits of technology

Under limited channel conditions, the energy efficiency of the wireless communication system is significantly improved, efficient data transmission and energy management between UAV and ground users are realized, and system energy consumption is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116471694B_ABST
    Figure CN116471694B_ABST
Patent Text Reader

Abstract

The present invention provides a scheduling control method and apparatus for a wireless communication system, including: minimizing the overall energy consumption by jointly optimizing the trajectory of the UAV and the access control strategy of the GU. Solving the problem of joint access control and trajectory planning through the MADDPG algorithm. The main factors affecting the energy consumption of data transmission in the wireless communication network are the access strategy of the UAV, trajectory planning, and channel conditions. When the GU has more energy, the access control strategy can select the active transmission mode; however, when the GU has less energy, the passive transmission mode can be selected. The present invention considers the actual situation more comprehensively. Through the joint optimization of the UAV access control and trajectory planning strategies, the proposed MADDPG transmission scheme enables the system to achieve maximum energy efficiency even under limited channel conditions. After simulation verification, compared with the benchmark scheme, the scheme proposed by the present invention achieves the best performance in terms of performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication, and more specifically, relates to a scheduling control method and device for a wireless communication system. Background Art

[0002] With the popularization of unmanned aerial vehicles (UAVs) in the Internet of Things (IoT), it has established a data collection channel for IoT users or sensors and is an indispensable part of the future IoT. Due to the mobility of ground users (GUs) and limited energy storage, a direct connection between a GU and a base station (BS) is usually difficult. Therefore, UAVs play an important role in assisting data collection and transmission from GUs to BSs. It can be used as a forwarding relay node to assist in the data transmission of GUs beyond the communication service range. However, due to the high complexity of distributed optimization, the lack of centralized coordination, and the unknown dynamics of the network environment, there are still some limitations in the joint control of the trajectory and transmission strategy of UAVs.

[0003] In currently studied UAV-assisted real-time wireless communication systems, in order to utilize their performance gains, trajectory planning is one of the most beneficial design problems, which can utilize the mobility of UAVs and dynamically reshape the network structure to support data transmission. By using dynamic programming to design the trajectory of UAVs, not only can the total energy consumption be reduced, but also the performance close to that of the exhaustive algorithm can be achieved with low complexity. There are also many existing works considering multi-UAV-assisted networks. By planning the flight trajectories of multiple UAVs, the data uploaded by IoT users has increased significantly. In addition, through joint optimization of bandwidth, power allocation, and the trajectory of UAVs, multi-UAV-assisted emergency communication has been explored. In particular, each UAV can first collect and cache user data, and then forward the data to the next UAV when they meet during flight. Coordination between different GUs is also a key design problem for efficient data collection and transmission. Since the coverage ranges of UAVs at different positions are different, it is necessary to divide GUs skillfully among different UAVs to balance interference and network coverage.

[0004] However, when the UAV performs access control on the GU, the data scheduling and energy transfer between the UAV and the GU are greatly interfered by the environment. Due to the time-varying channel conditions, it is difficult to maintain the stability of data transmission. Most of the current inventions on UAV-assisted networks consider the link switching between UAVs and how to optimize the UAV trajectories, while ignoring the importance of the access control strategy between the GU and the UAV. The UAV can also serve as an energy provider for some energy-starved GUs, providing energy for the GUs through radio frequency signals, which features wireless power transfer and low power consumption. When the UAV is an energy transmitter and the GU is a low-power sensor device with limited energy supply, it is difficult to control the consumed energy by selecting the data transmission mode and energy harvesting within the sensing time slot. The present invention aims to solve the access control strategy problem between the UAV and the GU, which is a high-dimensional control problem.

[0005] Secondly, most of the inventions only consider collecting GU data and completing data scheduling according to the planned UAV trajectories, without jointly considering the UAV trajectory planning and the access control strategy. In a dynamic environment, the efficiency of the GU-UAV access control strategy is related not only to the flight trajectory of the UAV but also to when to select to report data to the BS. It is a complex joint optimization problem to jointly consider planning the UAV flight trajectory and switching different transmission modes for data uploading according to the dynamic environment and its own state within the limited time when the UAV covers the GUs. The prior art does not combine the access control strategy with the UAV trajectory planning and cannot solve the problem of jointly optimizing the UAV trajectory. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a scheduling control method and device for a wireless communication system, aiming to solve the problem that the prior art does not jointly consider the UAV trajectory planning and the strategy of the UAV accessing the GU.

[0007] To achieve the above object, in the first aspect, the present invention provides a scheduling control method for a wireless communication system. The method is applied to a UAV-assisted wireless communication system, and the system includes: a base station BS, multiple UAVs, and multiple ground users GUs. The method includes the following steps:

[0008] Determine the energy efficiency of the wireless communication system. The energy efficiency is the average ratio of the total data volume received by the BS to the total energy consumed by the wireless communication system during the entire flight period of the UAV.

[0009] Determine the constraints of the wireless communication system; the constraints include: the distance between any two UAVs in any time slot is greater than a preset minimum distance, each GU accesses only one UAV in one time slot, only one UAV reports data to the BS in each time slot, the way for the GU to access the UAV is one of passive backscatter communication or active RF communication, the energy budget constraint for each GU in each time slot, and the amount of data reported by the UAV to the BS is determined by the distance between it and the BS and the channel conditions;

[0010] Determine the combinatorial optimization problem; the combinatorial optimization problem is used to design the scheduling strategy of the wireless communication system based on the constraints to maximize the energy efficiency; the scheduling strategy includes: the transmission control strategy of each GU, the flight trajectories of each UAV, and the transmission scheduling strategy of each UAV;

[0011] Define the combinatorial optimization problem as a Markov decision process MDP; where, the total reward of the MDP includes the long-term rewards of all UAVs, and the long-term reward of each UAV includes its self-reward under each decision step during its entire flight period. The self-reward includes: objective function reward, guidance reward, and penalty term; if a GU uploads data to the UAV, the UAV obtains a guidance reward. When the energy of the GU does not meet the requirements of its transmission control strategy, the value of the guidance reward is 0. If the distance between any two UAVs is less than the preset minimum distance, the UAV obtains a penalty term. If a UAV successfully reports data to the BS, the UAV obtains an objective function reward;

[0012] Solve the MDP to obtain the scheduling strategy of the wireless communication system when the energy efficiency is maximized.

[0013] In a possible example, each time slot t of the UAV includes: a flight sub-slot, a sensing sub-slot, and a reporting sub-slot, and the lengths of the three sub-slots are τ f , τ s , τ d ;

[0014] The constraints include:

[0015] ||l i (t + 1) - l i (t)|| ≤ υ max τ f ,

[0016] d i,j (t) ≥ d min ,

[0017] where, υ max τ f represents the maximum flight distance, d min represents the preset minimum distance, υmax represents the maximum flight speed, d i,j (t) represents the distance between the i-th UAV and the j-th UAV at the t-th time slot, the distance between the i-th UAV and the j-th UAV, l i (t) represents the position of the i-th UAV at the t-th time slot, l i (t + 1) represents the position of the i-th UAV at the (t + 1)-th time slot, i ≠ j.

[0018] In an optional example, the constraint conditions further include:

[0019]

[0020] where x m,i (t) ∈ {0, 1} represents the access control strategy of the m-th GU to the i-th UAV within the t-th time slot, x m,i (t) being 0 means the GU does not access the UAV, x m,i (t) being 1 means the GU accesses the UAV, represents the set of all GUs within the coverage range of the i-th UAV, and N represents the total number of UAVs.

[0021] In an optional example, the constraint conditions further include:

[0022] The data upload rate of the active radio frequency communication method is:

[0023]

[0024] where τ z is the sub-time slot allocated to the GU allowed to access control, p m (t) represents the transmission power of the m-th GU at the t-th time slot, h m,i represents the channel coefficient between the i-th UAV and the m-th GU, h m,i is composed of the channel coefficient under the line-of-sight and the non-line-of-sight between the UAV and the GU;

[0025] The data upload rate of the passive backscatter communication method is:

[0026]

[0027] where p A represents the fixed transmission power, Γ o is the constant coefficient of the antenna;

[0028] Let z m (t) ∈ {0, 1} represent the transmission control strategy of the m-th GU at the t-th time slot. When zm At (t)=0, the m-th GU will select the passive backscatter communication mode. When z m (t)=1, the m-th GU selects the active radio frequency communication mode.

[0029] In an optional example, to avoid scheduling interference between UAVs, the constraint conditions further include:

[0030]

[0031] Among them, y i (t)∈{0, 1} represents the transmission scheduling strategy of the i-th UAV in time slot t. Among them, y i (t)=1 means that the UAV reports data to the BS in time slot t;

[0032] When y i (t)=1:

[0033] O i (t)=τ d log(1 + p i,r (t)||g i || 2 )

[0034] Among them, O i (t) represents the data volume reported by the i-th UAV to the BS, p i,r (t) represents the transmission power used by the i-th UAV for information forwarding, and g i represents the channel condition between the UAV and the BS.

[0035] In an optional example, the constraint conditions further include:

[0036] When x m,i =1, let represent the energy collected by the m-th GU in the t-th time slot;

[0037] Each m-th GU needs to satisfy the following energy budget constraint in each time period:

[0038]

[0039] Among them, E m (t) represents the energy state of the m-th GU at the beginning of the t-th time slot, is the maximum battery capacity of the m-th GU, z n (t) represents the transmission control strategy of the n-th GU in the t-th time slot, and p m (t) represents the transmission power of the m-th GU in the t-th time slot.

[0040] In an optional example, the energy efficiency of the wireless communication system is:

[0041]

[0042] where Ξ represents the energy efficiency, represents the UAV time slot length, O i (t) represents the amount of data reported by the i-th UAV to the BS, y i (t) represents whether the i-th UAV plans to report data to the BS in a certain time slot, e i,o (t) represents the operating energy consumption of the UAV, e i,s (t) represents the sensing energy consumption of the UAV, e i,r (t) represents the reporting energy consumption of the UAV;

[0043] The sensing energy consumption e i,s (t) and the reporting energy consumption e i,r (t) of the UAV are specifically:

[0044]

[0045] e i,r (t) = y i (t)p i,r (t)τ d

[0046] where, represents the set of GUs allowed to access control by the i-th UAV,

[0047] In an optional example, the combined optimization problem is defined as an MDP, specifically:

[0048] The state of the wireless communication system in each time slot is represented as: s t = (s1(t), s2(t),..., s N (t)); where s i (t) represents the system state information observed by the i-th UAV; s i (t) = (χ i , ψ i ), where χ i = (E i , ζ m , Q i ) represents the energy storage and data buffering of the UAV and GUs, E i represents the set of energy queues of the UAV and the covered GUs, (ζ m , Q i ) is the set of all data buffers; ψi =(h i , g i ) represents the channel condition in the network. h i is the set of channel coefficients between the i-th UAV and all GUs allowed to access the i-th UAV, denoted as

[0049] Denote the actions of all UAVs as a t =(a1(t), a2(t),..., a N (t)), where the action represents the transmission control strategy of the GU, represents the access control strategy of the GU to the UAV, and y i =[y i (t)] represents the scheduling strategy of the UAV, representing the flight trajectory of the UAV;

[0050] The self-reward r i (t) of the i-th UAV is as follows:

[0051]

[0052] where γ and η are both adjustable parameters, and s m,i (t) represents the size of the sensing data uploaded from the m-th GU to the i-th UAV during the sub-slot τ z , r p (t) is the minimum distance index to avoid interference and collision between different UAVs; represents the guiding reward, and the objective function reward is denoted as represents the penalty term, and I(·) represents the indicator function;

[0053] The i-th UAV's long-term reward during the entire time period is

[0054] The total reward

[0055] In a second aspect, the present invention provides a scheduling control device for a wireless communication system. The device is applied to a UAV-assisted wireless communication system, and the system includes: a base station BS, multiple UAVs, and multiple ground users GUs; the device includes:

[0056] An energy efficiency determination unit for determining the energy efficiency of the wireless communication system; the energy efficiency is the average ratio of the total amount of data received by the BS to the total energy consumed by the wireless communication system during the entire flight period of the UAV;

[0057] A constraint determination unit for determining the constraints of a wireless communication system; the constraints include: the distance between any two UAVs in any time slot is greater than a preset minimum distance, each GU accesses only one UAV in one time slot, only one UAV reports data to the BS in each time slot, the way for the GU to access the UAV is one of passive backscatter communication or active radio frequency communication, the energy budget constraint of each GU in each time slot, and the amount of data reported by the UAV to the BS is determined by the distance between it and the BS and the channel conditions;

[0058] An optimization problem determination unit for determining a combinatorial optimization problem; the combinatorial optimization problem is used to design a scheduling strategy for the wireless communication system based on the constraints to maximize the energy efficiency; the scheduling strategy includes: the transmission control strategy of each GU, the flight trajectories of each UAV, and the transmission scheduling strategy of each UAV;

[0059] An MDP definition unit for defining the combinatorial optimization problem as a Markov decision process MDP; wherein, the total reward of the MDP includes the long-term rewards of all UAVs, and the long-term reward of each UAV includes its self-reward under each decision step during its entire flight period, and the self-reward includes: an objective function reward, a guidance reward, and a penalty term; if a GU uploads data to the UAV, the UAV obtains a guidance reward, and when the energy of the GU does not meet the requirements of its transmission control strategy, the value of the guidance reward is 0, if the distance between any two UAVs is less than the preset minimum distance, the UAV obtains a penalty term, and if a UAV successfully reports data to the BS, the UAV obtains an objective function reward;

[0060] A scheduling solution unit for solving the MDP to obtain the scheduling strategy of the wireless communication system when the energy efficiency is maximized.

[0061] In a third aspect, the present invention provides a scheduling control device for a wireless communication system, including: a memory and a processor;

[0062] The memory is used for storing a computer program;

[0063] The processor is used for implementing the method provided in the first aspect above when executing the computer program.

[0064] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects are obtained:

[0065] The present invention provides a scheduling control method and device for a wireless communication system, which takes into account the actual situation more comprehensively. To jointly optimize the trajectory planning and access control strategy of UAVs, a multi-agent reinforcement learning (Multi-Agent Deep Deterministic Policy Gradient, MADDPG) transmission scheme is adopted, enabling the system to achieve maximum energy efficiency even under limited channel conditions. Through simulation verification, compared with the benchmark scheme, the scheme proposed by the present invention achieves the best performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a flowchart of the scheduling control method for the wireless communication system provided by an embodiment of the present invention;

[0067] Figure 2 is an architecture diagram of a multi-UAV-assisted wireless communication system provided by an embodiment of the present invention;

[0068] Figure 3 is a time slot structure diagram of the working process of each UAV provided by an embodiment of the present invention;

[0069] Figure 4 is a convergence diagram of the reward value and a flight trajectory evaluation diagram during the training process provided by an embodiment of the present invention;

[0070] Figure 5 is a comparison diagram of the remaining data amounts of the GU and UAVs under the separate DDPG algorithm provided by an embodiment of the present invention;

[0071] Figure 6 is a comparison diagram of the remaining data amounts of the GU and UAVs under the MADDPG algorithm provided by an embodiment of the present invention;

[0072] Figure 7 is an architecture diagram of the scheduling control device for the wireless communication system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of them. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0074] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0075] In the description of the present invention, the meaning of "several" is more than one, the meaning of "multiple" is more than two, and understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0076] In the description of the present invention, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0077] The present invention can improve the service range of a wireless communication network. Since the task requirements of ground users GU are highly random, and the time-varying environment will cause obstacles to data transmission. To relieve the pressure of data link communication and improve the stability of the transmission process, it is necessary to increase the network coverage area and the flexibility of the network. Therefore, the concept of an unmanned aerial vehicle (UAV)-assisted computing network is proposed. Due to the flexible flight characteristics of UAVs, it is possible to perform temporary network deployment and information collection in areas such as sudden task requirements, emergency scenarios, and intelligent transportation.

[0078] The present invention formulates the trajectory planning and access control of multiple UAVs as a joint optimization problem. Since this problem has many variables and high complexity, traditional optimization algorithms consume a large amount of computing time to solve this problem and show poor performance. The present invention aims to solve this problem through a multi-agent deep reinforcement learning (DRL) method, considering a dynamic network environment that contains certain information about the spatial distribution and traffic demands of multiple GUs. The simulation results show that the trajectory planning and access control of UAVs can significantly improve the energy conversion efficiency of the unmanned aerial vehicles.

[0079] The present invention considers the problem of optimizing the trajectory planning and access control between UAVs and GUs. The objective of the present invention is to minimize the overall energy consumption by jointly optimizing the trajectories of the UAVs and the access control strategies of the GUs. To ensure satisfactory service coverage, different UAVs can negotiate the trajectory planning so that they do not collide in the same area. Therefore, according to the spatial distribution of the GUs and their traffic demands, the trajectories of the UAVs may have their own service areas. The UAVs responsible for more acquisition tasks need to fly close to the base station and report data to the base station. The present invention solves the problem of joint access control and trajectory planning through the MADDPG algorithm.

[0080] Figure 1 is the flowchart of the scheduling control method of the wireless communication system provided by the embodiment of the present invention; as Figure 1 shown, it includes the following steps:

[0081] S101, determine the energy efficiency of the wireless communication system; the energy efficiency is the average ratio of the total data volume received by the BS to the total energy consumed by the wireless communication system during the entire flight period of the UAV;

[0082] S102, determine the constraint conditions of the wireless communication system; the constraint conditions include: the distance between any two UAVs in any time slot is greater than the preset minimum distance, each GU accesses only one UAV in one time slot, only one UAV reports data to the BS in each time slot, the way for the GU to access the UAV is one of passive backscatter communication or active radio frequency communication, the energy budget constraint of each GU in each time slot, and the data volume reported by the UAV to the BS is determined by the distance between it and the BS and the channel conditions;

[0083] S103, determine the combined optimization problem; the combined optimization problem is used to design the scheduling strategy of the wireless communication system based on the constraint conditions to maximize the energy efficiency; the scheduling strategy includes: the transmission control strategies of each GU, the flight trajectories of each UAV, and the transmission scheduling strategies of each UAV;

[0084] S104, define the combined optimization problem as a Markov decision process MDP; where the total reward of the MDP includes the long-term rewards of all UAVs, and the long-term reward of each UAV includes its self-reward under each decision in its entire flight period, and the self-reward includes: objective function reward, guidance reward and penalty term; if a GU uploads data to the UAV, the UAV obtains a guidance reward, when the energy of the GU does not meet the requirements of its transmission control strategy, the value of the guidance reward is 0, if the distance between any two UAVs is less than the preset minimum distance, the UAV obtains a penalty term, and if a UAV successfully reports data to the BS, the UAV obtains an objective function reward;

[0085] S105, solve the MDP to obtain the scheduling strategy of the wireless communication system when the energy efficiency is maximized.

[0086] Specifically, the present invention considers a UAV-assisted wireless network system consisting of a BS, multiple UAVs, and GUs. First, the index of the UAV is denoted as The index of the GU is denoted as It is assumed that the GUs are spatially distributed beyond the direct communication range of the BS, so there is no direct link between the GUs and the BS. The UAV can receive the sensing data of the GUs and forward the collected data to the BS as a relay. Each GU can collect radio frequency energy from the beamforming signal of the UAV to charge its battery and maintain its operation, such as data transmission or processing. The workload of each GU can be transmitted to the UAV through active radio frequency (RF) or passive communication. Each channel is considered to be frequency-flat block fading, that is, the channel coefficient is constant within a time frame and may vary from frame to frame. Considering a dynamic network environment that contains certain information about the spatial distribution and traffic demand of the GUs. The present invention adopts the MADDPG algorithm to solve the joint access control and trajectory planning problems. The simulation results show that the joint trajectory optimization and access control strategy can better utilize multiple UAVs for data cooperative transmission and significantly improve the transmission energy efficiency of the system.

[0087] Due to limited channel capacity or poor channel quality, the BS cannot directly communicate with the GUs (multiple ground users). The goal of this solution is to improve the efficiency of its data collection and transmission by optimizing the trajectories of the UAVs. Each UAV has its own responsible collection area, which coordinates with each other and does not interfere with each other. At the same time, the UAV can also optimize its access control strategy to reduce the energy consumption of data transmission with the GU and improve the data throughput.

[0088] The present invention first mathematically models the optimization problems to be solved by each layer, and then derives the algorithm design of the present invention. Specifically as follows:

[0089] The present invention considers a UAV-assisted wireless network, in which a BS, multiple UAVs, and GUs are spatially distributed within the coverage range of the UAV, as Figure 2 shown. The set of UAVs is denoted as The set of all GUs is denoted as The present invention assumes that there is no direct link connection between all GUs and BSs due to the obstruction of surrounding objects on the ground. The UAV can fly above the GU, collect the sensing data of the GU, and forward the data information to the BS. Each GU can collect energy from the RF beamforming signal of the UAV to charge its battery and maintain its active operations, such as data sensing, transmission, and local processing. The sensing data of each GU can be uploaded to the relevant UAV through active radio communication or passive backscatter communication, depending on its energy state, channel conditions, and traffic requirements. After the UAV collects the sensing information of the GU, it forwards the information to the BS.

[0090] The present invention assumes that the trajectory planning of the UAV is implemented in a time-slot frame structure. Each time slot has a fixed length τ. It is further divided into three sub-slots for flying, sensing, and reporting, as Figure 3 shown. During the flying sub-time slot τ f , the UAV can fly to the preferred position and hover at that position during the sensing and reporting sub-time slots. During the sensing time slot τ s , a time-division protocol is considered to collect the sensing information of all GUs. In particular, each GU granted access will be assigned a time slot τ z . All GUs can upload their information to the UAV one by one through active or passive communication. In addition, each GU can collect RF energy when other GUs are actively transmitting. The third sub-time slot τ d is used for the UAV to report its information to the BS. The present invention assumes that the UAV-GU and UAV-BS channel coefficients are constant in each time slot and may change as the UAV adapts its trajectory.

[0091] The trajectory of each UAV-i can be defined as a set of positions on different time slots, i.e., each position is specified by 3D coordinates, i.e., l i (t) = (x i (t), y i (t), z i (t)). Let H B represent the height of the BS antenna. The present invention can assume that the position of the BS is l0(t) = (0, 0, H B ). Let d i,0 represent the distance between UAV-i and the BS. Assume that UAV-i moves in the direction of d i (t) ≤ υ max at a finite speed υ i (t). Therefore, the position of UAV-i in the next time slot is l i (t + 1) = l i (t) + υ i (t)τ f di (t), which is related to the flying sub - time - slot τ f , flying speed υ i (t) and direction d i (t). To avoid interference between different UAVs and ensure safety between different UAVs, the distance between UAV - i and UAV - j, that is, d i,j (t)=||l i (t)-l j (t)||, and the constraints are as follows:

[0092] ||l i (t + 1)-l i (t)||≤υ max τ f ,

[0093] d i,j (t)≥d min , (1)

[0094] where l j (t) represents the position of the j - th UAV at time - slot t, υ max τ f represents the maximum flying distance, and d min represents the minimum distance between UAVs to ensure safety.

[0095] Given the hovering position of the UAV in the sensing time - slot τ s , there may be multiple GUs within the coverage range of the same UAV. Note that the channel conditions of some GUs may be poor, so the data rate of information upload may be low. This means that the UAV must design an access control strategy to improve the energy efficiency of information upload to the UAV. Let represent the set of all GUs within the coverage range of UAV - i. Let represent the set of users allowed to upload sensing information to UAV - i. Due to insufficient energy or unsatisfactory channel conditions, the left - hand users may choose to retain their information upload in the current time - slot. When other UAVs come back, they can resume information transmission in a later period. Let x m,i (t)={0, 1} represent the access control strategy of GU - m to UAV - i in the t - th time - slot. Then, it can be obtained that The present invention further requires that to ensure that GU - m can only access one UAV in each time - slot.

[0096] For all GU - m in the set , consider using a time - division protocol to upload data for them. The length τ s of the sensing time - slot can be further divided into a length of The hourly slot. Each hourly slot can be used for radio frequency active transmission or backscatter passive transmission. For active radio frequency transmission, the received signal of UAV-i can be expressed as where p m represents the transmission power of GU-m, is the unit power of the information symbol, and v0 represents the noise signal. h m,i (t) represents the channel coefficient between the i-th UAV and the m-th GU in the current time slot. The present invention considers a realistic channel model composed of line-of-sight (LOS) and non-line-of-sight (NLOS) components. The channel coefficient can be modeled as where ψ m,i (t)=ω0(d m,i (t)) -α represents the large-scale fading, and the characteristics of the small-scale fading are as follows;

[0097]

[0098] The first term represents the LOS component, and the second term represents the NLOS component. The Rician factor K sets different weights for the LOS and NLOS components. Similarly, the present invention can define g i (t) as the channel vector from the multi-antenna UAV-i to the BS.

[0099] Therefore, the upload rate in active radio frequency transmission can be simplified as:

[0100]

[0101] The present invention assumes a normalized noise power. In passive data upload, GU-m relies on the radio frequency signal transmitted by UAV-i to backscatter information. Let represent the signal beamforming of UAV-i in the t-th hourly slot, where w m,i represents the normalized beamforming vector of UAV-i for GU-m p A represents the fixed transmission power, and s is a random symbol with unit power. After the backscattering of GU-m, the passive upload data rate can be approximated as:

[0102]

[0103] where Γ o is an antenna-specific constant coefficient. For simplicity, similar to the active transmission formula, the present invention assumes that UAV-i adopts the maximum ratio combining (MRC) scheme when detecting the information of GU-m. Therefore, the present invention has w m,i =h m,i / ||h m,i ||, and then Let \(z\) m (t) ∈ {0, 1} denote the transmission control strategy of GU-m in the \(t\)-th time slot. When \(z\) m (t) = 0, GU-m will choose backscatter communication, and when \(z\) m (t) = 1, it will choose RF active communication.

[0104] In each time slot, the UAV can collect data from the GU and then report the data to the BS. To avoid interference between UAVs, the present invention uses a binary variable \(y\) i (t) ∈ {0, 1} to indicate whether UAV-i plans to report its data to the BS. The present invention further requires to ensure that only one UAV can report to the BS within each time slot. Therefore, it can be expected that the data buffer of each UAV will be dynamically updated over time. Let \(s\) m,i (t) represent the size of the sensed data uploaded from GU-m to UAV-i during the sub-time slot \(\tau\) z . Given the transmission control strategy \(z\) m (t) of GU-m, the present invention has Let \(A\) m (t) represent the size of the sensed data arriving at GU-m at the beginning of the \(t\)-th time slot. For each GU-m, the present invention assumes that \(A\) m (t) ∈ [\(A\) m,min , \(A\) m,max is independently and identically distributed with an average value of \(\lambda\) m .

[0105] Let \((\zeta\) m (t), \(Q\) i (t)) represent the remaining data sizes in the buffers of GU-m and UAV-i, respectively. Therefore, the present invention can update the data queue as follows:

[0106]

[0107]

[0108] where \([X]\) + denotes the maximum operation, i.e., max{0, \(X\)}. The indicator \(y\) i (t) indicates whether UAV-i reports data to the BS, and \(O\) i (t) is the amount of data reported. When \(y\) i (t) = 1:

[0109] \(O\) i (t) = \(\tau\) d log(1 + \(p\) i,r (t)||\(g\) i ||2 ) (6)

[0110] where p i,r (t) represents the transmission power of UAV-i for information forwarding. Obviously, O i (t) depends on the distance d between UAV-i and the BS i,0 and the channel condition g i .

[0111] The present invention aims to maximize the energy efficiency of the UAV-assisted sensing network by jointly optimizing the trajectory, access control, and transmission scheduling strategies of the UAV, as well as the transmission strategy of the GU.

[0112] The total energy consumption per time slot includes the operating energy consumption of the UAV during flight and hovering, and the radio frequency energy consumption of the UAV during sensing and reporting. For simplicity, the present invention assumes that the operating energy consumption e i,o (t) is a constant, depending on the total time length of flight and hovering. The energy consumption of UAV sensing e i,s (t) depends on the signal beamforming in different sub-time slots when all GUs upload information through backscatter communication. Given a fixed beamforming power p A , the RF energy consumption e i,s (t) is related to the transmission strategy of the GU, i.e., where τ z is the fixed length of each sub-time slot. The energy consumption e i,r (t) = y i (t)p i,r (t)τ d during reporting can be simply modeled as a linear function of the transmission time τ d and the transmission power p i (t) when y i,r (t) = 1.

[0113] When GU-m is associated with UAV-i, i.e., x m,i = 1, its active radio frequency communication depends on the energy harvesting of UAV-i. Let represent the energy harvested by GU-m in the t-th time slot. Considering the linear energy harvesting model, the harvested energy can be estimated as follows:

[0114]

[0115] where μ is the energy conversion efficiency. When some other GU-n backscatters its information to UAV-i, i.e., z n (t) = 0, GU-m can obtain the radio frequency power s signal beamforming from UAV-i Therefore, for each time period of GU-m, the present invention has the following energy budget constraints:

[0116]

[0117] where E m (t) represents the energy state at the start of the t-th time slot, is the maximum battery capacity.

[0118] The present invention can define the energy efficiency Ξ as the time-average ratio between the total throughput received by the BS and the energy consumption of the UAV:

[0119]

[0120] Obviously, the energy efficiency depends on the access and transmission control strategies of the GUs, as well as the trajectory planning and scheduling strategies of the UAVs. Let represent the transmission control strategy of the GUs. Let represent the association and access control strategy of the GUs. Let and represent the trajectory planning and transmission scheduling strategies of the UAVs, respectively. Thus, the present invention can formulate the energy efficiency maximization problem as follows:

[0121]

[0122] The goal of the present invention is to optimize the trajectory access strategy x and reporting schedule y. The present invention also optimizes the transmission mode z of the GUs, which is related to the access control strategy of the UAVs in different time slots. For simplicity, the present invention can consider a fixed beamforming strategy in the present invention, that is, the amount of energy collected by each GU depends only on the channel conditions.

[0123] The inequality in (1) limits the minimum interference range between UAVs. The equalities in (2) and (3) represent the hybrid upload mode between the UAV and the GU. The constraints in (4)-(6) are the dynamics of the data buffers in the UAV and the GU. The constraints in (7) and (8) ensure that the energy is controllable within a certain range. In fact, the hovering power consumption e i,o (t) of the UAV is much greater than the sensing power e i,s (t) and the reporting power e i,r (t). Therefore, the power consumption of sensing and reporting can be ignored. Different transmission strategies of the GUs will significantly affect the trajectory planning and access control of the UAVs. Therefore, it is difficult to improve the energy conversion efficiency of the system by considering the control of the UAVs and the strategies of the GUs simultaneously. Another difficulty is that the UAVs should report information on the premise of avoiding interference, which will also affect the objective function.

[0124] The problem (9a) is a difficult combinatorial optimization problem. To simplify this problem, the present invention redefines (9a) as a Markov decision process (MDP), which jointly determines the strategy of the UAV and the transmission mode of the GU based on observations and past experience. Then, the present invention describes the states, actions, and rewards designed in this multi-UAV assisted network. Considering that the reconstructed MDP problem has multiple agents, and each agent needs to solve the combination of continuous variables and discrete variables, the present invention uses a multi-agent DRL algorithm to solve it. Multi-agent DRL combines deep neural networks (DNNs) and reinforcement learning (RL) in an environment where multiple agents interact. It can effectively coordinate the problems of large state spaces and dynamically changing action variables among agents.

[0125] Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is approximated as a combination of multiple single-agent DDPG agents running in parallel, i.e., a centralized training and decentralized execution scheme. Once the BS assigns the estimated actions to the UAVs, each UAV updates its own actions in a decentralized manner. Therefore, the trained actor- and critic-networks can be applied to the execution process of each drone.

[0126] The present invention first represents the state in a time slot as s t =(s1(t), s2(t),..., s N (t)). The system state s t in each time slot includes the observations of all UAVs in the network. The observation of each UAV includes energy storage, data buffer, and channel condition. The energy storage and data buffer of the UAV and GU are χ i =(E i , ξ m , Q i ), where E i is the set of energy queues of the UAV and the covered GU, and (ζ m , Q i ) is the set of all data buffers. Then, the channel condition in the network is represented as ψ i =(h i , g i ). Therefore, the present invention integerizes the system state as s i (t)=(χ i , ψ i ). In the present invention, it is assumed that all states can be measured at the beginning of the sensing slot.

[0127] Next, the present invention represents the actions of all UAVs as a t=(a1(t), a2(t),..., a N (t)). The action includes the GU transmission mode policy access control of the UAV scheduling policy y i =[y i (t)] and the trajectory

[0128] Finally, the present invention can represent the long-term reward of UAV-i as where is a discount factor Since the process of the UAV reporting information to the BS needs to be carried out on the premise that the UAV senses a certain amount of GU data. Therefore, the objective function is set such that the reward is sparse. Due to the sparsity of the objective function, the present invention introduces a guiding reward mechanism.

[0129] If the GU uploads data to the UAV, the system will obtain a guiding reward. To avoid interference and collision between different UAVs, the present invention adds a penalty term to the reward in, where I(·) is an indicator function. The present invention assumes that when the energy of the GU does not meet the requirements of its action decision, the reward value is 0. Therefore, under the condition of satisfying the energy queue constraint, the self-reward of UAV-i is defined as follows:

[0130]

[0131] where γ and η are adjustable parameters. The objective of the present invention is to select an optimal action to maximize the long-term return. represents the guiding reward, and the objective function reward is expressed as represents the penalty term.

[0132] Therefore, the total MDP reward

[0133] In addition, to evaluate the performance gain of the proposed algorithm, the present invention considers a wireless sensor network system with one BS, 2 UAVs and 6 GUs. For simplicity and intuitiveness, the present invention scales the x and y coordinates to the range of [-1, 1], assuming that the 6 GUs are randomly distributed outside the service range of the BS, so there is no direct link channel between the BS and the GUs. The UAV starts from a random starting position. More detailed parameters are listed in Table 1.

[0134] Table 1: Parameter settings in numerical simulation

[0135] Parameter Setting Training cycles per round 30 Path loss coefficient 2 Data size range of GU [5,15]M bits Maximum flight speed of UAV 25m / s Greedy parameter 0.05 Learning rate of Actor network <![CDATA[10 -3 > Learning rate of Critic network <![CDATA[10 -4 > Noise power -90dBm Initial data queue of GU [5,10]M bits

[0136] The present invention is in Figure 4The performance of the trajectory optimization algorithm was evaluated. Figure 4 In (a) and (b), the reward value function during the training process and the flight trajectory of the tested UAV after training are respectively shown. From Figure 4 As can be seen from (a), the training reward value of the present invention is increasing and finally gradually converges, which can verify the effectiveness of the algorithm during the training and learning process. In terms of the flight trajectory of the tested UAV, two UAVs take off from a random starting point respectively, and collect data from GUs along their trajectories according to the strategy of centralized training and distributed execution, as Figure 4 shown in (b). It can be seen that the UAVs cooperate with each other, have their own service areas, and there will be no conflict or interference between them.

[0137] It should be noted that the present invention can use the single DDPG algorithm or the MADDPG algorithm to solve the MDP to obtain the scheduling strategy of the communication system. Through experimental comparison, the scheduling strategy obtained by using the MADDPG algorithm can make the energy efficiency of the system higher. The specific comparative analysis is as follows:

[0138] The present invention Figure 5 evaluated the optimization performance of the single DDPG algorithm for solving the MDP to obtain the system access control strategy in Figure 5 In (a), it shows the schematic diagram of the remaining data volume of all GUs changing with time slots, and in (b), it shows the schematic diagram of the remaining data volume of the UAV changing with time slots. After determining the covered GUs, the UAV needs to allocate sensing strategies for each GU according to its specific state, so as to maximize the energy efficiency of the system.

[0139] The present invention evaluates the performance gain of the algorithm in this paper according to the observed data storage volumes of GUs and UAVs. The present invention compares the proposed method with the non-cooperative DDPG scheme. As Figure 6 shown, Figure 6 In (a), it shows the schematic diagram of the remaining data volume of all GUs changing with time slots, and in (b), it shows the schematic diagram of the remaining data volume of the UAV changing with time slots. For the convenience of simulation observation, we consider that when all the data of GUs are collected by the UAV and transmitted to the BS, all GUs generate new data volumes again. Compared with Figure 5 the single DDPG strategy in, the MADDPG algorithm applied in the present invention can transmit more data volumes within the same time period. Therefore, using the scheduling control strategy given by the method of the present invention makes the system have higher energy efficiency, and can also design access control strategies according to the task volumes and position situations of different GUs and report the tasks to the BS in time.

[0140] The UAV adaptive flight and acquisition solution proposed by the present invention can optimize the performance of emerging Internet of Things applications, improve the quality of service (reduce latency and energy consumption), and broaden the application scope of Internet of Things technologies. The multi-UAV optimization objectives for the multi-UAV assisted wireless communication network system: The objective of the present invention is to maximize the energy efficiency of the system by jointly optimizing the access strategies and flight trajectory control of multiple UAVs. The present invention uses the MADDPG algorithm to obtain the optimal solution of the model through training for the original stochastic optimization problem, comprehensively considering the influence of environmental factors and scheduling strategies in the model, reflecting the rationality of the solution, and ensuring the efficient operation of the system.

[0141] The main factors affecting the energy consumption of data transmission in the wireless communication network are the access strategy of the UAV, trajectory planning, and channel conditions. When the GU has more energy, the access control strategy allocates more active transmission time slots to the GU. However, when the GU has less energy, it becomes particularly important to allocate more passive transmission time slots to the GU. The present invention considers the actual situation more comprehensively. Through the joint optimization of the UAV access control and trajectory planning strategies, the proposed MADDPG transmission solution enables the system to achieve maximum energy efficiency even under limited channel conditions. After simulation verification, compared with the benchmark scheme, the scheme proposed by the present invention has the best performance.

[0142] Figure 7 It is the architecture diagram of the scheduling control device of the wireless communication system provided by the embodiment of the present invention, as Figure 7 shown, including:

[0143] The energy efficiency determination unit 710 is used to determine the energy efficiency of the wireless communication system; the energy efficiency is the average ratio of the total amount of data received by the BS to the total energy consumed by the wireless communication system during the entire flight period of the UAV;

[0144] The constraint condition determination unit 720 is used to determine the constraint conditions of the wireless communication system; the constraint conditions include: the distance between any two UAVs in any time slot is greater than the preset minimum distance, each GU accesses only one UAV in one time slot, only one UAV reports data to the BS in each time slot, the way for the GU to access the UAV is one of passive backscatter communication or active radio frequency active communication, the energy budget constraint of each GU in each time slot, and the amount of data reported by the UAV to the BS is determined by the distance between it and the BS and the channel conditions;

[0145] The optimization problem determination unit 730 is used to determine the combinatorial optimization problem; the combinatorial optimization problem is used to design the scheduling strategy of the wireless communication system based on the constraint conditions to maximize the energy efficiency; the scheduling strategy includes: the transmission control strategy of each GU, the flight trajectory of each UAV, and the transmission scheduling strategy of each UAV;

[0146] An MDP definition unit 740 is configured to define the combinatorial optimization problem as a Markov decision process (MDP). The total reward of the MDP includes the long-term rewards of all UAVs. The long-term reward of each UAV includes its self-reward at each decision step during its entire flight period. The self-reward includes: an objective function reward, a guidance reward, and a penalty term. If a GU uploads data to a UAV, the UAV obtains a guidance reward. When the energy of the GU does not meet the requirements of its transmission control strategy, the value of the guidance reward is 0. If the distance between any two UAVs is less than a preset minimum distance, the UAV obtains a penalty term. If a UAV successfully reports data to the BS, the UAV obtains an objective function reward.

[0147] A scheduling and solving unit 750 is configured to solve the MDP to obtain the scheduling strategy of the wireless communication system when the energy efficiency is maximized.

[0148] It can be understood that the detailed function implementation of each of the above units can be referred to the introduction in the foregoing method embodiments, and will not be elaborated here.

[0149] In addition, an embodiment of the present invention provides another scheduling control device for a wireless communication system, which includes: a memory and a processor;

[0150] The memory is used to store a computer program;

[0151] The processor is configured to implement the method in the above embodiments when executing the computer program.

[0152] In addition, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method in the above embodiments is implemented.

[0153] Based on the method in the above embodiments, an embodiment of the present invention provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.

[0154] Based on the method in the above embodiments, an embodiment of the present invention further provides a chip, which includes one or more processors and an interface circuit. Optionally, the chip may further include a bus. Wherein:

[0155] A processor may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The interface circuit can be used for sending or receiving data, instructions, or information. The processor can use the data, instructions, or other information received by the interface circuit for processing and can send the processed information through the interface circuit.

[0156] Optionally, the chip further includes a memory. The memory may include a read-only memory and a random access memory and provides operation instructions and data to the processor. A part of the memory may also include a non-volatile random access memory (NVRAM). Optionally, the memory stores executable software modules or data structures. The processor can execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions can be stored in the operating system). Optionally, the interface circuit can be used to output the execution result of the processor.

[0157] It should be noted that the respective functions corresponding to the processor and the interface circuit can be implemented through hardware design, can also be implemented through software design, or can be implemented in a combination of software and hardware. There is no limitation here. It should be understood that each step of the above method embodiment can be completed by the logic circuit in the hardware form or instructions in the software form in the processor.

[0158] It can be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments can be selectively executed according to the actual situation, can be partially executed, or can be fully executed. There is no limitation here.

[0159] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0160] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in the ASIC.

[0161] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0162] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A scheduling control method for a wireless communication system, characterized in that The method is applied to a drone-assisted wireless communication system, which includes: a base station BS, multiple drones UAV, and multiple ground users GU; the method includes the following steps: Determine the energy efficiency of the wireless communication system; the energy efficiency is the average ratio of the total data volume received by the BS to the total energy consumed by the wireless communication system during the entire flight period of the UAV; Determine the constraints of the wireless communication system; the constraints include: the distance between any two UAVs in any time slot is greater than a preset minimum distance, each GU accesses only one UAV in one time slot, only one UAV reports data to the BS in each time slot, the way for the GU to access the UAV is one of passive backscatter communication or active radio frequency communication, the energy budget constraint for each GU in each time slot, and the data volume reported by the UAV to the BS is determined by the distance between it and the BS and the channel conditions; Determine the combinatorial optimization problem; the combinatorial optimization problem is used to design the scheduling strategy of the wireless communication system based on the constraints to maximize the energy efficiency; the scheduling strategy includes: the transmission control strategy of each GU, the flight trajectories of each UAV, and the transmission scheduling strategy of each UAV; Define the combinatorial optimization problem as a Markov decision process MDP; where the total reward of the MDP includes the long-term rewards of all UAVs, and the long-term reward of each UAV includes its self-reward under each decision in its entire flight period, and the self-reward includes: objective function reward, guidance reward, and penalty term; if a GU uploads data to the UAV, the UAV obtains a guidance reward, when the energy of the GU does not meet the requirements of its transmission control strategy, the value of the guidance reward is 0, if the distance between any two drones is less than the preset minimum distance, the UAV obtains a penalty term, and if a UAV successfully reports data to the BS, the UAV obtains an objective function reward; Solve the MDP to obtain the scheduling strategy of the wireless communication system when the energy efficiency is maximized.

2. The method according to claim 1, characterized in that, Each time slot t of the UAV includes: a flight sub-time slot, a sensing sub-time slot, and a reporting sub-time slot, and the lengths of the three sub-time slots are The constraints include: d i,j (t) ≥ d min , Among them, represents the maximum flight distance, d min represents the preset minimum spacing, υ max represents the maximum flight speed, d i,j \(l_{ij}(t)\) represents the distance between the \(i\)-th UAV and the \(j\)-th UAV at the \(t\)-th time slot, the distance between the \(i\)-th UAV and the \(j\)-th UAV, l i \(l_{i}(t)\) represents the position of the \(i\)-th UAV at the \(t\)-th time slot, l i \(l_{i}(t + 1)\) represents the position of the \(i\)-th UAV at the \((t + 1)\)-th time slot, \(i\neq j\).

3. The method according to claim 2, wherein The constraints also include: where, x m,i (t) ∈ {0, 1} represents the access control policy of the m-th GU to the i-th UAV in the t-th time slot, and x m,i (t) being 0 means that the GU does not access the UAV, and x m,i (t) being 1 means that the GU accesses the UAV, represents the set of all GUs within the coverage range of the i-th UAV, and N represents the total number of UAVs.

4. The method according to claim 3, wherein The constraints also include: Data upload rate of the active radio frequency communication method is as follows: Among them, is the sub-slot allocated to the allowed access control GU, p m (t) represents the transmission power of the m-th GU in the t-th time slot, h m,i represents the channel coefficient between the i-th UAV and the m-th GU, h m,i is composed of the channel coefficient under the line-of-sight and the non-line-of-sight between the UAV and the GU; Data upload rate of passive backscatter communication mode is as follows: Among them, p A represents the fixed transmission power, and Γ o is the constant coefficient of the antenna; Let z m (t) ∈ {0, 1} represent the transmission control strategy of the m-th GU in the t-th time slot. When z m (t) = 0, the m-th GU will select the passive backscatter communication mode. When z m (t) = 1, the m-th GU selects the active radio frequency communication mode.

5. The method according to claim 4, wherein To avoid scheduling interference between UAVs, the constraints also include: where y i (t) ∈ {0, 1} represents the transmission scheduling strategy of the i-th UAV in time slot t, where y i (t) = 1 means that the UAV reports data to the BS in time slot t; When y i (t) = 1: O i (t) = τ d log(1 + p i,r (t) || g i || 2 ) Among them, O i (t) represents the amount of data reported by the i-th UAV to the BS, p i,r (t) represents the transmission power used by the i-th UAV for information forwarding, g i represents the channel condition between the UAV and the BS.

6. The method according to claim 3, wherein The constraints also include: When x m,i = 1, let denote the energy collected by the m-th GU in the t-th time slot; In each time period, the m-th GU needs to satisfy the following energy budget constraint: Among them, E m (t) represents the energy state at the start of the t-th time slot of the m-th GU, is the maximum battery capacity of the m-th GU, z n (t) represents the transmission control strategy of the n-th GU in the t-th time slot, p m (t) represents the transmit power of the m-th GU in the t-th time slot.

7. The method according to claim 5, wherein The energy efficiency of the wireless communication system is: where, Ξ represents the energy efficiency, represents the UAV time slot length, O i (t) represents the amount of data reported by the i-th UAV to the BS, y i (t) represents whether the i-th UAV plans to report data to the BS in a certain time slot, e i,o (t) represents the operating energy consumption of the UAV, e i,s (t) represents the sensing energy consumption of the UAV, e i,r (t) represents the reporting energy consumption of the UAV; The sensing energy consumption \(e\) of the UAV i,s \((t)\) and the reporting energy consumption \(e\) of the UAV i,r \((t)\) are specifically as follows: e i,r (t) = y i (t)p i,r (t)τ d Among them, represents the set of GUs allowed to access and control by the i-th UAV, 8. The method according to any one of claims 1 to 7, characterized in that, Define the combinatorial optimization problem as an MDP, specifically: Represent the state of the wireless communication system in each time slot as: s t =(s1(t), s2(t),..., s N (t)); where s i (t) represents the system state information observed by the i-th UAV; s i (t)=(χ i , ψ i ), where χ i =(E i , ξ m , Q i ) represents the energy storage and data buffering of the UAV and the GU, E i represents the set of energy queues of the UAV and the GUs covered, (ξ m , Q i ) is the set of all data buffers; ψ i =(h i , g i ) represents the channel conditions in the network, h i is the set of channel coefficients between the i-th UAV and all GUs allowed to access the i-th UAV, expressed as Represent the actions of all UAVs as a t =(a1(t), a2(t),..., a N (t)), where the action represents the transmission control strategy of the GU, represents the access control strategy of the GU to the UAV, y i =[y i (t)] represents the scheduling strategy of the UAV, represents the flight trajectory of the UAV; The self - reward r of the i - th UAV i (t) is as follows: where γ and η are both adjustable parameters, and s m,i (t) represents the size of the sensing data uploaded from the m-th GU to the i-th UAV during the sub-slot , and r p (t) is the minimum distance metric to avoid interference and collision between different UAVs; represents the guiding reward, and the objective function reward is expressed as represents the penalty term, and I(·) represents the indicator function; The long-term reward of the i-th UAV over the entire time period is the discount factor; The total reward 9. A scheduling control device for a wireless communication system, characterized in that, The device is applied to a drone-assisted wireless communication system, which includes: a base station BS, multiple drones UAV, and multiple ground users GU; the device includes: An energy efficiency determination unit, configured to determine the energy efficiency of the wireless communication system; the energy efficiency is the average ratio of the total data volume received by the BS to the total energy consumed by the wireless communication system during the entire flight period of the UAV; A constraint determination unit for determining the constraints of a wireless communication system; the constraints include: the distance between any two UAVs in any time slot is greater than a preset minimum distance, each GU accesses only one UAV in one time slot, only one UAV reports data to the BS in each time slot, the way for the GU to access the UAV is one of passive backscatter communication or active radio frequency communication, the energy budget constraint of each GU in each time slot, and the data volume reported by the UAV to the BS is determined by the distance between it and the BS and the channel conditions; An optimization problem determination unit for determining a combinatorial optimization problem; the combinatorial optimization problem is used to design the scheduling strategy of the wireless communication system based on the constraints to maximize the energy efficiency; the scheduling strategy includes: the transmission control strategy of each GU, the flight trajectories of each UAV, and the transmission scheduling strategy of each UAV; An MDP definition unit for defining the combinatorial optimization problem as a Markov decision process MDP; wherein, the total reward of the MDP includes the long-term rewards of all UAVs, and the long-term reward of each UAV includes its self-reward under each decision step during its entire flight period, and the self-reward includes: an objective function reward, a guidance reward, and a penalty term; if a GU uploads data to the UAV, the UAV obtains a guidance reward, and when the energy of the GU does not meet the requirements of its transmission control strategy, the value of the guidance reward is 0; if the distance between any two UAVs is less than the preset minimum distance, the UAV obtains a penalty term; if a UAV successfully reports data to the BS, the UAV obtains an objective function reward; A scheduling solution unit for solving the MDP to obtain the scheduling strategy of the wireless communication system when the energy efficiency is maximized.

10. A scheduling control device for a wireless communication system, characterized in that, Comprising: A memory and a processor; The memory is used for storing a computer program; The processor is used for implementing the method according to any one of claims 1-8 when executing the computer program.