Space-time-frequency parameter optimized air-ground hybrid communication method

By establishing UAV trajectory and channel models, combining TDD and FDMA technologies, and using reinforcement learning to optimize parameters, the problem of insufficient optimization of UAV communication system parameters in the time domain and frequency domain is solved, comprehensive coverage of ground nodes is achieved, and the effectiveness of emergency communications is improved.

CN119316845BActive Publication Date: 2025-10-17XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411432256.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-10-17
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

In the existing technology, drone communication systems fail to achieve full coverage in terms of time domain and frequency domain parameter optimization, resulting in poor communication performance of ground nodes, especially in disaster situations when the communication infrastructure is damaged and cannot be effectively restored.

Method used

A UAV trajectory model and channel model are established, and a hybrid communication architecture is adopted, combining TDD and FDMA technologies. Reinforcement learning is used to optimize the UAV trajectory, uplink and downlink ratios, and sub-bandwidth division. The total number of UEs covered by communication is optimized, and the PPO algorithm is used for parameter optimization.

Benefits of technology

It improves the effective coverage of drone network communications, ensures the efficient operation of emergency communication scenarios, and achieves comprehensive coverage of ground nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316845B_ABST
    Figure CN119316845B_ABST
Patent Text Reader

Abstract

The application discloses a space-time-frequency parameter optimization considering air-to-ground hybrid communication method, comprising: establishing a UAV trajectory and a channel model; based on the hybrid communication architecture, a communication link rule is established: each time slot is divided into uplink and downlink transmission stages based on TDD, the time length is divided according to the uplink and downlink ratio, based on FDMA, the total frequency bandwidth of the system is divided into the sub-frequency bandwidths of each UE according to the sub-bandwidth division ratio, the communication decision is determined according to the uplink and downlink data transmission rate, and the uplink and downlink data transmission rate is obtained based on the uplink and downlink ratio, the multiple sub-bandwidth division ratios and the channel model; based on the communication decision, the UAV trajectory, the uplink and downlink ratio and the multiple sub-bandwidth division ratios are used as optimization parameters, the total number of UEs covered by communication is used as an optimization target, and a problem model is established by adding constraint conditions; a reinforcement learning model is established, and a strategy optimization algorithm is introduced to optimize the optimization parameters to realize the optimization target. The application can obtain the best performance of the communication system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of communication, and particularly relates to a space-to-ground hybrid communication method considering optimization of space, time domain and frequency domain parameters. BACKGROUND

[0002] In recent years, the rapid development of unmanned aerial vehicles and communication technology has provided a solid technical foundation for low-altitude economy. In emergency rescue, the rapid and accurate transmission of information is the key to the entire communication. The damage of communication infrastructure caused by disasters such as earthquakes and heavy rains will lead to communication paralysis in the entire disaster area. In this case, unmanned aerial vehicles equipped with airborne communication base station systems can be quickly deployed to establish temporary communication networks and restore the communication capabilities of the disaster area, ensuring smooth rescue and command and dispatch.

[0003] Liu et al. studied the optimal aerial positioning and resource allocation of unmanned aerial vehicle base stations in Internet of Things (IoT) networks in the document "Liu Y, Liu K, Han J, Zhu L, Xiao Z, Xia X, Resource Allocation and 3D Placement for UAV-Enabled Energy-Efficient IoT Communications, IEEE Internet of Things Journal, 2020, PP(99): 1-1", aiming to minimize the total transmission power of IoT devices. They used the K-Means algorithm for clustering, the improved HD4M algorithm for subchannel allocation, and finally the alternating iteration method to jointly optimize the transmission power of IoT devices and the height of unmanned aerial vehicles.

[0004] Hattab et al. also conducted network resource optimization in the document "Hattab G, Cabric D, Energy-Efficient Massive IoT Shared Spectrum Access over UAV-enabled Cellular Networks, IEEE Transactions on Communications, 2020, PP(99): 1-1". They proposed a time division duplex transmission protocol to provide shared spectrum access and improve the average allocation of resources between IoT devices and users (UEs) through a random geometry protocol, but UEs will be subject to additional interference from IoT devices. Therefore, they optimized the nominal transmission power of IoT devices to maximize their energy efficiency while limiting interference to UEs.

[0005] In the Internet of Things network, Park et al. proposed a power splitting (PS) strategy in the document "Park G, Lee K, Optimization of the Trajectory, Transmit Power, and Power Splitting Ratio for Maximizing the Available Energy of a UAV Aded SWIPT System, Sensors, 2022, 22(23): 9081" to jointly optimize the trajectory, transmission power and power ratio of the UAV to maximize the average spectrum efficiency and ensure the minimum average available energy for ground user equipment.

[0006] The above methods optimize the UAV space, transmission power, time domain and frequency domain to varying degrees and in varying combinations, but ignore the simultaneous optimization of time domain and frequency domain parameters, which leaves room for improvement in system performance.

[0007] However, due to the limitations of UAV (Unmanned Aerial Vehicle) trajectories, time domain resources and frequency domain resources, the number of ground nodes for effective communication cannot achieve full coverage. Summary of the Invention

[0008] To address the aforementioned issues in the prior art, the present invention provides an air-to-ground hybrid communication method, device, and electronic device that optimizes parameters in the spatial, time, and frequency domains. The technical issues addressed by the present invention are achieved through the following technical solutions:

[0009] In a first aspect, an embodiment of the present invention provides an air-to-ground hybrid communication method that considers optimization of space, time domain, and frequency domain parameters, the method comprising:

[0010] For scenarios involving multiple UAVs and user nodes (UEs), a UAV trajectory model is established, and a channel model between the UAV and UE is established based on line-of-sight and non-line-of-sight links.

[0011] Based on the designed hybrid communication architecture, communication link rules are established, including: each time slot is divided into uplink and downlink transmission phases based on TDD technology, and the transmission duration is divided according to the uplink and downlink ratio; based on FDMA technology, the total system frequency bandwidth in each time slot is divided into sub-frequency bandwidths for each UE according to multiple sub-bandwidth division ratios matching the number of UEs; the communication decision between the UE and the UAV in each time slot is determined according to the uplink and downlink data transmission rates, which are calculated based on the uplink and downlink ratio, the multiple sub-bandwidth division ratios, and the channel model;

[0012] Based on the communication decision, taking the UAV trajectory, the uplink and downlink ratio, and the multiple sub-bandwidth division ratio as the optimization parameters, taking the total number of UEs covered by the communication as the optimization target, and adding multiple constraint conditions, a problem model is established;

[0013] A reinforcement learning model is established, and a preset strategy optimization algorithm is introduced to optimize the optimization parameters of the problem model, so as to realize the optimization target.

[0014] In an embodiment of the present application, the UAV trajectory model is represented as:

[0015]

[0016] In the scene, J UAVs are deployed above the ground, represented as I UEs are deployed on the ground, represented as The flight cycle of the UAV is divided into T equal time periods, represented as Each time period is a time slot, represented by a duration τ, and τ is small enough; the jth UAV is represented as UAV j , whose coordinates at time slot t are represented as And respectively represent the horizontal coordinates and vertical coordinates of the UAV j at time slot t, h j represents the flight height of the UAV j ; represents the initial coordinates of the UAV j ; represents the flight distance of the UAV j at time slot t; represents the flight angle of the UAV j at time slot t, and the value range is [0, 2π];

[0017] In the UAV trajectory model, the UAV flight coordinate constraint is represented as:

[0018]

[0019] Wherein, l max represents the maximum flight distance of the UAV in a time slot.

[0020] In an embodiment of the present application, the channel model between the UAV and the UE is established based on the line-of-sight link and the non-line-of-sight link, comprising:

[0021] The line-of-sight connection probability and the non-line-of-sight connection probability between the UAV j and the UE i in time slot t are determined, respectively represented as:

[0022]

[0023] where LoS and NLoS represent line-of-sight and non-line-of-sight, respectively; UE i represents the ith UE; represents the line-of-sight connection probability between the UAV j and the UE i in slot t, represents the non-line-of-sight connection probability between the UAV j and the UE i in slot t; a and b are environmental constants; represents the elevation angle between the UAV j and the UE i in slot t;

[0024] determines the LoS average path loss and the NLoS average path loss between the UAV j and the UE i in slot t, which are represented as:

[0025]

[0026] where, and represent the LoS average path loss and the NLoS average path loss between the UAV j and the UE i in slot t, respectively; f represents the channel transmission frequency of the UAV j and the UE i in Hz; represents the distance between the UAV j and the UE i in slot t; c represents the speed of light; η LoS and η NLoS represent the additional loss on the LoS path and the NLoS path, respectively;

[0027] According to the line-of-sight connection probability and the non-line-of-sight connection probability between the UAV j and the UE i in slot t, the LoS average path loss and the NLoS average path loss between the UAV j and the UE i in slot t, the average path loss between the UAV j and the UE i in slot t is determined and represented as:

[0028]

[0029] where, is the average path loss between the UAV jand UE i The average path loss between is used to characterize the channel status.

[0030] In one embodiment of the present invention, in the communication link rule,

[0031] UAV in time slot t j and UE i The calculation formula for the uplink data transmission rate is:

[0032]

[0033] UAV in time slot t j and UE i The calculation formula for the downlink data transmission rate is:

[0034]

[0035] in, is the UAV in time slot t j and UE i Uplink data transmission rate; is the UAV in time slot t j and UE i Downlink data transmission rate; k t is the uplink and downlink ratio in time slot t, k t τ represents the transmission duration of the uplink transmission phase in time slot t, (1-k t )τ represents the transmission duration of the downlink transmission phase in time slot t; w represents the size of the total frequency bandwidth of the system, and the sub-frequency bandwidth of I UE is expressed as a i is the sub-bandwidth allocation ratio of the i-th UE; P ij Indicates UE i to UAV j The transmission power; represents the UAV in time slot t j and UE i The average path loss between ji Indicates UAV j and UE i The transmission power; represents the UAV in time slot t j and UE i The average path loss between σ0 represents the power spectral density of the additive white Gaussian noise at the receiving end.

[0036] In one embodiment of the present invention, in the communication link rule, the communication decision is expressed as:

[0037]

[0038] wherein, denotes the communication decision of the UE i in the time slot t, denotes the maximum value of the uplink data transmission rate of the UE i with all UAVs; denotes the maximum value of the downlink data transmission rate of the UE i with all UAVs; iu,min denotes the minimum value of the uplink data transmission rate requirement; id,min denotes the minimum value of the downlink data transmission rate requirement; if denotes that the UE i establishes a communication link with the UAV and is within the communication coverage range; if denotes that the UE i does not establish a communication link with the UAV and is not within the communication coverage range.

[0039] In an embodiment of the present application, the expression of the problem model is:

[0040]

[0041] wherein, OP denotes the problem model; max denotes the maximum value; denotes the set of sub-bandwidth division ratios; denotes the set of UAV trajectories;

[0042] denotes the set of uplink / downlink ratios; s.t. the multiple items contained therein denote constraint conditions, wherein (a) denotes whether the UE i is covered; (b) is a constraint satisfied by the sub-bandwidth division ratio; (c)-(f) are constraint conditions for the UAV trajectory, (c) limits the flight angle θ of the UAV, (d) limits the flight distance l of the UAV within a period; (e) limits the coordinates of the UAV in the next time slot; (f) limits the flight area of the UAV; x t and y t denote the horizontal and vertical coordinates of the UAV in the time slot t; L and H denote the length and width of the flight area of the UAV.

[0043] In an embodiment of the present application, the reinforcement learning model comprises an MDP model, and the preset policy optimization algorithm comprises a PPO algorithm.

[0044] In an embodiment of the present application, the process of establishing the MDP model comprises:

[0045] According to the UE i and the UAV juplink data transmission rate, downlink data transmission rate and channel state, a state space is constructed and represented as:

[0046]

[0047] According to the UAV flight angle, UAV flight distance, uplink / downlink ratio and sub-bandwidth division ratio in the time slot t, an action space is constructed and represented as:

[0048]

[0049] The reward function is set as:

[0050]

[0051] Wherein, N t is the total number of UEs covered by communication in time slot t; N t is the total number of UEs covered by communication in time slot t-1.

[0052] In a second aspect, the embodiments of the present application provide a space-time-frequency parameter optimization considering air-ground hybrid communication device, the device comprises:

[0053] The UAV trajectory model and channel model establishment module is used to establish a UAV trajectory model for a scenario containing multiple UAVs and user nodes UEs, and establish a channel model between the UAV and the UE based on a line-of-sight link and a non-line-of-sight link;

[0054] The communication link rule establishment module is used to establish a communication link rule based on the designed hybrid communication architecture, including: each time slot is divided into uplink and downlink transmission stages based on TDD technology, and the transmission duration is divided according to the uplink / downlink ratio; based on FDMA technology, the total frequency bandwidth in each time slot is divided into sub-frequency bandwidths of each UE according to a plurality of sub-bandwidth division ratios matched with the number of UEs; the communication decision between the UE and the UAV in each time slot is determined according to the uplink / downlink data transmission rate, which is calculated based on the uplink / downlink ratio, the plurality of sub-bandwidth division ratios and the channel model;

[0055] The problem model establishment module is used to establish a problem model based on the communication decision, taking the UAV trajectory, the uplink / downlink ratio and the plurality of sub-bandwidth division ratios as optimization parameters, taking the total number of UEs covered by communication as an optimization target, and adding a plurality of constraint conditions.

[0056] The parameter optimization module is used to establish a reinforcement learning model, and introduce a preset policy optimization algorithm to optimize the optimization parameters of the problem model, so as to realize the optimization target.

[0057] In a third aspect, the embodiments of the present application provide an air-to-ground communication system including a plurality of unmanned aerial vehicles (UAVs) and user nodes (UEs), and the air-to-ground communication system implements air-to-ground communication by using the air-to-ground hybrid communication method considering the optimization of spatial, time domain and frequency domain parameters.

[0058] The present application has the following beneficial effects:

[0059] In the air-to-ground hybrid communication method considering the optimization of spatial, time domain and frequency domain parameters provided by the embodiments of the present application, based on the air-to-ground communication system, a UAV trajectory model and a channel model are established, a hybrid communication architecture is proposed, the communication link rules between the ground user nodes and the UAVs are defined, the optimization of parameters in the aspects of space, time domain and frequency domain is considered, the UAV trajectory, the uplink and downlink ratio and the multiple sub-bandwidth division ratio are taken as the optimization parameters, the total number of UEs covered by communication is taken as the optimization target, multiple constraint conditions are added, a problem model is established, the optimization of the optimization parameters is realized by using a policy optimization algorithm through the establishment of a reinforcement learning model, the optimization target is achieved, and the best performance of the communication system is obtained. The present application improves the number of effective coverage UE nodes of the UAV network communication, and provides a new method for efficient operation of emergency communication and the like. BRIEF DESCRIPTION OF DRAWINGS

[0060] Figure 1 A flowchart of the air-to-ground hybrid communication method considering the optimization of spatial, time domain and frequency domain parameters provided by the embodiments of the present application is shown in the figure.

[0061] Figure 2 A scene diagram of the embodiments of the present application including a plurality of unmanned aerial vehicles (UAVs) and user nodes (UEs) is shown in the figure.

[0062] Figure 3 A structure diagram of the hybrid communication architecture provided by the embodiments of the present application is shown in the figure.

[0063] Figure 4 A structure diagram of the air-to-ground hybrid communication device considering the optimization of spatial, time domain and frequency domain parameters provided by the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0064] The present application will be further described in detail below in combination with specific embodiments, but the embodiments of the present application are not limited thereto.

[0065] In order to optimize the UAV running trajectory, the uplink and downlink time slot ratio and the bandwidth allocation parameters in the communication process between the ground multiple nodes and the UAVs, and achieve the purpose of optimizing the performance of the communication system, the embodiments of the present application provide an air-to-ground hybrid communication method, device and air-to-ground communication system considering the optimization of spatial, time domain and frequency domain parameters, which can be used in the field of social public safety, such as emergency communication.

[0066] It should be noted that the execution subject of the space, time domain and frequency domain parameter optimization considering space, time domain and frequency domain parameter optimization of the air-to-ground hybrid communication method provided by the embodiment of the application can be a space, time domain and frequency domain parameter optimization considering space, time domain and frequency domain parameter optimization of the air-to-ground hybrid communication device, which can run in an electronic device. The electronic device can be a server of a back-end control center, a cloud server, a UAV edge computing device, and the like, but is not limited thereto.

[0067] In a first aspect, the embodiment of the application provides a space, time domain and frequency domain parameter optimization considering space, time domain and frequency domain parameter optimization of an air-to-ground hybrid communication method, as shown in the method can include the following steps: Figure 1

[0068] S1, for a scene containing multiple unmanned aerial vehicles (UAVs) and user nodes (UEs), a UAV trajectory model is established, and a channel model between the UAVs and the UEs is established based on a line-of-sight link and a non-line-of-sight link;

[0069] Figure 2 is a schematic diagram of a scene containing multiple unmanned aerial vehicles (UAVs) and user nodes (UEs); in the scene, J UAVs are deployed above the ground, denoted as The purpose is to use UAVs as air base stations to quickly restore communication; I UEs are deployed on the ground, which can be communication equipment of users in the disaster area, denoted as The flight cycle of the UAV is divided into T equal time periods, denoted as Wherein each time period is a time slot, denoted by duration τ, assuming that τ is small enough and equal in all time periods, the position of the UAV remains relatively constant in each time period; the jth UAV is denoted as UAV j , whose coordinates at time slot t are denoted as and respectively represent the horizontal coordinate and the vertical coordinate of the UAV j at time slot t, h j represents the flight height of the UAV j , and the flight height is constant; represents the initial coordinates of the UAV j ; represents the flight distance of the UAV j at time slot t; represents the flight angle of the UAV j at time slot t, the value range is [0, 2π]; the maximum flight speed of the UAV is v max , then the maximum flight distance of the UAV in each time slot t can be represented as l max = v max τ.

[0070] ​Therefore, the horizontal and vertical coordinates of the UAV in the time slot t, i.e., the UAV trajectory model, can be expressed as:

[0071]

[0072] wherein the UAV flight coordinate constraint is expressed as:

[0073]

[0074] wherein l max represents the maximum flight distance of the UAV in a time slot.

[0075] In the embodiment of the present application, the channel model covers two types of line-of-sight (LoS) and non-line-of-sight (NLoS), and specifically, the channel model between the UAV and the UE is established based on the line-of-sight link and the non-line-of-sight link, including the following steps:

[0076] 1) determining the line-of-sight connection probability and the non-line-of-sight connection probability between the UAV j and the UE i in the time slot t, which are respectively expressed as:

[0077]

[0078] wherein LoS and NLoS respectively represent line-of-sight and non-line-of-sight; UE i represents the i-th UE; represents the line-of-sight connection probability between the UAV j and the UE i in the time slot t, represents the non-line-of-sight connection probability between the UAV j and the UE i in the time slot t; and α and β are environmental constants; represents the elevation angle between the UAV j and the UE i in the time slot t;

[0079] 2) determining the LoS average path loss and the NLoS average path loss between the UAV j and the UE i in the time slot t, which are respectively expressed as:

[0080]

[0081] wherein, and respectively represent the LoS average path loss and the NLoS average path loss between the UAV j and the UE i in the time slot t; and f represents the distance between the UAV j and the UE iThe channel transmission frequency of the channel is Hz. The distance between the UAV j and the UE i ; c represents the speed of light; η LoS and η NLoS respectively represent the additional loss on the LoS path and the additional loss on the NLoS path, η LoS and η NLoS are known values.

[0082] 3) According to the line-of-sight connection probability, the non-line-of-sight connection probability between the UAV j and the UE i in the time slot t, the LoS average path loss, the NLoS average path loss between the UAV j and the UE i in the time slot t, the average path loss between the UAV j and the UE i in the time slot t is determined, which is represented as:

[0083]

[0084] Wherein, is the average path loss between the UAV j and the UE i in the time slot t, which is used to characterize the channel state.

[0085] S2, based on the designed hybrid communication architecture, establishes a communication link rule, including: each time slot is divided into uplink and downlink transmission stages based on TDD technology, and the transmission duration is divided according to the uplink and downlink ratio; based on the FDMA technology, the total frequency bandwidth in each time slot is divided into the sub-frequency bandwidths of each UE according to a plurality of sub-bandwidth division ratios matched with the number of UEs; the communication decision between the UE and the UAV in each time slot is determined according to the uplink and downlink data transmission rate, and the uplink and downlink data transmission rate is calculated based on the uplink and downlink ratio, the plurality of sub-bandwidth division ratios and the channel model.

[0086] In the embodiment of the application, a hybrid communication architecture, referred to as HybridComm, is designed, which is divided into two layers of frameworks, the upper layer adopts TDD technology to realize full-duplex communication, and the lower layer adopts FDMA to realize multiple access of multiple ground nodes. The specific block diagram is shown in 3, and the communication link rule is established according to the hybrid communication architecture as follows.

[0087] Specifically, the full-duplex communication function is realized by TDD technology. In the TDD technology, each time slot is divided into uplink and downlink signal transmission stages. Figure 3 For example, the time slot t has a time duration of τ, and the user UE iThe uplink signal transmission phase to the UAV, the transmission duration is kτ, the UAV to UE i The downlink signal transmission phase, the transmission duration is (1-k)τ, the uplink and downlink ratio k is between 0 and 1, the uplink and downlink ratio k of the time slot t is dynamically adjusted according to the data volume of the uplink and downlink in the time slot t-1, and this part is realized by establishing a problem model and optimizing parameters by using reinforcement learning. For details, please understand S3-S4 in the following text.

[0088] The ground multi-node multiple access function is realized by the FDMA technology. In the FDMA technology, the simultaneous communication of multiple users is realized by dividing the available frequency spectrum into multiple independent frequency channels. Specifically, the system total frequency bandwidth is divided according to the division ratio of a plurality of sub-bandwidths matched with the number of UEs, each user is allocated a fixed sub-frequency bandwidth, which can be used continuously during the call process, and the signals between different users do not interfere with each other. FDMA supports parallel communication, allows multiple users to transmit data at the same time, and can multiplex the same frequency in different geographical areas to improve the frequency spectrum utilization. Therefore, the entire system total frequency bandwidth is divided into I non-overlapping sub-frequency bandwidths, denoted as Wherein, w represents the size of the system total frequency bandwidth, a i represents the sub-bandwidth division ratio of the i-th UE, and the sum of all sub-bandwidth division ratios is 1.

[0089] Due to factors such as distance and environment, the network throughput is different when the UAV establishes a communication link with different UEs, which directly affects the network communication quality. In the present application, the communication decision is represented by If the maximum transmission rate between the UE i and the J UAVs is greater than or equal to a certain threshold, a communication link is established with the UAV with the maximum transmission rate, denoted as If the maximum transmission rate between the UE i and the J UAVs is less than a certain threshold, no communication link is established, denoted as

[0090] According to the related settings of the hybrid communication architecture described in the foregoing, the uplink data transmission rate of the UAV j and the UE i in the time slot t is calculated as follows:

[0091]

[0092] The downlink data transmission rate of the UAV j and the UE i in the time slot t is calculated as follows:

[0093]

[0094] Wherein, the uplink data transmission rate of the UAV j and the UE i in slot t; the downlink data transmission rate of the UAV j and the UE i in slot t; k t is the uplink / downlink ratio in slot t, k t τ represents the transmission duration of the uplink transmission phase in slot t, (1-k t )τ represents the transmission duration of the downlink transmission phase in slot t; w represents the size of the total frequency bandwidth of the system, and the sub-frequency bandwidth of the I UEs is represented as a i is the sub-bandwidth division ratio of the i-th UE; P ij represents the transmission power of the UE i to the UAV j ; represents the average path loss between the UAV j and the UE i in slot t; P ji represents the transmission power of the UAV j and the UE i ; represents the average path loss between the UAV j and the UE i in slot t, which is equal to σ0 represents the power spectral density of the additive white Gaussian noise at the receiving end.

[0095] The communication decision is represented as:

[0096]

[0097] wherein, represents the communication decision of the UE i in slot t, represents the maximum value of the uplink data transmission rate of the UE i and all UAVs; represents the maximum value of the downlink data transmission rate of the UE i and all UAVs; c iu,min represents the minimum value of the uplink data transmission rate requirement; c id,min represents the minimum value of the downlink data transmission rate requirement; if represents that the UE i establishes a communication link with the UAV and is within the communication coverage range; if represents that the UE i does not establish a communication link with the UAV and is not within the communication coverage range.

[0098] S3, based on the communication decision, taking the UAV trajectory, the uplink-downlink ratio, and the multiple sub-bandwidth division ratios as optimization parameters, taking the total number of UEs in communication coverage as an optimization target, and adding multiple constraint conditions to establish a problem model;

[0099] The embodiment of the application optimizes the total number of UEs in communication coverage, which is a key indicator depending on the data transmission rate between the UAV and the UE. To achieve maximum UE coverage, a holistic approach is needed, and the embodiment of the application optimizes several key parameters simultaneously, including the UAV trajectory, the uplink-downlink ratio in TDD technology, and the sub-bandwidth division ratio in FDMA technology.

[0100] Specifically, according to the foregoing steps, the UE i transmission rate between the UAV j , thereby obtaining a function of the total number of UEs in communication coverage and constructing a problem model taking the function as an objective function, and considering constraint conditions such as the value range of the communication decision, the sum of all sub-bandwidth division ratios, the flight angle range of the UAV, the flight distance range of the UAV in a time slot, the value range of the horizontal and vertical coordinates of the UAV in a time slot, and the flight range of the UAV in the entire cycle.

[0101] The expression of the problem model is:

[0102]

[0103] Wherein, OP represents the problem model; max represents the maximum value; represents the set of sub-bandwidth division ratios; represents the set of UAV trajectories;

[0104] represents the set of uplink-downlink ratios; s.t. contains multiple items representing constraint conditions, wherein (a) represents whether the UE i is covered; (b) is a constraint satisfied by the sub-bandwidth division ratio; (c)-(f) are constraint conditions for the UAV trajectory, (c) limits the flight angle θ of the UAV, (d) limits the flight distance l of the UAV in a cycle; (e) limits the coordinates of the UAV in the next time slot; (f) limits the flight area of the UAV; x t and y t represent the horizontal and vertical coordinates of the UAV in a time slot; L and H represent the length and width of the flight area of the UAV.

[0105] S4, a reinforcement learning model is established, and a preset policy optimization algorithm is introduced to optimize the optimization parameters of the problem model to achieve the optimization target.

[0106] The embodiment of the application can utilize an existing reinforcement learning algorithm, construct a reinforcement learning model, and introduce a related policy optimization algorithm to optimize the above-mentioned optimization parameters. For example, in an optional implementation, the reinforcement learning model can include an MDP (Markov Decision Process) model, and the preset policy optimization algorithm can include a PPO (Proximal Policy Optimization) algorithm.

[0107] Accordingly, specifically, the process of establishing the MDP model can include:

[0108] 1) According to the uplink data transmission rate, the downlink data transmission rate, and the channel state of the UE i and the UAV j in the time slot t, a state space is constructed, denoted as:

[0109]

[0110] 2) According to the flight angle, the flight distance, the uplink / downlink ratio, and the sub-bandwidth division ratio of the UAV in the time slot t, an action space is constructed, denoted as:

[0111]

[0112] 3) The reward function is set as:

[0113]

[0114] wherein the reward is an index for measuring the action effect, and the flight position, the uplink / downlink ratio, and the sub-bandwidth division ratio of the UAV in each time period are random. When designing the reward and punishment function, in order to quickly converge and find the optimal solution, the application gives a reward when the coverage rate of adjacent time periods increases, gives a punishment when the coverage rate decreases, gives a reward when the coverage rate is unchanged, and gives a punishment when the coverage rate increases. In addition, in order to ensure uninterrupted communication of the ground rescue equipment, the same punishment is applied when there is no coverage. t Nt is the total number of communication-covered UEs in the time slot t; Nt-1 is the total number of communication-covered UEs in the time slot t-1. t Nt is the total number of communication-covered UEs in the time slot t; Nt-1 is the total number of communication-covered UEs in the time slot t-1.

[0115] After establishing the MDP model, the PPO algorithm is introduced to optimize the optimization parameters of the problem model. The PPO algorithm is a prominent policy gradient method widely used in DRL (Deep Reinforcement Learning), which aims to enhance training stability by limiting the difference between the updated policy and its previous version.

[0116] In the embodiment of the present invention, the process of introducing the PPO algorithm to optimize the optimization parameters of the problem model includes:

[0117] Step 1: Input parameters, including the policy network parameter set θ π , value network parameter set φ V ;

[0118] Step 2: For each value of K, execute steps 3 to 7, where K represents the number of training cycles of the PPO algorithm and is a positive integer.

[0119] Step 3, by running the policy π in the environment K =π(θ K ) to collect the trajectory set Among them, the input state space and the output action space are the mean and standard deviation of each dimension. Each action can be sampled from a multidimensional Gaussian distribution. The trajectory set refers to the interaction between the agent and the environment under the guidance of the current strategy, collecting a series of state, action, reward and next state data, τ i ={(s 1 ,a 1 ,r 1 ,s 2 ),(s 2 ,a 2 ,r 2 ,s 2+1 ),...,(s T ,a T ,r T ,s T+1 )};

[0120] Regarding states, actions, and rewards, please refer to the previous state space, action space, reward function, and the existing PPO algorithm for understanding. I will not explain them in detail here.

[0121] Step 4: Calculate the cumulative return

[0122] Where γ is the discount factor;

[0123] Step 5: Based on the current value function Calculating odds estimates The calculation formula is: in, represents the cumulative reward after selecting action a in state s; where δ t is the time difference error at time t, expressed as s t+1 is the state value at time t+1, s t is the state value at time t; λ is the GAE parameter;

[0124] Step 6, update the policy by maximizing the PPO-Clip objective, which is expressed as:

[0125]

[0126] Among them, π θ (a t |s t ) represents the state s under the current strategy t Take action a t probability; Indicates the state s under the old policy t Take action a t probability; is the advantage function estimate; ε is the hyperparameter of the shear range; Represents the probability ratio of the policy to be pruned to avoid excessive policy updates that affect stability.

[0127] Step 7, fit the value function through mean square error regression, expressed as:

[0128]

[0129] Among them, V φ (s t ) is the state s t The value function of .

[0130] In general, the current strategy Next, we collect the state, action, reward, and next state value of the MDP model in time slot t, and use the value network To calculate the advantage estimate for each state-action pair Finally, the policy network parameter set θ is completed π and the value network parameter set φ V value.

[0131] Through the above steps, the optimization parameters of each time slot can be continuously optimized within a flight cycle to achieve the optimization goal.

[0132] In the air-to-ground hybrid communication method that considers the optimization of space, time domain and frequency domain parameters provided in the embodiment of the present invention, a UAV trajectory model and a channel model are established based on the air-to-ground communication system, and a hybrid communication architecture is proposed. The communication link rules between ground user nodes and UAVs are defined, and parameter optimization in space, time domain and frequency domain is considered. The UAV trajectory, uplink and downlink ratio, and the ratio of multiple sub-bandwidth divisions are used as optimization parameters, and the total number of UEs covered by the communication is used as the optimization target. Multiple constraints are added to establish a problem model, and the optimization parameters are optimized by using a strategy optimization algorithm through the establishment of a reinforcement learning model to achieve the optimization target and obtain the optimal performance of the communication system. The present invention increases the number of UE nodes effectively covered by UAV network communication, and provides a new method for the efficient operation of scenarios such as emergency communications.

[0133] In the second aspect, corresponding to the above method embodiment, the embodiment of the present invention further provides an air-to-ground hybrid communication device considering the optimization of space, time domain and frequency domain parameters, such as Figure 4 As shown, the device includes:

[0134] The UAV trajectory model and channel model establishment module is used to establish the UAV trajectory model for scenarios containing multiple UAVs and user nodes UE, and to establish the channel model between the UAV and UE based on line-of-sight links and non-line-of-sight links;

[0135] A communication link rule establishment module is configured to establish communication link rules based on the designed hybrid communication architecture, including: each time slot is divided into uplink and downlink transmission phases based on TDD technology, and the transmission duration is divided according to the uplink and downlink ratio; based on FDMA technology, the total system frequency bandwidth in each time slot is divided into sub-frequency bandwidths for each UE according to multiple sub-bandwidth division ratios matching the number of UEs; the communication decision between the UE and the UAV in each time slot is determined according to the uplink and downlink data transmission rates, which are calculated based on the uplink and downlink ratio, the multiple sub-bandwidth division ratios, and the channel model;

[0136] A problem model building module is used to build a problem model based on the communication decision, using the UAV trajectory, uplink and downlink ratio, and multiple sub-bandwidth division ratios as optimization parameters, the total number of UEs covered by the communication as the optimization target, and adding multiple constraints;

[0137] The parameter optimization module is used to establish a reinforcement learning model and introduce a preset strategy optimization algorithm to optimize the optimization parameters of the problem model to achieve the optimization goal.

[0138] For the specific processing procedures of each module of the device, please refer to the relevant content of the first aspect and will not be elaborated here.

[0139] The space, time domain and frequency domain parameter optimization considering air-to-ground hybrid communication device provided by the embodiment can improve the effective coverage of the UAV network communication UE node quantity, and provides a new scheme for efficient operation of emergency communication and the like.

[0140] In a third aspect, corresponding to the method embodiments, the embodiment of the present application also provides an air-to-ground communication system including a plurality of unmanned aerial vehicles (UAVs) and user nodes (UEs), wherein the air-to-ground communication system adopts the space, time domain and frequency domain parameter optimization considering air-to-ground hybrid communication method of the first aspect to realize air-to-ground communication. For specific communication methods, please refer to the content of the first aspect, which will not be repeated here.

[0141] The air-to-ground communication system provided by the embodiment adopts the space, time domain and frequency domain parameter optimization considering air-to-ground hybrid communication method to realize air-to-ground communication, which can obtain the best performance of the system and improve the effective coverage of the UAV network communication UE node quantity.

[0142] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in the present specification.

[0143] Each embodiment in the present specification is described in a relevant manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments. Especially, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the part of the method embodiment.

[0144] The above only describes the preferred embodiments of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application is included in the protection scope of the present application.

Claims

1. An air-to-ground hybrid communication method considering optimization of space, time and frequency domain parameters, characterized in that: include: For scenarios involving multiple UAVs and user nodes (UEs), a UAV trajectory model is established, and a channel model between the UAV and UE is established based on line-of-sight and non-line-of-sight links. Based on the designed hybrid communication architecture, communication link rules are established, including: each time slot is divided into uplink and downlink transmission phases based on TDD technology, and the transmission duration is divided according to the uplink and downlink ratio; based on FDMA technology, the total system frequency bandwidth in each time slot is divided into sub-frequency bandwidths for each UE according to multiple sub-bandwidth division ratios matching the number of UEs; the communication decision between the UE and the UAV in each time slot is determined according to the uplink and downlink data transmission rates, which are calculated based on the uplink and downlink ratio, the multiple sub-bandwidth division ratios, and the channel model; Based on the communication decision, a problem model is established with the UAV trajectory, uplink and downlink ratio, and multiple sub-bandwidth division ratios as optimization parameters, the total number of UEs covered by the communication as the optimization target, and multiple constraints added; Establishing a reinforcement learning model and introducing a preset strategy optimization algorithm to optimize the optimization parameters of the problem model to achieve the optimization goal; Among them, in the communication link rule, the communication decision is expressed as: ; In the scenario, there are J UAVs deployed above the ground, represented by ; There are 1 UE deployed on the ground, represented by ; The UAV flight cycle is divided into T equal time periods, expressed as , where each time period is a time slot, with duration Indicates that, and Small enough; , , indicating time slot middle communication decisions, express With all The maximum uplink data transmission rate; express With all The maximum value of the downlink data transmission rate; Indicates the minimum value required for uplink data transmission rate; Indicates the minimum value required for downlink data transmission rate; if ,express and Establish a communication link within the communication coverage area; if ,express and The communication link has not been established and the device is not within the communication coverage area.

2. The air-to-ground hybrid communication method considering spatial, time and frequency domain parameter optimization according to claim 1, characterized in that: The UAV trajectory model is expressed as: ; Among them, UAVs are represented as , which is in the time slot The coordinates are expressed as ; and Respectively In the time slot The horizontal and vertical coordinates of express Flight altitude; express The initial coordinates of express In the time slot Flight distance; express In the time slot The flight angle range is ; In the UAV trajectory model, the UAV flight coordinate constraint is expressed as: ; ; in, Indicates the maximum flight distance of the drone in a time slot.

3. The air-to-ground hybrid communication method considering space, time and frequency domain parameter optimization according to claim 2, characterized in that: The establishing of a channel model between the UAV and the UE based on the line-of-sight link and the non-line-of-sight link includes: Determine time slot middle and The line-of-sight connection probability and the non-line-of-sight connection probability are expressed as: ; ; Among them, LoS and NLoS represent line-of-sight and non-line-of-sight respectively; Indicates the indivual ; Indicates time slot middle and The probability of line-of-sight connection between Indicates time slot middle and The probability of non-line-of-sight connection between and is an environmental constant; Indicates time slot middle and The elevation angle between Determine time slot middle and The average LoS path loss and the average NLoS path loss between them are expressed as: ; ; in, and Represents time slots middle and The average LoS path loss and the average NLoS path loss between them; express and The channel transmission frequency, in Hz; Indicates time slot middle and the distance between them; represents the speed of light; and They represent the additional loss on the LoS path and the additional loss on the NLoS path respectively; According to time slot middle and The probability of line-of-sight connection, the probability of non-line-of-sight connection, and the time slot middle and The average LoS path loss and the average NLoS path loss between them determine the time slot middle and The average path loss between is expressed as: ; in, It is a time slot middle and The average path loss between is used to characterize the channel status.

4. The air-to-ground hybrid communication method considering space, time and frequency domain parameter optimization according to claim 3, characterized in that: In the communication link rules, Time Slot middle and The calculation formula for the uplink data transmission rate is: ; Time Slot middle and The calculation formula for the downlink data transmission rate is: ; in, Time slot middle and Uplink data transmission rate; Time slot middle and Downlink data transmission rate; Time slot The ratio of upper and lower levels, Indicates time slot The transmission duration of the uplink transmission phase, Indicates time slot The transmission duration of the downlink transmission phase; Represents the size of the total frequency bandwidth of the system, and the sub-frequency bandwidth of I UE is expressed as , For the The sub-bandwidth allocation ratio of each UE; express arrive The transmission power; Indicates time slot middle and The average path loss between express and The transmission power; Indicates time slot middle and The average path loss between ; Represents the power spectral density of the additive white Gaussian noise at the receiver.

5. The air-to-ground hybrid communication method considering space, time and frequency domain parameter optimization according to claim 4, characterized in that: The expression of the problem model is: ; Among them, OP represents the problem model; max represents finding the maximum value; , represents the set of sub-bandwidth division ratios; ,express A collection of trajectories; A set representing the uplink and downlink ratios; The multiple items contained represent constraints, where (a) represents (b) is the constraint satisfied by the sub-bandwidth division ratio; (c) to (f) are the constraints of the drone trajectory, (c) limits the drone's flight angle (d) Limit the flight distance of the drone within the cycle ;(e) limit the coordinates of the drone in the next time slot;(f) limit the flight area of ​​the drone; and Indicates that the drone is in the time slot The horizontal and vertical coordinates of and Indicates the length and width of the drone's flight area.

6. The air-to-ground hybrid communication method considering space, time and frequency domain parameter optimization according to claim 5, characterized in that: The reinforcement learning model includes an MDP model, and the preset policy optimization algorithm includes a PPO algorithm.

7. The air-to-ground hybrid communication method considering space, time and frequency domain parameter optimization according to claim 6, characterized in that: The process of establishing an MDP model includes: According to time slot middle and The uplink data transmission rate, downlink data transmission rate and channel state are used to construct the state space, which is expressed as: ; According to time slot The UAV flight angle, UAV flight distance, uplink and downlink ratio, and sub-bandwidth division ratio are used to construct the action space, which is expressed as: ; Set the reward function to: ; in, Time slot The total number of UEs covered by communication; Time slot The total number of UEs covered by communication.

8. An air-to-ground hybrid communication device considering optimization of space, time and frequency domain parameters, characterized in that: include: The UAV trajectory model and channel model establishment module is used to establish the UAV trajectory model for scenarios containing multiple UAVs and user nodes UE, and to establish the channel model between the UAV and UE based on line-of-sight links and non-line-of-sight links; A communication link rule establishment module is configured to establish communication link rules based on the designed hybrid communication architecture, including: each time slot is divided into uplink and downlink transmission phases based on TDD technology, and the transmission duration is divided according to the uplink and downlink ratio; based on FDMA technology, the total system frequency bandwidth in each time slot is divided into sub-frequency bandwidths for each UE according to multiple sub-bandwidth division ratios matching the number of UEs; the communication decision between the UE and the UAV in each time slot is determined according to the uplink and downlink data transmission rates, which are calculated based on the uplink and downlink ratio, the multiple sub-bandwidth division ratios, and the channel model; A problem model building module is used to build a problem model based on the communication decision, using the UAV trajectory, uplink and downlink ratio, and multiple sub-bandwidth division ratios as optimization parameters, the total number of UEs covered by the communication as the optimization target, and adding multiple constraints; A parameter optimization module is used to establish a reinforcement learning model and introduce a preset strategy optimization algorithm to optimize the optimization parameters of the problem model to achieve the optimization goal; Among them, in the communication link rule, the communication decision is expressed as: ; In the scenario, there are J UAVs deployed above the ground, represented by ; There are 1 UE deployed on the ground, represented by ; The UAV flight cycle is divided into T equal time periods, expressed as , where each time period is a time slot, with duration Indicates that, and Small enough; , , indicating time slot middle communication decisions, express With all The maximum uplink data transmission rate; express With all The maximum value of the downlink data transmission rate; Indicates the minimum value required for uplink data transmission rate; Indicates the minimum value required for downlink data transmission rate; if ,express and Establish a communication link within the communication coverage area; if ,express and The communication link has not been established and the device is not within the communication coverage area.

9. An air-to-ground communication system, characterized in that: The air-to-ground communication system comprises multiple unmanned aerial vehicles (UAVs) and user nodes (UEs). The air-to-ground communication system realizes air-to-ground communication by adopting the air-to-ground hybrid communication method considering the optimization of space, time domain and frequency domain parameters as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle (UAV)-based air-to-ground hybrid communication method, device and communication system

    CN119316844A