Unmanned aerial vehicle communication situation awareness method based on deep reinforcement learning

CN120474595BActive Publication Date: 2026-09-22BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510429669.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2025-04-08
Publication Date
2026-09-22
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

所述环境信息有助于解决未知环境中的抗截获安全传输问题

Benefits of technology

[0089]1.与传统空基应急通信系统的路径规划与资源优化技术相比,本发明公开的一种基于深度强化学习的无人机通信态势感知方法,无人机充分发挥在应急通信中的作用,实现有限资源条件下的系统感知、计算和通信资源的合理分配,有助于紧急情况下敏捷抗截获的通信。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474595B_ABST
    Figure CN120474595B_ABST
Patent Text Reader

Abstract

The unmanned aerial vehicle communication situation awareness method based on deep reinforcement learning belongs to the field of satellite communication. For unmanned aerial vehicle assisted satellite communication, an integrated framework based on perception, transmission, calculation and navigation is constructed. The framework deploys a multifunctional unmanned aerial vehicle with integrated sensing and computing. It enhances physical layer security through target perception and improves transmission efficiency using computing power. The mobile trajectory, sensing and computing power and strategy of the unmanned aerial vehicle are jointly designed to maximize data transmission efficiency while meeting safety requirements. The situation awareness-based unmanned aerial vehicle trajectory and multi-domain resource joint planning method uses deep reinforcement learning to adaptively plan the trajectory of the unmanned aerial vehicle. In an unknown eavesdropping environment, it considers data calculation and uploading, gradually constructs an eavesdropping environment situation map, balances the transmission, calculation and perception tasks of the unmanned aerial vehicle under energy constraints, and realizes the cooperation among the transmission, calculation and perception tasks of the unmanned aerial vehicle, thereby improving the communication security and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a UAV communication situational awareness method based on deep reinforcement learning, belonging to the field of satellite communication. Background Technology

[0002] High-orbit satellite communication boasts wide coverage, high reliability, and large communication capacity, making it a common method for emergency communication. However, it is easily obstructed by obstacles, resulting in blind spots. Unmanned aerial vehicles (UAVs), with their high mobility and flexible communication networking capabilities, can serve as aerial relays to assist communication between satellites and ground users. Furthermore, the highly open nature of satellite channels makes them vulnerable to potential eavesdropping threats; therefore, secure transmission is a significant challenge for high-orbit satellite communication.

[0003] In recent years, advanced beamforming techniques have been developed to achieve physical layer security for high-orbit satellite communications, such as precoding, interference alignment, or artificial noise. However, most research assumes that potential eavesdroppers have been identified and that the serving base station is somehow aware of some of their Channel State Information (CSI). In reality, however, potential eavesdroppers typically do not interact with the base station, making their presence difficult to detect. This makes obtaining relevant CSI extremely challenging, becoming a key constraint on secure wireless network transmission.

[0004] When some base stations are damaged or out of service, both terrestrial and satellite communication networks alone will struggle to provide communication coverage to ground users in areas inaccessible by vehicles or in satellite-denied scenarios due to mountainous or forested terrain. Drones, however, due to their mobility, flexibility, and ease of deployment, can combine with terrestrial and satellite networks to penetrate communication blind spots, providing communication coverage in complex terrain areas and becoming a communication hub between the outside world and emergency zones. Firstly, during the drone's approach and data collection phase, based on emergency communication needs, drones carry various types of sensor equipment, such as cameras and infrared sensors, and can quickly fly to the target area to collect various data in real time, such as disaster images, personnel needs, and communication information. Simultaneously, drones can utilize their sensing capabilities based on integrated communication and sensing signals to achieve proactive perception of target drones or the environment. This effectively extracts the impact of dynamic targets such as drones and static targets such as the environment on wireless signal characteristics, assisting in security monitoring and eavesdropping avoidance. Once the drone reaches the communication coverage area, it will transmit the collected data to the receiving end via an air link. The data will then be forwarded to the emergency communication processing system to quickly establish an emergency private network, providing timely and reliable communication support and rescue services to the affected areas, and creating a situational awareness map. Summary of the Invention

[0005] In the context of UAV-assisted satellite communication in unreliable and complex channel environments, this invention aims to provide a UAV communication situational awareness method based on deep reinforcement learning. This method utilizes onboard sensing devices on the UAV and integrates sensing capabilities into existing wireless networks to acquire the necessary environmental information for the target area. This environmental information helps address the problem of secure transmission against interception in unknown environments. By employing deep reinforcement learning for adaptive trajectory planning of the UAV, the method enables coordination between UAV transmission, computation, and sensing tasks, thereby improving communication security and stability.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] This invention discloses a UAV communication situational awareness method based on deep reinforcement learning. For UAV-assisted satellite communication in unreliable and complex channel environments, it constructs a comprehensive framework integrating perception, transmission, computation, and navigation. This framework deploys a multi-functional UAV with integrated sensing, computation, and communication capabilities. It enhances physical layer security through target perception and improves transmission efficiency using computational power. Simultaneously, it jointly designs the UAV's movement trajectory, sensing and computational power, and strategies. While meeting security requirements, it maximizes data transmission efficiency. The situational awareness-based UAV trajectory and resource joint planning method utilizes deep reinforcement learning for adaptive trajectory planning. In unknown eavesdropping environments, it balances data computation and uploading while gradually constructing an eavesdropping environment situation map. Based on the eavesdropping environment situation map, it balances the UAV's transmission, computation, and perception tasks under energy constraints, achieving synergy among these tasks and improving communication security and stability.

[0008] The UAV communication situational awareness method disclosed in this invention includes the following steps:

[0009] Step 1: Establish a multi-domain fusion model of the aerospace emergency communication system, divide the communication area of ​​the high-orbit satellite in the aerospace emergency communication system model into a grid, and obtain the distance that the UAV moves between time slots.

[0010] A multi-domain fusion model of an aerospace emergency communication system is established. This model includes a high-orbit satellite and an auxiliary communication UAV. The ground area served by the high-orbit satellite has communication blind spots due to terrain. An airborne eavesdropper, Eve, exists within the high-orbit satellite's communication area, but its location is obtained through radar sensing. When the channel quality in the communication blind spot is severely compromised, preventing direct satellite-to-ground communication, the UAV acts as a relay, collecting disaster information in the blind spot and then flying to the satellite's coverage area to upload data. During the UAV's upload process, Eve eavesdrops on the information. Throughout the flight and upload process, the UAV sends sensing signals to determine Eve's location and simultaneously performs onboard computation to compress and upload the data. The UAV is configured with a uniform linear array of NT antennas for communication and sensing, while Eve uses a single antenna for eavesdropping. The positions of the satellite and Eve are fixed; Eve's position is unknown to the UAV; the UAV moves in a two-dimensional plane.

[0011] The high-orbit satellite communication area is divided into I×I grids, each grid containing two attributes: satellite channel status and eavesdropping channel status. Satellite channel status is determined by geographical location and environment, exhibiting a degree of randomness. Eavesdropping channel status is determined by the location of the Eve (or similar feature), with grid areas closer to the Eve being more susceptible to eavesdropping. UAVs (Unmanned Aerial Vehicles) all start from a fixed starting point and move in grid units, advancing one grid per time slot or remaining in one position, returning to the starting point after one complete flight to form a closed loop. The total flight time slots of the UAV are represented as... Where T end The time to return to the starting point; the UAV performs sensing calculations and transmissions on each grid at time intervals ΔT; in time slot t, the UAV is located at the lth time slot. t =[x t ,y t ] Grid; Eve's position l e =[x e ,y e High-orbit satellite position l s =[x s ,y s The height difference between the drone and the flight plane is H; Indicates the lth t and l t+1 The distance between grids, and satisfying

[0012]

[0013] Among them x∈[1,I], y∈[1,I].

[0014] Step 2: Establish a transmission model based on statistical CSI to model the channel; obtain the security interruption probability based on the transmission model.

[0015] Satellite service areas are affected by the environment; therefore, satellite communication channels are modeled as uncertainties. The UAV in the lth... t The channel state of the grid is

[0016]

[0017] in Here, f represents the path loss from the UAV to the satellite, and f is the operating frequency. For drones in the l t Distance between the grid and the satellite; |g s | 2 The channel fading factor characterizes the impact of complex environmental obstructions on communication between UAVs and satellites; these impacts include scattering effects, masking effects, and a combination of scattering and masking effects; |g s | 2 The probability density function PDF satisfies

[0018]

[0019] Where 1F1(m; 1; δx) is the confluence hypergeometry function, and Nakagami-m fading parameter Real number m∈(0,∞), μ is the average power of the line-of-sight component, and 2κ is the average power of the multipath component; superimposing equation (3) yields the cumulative distribution function CDF, expressed as follows:

[0020]

[0021] in It is an increasing power function. It is a gamma function; Represents the incomplete gamma function; integrating (2) yields |g i | 2 cumulative distribution function

[0022] The satellite channel status of the target area is G t ={g 1,1 ,…,g i,j ,…,g I,I}, g i,j Indicate l t = Satellite channel state at position [i,j];

[0023] l t Location eavesdropping channel state is represented as

[0024] (5) For the path loss of eavesdroppers, Let θ be the distance between the drone and Eve. t∈[-π / 2,π / 2] is the azimuth angle from the UAV to Eve, a(θ) t ) represents the drone-eavesdropping link steering vector;

[0025] In time slot t, the UAV transmit signal is

[0026] x(t)=w(t)s com (t)+w e (t)s rad (t) (6)

[0027] Where s com (t) and s rad (t) represent the communication and sensing parts of the integrated sensing signal, respectively. and This is the precoding matrix;

[0028] The signal received by the satellite is represented as

[0029]

[0030] in For l t Location communication channel status, It is Gaussian white noise;

[0031] Eve received the signal as

[0032]

[0033] in For l t Location eavesdropping channel status, It is Gaussian white noise;

[0034] Define an auxiliary variable r(t) as the lower bound of the safe rate achieved by the satellite for the UAV in time slot t. Then the achieved safe throughput C(t) is expressed as:

[0035] C(t) = I C ([C s (t)-C e (t)] + ≥r(t))·r(t) (9)

[0036] Among them I C (·) is an indicator function, when [C s (t)-C e (t)] + When I ≥ r(t) C (·)=1, r(t) is the minimum expected communication rate per time slot link, C s (t) and C e(t) represents the transmission rate from the drone to the satellite and the eavesdropping rate, respectively, expressed as...

[0037]

[0038] as well as

[0039]

[0040] Where B is the bandwidth;

[0041] The probability of safe interruption from drone to satellite meets the following conditions.

[0042] Pr{[C s (t)-C e (t)] + ≤r(t)}≤ε (12)

[0043] Where ε≤0.1 is the maximum allowable interruption probability.

[0044] Step 3: Use radar signals to sense Eve's location and estimate the state of the eavesdropping channel.

[0045] Eve's location parameters were estimated based on the radar echo signal; the potential eavesdropper's location uncertainty model is as follows:

[0046]

[0047] In the formula Δl is the location estimate obtained when detecting a potential eavesdropper. e (t) represents the position measurement error; In the actual location l e The distribution at (t) follows a Gaussian distribution, i.e.

[0048]

[0049] Where δ is the measurement standard deviation; the measurement error of the t-th time slot satisfies

[0050]

[0051] Among them G MF For MF gain;

[0052] The distribution measured from multiple time slots is accumulated as follows:

[0053]

[0054] in satisfy To sum the measurement standard deviation of the distribution function; set the measurement accuracy threshold to Ω, when At that time, Eve's positional awareness is sufficient;

[0055] The estimated eavesdropping channel state is in This indicates that the estimated location of Eve is In this case, l t =The state of the eavesdropping channel at position [i,j].

[0056] Step 4: Compress the data file and transmit the compressed file to the satellite.

[0057] Each UAV mission collects N data files to be transmitted from communication blind spots, and these data files are independent of each other. During this mission, the UAV needs to upload all the data it carries to the satellite. Each data file is associated with an onboard computing task, and the nth onboard computing task is represented by (d...). an ,d bn ,c n ) description, where c n To calculate the number of CPU revolutions required for this task, d an and d bn These represent the data sizes before and after the calculation; define the indicator variable x. n (t)∈{0,1} represents the computation strategy for the nth task in the t-th time slot, x n (t) = 1 indicates that the data for the nth task has been processed; otherwise, x n (t) = 0; d is defined for simplicity. a =[d a1 ,…,d aN ] T ,d b =[d b1 ,…,d bN ] T and c = [c1,…,c N ] T ;

[0058] Utilizing the drone's computing capabilities, data files are compressed during flight uploads to reduce data transmission volume; t time slot, drone position l t Carrying compressed data Uncompressed data volume Its satisfaction

[0059]

[0060] as well as

[0061]

[0062] in The amount of data uploaded from the compressed queue, The amount of data uploaded to the uncompressed queue, ΔT is the time slot length; This represents the reduction in uncompressed queue data after calculation. To calculate the increase in data in the compressed queue, x(t) = [x1(t), ..., x...]. N (t)] T The calculation strategy for the t-th time slot.

[0063] Step 5: Analyze the various constraints and establish a situational awareness optimization problem.

[0064] Analyze the power consumption constraints on the UAV; power consumption includes three parts: communication power consumption, computing power consumption, and sensing power consumption; the total power of the UAV mission in each time slot is constrained, with the total power P... t satisfy

[0065]

[0066] Where P w (t)=|w(t)| 2 For transmission power, P e (t)=|w e (t)| 2 To sense power, To calculate power, To calculate the power consumption factor, f c (t) represents the calculated frequency, P max The maximum power allowed per time slot;

[0067] To jointly optimize limited airborne resources and address the issue of efficient and secure data transmission in satellite coverage blind spots under eavesdropping conditions, a situational awareness optimization problem is established. This problem aims to ensure data is uploaded to the satellite in the shortest possible time while maintaining secure transmission, and simultaneously allows the UAV to return to its starting point for the next mission. Therefore, the objective function and corresponding constraints of the overall situational awareness optimization problem are expressed by the following formula.

[0068]

[0069] st(9),(18), (20a)

[0070] c T x(t)≤ΔTf c (t), t∈[1,T] s (20b)

[0071] 0≤f c (t)≤f max ,t∈[1,T s (20c)

[0072] x n (t)∈{0,1},t∈[1,T s (20d)

[0073] L I ,f c ,P w ,P s ,T s ,X I ∈U (20e) where To perceive the path, For the sensing phase, calculate the frequency, transmission power, and calculation strategy; T s For the perception strategy, i.e., the moment when the UAV's perception ends, it satisfies To calculate the frequency, For transmission power variables, To sense power variables.

[0074] Step Six: Applying Deep Reinforcement Learning to the Situational Awareness Optimization Problem The solution is then performed to obtain the eavesdropping environment situation.

[0075] Due to situational awareness optimization problem The objective function in the problem has the characteristic of long-term accumulation, which makes the situational awareness optimization problem... The problem is transformed into a Markov decision problem, and the DRL method is used to solve the specific problem.

[0076] Optimize situational awareness problem Transformed into a Markov decision process, i.e., a tuple For state space, For the action space, For reward space, γ represents the transition probability and the reward discount factor; at time t, the UAV first observes the current environmental information to obtain the state. Then, take action. Acting on the environment, and then using the transition probability Transition to the next state s t+1 And generate a reward r after the action is performed. t ;

[0077]

[0078] G t These represent the eavesdropping channel status and satellite channel status of the target area sensed at time t, respectively.

[0079] a t ={l t+1 ,P w(t+1),P s (t+1),f c (t+1),x(t+1),b t} (twenty two)

[0080] Where b t ∈{0,1} indicates whether the perception phase has ended, b t =1 means that the perception phase ends at the current moment; as can be seen from (22), l t+1 ,x(t+1),b t P is a discrete variable. w (t+1),P s (t+1),f c (t+1) is a continuous variable; therefore, the agent's output action is a mixed high-dimensional vector, which increases the complexity of the search space for finding the optimal action.

[0081] r t =R(t)+X1(t)+X2(t) (23)

[0082] The constant penalty terms X1(t) and R(t) are expressed as follows:

[0083] R(t)=ξ(P s (t),t)-ξ(P s (t-1),t-1) (24)

[0084] The penalty term X2(t) with lag is expressed as

[0085]

[0086] The trained Actor network is independently deployed on the drone, and performs trajectory planning and resource allocation in real time according to the environment; finally, a situational awareness map is obtained, achieving the minimum completion time under the constraint of the lower limit of data completion.

[0087] Step Seven: Based on Steps Five and Six, construct a joint planning method for UAV trajectory and resources based on situational awareness. This method utilizes deep reinforcement learning to perform adaptive trajectory planning for the UAV, gradually constructing a situational awareness map of the eavesdropping environment while simultaneously handling data computation and uploading under unknown eavesdropping conditions. Based on this situational awareness map, the UAV's transmission, computation, and sensing tasks are balanced under energy constraints, achieving coordination among these tasks.

[0088] Beneficial effects:

[0089] 1. Compared with the path planning and resource optimization techniques of traditional airborne emergency communication systems, the present invention discloses a UAV communication situational awareness method based on deep reinforcement learning. This method enables UAVs to fully play their role in emergency communication, realize the rational allocation of system perception, computing and communication resources under limited resource conditions, and help to achieve agile and anti-interception communication in emergency situations.

[0090] 2. Compared with the multi-domain resource allocation technology of traditional airborne emergency communication systems, this invention discloses a UAV communication situational awareness method based on deep reinforcement learning. This method utilizes deep reinforcement learning to construct a joint planning method for UAV trajectory and multi-domain resources based on situational awareness. The joint planning method uses deep reinforcement learning to perform adaptive trajectory planning for the UAV, gradually constructing an eavesdropping environment situation map while simultaneously handling data computation and uploading in an unknown eavesdropping environment. Based on the eavesdropping environment situation map, the UAV's transmission, computation, and sensing tasks are balanced under energy constraints, achieving coordination among these tasks.

[0091] 3. Compared with the multi-domain resource allocation technology of traditional airborne emergency communication systems, the UAV communication situational awareness method disclosed in this invention, based on deep reinforcement learning, implements a specialized reward and punishment mechanism for different dimensions of multi-dimensional actions in the multi-domain resource planning of UAVs based on reward decomposition DDPG. On the basis of overall performance improvement, it better reflects the independent contribution of each action dimension to the system performance and successfully and gradually constructs the eavesdropping environment situation map. Attached Figure Description

[0092] Figure 1 This is a flowchart of the UAV communication situational awareness method based on deep reinforcement learning in this invention;

[0093] Figure 2 This is a model diagram of an integrated UAV-assisted satellite communication system that combines perception, transmission, computation, and navigation in the UAV communication situational awareness method based on deep reinforcement learning of this invention.

[0094] Figure 3 This invention describes the environmental situation assessment process in the UAV communication situational awareness method based on deep reinforcement learning, wherein: Figure 3 In the middle (a), the initial unknown state of the channel environment (step = 0) is represented. Figure 3 (b) represents the initial intermediate state of the channel environment (step = 3). Figure 3 (c) represents the initial intermediate state of the channel environment (step = 5). Figure 3 In the middle (d), the initial complete state of the channel environment (step=7) is shown. Detailed Implementation

[0095] The present invention will now be described in detail with reference to the accompanying drawings and embodiments, while also discussing the technical problems solved by the present invention and its beneficial effects. It should be noted that the described embodiments are intended to facilitate understanding of the present invention and do not constitute any limitation thereof.

[0096] Example 1

[0097] This embodiment describes the application of a deep reinforcement learning-based UAV communication situational awareness method in UAV-assisted satellite communication in an emergency rescue environment in a primeval forest. During a geological exploration mission, an exploration team ventured into a primeval forest for geological surveys. However, due to the complex terrain and dense vegetation, traditional communication methods were obstructed, causing the exploration team to lose contact with the outside world. Satellite communication was ineffective due to forest cover, and ground base station signals could not provide coverage, posing a severe challenge to the rescue operation. To restore communication as quickly as possible, the rescue team decided to deploy UAVs deep into the forest, utilizing their mobility, flexibility, and rapid deployment capabilities, combined with ground and satellite networks, to provide coverage in communication blind spots, becoming a crucial communication hub between the outside world and the trapped personnel. The rescue UAV, equipped with high-precision cameras, infrared sensors, lidar, and other equipment, quickly flew to the target area to conduct comprehensive data collection. The UAV could acquire real-time images through high-resolution imaging, monitor the exploration team's movement trajectory, and use infrared sensors to detect vital signs of personnel, effectively locating them even in low light or complex environments. In addition, the UAV also collected meteorological data for the area, such as temperature, humidity, wind speed, and air quality, providing support for subsequent rescue decisions. Once the drone enters the communication coverage area, it transmits the collected data to the ground receiver via a high-bandwidth air link, and then forwards it to the emergency communication processing system. In the system's backend, the rescue team analyzes the transmitted data through a visualization platform to create a precise situational awareness map, including information such as the exploration team's current location, environmental risk assessment, and feasible rescue routes, providing a scientific basis for command and dispatch. Furthermore, the drone collaborates with other rescue drones or relay stations through a distributed network to expand communication coverage and improve the response efficiency of rescue operations.

[0098] like Figure 1 As shown in the figure, this embodiment discloses a UAV communication situational awareness method based on deep reinforcement learning. The specific implementation steps are as follows:

[0099] Step 1: Establish a multi-domain fusion model of the aerospace emergency communication system based on synergy and computing. Divide the communication area of ​​the high-orbit satellite in the model into a grid to obtain the distance the UAV moves between time slots.

[0100] Establish a model of a space-based emergency communication system that integrates multiple domains of sensing and computing; the system model diagram is as follows. Figure 2As shown, the model includes a high-orbit satellite and an auxiliary communication drone (UAV). The ground area served by the high-orbit satellite has communication blind spots due to terrain. An aerial eavesdropper, Eve, exists within the high-orbit satellite's communication area, but its location can be obtained through radar sensing. When the channel quality in the communication blind spot is severely compromised, preventing direct satellite-to-ground communication, the UAV acts as a relay, collecting disaster information in the blind spot and then flying to the satellite's communication coverage area to upload the data. During the UAV's upload process, Eve eavesdrops on the information. Throughout the flight and upload process, the UAV sends sensing signals to determine Eve's location and simultaneously performs onboard computation to compress and upload the data. The UAV is configured with a uniform linear array of NT antennas for communication and sensing, while Eve uses a single antenna for eavesdropping. The positions of the satellite and Eve are fixed; Eve's position is unknown to the drone; the UAV moves in a two-dimensional plane.

[0101] The high-orbit satellite communication area is divided into a 10×10 grid, with each grid containing two attributes: satellite channel status and eavesdropping channel status. The UAV starts at (0m, 0m, Hm) and travels at a constant speed of 20m / s, hovering for 5 seconds at each point to perform data transmission and computational tasks. The satellite channel status is determined by geographical location and environment, exhibiting a degree of randomness; the eavesdropping channel status is determined by the location of Eve, with grid areas closer to Eve being more susceptible to eavesdropping. All UAVs start from a fixed starting point and move in grid units, advancing one grid or remaining in one position per time slot, returning to the starting point after one complete flight to form a closed loop. The total flight time slots of the UAV are represented as... Where T end The time to return to the starting point; the UAV performs sensing calculations and transmissions on each grid at time intervals ΔT; in time slot t, the UAV is located at the lth time slot. t =[x t ,y t ] Grid; Eve's position l e =[x e ,y e High-orbit satellite position l s =[x s ,y s The height difference between the drone and the flight plane is H; Indicates the lth t and l t+1 The distance between grids, and satisfying

[0102]

[0103] Where x∈[1,I], y∈[1,I];

[0104] Step 2: Establish a transmission model based on statistical CSI to model the channel and obtain the security interruption probability;

[0105] Satellite service areas are affected by the environment; therefore, satellite communication channels are modeled as uncertainties. The UAV in the lth... t The channel state of the grid is

[0106]

[0107] in Here, f represents the path loss from the UAV to the satellite, and f is the operating frequency. For drones in the l t Distance between the grid and the satellite; |g s | 2 The channel fading factor characterizes the impact of complex environmental obstructions on communication between UAVs and satellites; these impacts include scattering effects, masking effects, and a combination of scattering and masking effects; |g s | 2 The probability density function PDF satisfies

[0108]

[0109] Where 1F1(m; 1; δx) is the confluence hypergeometry function, and Nakagami-m fading parameter Real number m∈(0,∞), μ is the average power of the line-of-sight component, and 2κ is the average power of the multipath component; superimposing equation (3) yields the cumulative distribution function CDF, expressed as follows:

[0110]

[0111] in It is an increasing power function. It is a gamma function; Represents the incomplete gamma function; integrating (2) yields |g i | 2 cumulative distribution function

[0112] The satellite channel status of the target area is G t ={g 1,1 ,…,g i,j ,…,g I,I}, g i,j Indicate l t = Satellite channel state at position [i,j];

[0113] l t Location eavesdropping channel state is represented as

[0114]

[0115] in For the path loss of eavesdroppers, Let θ be the distance between the drone and Eve. t ∈[-π / 2,π / 2] is the azimuth angle from the UAV to Eve, a(θ) t ) represents the drone-eavesdropping link steering vector;

[0116] In time slot t, the UAV transmit signal is

[0117] x(t)=w(t)s com (t)+w e (t)s rad (t) (6)

[0118] Where s com (t) and s rad (t) represent the communication and sensing parts of the integrated sensing signal, respectively. and This is the precoding matrix;

[0119] The signal received by the satellite is represented as

[0120] (7)

[0121] in For l t Location communication channel status, It is Gaussian white noise;

[0122] Eve received the signal as

[0123]

[0124] in For l t Location eavesdropping channel status, It is Gaussian white noise;

[0125] Define an auxiliary variable r(t) as the lower bound of the safe rate achieved by the satellite for the UAV in time slot t. Then the achievable safe throughput C(t) is expressed as:

[0126] C(t) = I C ([C s (t)-C e (t)] + ≥r(t))·r(t) (9)

[0127] Among them I C (·) is an indicator function, when [C s (t)-C e (t)]+ When I ≥ r(t) C (·)=1, r(t) is the minimum expected communication rate per time slot link, C s (t) and C e (t) represents the transmission rate from the drone to the satellite and the eavesdropping rate, respectively, expressed as...

[0128]

[0129] as well as

[0130]

[0131] Where B is the bandwidth;

[0132] The probability of safe interruption from drone to satellite meets the following conditions.

[0133] Pr{[C s (t)-C e (t)] + ≤r(t)}≤ε (12)

[0134] Where ε≤0.1 is the maximum allowable interruption probability;

[0135] Step 3: Use radar signals to sense Eve's location and estimate the state of the eavesdropping channel;

[0136] The location of Eve is determined using radar signals. The radar echo signal received by the UAV is represented as follows:

[0137]

[0138] in For round-trip channel states, satisfying

[0139]

[0140] This is Eve's radar cross-section. It is noise.

[0141] Eve's location parameters were estimated based on the radar echo signal; the potential eavesdropper's location uncertainty model is as follows:

[0142]

[0143] In the formula Δl is the location estimate obtained when detecting a potential eavesdropper. e (t) represents the position measurement error; the main consideration is that the error in estimating the position is caused by noise, therefore Δl e (t) follows a Gaussian distribution. Therefore, from Δl e(t) Calculation results of Eve's location affected by the event In the actual location l e The region around (t) follows a Gaussian distribution, i.e.

[0144]

[0145] Where δ is the measurement standard deviation; the measurement error of the t-th time slot satisfies

[0146]

[0147] Among them G MF For MF gain;

[0148] The distribution measured from multiple time slots is accumulated as follows:

[0149]

[0150] in satisfy The standard deviation of the cumulative distribution function is used; since each measured location is distributed near the true location, the cumulative probability distribution of multiple measurements should have a maximum value at the true location of the radiation source. A threshold of measurement accuracy is set to Ω. At that time, Eve's positional awareness is sufficient;

[0151] The estimated eavesdropping channel state is in This indicates that the estimated location of Eve is In this case, l t =The state of the eavesdropping channel at position [i,j];

[0152] Step 4: Compress the data file and transmit the compressed file to the satellite;

[0153] Each UAV mission collects N data files to be transmitted from communication blind spots, and these data files are independent of each other. During this mission, the UAV needs to upload all the data it carries to the satellite. Each data file is associated with an onboard computing task, and the nth onboard computing task is represented by (d...). an ,d bn ,c n ) description, where c n To calculate the number of CPU revolutions required for this task, d an and d bn These represent the data sizes before and after the calculation; define the indicator variable x. n (t)∈{0,1} represents the computation strategy for the nth task in the t-th time slot, x n (t) = 1 indicates that the data for the nth task has been processed; otherwise, xn (t) = 0; d is defined for simplicity. a =[d a1 ,…,d aN ] T ,d b =[d b1 ,…,d bN ] T and c = [c1,…,c N ] T ;

[0154] Utilizing the drone's computing capabilities, data files are compressed during flight upload to reduce data transmission volume; it is assumed that the drone will prioritize transmitting already compressed data; time slot t, drone position l t Carrying compressed data Uncompressed data volume Its satisfaction

[0155]

[0156] as well as

[0157]

[0158] in The amount of data uploaded from the compressed queue, The amount of data uploaded to the uncompressed queue, ΔT is the time slot length; This represents the reduction in uncompressed queue data after calculation. To calculate the increase in data in the compressed queue, x(t) = [x1(t), ..., x...]. N (t)] T The calculation strategy for the t-th time slot;

[0159] Step 5: Analyze the various constraints and establish a situational awareness optimization problem.

[0160] First, we analyze the power consumption constraints on the UAV. Power consumption includes three parts: communication power consumption, computing power consumption, and sensing power consumption. The total power of the UAV mission in each time slot is constrained, with the total power P... t satisfy

[0161]

[0162] Where P w (t)=|w(t)| 2 For transmission power, P e (t)=|w e (t)| 2 To sense power, To calculate power, To calculate the power consumption factor, f c (t) represents the calculated frequency, P max The maximum power allowed per time slot;

[0163] To jointly optimize limited airborne resources and address the issue of efficient and secure data transmission in satellite coverage blind spots under eavesdropping conditions, an optimization problem is established to ensure that data is uploaded to the satellite in the shortest possible time while the UAV returns to its starting point to perform the next mission, all under secure transmission conditions. Therefore, the objective function of the overall problem and its corresponding constraints are expressed by the following mathematical formula.

[0164]

[0165] st(9),(18), (22a)

[0166] c T x(t)≤ΔTf c (t), t∈[1,T] s (22b)

[0167] 0≤f c (t)≤f max ,t∈[1,T s (22c)

[0168] x n (t)∈{0,1},t∈[1,T s (22d)

[0169] L I ,f c ,P w ,P s ,T s ,X I ∈U (22e) where To perceive the path, For the sensing phase, calculate the frequency, transmission power, and calculation strategy; T s For the perception strategy, i.e., the moment when the UAV's perception ends, it satisfies To calculate the frequency, For transmission power variables, To sense power variables;

[0170] Step Six: Solve the situational awareness optimization problem using deep reinforcement learning methods to obtain the situational awareness of the eavesdropping environment;

[0171] because The objective function in the problem has the characteristic of long-term accumulation. This invention transforms the optimization problem into a Markov decision problem and uses the DRL method to solve the specific problem.

[0172] 1) Problem transformation based on Markov decision processes;

[0173] First of all Transformed into a Markov decision process, i.e., a tuple For state space, For the action space, For reward space, γ represents the transition probability and the reward discount factor; at time t, the UAV first observes the current environmental information to obtain the state. Then, take action. After acting on the environment, the transition probability is used. Transition to the following state s t+1 And generate a reward r after the action is performed. t Specifically, the state, action, and reward function of a Markov process can be represented as follows;

[0174] 1. State Space In the t-th time slot, the system state comprises the distance of the UAV from the origin, the estimated state of the eavesdropping channel, the satellite channel state, airborne data, and the perception standard deviation, and can be represented as follows:

[0175]

[0176] G t These represent the eavesdropping channel status and satellite channel status of the target area sensed at time t, respectively.

[0177] 2. Motion space During the situational awareness phase, the UAV needs to adaptively plan its flight path, detect the location of the eavesdropper, and simultaneously complete data transmission and onboard compression tasks. Therefore, in the t-th time slot, the agent needs to adjust its flight path based on the current environmental state s. t The decision-making process involves determining the UAV's position in the next time slot, the power used for perception computing transmission, the computing strategy, and the perception strategy, and then executing these in the next time slot. Therefore, the action space needs to be mapped to high-dimensional parameters, and the agent's action in the t-th time slot can be represented as...

[0178] a t ={l t+1 ,P w (t+1),P s (t+1),f c (t+1),x(t+1),b t} (twenty four)

[0179] Where b t ∈{0,1} indicates whether the perception phase has ended, b t =1 means that the perception phase ends at the current moment; as can be seen from (24), lt+1 ,x(t+1),b t P is a discrete variable. w (t+1),P s (t+1),f c (t+1) is a continuous variable; therefore, the agent's output action is a mixed high-dimensional vector, which increases the complexity of the search space for finding the optimal action.

[0180] 3. Reward Space Due to the complexity of the mixed high-dimensional action space, the rewards from actions in different dimensions are prone to compensation, making it difficult for the agent to analyze the merits of actions in different dimensions based on system rewards, and training the agent may become more difficult. Therefore, this invention sets up a specialized reward and punishment mechanism for actions in different dimensions based on system rewards, namely the "basic reward + individual reward" mode, so that the agent can learn and understand the impact of actions in various dimensions more effectively.

[0181] Specifically, a basic reward function is first set to evaluate the overall system performance. The basic reward reflects the agent's contribution to the overall system goal. To detect the eavesdropper's location and complete the data transmission task, the drone needs to choose a location that is more conducive to shortening the overall task completion time. Therefore, the basic reward is defined as the estimated task completion time gain based on the current channel state. That is, assuming the drone's perception phase ends at the current moment, the optimal time result for the anti-interception transmission phase is estimated based on the currently obtained channel state. The difference between the estimated result at this moment and the estimated result at the previous moment is used as the reward for that moment, expressed as:

[0182] R(t)=ξ(P s (t),t)-ξ(P s (t-1),t-1) (25)

[0183] Where ξ(P) s ,t) represents the shortest task completion time for the anti-interception transmission phase that can be achieved under the environmental situation perceived by the UAV in time slot t;

[0184] For drone trajectory planning dimension l t+1 and perception strategy dimension b t This invention sets up an "individual reward," wherein when the agent's trajectory planning action exceeds the boundary of the target area, a constant penalty term X1(t) is set to regulate its trajectory planning range; for the perception policy dimension b t Set a penalty term X2 with lag, represented as

[0185]

[0186] Where C2 > 0 is a constant, that is, if the base reward in time slot t+1 is positive, then the perception policy b in time slot t is positive. t This shortens the task completion time and is rewarded accordingly; if the base reward for slot t+1 is negative, then the perception strategy b for slot t... t Errors prolong the task completion time, and are therefore penalized; thus, the final reward function can be expressed as:

[0187] r t =R(t)+X1(t)+X2(t) (27)

[0188] Based on this, the present invention implements a specialized reward and punishment mechanism for different dimensions in multi-dimensional actions, which better reflects the independent contribution of each action dimension to the system performance on the basis of overall performance improvement.

[0189] 2) A multi-domain resource planning algorithm for unmanned aerial vehicles (UAVs) based on reward decomposition DDPG;

[0190] The DDPG algorithm is a reinforcement learning algorithm that performs well in continuous action control tasks, and is an improvement on the deterministic policy gradient (DPG) algorithm. DDPG combines deep neural networks and Q-Learning to solve optimal decision problems in continuous control tasks. The algorithm is based on the Actor-Critic deep reinforcement learning framework and includes two types of training networks: Actor networks and Critic networks. The input network parameters include the perceptual accuracy threshold Ω, initialization of the Actor and Critic network structures and parameters, and the agent replay buffer. Training rounds M, maximum number of steps allowed per episode T max The output is the parameters of the trained Actor network.

[0191] The trained Actor network can be deployed independently on drones to perform trajectory planning and resource allocation in real time based on the environment; ultimately, a situational awareness map is obtained, achieving the minimum completion time under the constraint of the lower limit of data completion.

[0192] Step 7: Based on Steps 5 and 6, construct a joint planning method for UAV trajectory and resources based on situational awareness; use deep reinforcement learning to perform adaptive trajectory planning for UAVs, and gradually construct an eavesdropping environment situation map while taking into account data calculation and uploading in an unknown eavesdropping environment; according to the eavesdropping environment situation map, balance the UAV's transmission, calculation, and perception tasks under energy constraints to achieve coordination among the UAV's transmission, calculation, and perception tasks.

[0193] The deep reinforcement learning framework based on Actor-critic takes a perceptual accuracy threshold Ω as input to the network, initializes the Actor and Critic network structures and parameters, and includes an agent replay buffer. Training rounds M, maximum number of steps allowed per episode T max For each episode, first initialize the system state s0, when t≤T max At that time, select the action according to the ∈-greedy strategy. Perform action a t And transition to the next system state s t+1 Then, the estimated task completion time gain ξ(P) based on the current channel state is obtained. s ,t), where the corresponding dimension of the action receives the corresponding reward / penalty r. t Storage experience groups (s) t ,a t ,r t ,s t+1 ) to the experience replay pool from A mini-batch of experiences is randomly selected. The Critic network parameters θ are updated by minimizing the loss, and the Actor network parameters φ are updated using policy gradient descent. The target network parameters are updated every G steps. This invention trains the proposed model using the Adam optimizer on episode = 500 epochs with a learning rate of 0.001, where a maximum of 100 training steps are allowed per epoch. In each training step, the mini-batch size is 32. The trained Actor network parameters are output.

[0194] Tables 1 to 3 illustrate the assessment process of the eavesdropping environment during the UAV's perception process. The initial state, step 2, and step 7 of a specific perception process are selected, with higher values ​​indicating better resistance to interception of the transmission channel. The environmental situation assessment process is as follows: Figure 3 As shown.

[0195] Table 1 Initial state of the channel environment sensing process

[0196]

[0197] Table 2. Status of Step 2 in the Channel Environment Sensing Process

[0198]

[0199] Table 3. Status of Step 7 in the Channel Environment Sensing Process

[0200]

[0201] In the second step of the perception process, it can be seen that the drone can only make a preliminary judgment on the approximate location of the eavesdropper, and only partial information is available for the initial assessment of the eavesdropping environment. The state of the eavesdropping channel remains relatively unclear. As the perception process progresses, by the seventh step, it can be observed that the location of the eavesdropper has gradually been determined, and the state of the eavesdropping channel has been further improved. The drone has a more accurate understanding of the channel's condition and potential threats. Finally, the perception of the eavesdropping environment is relatively complete. The results show that, under the proposed algorithm, the drone can gradually improve its understanding of the eavesdropping environment during the perception process, thereby effectively responding to potential security threats and providing protection for the subsequent anti-interception transmission of data.

[0202] This example describes a UAV communication situational awareness method based on deep reinforcement learning. Addressing the unreliable and uncertain channel environment in emergency scenarios, it proposes an integrated aerospace emergency communication system framework encompassing sensing, transmission, computation, and navigation. A multi-functional UAV integrating sensing, computation, and communication is deployed, leveraging sensing to enhance physical layer security and utilizing computation to improve transmission efficiency. By obtaining a situational awareness map through sensing, data transmission can be completed in the shortest possible time while meeting security requirements.

[0203] The above detailed description further illustrates the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for UAV communication situational awareness based on deep reinforcement learning, characterized in that: Includes the following steps, Step 1: Establish a multi-domain fusion model of the aerospace emergency communication system, divide the communication area of ​​the high-orbit satellite in the aerospace emergency communication system model into a grid, and obtain the distance that the UAV moves between time slots; Step 2: Establish a transmission model based on statistical CSI to model the channel; The probability of secure interruption is obtained based on the transmission model; Step 3: Use radar signals to sense Eve's location and estimate the state of the eavesdropping channel; Step 4: Compress the data file and transmit the compressed file to the satellite; Step 5: Analyze the various constraints and establish a situational awareness optimization problem. ; Step Six: Applying Deep Reinforcement Learning to the Situational Awareness Optimization Problem The solution is then performed to obtain the eavesdropping environment situation; Step 7: Based on Steps 5 and 6, construct a joint planning method for UAV trajectory and resources based on situational awareness; utilize deep reinforcement learning to perform adaptive trajectory planning for UAVs, and gradually construct an eavesdropping environment situation map while balancing data computation and uploading in an unknown eavesdropping environment; based on the eavesdropping environment situation map, balance the UAV's transmission, computation, and perception tasks under energy constraints to achieve coordination among the UAV's transmission, computation, and perception tasks.

2. The UAV communication situational awareness method based on deep reinforcement learning as described in claim 1, characterized in that: In step one, A multi-domain fusion model of an aerospace emergency communication system was established. This model includes a high-orbit satellite and an auxiliary communication UAV. Due to terrain, the ground area served by the high-orbit satellite has communication blind spots. An unidentified aerial eavesdropper, Eve, exists within the high-orbit satellite's communication area, but its location is obtained through radar sensing. When the channel quality in the communication blind spot is severely compromised, preventing direct satellite-to-ground communication, the UAV acts as a relay. After collecting disaster information in the communication blind spot, it flies to the satellite's communication coverage area to upload the data. During the UAV's upload process, Eve eavesdrops on the information. Throughout the entire flight and upload process, the UAV sends sensing signals to determine Eve's location and simultaneously performs onboard computation to compress and upload the data. The UAV is configured with a uniform linear array of NT antennas for communication and sensing, while Eve uses a single antenna for eavesdropping; the positions of the satellite and Eve are fixed; Eve's position is unknown to the UAV; the UAV moves in a two-dimensional plane. The high-orbit satellite communication area is divided into an I×I grid. Each grid contains two attributes: satellite channel status and eavesdropping channel status. The satellite channel status is determined by geographical location and environment, exhibiting a degree of randomness. The eavesdropping channel status is determined by the location of Eve; grid areas closer to Eve are more susceptible to eavesdropping. UAVs all start from a fixed starting point and move in grid units, advancing one grid per time slot or remaining in the same position, returning to the starting point after one complete flight to form a closed loop. The total flight time slots of the UAV are represented as... ,in The time to return to the starting point; the UAV on each grid at time intervals Perform sensing, computing, and transmission; The drone is located in the time slot. Grid; Eve position ; High-orbit satellite position The height difference between the drone's flight plane and the flight plane is ; Indicates the first and The distance between grids, and satisfying (1) in .

3. The UAV communication situational awareness method based on deep reinforcement learning as described in claim 2, characterized in that: In step two, Satellite service areas are affected by the environment; therefore, satellite communication channels are modeled as uncertainties. The UAV in the [missing information]... The channel state of the grid is (2) in For the path loss from the drone to the satellite, For operating frequency, For drones in the Distance between the grid and the satellite; The channel fading factor characterizes the impact of complex environmental obstructions on communication between UAVs and satellites; the impact includes scattering effects, masking effects, and a combination of scattering and masking effects. The probability density function PDF satisfies (3) in, For the confluence hypergeometry function, Nakagami-m fading parameters real numbers , The average power of the line-of-sight component, Let be the average value of the multipath power; superimposing equation (3) yields the cumulative distribution function CDF, expressed as follows: (4) in It is an ascending power function. It is a gamma function; Indicates the incomplete gamma function; integral over (2) yields cumulative distribution function The satellite channel status of the target area is... , express Satellite channel status at the location; Location eavesdropping channel state is represented as (5) in For the path loss of eavesdroppers, The distance between the drone and Eve. The azimuth angle of the drone to Eve. The guide vector for the drone-eavesdropping link; In the time slot UAV transmission signal is (6) in and These are the communication and sensing parts of the integrated sensing signal, respectively. and This is the precoding matrix; The signal received by the satellite is represented as (7) in for Location communication channel status, It is Gaussian white noise; Eve received the signal as (8) in for Location eavesdropping channel status, It is Gaussian white noise; Define an auxiliary variable As a satellite in The time slot represents the lower bound of the safe rate achieved by the drone, and thus the safe throughput. Represented as (9) in For indicator functions, when hour , The minimum expected communication rate per time slot link, and These represent the transmission rate from the drone to the satellite and the eavesdropping rate, respectively, denoted as... (10) as well as (11) in For bandwidth; The probability of safe interruption from drone to satellite meets the following conditions. (12) in This represents the maximum permissible interruption probability.

4. The UAV communication situational awareness method based on deep reinforcement learning as described in claim 3, characterized in that: The method for implementing step three is as follows: Eve's location parameters were estimated based on the radar echo signal; the potential eavesdropper's location uncertainty model is as follows: (13) In the formula To obtain a location estimate when detecting a potential eavesdropper. This is for position measurement error; In real location It follows a Gaussian distribution, that is (14) in To measure the standard deviation; the first The measurement error of each time slot satisfies , (15) in For MF gain; The distribution measured from multiple time slots is accumulated as follows: (16) in satisfy To sum the measurement standard deviation of the cumulative distribution function; set the measurement accuracy threshold to . ,when At that time, Eve's positional awareness is sufficient; The estimated eavesdropping channel state is ,in This indicates that the estimated location of Eve is In this case, The eavesdropping channel status at the location.

5. The UAV communication situational awareness method based on deep reinforcement learning as described in claim 4, characterized in that: In step four, Each drone mission collects information on communication blind spots. There are 10 data files to be transmitted, and each data file is independent of the others; The drone needs to upload all the data it carries to the satellite during this mission; and associate each data file with an onboard computing task. For individual airborne computing tasks Description, in which To calculate the number of CPU revolutions required for this task, and Define the data size before and after the calculation; define indicator variables. Indicates the first The first time slot Individual task computation strategy Indicates the first The data for each task was processed and calculated, and vice versa. ; for a concise definition and ; By utilizing the computing capabilities of drones, data files are compressed during flight uploads to reduce the amount of data transmitted. Time slot, drone location Carrying compressed data Uncompressed data volume Its satisfaction (17) as well as (18) in The amount of data uploaded from the compressed queue, The amount of data uploaded to the uncompressed queue, , The time slot length; This represents the reduction in uncompressed queue data after calculation. This represents the increase in the amount of data in the compressed queue after calculation. For the first The calculation strategy for each time slot.

6. The UAV communication situational awareness method based on deep reinforcement learning as described in claim 5, characterized in that: Step five is implemented as follows: Analyze the power consumption constraints on the UAV; power consumption includes three parts: communication power consumption, computing power consumption, and sensing power consumption; the total power of the UAV mission in each time slot is constrained, and the total power... satisfy (19) in For transmission power, To sense power, To calculate power, To calculate the power consumption factor, To calculate the frequency, The maximum power allowed per time slot; To jointly optimize limited airborne resources and address the issue of efficient and secure data transmission in satellite coverage blind spots under eavesdropping conditions, a situational awareness optimization problem is established. This problem aims to ensure data is uploaded to the satellite in the shortest possible time while maintaining secure transmission, and simultaneously allows the UAV to return to its starting point for the next mission. Therefore, the objective function and corresponding constraints of the overall situational awareness optimization problem are expressed by the following formula. (20) (20a) (20b) (20c) (20d) (20e) in To perceive the path, Calculate the frequency, transmission power, and calculation strategy for the sensing phase; For the perception strategy, i.e., the moment when the UAV's perception ends, it satisfies ; To calculate the frequency, For transmission power variables, To sense power variables.

7. The UAV communication situational awareness method based on deep reinforcement learning as described in claim 6, characterized in that: In step six, Due to situational awareness optimization problem The objective function in the problem has the characteristic of long-term accumulation, which makes the situational awareness optimization problem... The problem is transformed into a Markov decision problem, and the DRL method is used to solve the specific problem. Optimize situational awareness problem Transformed into a Markov decision process, i.e., a tuple , For state space, For the action space, For reward space, For the transition probability, As a reward discount factor; in At any given moment, the drone first observes the current environmental information to obtain its status. Then, take action. Acting on the environment, and then using the transition probability Transition to the next state And generate a reward after the action is performed. ; (21) , They are respectively The target area's eavesdropping channel status and satellite channel status are constantly monitored; (22) in Indicate whether the perception phase has ended. That is, the perception phase ends at the current moment; as can be seen from (22), For discrete variables, Since these are continuous variables, the agent's output actions are mixed high-dimensional vectors, increasing the complexity of the search space for finding the optimal action. (23) Constant penalty term , Represented as: (24) Penalties with lag Represented as (25) The trained Actor network is independently deployed on the drone, performing trajectory planning and resource allocation in real time according to the environment; ultimately, a situational awareness map is obtained, achieving the minimum completion time under the constraint of the lower limit of data completion.