Unmanned aerial vehicle communication situation awareness method based on deep reinforcement learning
Through the deep reinforcement learning drone communication situational awareness method, an integrated framework of perception, transmission, computing and navigation is built, which solves the detection and transmission security problems of unknown eavesdroppers in drone-assisted satellite communications, and realizes efficient and secure data transmission in complex environments.
Patent Information
- Application Number
- CN202510429669.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-28
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In drone-assisted satellite communication, the prior art is difficult to effectively detect and deal with unknown eavesdroppers, resulting in insufficient communication security and stability, especially in complex environments, which makes it difficult to achieve secure transmission of communication blind spots.
Adopting a drone communication situational awareness method based on deep reinforcement learning, by building a comprehensive framework that integrates perception, transmission, computing and navigation, using the multi-functional sensing equipment and computing capabilities of the drone, combining deep reinforcement learning for adaptive trajectory planning, realize the coordination of drone transmission, computing, and perception tasks, and improve communication security and stability.
In an unknown eavesdropping environment, the coordination between drone transmission, computing and perception tasks is achieved, which improves communication security and stability, and ensures efficient and secure data transmission under energy constraints.
Smart Images

Figure CN120474595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for UAV communication situation awareness based on deep reinforcement learning, and belongs to the field of satellite communications. Background Art
[0002] High-orbit satellite communications, with their wide coverage, high reliability, and large communication capacity, are a common method for emergency communications. However, communications are easily obstructed by obstacles, resulting in blind spots. Drones, with their advantages of high maneuverability and flexible communication networking, can serve as aerial relays to facilitate communication between satellites and users on the ground. Furthermore, the highly open nature of satellite channels makes them susceptible to potential eavesdropping threats, making secure transmission a key concern for high-orbit satellite communications.
[0003] In recent years, advanced beamforming techniques have been developed to achieve physical layer security for high-orbit satellite communications, such as precoding, interference alignment, and artificial noise. However, most research assumes that potential eavesdroppers have been identified and that the serving base station somehow knows some of their channel state information (CSI). In reality, however, potential eavesdroppers typically do not interact with the base station, making their presence difficult to detect. This makes obtaining relevant CSI extremely challenging, and poses a key constraint on secure wireless network transmission.
[0004] When some base stations are damaged and out of service, both ground and satellite networks alone will struggle to provide coverage for users on the ground, especially in areas inaccessible to vehicles and in satellite-denied scenarios due to obstruction by mountainous and forested terrain. However, drones, due to their inherent mobility, flexibility, and ease of deployment, can be integrated with ground and satellite networks to penetrate communication blind spots, providing coverage in complex terrain and becoming a communication hub between the outside world and emergency response areas. First, during the approaching data collection phase, drones, equipped with various sensor devices such as cameras and infrared sensors, can rapidly fly to the target area to collect real-time data and information, including disaster images, personnel needs, and communication information. Furthermore, drones leverage their perception capabilities to proactively sense target drones or the environment based on integrated communication and perception signals. These drones effectively extract the impact of dynamic targets like drones and static objects like the environment on wireless signal characteristics, assisting in security monitoring and eavesdropping avoidance. When the drone reaches the communication coverage area, it will enter the data transmission stage, transmit the collected data to the receiving end through the air link, and further forward the data to the emergency communication processing system to quickly establish an emergency dedicated network, provide timely and reliable communication support and rescue services to the affected areas, and establish a situational awareness map. Summary of the Invention
[0005] Based on a UAV-assisted satellite communication scenario in an unreliable and complex channel environment, the present invention aims to provide a UAV communication situational awareness method based on deep reinforcement learning. This method utilizes UAV-mounted sensing equipment and the integration of sensing capabilities into existing wireless networks to acquire the required environmental information about the target area. This environmental information helps address the issue of secure transmission against interception in unknown environments. Using deep reinforcement learning methods to perform adaptive UAV trajectory planning, it is possible to achieve coordination between the UAV's transmission, computing, and sensing tasks, improving communication security and stability.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] The present invention discloses a method for UAV communication situational awareness based on deep reinforcement learning. It aims at UAV-assisted satellite communications in unreliable and complex channel environments and builds a comprehensive framework based on the integration of perception, transmission, computing and navigation. The framework deploys a multifunctional UAV with integrated sensing and computing, which enhances physical layer security through target perception and improves transmission efficiency by using computing power. At the same time, the UAV's movement trajectory, sensing and computing power and strategy are jointly designed. While meeting security requirements, data transmission efficiency is maximized. The UAV trajectory and resource joint planning method based on situational awareness uses deep reinforcement learning methods to perform UAV adaptive trajectory planning. While taking into account data calculation and upload in an unknown eavesdropping environment, it gradually builds an eavesdropping environment situation map. According to the eavesdropping environment situation map, the UAV transmission, computing and perception tasks are balanced under energy constraints, and the coordination between the UAV transmission, computing and perception tasks is achieved, thereby improving communication security and stability.
[0008] The present invention discloses a method for UAV communication situation awareness based on deep reinforcement learning, comprising the following steps:
[0009] Step 1: Establish an aerospace emergency communication system model with multi-domain fusion of telemetry and computing, divide the communication area of the high-orbit satellite in the aerospace emergency communication system model into grids, and obtain the distance the UAV moves between time slots.
[0010] A space emergency communication system model integrating inter-sensory computing and multi-domain is established. The space emergency communication system model includes a high-orbit satellite and an auxiliary communication UAV. The ground area served by the high-orbit satellite has a communication blind spot due to the influence of terrain. There is an aerial eavesdropper Eve with an unknown position in the high-orbit satellite communication area, but the position information of the aerial eavesdropper is obtained through radar perception. When the channel quality in the communication blind spot is severely damaged and direct satellite-to-ground communication is impossible, the UAV acts as a relay, collects disaster information in the communication blind spot, and then flies to the satellite communication coverage area to upload data. During the UAV uploading process, Eve eavesdrops on the information. During the entire flight upload process, the UAV sends a perception signal to determine Eve's position information, and performs on-board computing to compress the uploaded data. The UAV is equipped with a uniform linear array composed of NT antennas for communication and perception, and Eve uses a single antenna for eavesdropping. The positions of the satellite and Eve are fixed. Eve's position is unknown to the UAV. The UAV moves on a two-dimensional plane.
[0011] The high-orbit satellite communication area is divided into I×I grids, each of which contains two attributes: satellite channel status and eavesdropping channel status. The satellite channel status is determined by the geographical location and environment and has a certain degree of randomness. The eavesdropping channel status is determined by Eve's position, and the grid area close to Eve is more susceptible to eavesdropping. UAVs start from a fixed starting point and move in grid units. They advance one grid per time slot or maintain the same position. After flying one circle, they return to the starting point to form a closed-loop path. The total flight time slot of the UAV is expressed as Where T end is the time to return to the starting point; UAV performs perception calculation and transmission on each grid at a time interval of ΔT; in time slot t, the UAV is located at the lth t =[x t ,y t ] grid;Eve position l e =[x e ,y e ]; High-orbit satellite position l s =[x s ,y s ], the height difference with the UAV flight plane is H; Indicates the first t and l t+1 The distance between grids, and satisfy
[0012]
[0013] Among them x∈[1,I], y∈[1,I].
[0014] Step 2: Establish a transmission model based on statistical CSI to model the channel; obtain the safety interruption probability based on the transmission model.
[0015] The satellite service area will be affected by the environment, so the satellite communication channel is modeled as an uncertainty model. t The channel state of the grid is
[0016]
[0017] in is the path loss from the UAV to the satellite, f is the operating frequency, For drones in the first t The distance between the grid and the satellite; |g s | 2 is the channel fading factor, which characterizes the impact of complex environment blocking on the communication between the UAV and the satellite; the impact includes scattering effect, masking effect, and scattering and masking effect; |g s | 2 The probability density function PDF satisfies
[0018]
[0019] Among them, 1F1(m; 1; δx) is the confluent hypergeometric function, and the Nakagami-m fading parameter is Real number m∈(0,∞), μ is the average power of the line-of-sight component, 2κ is the average value of the multipath power; superimposing equation (3) yields the cumulative distribution function CDF, which is expressed as follows
[0020]
[0021] in is a raised power function, is the gamma function; represents the incomplete gamma function; integrating (2) yields |g i | 2 Cumulative distribution function of
[0022] The satellite channel status in the target area is G t ={g 1,1 ,…,g i,j ,…,g I,I}, g i,j Indicates l t = Satellite channel status at position [i, j];
[0023] l t The location eavesdropping channel state is expressed as
[0024] (5) is the path loss of the eavesdropper, is the distance between the drone and Eve, θ t∈[-π / 2,π / 2] is the azimuth angle from the UAV to Eve, a(θ t ) is the drone-eavesdropper link steering vector;
[0025] At time slot t, the UAV transmits a signal:
[0026] x(t)=w(t)s com (t)+w e (t)s rad (t) (6)
[0027] where s com (t) and s rad (t) are the communication part and the perception part of the synaesthesia integrated signal, and is the precoding matrix;
[0028] The signal received by the satellite is expressed as
[0029]
[0030] in for l t Position communication channel status, is Gaussian white noise;
[0031] Eve receives the signal
[0032]
[0033] in for l t Location eavesdropping channel status, is Gaussian white noise;
[0034] Define an auxiliary variable r(t) as the lower bound of the safe rate that the satellite can reach for the UAV in the tth time slot, and the achieved safe throughput C(t) is expressed as
[0035] C(t)=I C ([C s (t)-C e (t)] + ≥r(t))·r(t) (9)
[0036] Among them I C (·) is the indicator function, when [C s (t)-C e (t)] + When ≥r(t) C (·)=1, r(t) is the minimum expected communication rate of the link per time slot, C s (t) and C e(t) are the transmission rate from the UAV to the satellite and the eavesdropping rate, respectively, expressed as
[0037]
[0038] as well as
[0039]
[0040] Where B is the bandwidth;
[0041] The probability of safe interruption from drone to satellite meets
[0042] Pr{[C s (t)-C e (t)] + ≤r(t)}≤ε (12)
[0043] Where ε≤0.1 is the maximum allowable interruption probability.
[0044] Step 3: Use radar signals to perceive Eve’s location and estimate the state of the eavesdropping channel.
[0045] The position parameters of Eve are estimated based on the radar echo signal; the uncertainty model of the potential eavesdropper's position is:
[0046]
[0047] In the formula is the location estimate obtained when sensing a potential eavesdropper, Δl e (t) is the position measurement error; In real location e (t) satisfies the Gaussian distribution, that is,
[0048]
[0049] Where δ is the measurement standard deviation; the measurement error of the tth time slot satisfies
[0050]
[0051] Among them G MF is the MF gain;
[0052] The distribution measured over multiple time slots is accumulated as
[0053]
[0054] in satisfy is the measurement standard deviation of the cumulative distribution function; the threshold of measurement accuracy is set to Ω, when When , Eve’s position perception is sufficient;
[0055] The estimated eavesdropping channel state is in Indicates that Eve's estimated position is In this case, t = the eavesdropping channel state at position [i, j].
[0056] Step 4: Compress the data file and transmit the compressed file to the satellite.
[0057] In each mission, the UAV collects N data files to be transmitted in the communication blind area. Each data file is not related to each other. The UAV needs to upload all the data it carries to the satellite in this mission. Each data file is associated with an onboard computing task. The nth onboard computing task is used (d an ,d bn ,c n ) description, where c n To calculate the number of CPU revolutions required for this task, d an and d bn are the data sizes before and after calculation respectively; define indicator variable x n (t)∈{0,1} represents the calculation strategy of the nth task in the tth time slot, x n (t) = 1 means that the data of the nth task has been calculated and processed, otherwise x n (t) = 0; for simplicity, define d a =[d a1 ,…,d aN ] T ,d b =[d b1 ,…,d bN ] T and c=[c1,…,c N ] T ;
[0058] Using the computing function of the drone, the data files are compressed during the flight upload process to reduce the amount of data transmission; t time slot, drone position l t , carrying compressed data Uncompressed data volume Its satisfaction
[0059]
[0060] as well as
[0061]
[0062] in is the amount of data uploaded from the compressed queue, is the amount of data uploaded from the uncompressed queue, ΔT is the time slot length; To calculate the amount of data reduction in the uncompressed queue, The amount of data increase in the compressed queue after calculation is x(t)=[x1(t),…,x N (t)] T is the calculation strategy for the tth time slot.
[0063] Step 5: Analyze each constraint and establish a situational awareness optimization problem
[0064] Analyze the power consumption constraints on the drone; the power consumption includes communication power consumption, computing power consumption and perception power consumption; the total power of the drone in each time slot is constrained, the total power P t satisfy
[0065]
[0066] Among them, P w (t)=|w(t)| 2 is the transmission power, P e (t)=|w e (t)| 2 is the perceived power, To calculate power, To calculate the power consumption factor, f c (t) is the calculation frequency, P max The upper limit of power allowed for each time slot;
[0067] In order to jointly optimize limited onboard resources to solve the problem of efficient and secure data transmission in satellite coverage blind spots under eavesdropping environments, a situational awareness optimization problem is established, so that data can be uploaded to the satellite in the shortest time under secure transmission conditions, and the drone returns to the starting point to perform the next mission; therefore, the objective function of the overall situational awareness optimization problem and the corresponding constraints are expressed as follows
[0068]
[0069] st(9),(18), (20a)
[0070] c T x(t)≤ΔTf c (t),t∈[1,T s ] (20b)
[0071] 0≤f c (t)≤f max ,t∈[1,T s ] (20c)
[0072] x n (t)∈{0,1},t∈[1,T s ] (20d)
[0073] L I ,f c ,P w ,P s ,T s ,X I ∈U (20e) in For the perception path, Calculate frequency, transmission power and calculation strategy for the perception stage; T s is the perception strategy, that is, the drone perception end time, satisfying To calculate the frequency, is the transmission power variable, is the perceived power variable.
[0074] Step 6: Use deep reinforcement learning methods to optimize situational awareness Solve the problem and get the eavesdropping environment situation.
[0075] Due to the situational awareness optimization problem The objective function in the problem has the characteristics of long-term accumulation, which makes the situation awareness optimization problem Convert it into a Markov decision problem and use the DRL method to solve the specific problem;
[0076] Optimizing situational awareness Transformed into a Markov decision process, i.e., the tuple is the state space, is the action space, For reward space, is the transfer probability, γ is the reward discount factor; at time t, the drone first observes the current environment information to obtain the state Then take action Act on the environment, and then use the transition probability Transition to the next state s t+1 , and generates the reward r after the action is executed t ;
[0077]
[0078] G t are the eavesdropping channel state and satellite channel state of the target area sensed at time t respectively;
[0079] a t ={l t+1,P w (t+1),P s (t+1),f c (t+1),x(t+1),b t} (twenty two)
[0080] where b t ∈{0,1} indicates whether the perception phase is over, b t =1 means that the perception phase ends at the current moment; from (22), we can see that l t+1 ,x(t+1),b t is a discrete variable, P w (t+1),P s (t+1),f c (t+1) is a continuous variable; therefore, the agent's output action is a mixed high-dimensional vector, which increases the complexity of the search space for finding the optimal action;
[0081] r t =R(t)+X1(t)+X2(t) (23)
[0082] The constant penalty term X1(t), R(t) is expressed as:
[0083] R(t)=ξ(P s (t),t)-ξ(P s (t-1),t-1) (24)
[0084] The penalty term X2(t) with hysteresis is expressed as
[0085]
[0086] The trained Actor network is independently deployed on the drone, performing trajectory planning and resource allocation in real time based on the environment; ultimately, a situational awareness map is obtained, achieving the minimum completion time under the constraint of the lower limit of data completion volume.
[0087] Step 7: Based on Steps 5 and 6, a situational awareness-based joint planning method for drone trajectories and resources is constructed. This method utilizes deep reinforcement learning to perform adaptive drone trajectory planning. While balancing data computation and upload in unknown eavesdropping environments, it gradually constructs a situational map of the eavesdropping environment. Based on this situational map, the drone's transmission, computation, and perception tasks are balanced within energy constraints, achieving coordination among these tasks.
[0088] Beneficial effects:
[0089] 1. Compared with the path planning and resource optimization technologies of traditional air-based emergency communication systems, the present invention discloses a UAV communication situational awareness method based on deep reinforcement learning. UAVs fully play their role in emergency communications, realize the rational allocation of system perception, computing and communication resources under limited resource conditions, and contribute to agile and anti-interception communications in emergency situations.
[0090] 2. Compared to the multi-domain resource allocation technology of traditional airborne emergency communication systems, the present invention discloses a method for UAV communication situational awareness based on deep reinforcement learning. This method utilizes deep reinforcement learning to construct a joint planning method for UAV trajectories and multi-domain resources based on situational awareness. This method utilizes deep reinforcement learning to perform adaptive UAV trajectory planning, balancing data computation and upload in unknown eavesdropping environments while gradually constructing a situational map of the eavesdropping environment. Based on this situational map, the UAV's transmission, computation, and perception tasks are balanced under energy constraints, achieving coordination among these tasks.
[0091] 3. Compared with the multi-domain resource allocation technology of traditional air-based emergency communication systems, the present invention discloses a UAV communication situational awareness method based on deep reinforcement learning. In the UAV multi-domain resource planning based on reward decomposition DDPG, a specialized reward and punishment mechanism is implemented for different dimensions in multi-dimensional actions. On the basis of overall performance improvement, it better reflects the independent contribution of each action dimension to system performance and successfully and gradually constructs an eavesdropping environment situation map. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 This is a flowchart of the UAV communication situation awareness method based on deep reinforcement learning of the present invention;
[0093] Figure 2 This is a model diagram of a UAV-assisted satellite communication system that integrates perception, transmission, calculation, and navigation in the UAV communication situation awareness method based on deep reinforcement learning of the present invention;
[0094] Figure 3 This is the environmental situation assessment process in the UAV communication situation awareness method based on deep reinforcement learning of the present invention, wherein: Figure 3 (a) is the initial unknown state of the channel environment (step = 0), Figure 3 (b) is the initial intermediate state of the channel environment (step = 3), Figure 3 (c) is the initial intermediate state of the channel environment (step = 5), Figure 3 (d) is the initial complete state of the channel environment (step = 7). DETAILED DESCRIPTION
[0095] The present invention will be described in detail below with reference to the accompanying drawings and embodiments, and the technical problems solved by the technical solution of the present invention and the beneficial effects thereof will be discussed. It should be noted that the described embodiments are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.
[0096] Example 1
[0097] This example describes the application of a deep reinforcement learning-based drone communication situational awareness method for drone-assisted satellite communication in a primeval forest emergency rescue environment. During a geological survey, a team conducted a geological survey deep within a primeval forest. However, due to the complex terrain and dense vegetation in the area, traditional communication methods were blocked, resulting in the team losing contact with the outside world. Satellite communications were obstructed by the forest, and ground base station signals were unable to reach them, posing a significant challenge to the rescue effort. To quickly restore communications, the rescue team decided to deploy drones deep into the forest. Leveraging their mobility, flexibility, and rapid deployment capabilities, combined with terrestrial and satellite networks, they provided coverage in communication-blind areas, becoming a critical communication hub between the outside world and the trapped personnel. Equipped with a variety of equipment, including high-precision cameras, infrared sensors, and lidar, the rescue drones rapidly flew to the target area and collected comprehensive data. The drones used high-resolution imaging to capture real-time footage, monitor the team's progress, and used infrared sensors to detect vital signs, enabling effective positioning even in low-light and complex environments. Furthermore, the drones collected meteorological data in the area, such as temperature, humidity, wind speed, and air quality, to support subsequent rescue decisions. Once a drone enters a communication coverage area, it transmits the collected data via a high-bandwidth airlink to a ground receiver, which then forwards it to the emergency communication processing system. Backstage, the rescue team uses a visualization platform to analyze the transmitted data and create a precise situational awareness map, including information such as the exploration team's current location, environmental risk assessments, and feasible rescue routes, providing a scientific basis for command and dispatch. Furthermore, the drone collaborates with other rescue drones or relay stations via a distributed network, expanding communication coverage and improving the efficiency of rescue response efforts.
[0098] like Figure 1 As shown, this embodiment discloses a method for UAV communication situation awareness based on deep reinforcement learning, and the specific implementation steps are as follows:
[0099] Step 1: Establish a multi-domain integrated aerospace emergency communication system model, divide the communication area of the high-orbit satellite in the model into grids, and obtain the distance the UAV moves between time slots;
[0100] Establish a multi-domain integrated aerospace emergency communication system model; the system model is shown in the figure below. Figure 2As shown in the figure, the model includes a high-orbit satellite and an auxiliary communication UAV. The ground area served by the high-orbit satellite has a communication blind spot due to the influence of terrain. There is an aerial eavesdropper Eve with an unknown position in the high-orbit satellite communication area, but the position information of the aerial eavesdropper can be obtained through radar perception. When the channel quality in the communication blind spot is severely damaged and direct satellite-to-ground communication is impossible, the UAV acts as a relay, collects disaster information in the communication blind spot, and then flies to the satellite communication coverage area to upload data. During the UAV uploading process, Eve eavesdrops on the information. During the entire flight upload process, the UAV sends a perception signal to determine Eve's position information, and at the same time performs on-board calculations to compress the uploaded data. The UAV is equipped with a uniform linear array composed of NT antennas for communication and perception, and Eve uses a single antenna for eavesdropping. The positions of the satellite and Eve are fixed. Eve's position is unknown to the UAV. The UAV moves on a two-dimensional plane.
[0101] The high-orbit satellite communication area is divided into 10×10 grids, each containing two attributes: satellite channel status and eavesdropping channel status. The UAV starts at (0m, 0m, H m) and flies at a constant speed of 20m / s, hovering for 5 seconds at each point to perform data transmission and computation tasks. The satellite channel status is determined by the geographic location and environment and has a certain degree of randomness. The eavesdropping channel status is determined by Eve's position, and grid areas close to Eve are more susceptible to eavesdropping. UAVs start from a fixed starting point and move in grid units. They advance one grid per time slot or maintain the same position. After flying one circle, they return to the starting point to form a closed-loop path. The total flight time slot of the UAV is expressed as Where T end is the time to return to the starting point; UAV performs perception calculation and transmission on each grid at a time interval of ΔT; in time slot t, the UAV is located at the lth t =[x t ,y t ] grid;Eve position l e =[x e ,y e ]; High-orbit satellite position l s =[x s ,y s ], the height difference from the UAV flight plane is H; Indicates the first t and l t+1 The distance between grids, and satisfy
[0102]
[0103] Where x∈[1,I], y∈[1,I];
[0104] Step 2: Establish a transmission model based on statistical CSI, model the channel, and obtain the safety interruption probability;
[0105] The satellite service area will be affected by the environment, so the satellite communication channel is modeled as an uncertainty model. t The channel state of the grid is
[0106]
[0107] in is the path loss from the UAV to the satellite, f is the operating frequency, For drones in the first t The distance between the grid and the satellite; |g s | 2 is the channel fading factor, which characterizes the impact of complex environment blocking on the communication between the UAV and the satellite; the impact includes scattering effect, masking effect, and scattering and masking effect; |g s | 2 The probability density function PDF satisfies
[0108]
[0109] Among them, 1F1(m; 1; δx) is the confluent hypergeometric function, and the Nakagami-m fading parameter is Real number m∈(0,∞), μ is the average power of the line-of-sight component, 2κ is the average value of the multipath power; superimposing equation (3) yields the cumulative distribution function CDF, which is expressed as follows
[0110]
[0111] in is a raised power function, is the gamma function; represents the incomplete gamma function; integrating (2) yields |g i | 2 Cumulative distribution function of
[0112] The satellite channel status in the target area is G t ={g 1,1 ,…,g i,j ,…,g I,I}, g i,j Indicates l t = Satellite channel status at position [i, j];
[0113] l t The location eavesdropping channel state is expressed as
[0114]
[0115] in is the path loss of the eavesdropper, is the distance between the drone and Eve, θ t ∈[-π / 2,π / 2] is the azimuth angle from the UAV to Eve, a(θ t ) is the drone-eavesdropper link steering vector;
[0116] At time slot t, the UAV transmits a signal:
[0117] x(t)=w(t)s com (t)+w e (t)s rad (t) (6)
[0118] where s com (t) and s rad (t) are the communication part and the perception part of the synaesthesia integrated signal, and is the precoding matrix;
[0119] The signal received by the satellite is expressed as
[0120] y(t)=g t w(t)s com (t)+g lt w e (t)s rad (t)+n (7)
[0121] in for l t Position communication channel status, is Gaussian white noise;
[0122] Eve receives the signal
[0123]
[0124] in for l t Location eavesdropping channel status, is Gaussian white noise;
[0125] Define an auxiliary variable r(t) as the lower bound of the safe rate that the satellite can reach for the UAV in the tth time slot, and the achievable safe throughput C(t) is expressed as
[0126] C(t)=I C ([C s (t)-C e (t)] + ≥r(t))·r(t) (9)
[0127] Among them I C (·) is the indicator function, when [C s (t)-C e (t)] + When ≥r(t) C (·)=1, r(t) is the minimum expected communication rate of the link per time slot, C s (t) and C e (t) are the transmission rate from the UAV to the satellite and the eavesdropping rate, respectively, expressed as
[0128]
[0129] as well as
[0130]
[0131] Where B is the bandwidth;
[0132] The probability of safe interruption from drone to satellite meets
[0133] Pr{[C s (t)-C e (t)] + ≤r(t)}≤ε (12)
[0134] Where ε≤0.1 is the maximum permissible interruption probability;
[0135] Step 3: Use radar signals to perceive Eve’s location and estimate the state of the eavesdropping channel;
[0136] The radar signal is used to perceive Eve’s position. The radar echo signal received by the drone is expressed as
[0137]
[0138] in For the round-trip channel state, satisfy
[0139]
[0140] For Eve's radar cross section, For noise.
[0141] The position parameters of Eve are estimated based on the radar echo signal; the uncertainty model of the potential eavesdropper's position is:
[0142]
[0143] In the formula is the location estimate obtained when sensing a potential eavesdropper, Δl e(t) is the position measurement error; the error in the estimated position is mainly caused by noise, so Δl e (t) satisfies Gaussian distribution. Then Δl e (t) The calculation results of Eve position affected In real location e (t) satisfies the Gaussian distribution, that is,
[0144]
[0145] Where δ is the measurement standard deviation; the measurement error of the tth time slot satisfies
[0146]
[0147] Among them G MF is the MF gain;
[0148] The distribution measured over multiple time slots is accumulated as
[0149]
[0150] in satisfy is the measurement standard deviation of the cumulative distribution function; because the position measured each time is distributed near the true position, the cumulative probability distribution of the cumulative distribution of multiple measurements should have the highest value at the true position of the radiation source. Set the threshold of measurement accuracy to Ω, when When , Eve’s position perception is sufficient;
[0151] The estimated eavesdropping channel state is in Indicates that Eve's estimated position is In this case, t = the eavesdropping channel state at position [i, j];
[0152] Step 4: compress the data file and transmit the compressed file to the satellite;
[0153] In each mission, the UAV collects N data files to be transmitted in the communication blind area. Each data file is not related to each other. The UAV needs to upload all the data it carries to the satellite in this mission. Each data file is associated with an onboard computing task. The nth onboard computing task is used (d an ,d bn ,c n ) description, where c n To calculate the number of CPU revolutions required for this task, d an and d bn are the data sizes before and after calculation respectively; define indicator variable x n(t)∈{0,1} represents the calculation strategy of the nth task in the tth time slot, x n (t) = 1 means that the data of the nth task has been calculated and processed, otherwise x n (t) = 0; for simplicity, define d a =[d a1 ,…,d aN ] T ,d b =[d b1 ,…,d bN ] T and c=[c1,…,c N ] T ;
[0154] Using the computing function of the drone, the data files are compressed during the flight upload process to reduce the amount of data transmission; it is assumed that the drone will give priority to transmitting the compressed data; at time slot t, the drone position l t , carrying compressed data Uncompressed data volume Its satisfaction
[0155]
[0156] as well as
[0157]
[0158] in is the amount of data uploaded from the compressed queue, is the amount of data uploaded from the uncompressed queue, ΔT is the time slot length; To calculate the amount of data reduction in the uncompressed queue, The amount of data increase in the compressed queue after calculation is x(t)=[x1(t),…,x N (t)] T is the calculation strategy for the tth time slot;
[0159] Step 5: Analyze each constraint and establish a situational awareness optimization problem
[0160] First, the power consumption constraints on the UAV are analyzed; the power consumption includes communication power consumption, computing power consumption and perception power consumption; the total power of the UAV task execution in each time slot is constrained, the total power P t satisfy
[0161]
[0162] Among them, P w (t)=|w(t)| 2is the transmission power, P e (t)=|w e (t)| 2 is the perceived power, To calculate power, To calculate the power consumption factor, f c (t) is the calculation frequency, P max The upper limit of power allowed for each time slot;
[0163] In order to jointly optimize limited onboard resources to solve the problem of efficient and secure data transmission in satellite coverage blind spots under eavesdropping environments, an optimization problem is established to upload data to the satellite in the shortest time under secure transmission conditions, while the drone returns to the starting point to perform the next mission; therefore, the objective function of the overall problem and the corresponding constraints are expressed as follows
[0164]
[0165] st(9),(18), (22a)
[0166] c T x(t)≤ΔTf c (t),t∈[1,T s ] (22b)
[0167] 0≤f c (t)≤f max ,t∈[1,T s ] (22c)
[0168] x n (t)∈{0,1},t∈[1,T s ] (22d)
[0169] L I ,f c ,P w ,P s ,T s ,X I ∈U (22e) in For the perception path, Calculate frequency, transmission power and calculation strategy for the perception stage; T s is the perception strategy, that is, the drone perception end time, satisfying To calculate the frequency, is the transmission power variable, is the perceived power variable;
[0170] Step 6: Use deep reinforcement learning methods to solve the situational awareness optimization problem and obtain the eavesdropping environment situation;
[0171] because The objective function in has the characteristics of long-term accumulation. The present invention transforms the optimization problem into a Markov decision problem and adopts the DRL method to solve the specific problem.
[0172] 1) Problem transformation based on Markov decision process;
[0173] First, Transformed into a Markov decision process, i.e., the tuple is the state space, is the action space, For reward space, is the transfer probability, γ is the reward discount factor; at time t, the drone first observes the current environment information to obtain the state Then take action After acting on the environment, the transition probability Transfer to the following state s t+1 , and generates the reward r after the action is executed t ; Specifically, the state, action and reward functions of the Markov process can be expressed as follows;
[0174] 1. State Space At the tth time slot, the state of the system includes the distance between the UAV and the origin, the eavesdropping channel estimation state, the satellite channel state, the airborne data, and the perception standard deviation, which can be expressed as
[0175]
[0176] G t are the eavesdropping channel state and satellite channel state of the target area sensed at time t respectively;
[0177] 2. Action Space In the situation awareness phase, the UAV needs to adaptively plan the flight path, perceive the eavesdropper's position, and complete the data transmission and airborne compression tasks at the same time; therefore, in the tth time slot, the intelligent agent needs to t The decision is made on the drone's position in the next time slot, the power used for perception, computing, and transmission, the computing strategy, and the perception strategy, and then executed in the next time slot. Therefore, the action space needs to be mapped into high-dimensional parameters. The action of the agent in the tth time slot can be expressed as
[0178] a t ={l t+1 ,P w (t+1),P s (t+1),f c (t+1),x(t+1),b t} (twenty four)
[0179] where b t ∈{0,1} indicates whether the perception phase is over, b t =1 means that the perception phase ends at the current moment; from (24), we can see that l t+1 ,x(t+1),b t is a discrete variable, P w (t+1),P s (t+1),f c (t+1) is a continuous variable; therefore, the agent's output action is a mixed high-dimensional vector, which increases the complexity of the search space for finding the optimal action;
[0180] 3. Reward Space Due to the complexity of mixed high-dimensional action spaces, rewards from actions in different dimensions can easily lead to compensation. This makes it difficult for agents to analyze the pros and cons of actions in different dimensions based on system rewards, making training agents more difficult. Therefore, this invention sets up specialized reward and punishment mechanisms for actions in different dimensions based on system rewards, namely the "basic reward + individual reward" model, so that agents can more effectively learn and understand the impact of actions in each dimension.
[0181] Specifically, we first set up a basic reward function to evaluate the overall system performance. The basic reward reflects the contribution of the agent to the overall goal of the system. In order to perceive the eavesdropper's position and complete the data transmission task, the drone needs to choose a location that is more conducive to shortening the overall task completion time. Therefore, the basic reward is defined as the estimated task completion time gain based on the current channel state. That is, assuming that the drone perception phase ends at the current moment, then according to the current channel state, the optimal time result of the anti-interception transmission phase is estimated. The difference between the estimated result at this moment and the estimated result at the previous moment is used as the reward at this moment, which is expressed as
[0182] R(t)=ξ(P s (t),t)-ξ(P s (t-1),t-1) (25)
[0183] where ξ(P s ,t) is the shortest task completion time of the anti-interception transmission phase that can be achieved under the environmental situation perceived by the UAV in time slot t;
[0184] For the UAV trajectory planning dimension l t+1 and perceived strategy dimension b t , the present invention sets "individual reward", in which when the trajectory planning action of the intelligent agent exceeds the boundary of the target area, a constant penalty term X1(t) is set to regulate its trajectory planning range; for the perception strategy dimension b t, set the penalty term X2 with hysteresis, expressed as
[0185]
[0186] Where C2>0 is a constant, that is, if the basic reward of time slot t+1 is positive, then the perception strategy b of time slot t t It plays a role in shortening the task completion time and is rewarded; if the basic reward of the t+1 time slot is negative, the perception strategy b of the t time slot t Errors prolong the task completion time and are punished; therefore, the final reward function can be expressed as
[0187] r t =R(t)+X1(t)+X2(t) (27)
[0188] Based on this, the present invention implements a specialized reward and punishment mechanism for different dimensions in multi-dimensional actions, which better reflects the independent contribution of each action dimension to system performance on the basis of overall performance improvement;
[0189] 2) UAV multi-domain resource planning algorithm based on reward decomposition DDPG;
[0190] The DDPG algorithm is a reinforcement learning algorithm that performs well in continuous action control tasks and is an improvement on the deterministic policy gradient (DPG) algorithm. The DDPG algorithm combines deep neural networks and the Q-Learning algorithm to solve the optimal decision-making problem in continuous control tasks. The algorithm is based on the actor-critic deep reinforcement learning framework and contains two types of training networks: the actor network and the critic network. The input network parameters include the perception accuracy threshold Ω, the initialization of the actor and critic network structures and parameters, and the agent playback buffer. Training round M, each episode allows a maximum number of steps T max ;The output is the trained Actor network parameters;
[0191] The trained Actor network can be independently deployed on drones to perform trajectory planning and resource allocation in real time based on the environment; ultimately, a situational awareness map is obtained, achieving the minimum completion time under the constraint of the lower limit of data completion volume.
[0192] Step 7: Based on steps 5 and 6, a joint planning method for drone trajectories and resources based on situational awareness is constructed. Deep reinforcement learning methods are used to perform adaptive drone trajectory planning, balancing data computation and upload in unknown eavesdropping environments while gradually building a situational map of the eavesdropping environment. Based on the situational map of the eavesdropping environment, the drone's transmission, computation, and perception tasks are balanced under energy constraints, achieving coordination among these tasks.
[0193] Actor-critic based deep reinforcement learning framework, the input network parameters include perception accuracy threshold Ω, initialization of Actor and Critic network structure and parameters, and agent playback buffer Training round M, each episode allows a maximum number of steps T max For each episode, first initialize the system state s0, when t≤T max When ∈-greedy strategy is used to select an action Execute action a t And transfer to the next system state s t+1 , then we get the estimated task completion time gain ξ(P s ,t), the action in the corresponding dimension gets the corresponding reward / penalty r t Storage experience group (s t ,a t ,r t ,s t+1 ) to the experience replay pool from Mini-batches of experiences are randomly selected from [1, 2]. The critic network parameters θ are updated by minimizing the loss, and the actor network parameters φ are updated using policy gradient descent, with the target network parameters updated every G steps. The proposed model is trained using the Adam optimizer for 500 episodes with a learning rate of 0.001, allowing a maximum of 100 training steps per episode. In each training step, the mini-batch size is 32. The trained actor network parameters are output.
[0194] Tables 1 to 3 show the evaluation process of the eavesdropping environment during the UAV perception process, and respectively select the initial state, the second step and the seventh step of the eavesdropping environment perception results in a certain perception process. The larger the value, the better the anti-interception transmission channel state. Figure 3 shown.
[0195] Table 1 Initial state of the channel environment perception process
[0196]
[0197] Table 2 Status of step 2 of the channel environment perception process
[0198]
[0199] Table 3 Status of step 7 of the channel environment perception process
[0200]
[0201] In the perception results of step 2, we can see that the drone can only make a preliminary assessment of the eavesdropper's approximate location. The initial assessment of the eavesdropping environment only provides partial information, and the state of the eavesdropping channel remains relatively ambiguous. As the perception process progresses, by step 7, we can observe that the eavesdropper's location has been gradually determined, and the state of the eavesdropping channel has been further refined. The drone has a more accurate understanding of the channel's conditions and potential threats. Finally, the perception of the eavesdropping environment has become relatively complete. These results demonstrate that the proposed algorithm enables the drone to gradually improve its understanding of the eavesdropping environment during the perception process, effectively addressing potential security threats and ensuring the subsequent anti-interception transmission of data.
[0202] This example describes a deep reinforcement learning-based drone communication situational awareness method. Addressing the unreliable and uncertain channel environments encountered in emergency scenarios, it proposes a framework for an aerospace emergency communication system integrating perception, transmission, computing, and navigation. By deploying multifunctional drones with intersensory computing, the system enhances physical layer security through perception and improves transmission efficiency through computing. This method generates a situational awareness map through perception, enabling data transmission to be completed in the shortest possible time while meeting safety requirements.
[0203] The above specific description further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A UAV communication situational awareness method based on deep reinforcement learning, characterized by: The following steps are included: Step 1: Establish a multi-domain integrated aerospace emergency communication system model using synaesthesia and computing. Grid-divide the communication area of the high-orbit satellite in the aerospace emergency communication system model to obtain the distance the UAV moves between time slots. Step 2: Establish a transmission model based on statistical CSI to model the channel; Obtain the security interruption probability based on the transmission model; Step 3: Use radar signals to perceive Eve’s location and estimate the state of the eavesdropping channel; Step 4: compress the data file and transmit the compressed file to the satellite; Step 5: Analyze each constraint and establish the situation awareness optimization problem P1; Step 6: Use deep reinforcement learning to solve the situational awareness optimization problem P1 and obtain the eavesdropping environment situation; Step 7: Based on steps 5 and 6, a joint planning method for UAV trajectories and resources based on situational awareness is constructed. Deep reinforcement learning methods are used to perform adaptive trajectory planning for UAVs, taking into account data calculation and upload in unknown eavesdropping environments, and gradually constructing an eavesdropping environment situation map. Based on the eavesdropping environment situation map, the UAV transmission, computing, and perception tasks are balanced under energy constraints to achieve coordination among the UAV transmission, computing, and perception tasks.
2. The method for UAV communication situation awareness based on deep reinforcement learning according to claim 1, characterized in that: In step one, A space emergency communication system model integrating inter-sensing and multi-domain computing was established. The space emergency communication system model includes a high-orbit satellite and an auxiliary communication UAV. The ground area served by the high-orbit satellite has a communication blind spot due to the influence of terrain. There is an aerial eavesdropper Eve with an unknown position in the high-orbit satellite communication area, but the aerial eavesdropper's position information is obtained through radar perception. When the channel quality in the communication blind spot is severely impaired and direct satellite-to-ground communication is impossible, the UAV acts as a relay, collecting disaster information in the communication blind spot and then flying to the satellite communication coverage area to upload data. During the UAV uploading process, Eve eavesdrops on the information. During the entire flight upload process, the UAV sends perception signals to determine Eve's position information, and at the same time performs on-board computing to compress the uploaded data. The UAV is equipped with a uniform linear array of NT antennas for communication and sensing, while Eve uses a single antenna for eavesdropping. The satellite and Eve are fixed in position; Eve's position is unknown to the UAV; and the UAV moves in a two-dimensional plane. The high-orbit satellite communication area is divided into I×I grids, each of which contains two attributes: satellite channel status and eavesdropping channel status. The satellite channel status is determined by the geographical location and environment and has a certain degree of randomness. The eavesdropping channel status is determined by Eve's position, and the grid area close to Eve is more susceptible to eavesdropping. UAVs start from a fixed starting point and move in grid units. They advance one grid per time slot or maintain the same position. After flying one circle, they return to the starting point to form a closed-loop path. The total flight time slot of the UAV is expressed as T = {1,…,t,…,T end }, where T end is the time to return to the starting point; UAV performs perception calculation and transmission on each grid at a time interval of ΔT; in time slot t, the UAV is located at the lth t =[x t ,y t ] grid;Eve position l e =[x e ,y e ]; High-orbit satellite position l s =[x s ,y s ], the height difference from the UAV flight plane is H; Indicates the first t and l t+1 The distance between grids, and satisfy Among them x∈[1,I], y∈[1,I].
3. The method for UAV communication situation awareness based on deep reinforcement learning according to claim 2, characterized in that: In step 2, The satellite service area will be affected by the environment, so the satellite communication channel is modeled as an uncertainty model. t The channel state of the grid is in is the path loss from the UAV to the satellite, f is the operating frequency, For drones in the first t The distance between the grid and the satellite; |g s | 2 is the channel fading factor, which characterizes the impact of complex environment blocking on the communication between the UAV and the satellite; the impact includes scattering effect, masking effect, and scattering and masking effect; |g s | 2 The probability density function PDF satisfies Among them, 1F1(m; 1; δx) is the confluent hypergeometric function, and the Nakagami-m fading parameter is Real number m∈(0,∞), μ is the average power of the line-of-sight component, 2κ is the average value of the multipath power; superimposing equation (3) yields the cumulative distribution function CDF, which is expressed as follows in is a raised power function, is the gamma function; represents the incomplete gamma function; integrating (2) yields |g i | 2 Cumulative distribution function of The satellite channel status in the target area is G t ={g 1,1 ,…,g i,j ,…,g I,I }, g i,j Indicates l t = Satellite channel status at position [i, j]; l t The location eavesdropping channel state is expressed as in is the path loss of the eavesdropper, is the distance between the drone and Eve, θ t ∈[-π / 2,π / 2] is the azimuth angle from the UAV to Eve, a(θ t ) is the drone-eavesdropper link steering vector; At time slot t, the UAV transmits a signal: x(t)=w(t)s com (t)+w e (t)s rad (t) (6) where s com (t) and s rad (t) are the communication part and the perception part of the synaesthesia integrated signal, and is the precoding matrix; The signal received by the satellite is expressed as y(t)=g t w(t)s com (t)+g lt w e (t)s rad (t)+n (7) in for l t Position communication channel state, n~CN(0,σ 2 ) is Gaussian white noise; Eve receives the signal in for l t Location eavesdropping channel status, is Gaussian white noise; Define an auxiliary variable r(t) as the lower bound of the safe rate that the satellite can reach for the UAV in the tth time slot, and the achieved safe throughput C(t) is expressed as C(t)=I C ([C s (t)-C e (t)] + ≥r(t))·r(t) (9) Among them I C (·) is the indicator function, when [C s (t)-C e (t)] + When ≥r(t) C (·)=1, r(t) is the minimum expected communication rate of the link per time slot, C s (t) and C e (t) are the transmission rate from the UAV to the satellite and the eavesdropping rate, respectively, expressed as as well as Where B is the bandwidth; The probability of safe interruption from drone to satellite meets Pr{[C s (t) -C e (t)] + ≤r(t)}≤ε (12) Where ε≤0.1 is the maximum allowable interruption probability.
4. The method for UAV communication situation awareness based on deep reinforcement learning according to claim 3, characterized in that: The implementation method of step three is: The position parameters of Eve are estimated based on the radar echo signal; the uncertainty model of the potential eavesdropper's position is: In the formula is the location estimate obtained when sensing a potential eavesdropper, Δl e (t) is the position measurement error; In real location e (t) satisfies the Gaussian distribution, that is, Where δ is the measurement standard deviation; the measurement error of the tth time slot satisfies Among them G MF is the MF gain; The distribution measured over multiple time slots is accumulated as in satisfy is the measurement standard deviation of the cumulative distribution function; the threshold of measurement accuracy is set to Ω, when When , Eve’s position perception is sufficient; The estimated eavesdropping channel state is in Indicates that Eve's estimated position is In this case, t = the eavesdropping channel state at position [i, j].
5. The method for UAV communication situation awareness based on deep reinforcement learning according to claim 4, characterized in that: In step four, In each mission, the UAV collects N data files to be transmitted in the communication blind area. Each data file is not related to each other. The UAV needs to upload all the data it carries to the satellite in this mission. Each data file is associated with an onboard computing task. The nth onboard computing task is used (d an ,d bn ,c n ) description, where c n To calculate the number of CPU revolutions required for this task, d an and d bn are the data sizes before and after calculation respectively; define indicator variable x n (t)∈{0,1} represents the calculation strategy of the nth task in the tth time slot, x n (t) = 1 means that the data of the nth task has been calculated and processed, otherwise x n (t) = 0; for simplicity, define d a =[d a1 ,…,d aN ] T ,d b =[d b1 ,…,d bN ] T and c=[c1,…,c N ] T ; Using the computing function of the drone, the data files are compressed during the flight upload process to reduce the amount of data transmission; t time slot, drone position l t , carrying compressed data Uncompressed data volume Its satisfaction as well as in is the amount of data uploaded from the compressed queue, is the amount of data uploaded from the uncompressed queue, ΔT is the time slot length; To calculate the amount of data reduction in the uncompressed queue, The amount of data increase in the compressed queue after calculation is x(t)=[x1(t),…,x N (t)] T is the calculation strategy for the tth time slot.
6. The method for UAV communication situation awareness based on deep reinforcement learning according to claim 5, characterized in that: The implementation method of step five is: Analyze the power consumption constraints on the drone; the power consumption includes communication power consumption, computing power consumption and perception power consumption; the total power of the drone in each time slot is constrained, the total power P t satisfy P t =P w (t)+P e (t)+P c (t)≤P max ,t∈T (19) Among them, P w (t)=|w(t)| 2 is the transmission power, P e (t)=|w e (t)| 2 is the perceived power, To calculate power, To calculate the power consumption factor, f c (t) is the calculation frequency, P max The upper limit of power allowed for each time slot; In order to jointly optimize limited onboard resources to solve the problem of efficient and secure data transmission in satellite coverage blind spots under eavesdropping environments, a situational awareness optimization problem is established, so that data can be uploaded to the satellite in the shortest time under secure transmission conditions, and the drone returns to the starting point to perform the next mission; therefore, the objective function of the overall situational awareness optimization problem and the corresponding constraints are expressed as follows st(9),(18), (20a) c T x(t)≤ΔTf c (t),t∈[1,T s ] (20b) 0≤f c (t)≤f max ,t∈[1,T s ] (20c) x n (t)∈{0,1},t∈[1,T s ] (20d) L I ,f c ,P w ,P s ,T s ,X I ∈U (20e) in For the perception path, Calculate frequency, transmission power and calculation strategy for the perception stage; T s is the perception strategy, that is, the drone perception end time, satisfying To calculate the frequency, is the transmission power variable, is the perceived power variable.
7. The method for UAV communication situation awareness based on deep reinforcement learning according to claim 6, characterized in that: In step six, Since the objective function in the situation awareness optimization problem P1 has the characteristics of long-term accumulation, the situation awareness optimization problem P1 is transformed into a Markov decision problem, and the DRL method is used to solve the specific problem; The situational awareness optimization problem P1 is transformed into a Markov decision process, that is, a tuple (S, A, R, P, γ), where S is the state space, A is the action space, R is the reward space, P is the transition probability, and γ is the reward discount factor. At time t, the drone first observes the current environment information to obtain the state st∈S, then takes action at∈A to act on the environment, and then takes the transition probability P(s t+1 ∣s t ,a t )Transfer to the next state s t+1 , and generates the reward r after the action is executed t ; G t are the eavesdropping channel state and satellite channel state of the target area sensed at time t respectively; a t ={l t+1 ,P w (t+1),P s (t+1),f c (t+1),x(t+1),b t } (22) where b t ∈{0,1} indicates whether the perception phase is over, b t =1 means that the perception phase ends at the current moment; from (22), we can see that l t+1 ,x(t+1),b t is a discrete variable, P w (t+1),P s (t+1),f c (t+1) is a continuous variable; therefore, the agent's output action is a mixed high-dimensional vector, which increases the complexity of the search space for finding the optimal action; r t =R(t)+X1(t)+X2(t) (23) The constant penalty term X1(t), R(t) is expressed as: R(t)=ξ(P s (t),t)-ξ(P s (t-1),t-1) (24) The penalty term X2(t) with hysteresis is expressed as The trained Actor network is independently deployed on the drone, performing trajectory planning and resource allocation in real time based on the environment; ultimately, a situational awareness map is obtained, achieving the minimum completion time under the constraint of the lower limit of data completion volume.
8. The method for UAV communication situation awareness based on deep reinforcement learning according to claim 7, characterized in that: The specific implementation method of step seven is: based on steps five and six, construct a joint planning method for drone trajectories and resources based on situational awareness; the joint planning method for drone trajectories and resources uses deep reinforcement learning methods to perform adaptive drone trajectory planning, taking into account data calculation and uploading in an unknown eavesdropping environment while gradually building an eavesdropping environment situation map; according to the eavesdropping environment situation map, balance the drone's transmission, computing, and perception tasks under energy constraints to achieve coordination between the drone's transmission, computing, and perception tasks.