A full-duplex unmanned aerial vehicle-assisted integrated communication and sensing covert communication optimization method, system, medium and device

CN122802930APending Publication Date: 2026-09-22ANQING NORMAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610823628.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0009]本申请提供一种全双工无人机辅助的通感一体化隐蔽通信优化方法、系统、介质及设备,旨在解决背景技术中提出的现有技术在全双工无人机辅助的通感一体化场景下,尚缺乏兼顾高隐蔽性、高传输速率及强鲁棒性的协同优化方案的问题

Benefits of technology

[0020]本申请通过联合优化通信与感知功率、无人机三维轨迹及人工噪声发射,突破了传统技术中通信速率、雷达感知性能与隐蔽性相互制约的瓶颈,在满足雷达探测需求与严格隐蔽性约束的前提下,实现了平均隐蔽通信速率的最大化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802930A_ABST
    Figure CN122802930A_ABST
Patent Text Reader

Abstract

The application discloses a full-duplex unmanned aerial vehicle assisted integrated sensing and communication optimization method, system, medium and equipment, belonging to the field of wireless communication. The communication optimization method constructs a cooperative countermeasure architecture containing a full-duplex ground base station and a full-duplex air mobile unmanned aerial vehicle, deduces communication and sensing performance parameters by establishing air-ground and ground-ground channel models; for the uncertainty of the position of an eavesdropper, a robust concealment constraint boundary in the worst case is constructed, a soft actor-critic (SAC) deep reinforcement learning algorithm is used to convert a multi-constraint coupling problem into a Markov decision process for solving, and joint optimization of base station power distribution, unmanned aerial vehicle three-dimensional trajectory and artificial noise emission is realized. The communication optimization method significantly improves the concealment communication rate and robustness of the system in a dynamic electromagnetic environment without the need for heavy hardware interference cancellation equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless communication technology, specifically to a method, system, medium, and device for optimizing full-duplex unmanned aerial vehicle-assisted integrated sensing and covert communication. Background Technology

[0002] With the evolution of sixth-generation mobile communication (6G) and integrated air-ground networks, the Integrated Sensing and Communication Capabilities (ISAC) technology, through the unified reuse of spectrum resources and radio frequency hardware, achieves deep integration of communication and sensing functions within the same waveform and time slot, and has become a core enabling technology for wireless networks. Unmanned aerial vehicles (UAVs), with their high mobility, wide-area coverage, and strong line-of-sight (LoS) transmission advantages, can be widely applied in scenarios such as emergency communication, low-altitude security, and disaster monitoring when combined with ISAC technology. Furthermore, the introduction of full-duplex technology gives UAVs the ability to transmit and receive simultaneously on the same frequency, enabling them to transmit detection or jamming waveforms while receiving legitimate signals, providing technical support for improving the environmental adaptability and proactive defense capabilities of air-based networks.

[0003] However, the inherent broadcast characteristics and strong line-of-sight links of UAV communication make them highly vulnerable to signal interception and detection by external nodes, posing security risks of communication exposure and data leakage. Traditional upper-layer encryption technologies can only ensure the confidentiality of information content, but cannot conceal the physical existence of the communication link, making them susceptible to passive eavesdropping attacks based on energy detection. Covert communication technology, aiming to hide communication behavior from the physical layer through low probability of detection (LPD) transmission, is a key means to solve the above problems. Combined with full-duplex technology, UAVs can actively emit artificial noise to mask legitimate signals, thereby building a highly covert security protection system while ensuring communication functionality.

[0004] Chinese invention patent application CN118042454A discloses a physical layer secure transmission method and system in a sensor-integrated drone network. It uses sensor-integrated signals to sense the location of eavesdroppers and combines trajectory planning and artificial noise beamforming to improve communication security. However, relying on traditional physical layer encryption only protects the confidentiality of the content and cannot hide the energy characteristics of the communication signal itself, making the drone's communication behavior easily detected and located passively.

[0005] Chinese invention patent CN117880817B discloses a method, apparatus, and electronic device for determining the trajectory and beamforming vector of a UAV. It maximizes the average covert communication rate by jointly optimizing the UAV trajectory and beamforming vector. However, it relies on environmental background noise to achieve covertness, requiring the system to reduce transmission power or adjust direction to meet covert constraints, thus limiting the legal communication rate and making it difficult to balance high covertness with high transmission performance.

[0006] Chinese invention patent CN118869037B discloses a trajectory and power design method and system for a covert relay communication network for unmanned aerial vehicles (UAVs). It achieves covert relay transmission by jointly optimizing the UAV trajectory and user / UAV power, utilizing environmental background noise and jammer signals. However, its reliance on passive environmental noise or external jammers to mask communication behavior necessitates a significant reduction in transmission power to meet covert constraints under strong line-of-sight links. This results in a significant decrease in legitimate communication rates and transmission reliability, making it difficult to balance high covertness with high system efficiency.

[0007] Chinese invention patent application CN118432672A discloses a covert transmission strategy for a NOMA-RIS-assisted integrated sensing system. It introduces NOMA technology and a reconfigurable smart surface (RIS), utilizing RIS to adjust signal phase to constructively superimpose legitimate signals and disrupt eavesdropping signals. Simultaneously, it combines beamforming optimization to maximize covert communication rates. However, this scheme relies on the ideal assumption that the eavesdropper's location is completely known, depends on complex channel state information (CSI) acquisition, lacks robust modeling for uncertain eavesdropper locations, and is difficult to adapt to real-world scenarios with dynamic changes in eavesdropper locations or imperfect information acquisition in complex electromagnetic environments.

[0008] In summary, existing technologies lack a collaborative optimization scheme that balances high concealment, high transmission rate, and strong robustness in full-duplex drone-assisted sensory integration scenarios. Therefore, this application provides a method, system, medium, and device for optimizing concealed communication in full-duplex drone-assisted sensory integration to address the aforementioned problems. Summary of the Invention

[0009] This application provides a method, system, medium, and device for optimizing full-duplex drone-assisted sensor-integrated covert communication, aiming to solve the problem mentioned in the background art that the existing technology lacks a collaborative optimization scheme that takes into account high concealment, high transmission rate, and strong robustness in full-duplex drone-assisted sensor-integrated scenarios.

[0010] To achieve the above objectives, this application provides the following technical solution: an optimization method for full-duplex unmanned aerial vehicle (UAV)-assisted integrated sensing and covert communication, which is applied to a full-duplex UAV-assisted integrated sensing and covert communication system, the communication system including a full-duplex ground base station, a full-duplex aerial mobile UAV, an aerial radar-sensing target, and a ground eavesdropper with an uncertain location; the method includes: Air-to-ground and ground-to-ground channel models were established separately. Channel characteristic parameters of legitimate communication links, sensing links, and eavesdropping links were derived. The communication signal-to-interference-plus-noise ratio (SIR) of the full-duplex airborne mobile UAV receiver, the sensing SIR of the full-duplex ground base station receiver, and the target covert communication rate were calculated. Based on the signal detection mechanism of ground eavesdroppers, a binary hypothesis detection model is constructed. Considering the uncertainty of the location of ground eavesdroppers, the worst-case channel gain extremum is solved through geometric relationships, and a robust constraint boundary is constructed to ensure the concealment of the system. With the goal of maximizing the average covert communication rate, a multi-constraint coupled optimization problem is constructed by integrating the three-dimensional flight maneuver constraints of full-duplex aerial mobile UAVs, the transmit power constraints of full-duplex ground base stations, the performance constraints of aerial radar sensing targets, and the robust covertness constraint boundary. Define the state space and action space, design a reward function that includes a penalty term for constraint violation, transform the multi-constraint optimization problem into a Markov decision process, and use a deep reinforcement learning algorithm to train and obtain the optimal control policy; The optimal control strategy, once trained, is deployed to the communication system to generate and execute full-duplex ground base station power allocation commands, full-duplex aerial mobile UAV flight trajectory commands, and jamming signal transmission commands in real time, thereby achieving closed-loop control.

[0011] Preferably, the air-to-ground channel model adopts a probabilistic line-of-sight channel model, and its large-scale average channel power gain is dynamically calculated based on the three-dimensional distance, elevation angle and environmental parameters between the full-duplex airborne mobile UAV and the full-duplex ground base station. The ground-to-ground channel model adopts the Rayleigh fading channel model, and its channel gain follows an exponential distribution.

[0012] Preferably, the robust constraint boundaries for constructing the system's concealment include: Obtain the estimated center and uncertainty radius of the eavesdropper's location; The worst-case extreme distances between the ground eavesdropper and the full-duplex ground base station, and between the ground eavesdropper and the full-duplex aerial mobile drone are calculated based on geometric relationships, and the worst-case extreme channel gain is calculated based on the extreme distances. Calculate the minimum interference to leakage power ratio based on the channel gain extremum; Construct robust constraints that ensure the average minimum detection error probability is not lower than a preset concealment threshold.

[0013] Preferably, the geometric relationship is expressed using the triangle inequality.

[0014] Preferably, the deep reinforcement learning algorithm employs the soft actor-critic algorithm, and the decision optimization problem is modeled as a Markov decision process; The state space includes the position state, perception performance state, and stealth state of the full-duplex aerial mobile UAV. The action space includes the power allocation parameters of the full-duplex ground base station and the flight control parameters of the full-duplex aerial mobile UAV; The reward function includes a communication rate gain term and a constraint violation penalty term.

[0015] Preferably, the soft actor-critic algorithm includes a double-Q network structure, reparameterization techniques, and a target network soft update mechanism.

[0016] Preferably, the deployment execution optimal control strategy includes: Within each control time slot, the optimal control strategy is input based on the current system state, and the full-duplex ground base station power allocation parameters and the full-duplex aerial mobile UAV flight speed vector parameters are output. Control commands are sent to full-duplex ground base stations and full-duplex aerial mobile drones via the control link; the full-duplex ground base stations adjust the power distribution of communication signals and sensing signals according to the commands. Full-duplex aerial mobile drones perform spatial maneuvers and emit artificial noise interference signals according to instructions.

[0017] A full-duplex UAV-assisted sensor-integrated covert communication optimization system, the system being used to execute the aforementioned full-duplex UAV-assisted sensor-integrated covert communication optimization method, the communication optimization system comprising: The channel modeling module establishes air-to-ground and ground-to-ground channel models respectively, derives the channel characteristic parameters of legitimate communication links, sensing links, and eavesdropping links, and calculates the communication signal-to-interference-plus-noise ratio (SIR) of the full-duplex airborne mobile UAV receiver, the sensing SIR of the full-duplex ground base station receiver, and the target covert communication rate. The constraint construction module constructs a binary hypothesis detection model based on the signal detection mechanism of ground eavesdroppers. It also solves the worst-case channel gain extremum through geometric relationships to address the uncertainty of the ground eavesdropper's location, thus constructing a robust constraint boundary to ensure the system's concealment. The optimization problem construction module aims to maximize the average covert communication rate. It integrates the three-dimensional flight maneuver constraints of full-duplex aerial mobile UAVs, the transmit power constraints of full-duplex ground base stations, the performance constraints of aerial radar sensing targets, and the robust covert constraint boundary to construct a multi-constraint coupled optimization problem. The intelligent solution module defines the state space and action space, designs a reward function that includes a penalty term for constraint violation, transforms the multi-constraint optimization problem into a Markov decision process, and uses a deep reinforcement learning algorithm to train and obtain the optimal control strategy. The strategy execution module deploys the trained optimal control strategy to the communication system, and generates and executes full-duplex ground base station power allocation commands, full-duplex aerial mobile UAV flight trajectory commands, and jamming signal transmission commands in real time to achieve closed-loop control.

[0018] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described full-duplex UAV-assisted sensor-integrated covert communication optimization method.

[0019] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-described full-duplex UAV-assisted sensor-integrated covert communication optimization method.

[0020] This application overcomes the bottleneck of mutual constraints between communication rate, radar perception performance and stealth in traditional technologies by jointly optimizing communication and sensing power, UAV three-dimensional trajectory and artificial noise emission. Under the premise of meeting radar detection requirements and strict stealth constraints, it maximizes the average stealth communication rate.

[0021] This application addresses the complex scenario where the location of the eavesdropper is uncertain. It constructs a robust concealment constraint model under worst-case conditions and utilizes a full-duplex UAV to emit artificial noise to actively mask base station signal leakage. This achieves low probability of detection (LPD) transmission at the physical layer, effectively solving the problem of a sharp drop in security due to eavesdropper location estimation errors in existing technologies and significantly improving the system's survivability in dynamic electromagnetic environments.

[0022] This application employs a soft actor-critic (SAC) algorithm based on the maximum entropy framework, which efficiently solves the problem of joint optimization in high-dimensional non-convex and strongly coupled environments, avoiding getting trapped in local optima. At the same time, it compensates for residual self-interference in the full-duplex mode of UAVs through refined modeling at the algorithm level, ensuring the smoothness and engineering feasibility of the optimized trajectory without relying on expensive heavy hardware elimination equipment. Attached Figure Description

[0023] Figure 1 System architecture diagram of a full-duplex drone-assisted integrated sensory covert communication system; Figure 2 Overall control flowchart for an optimized method of sensor-integrated covert communication assisted by full-duplex UAVs; Figure 3 The flowchart for maximum entropy network training control of the maximum entropy deep reinforcement learning SAC algorithm for synesthetic covert collaboration. Figure 4 Flowchart for calculating worst-case robustness concealment extrema; Figure 5 Comparison of UAV 3D flight trajectories optimized for SAC algorithm, PPO, and TD3 algorithm; Figure 6 A comparison chart showing the evolution of average communication rate over time steps; Figure 7 Concealment threshold Simulation graph showing the relationship between the average covert communication rate and the actual rate. Figure 8 Radius of uncertainty of the eavesdropper's location Simulation graph showing the relationship between the average covert communication rate and the actual rate. Figure 9 Self-interference residual coefficient Simulation graph showing the relationship between the average covert communication rate and the actual rate. Figure 10 To sense the SINR threshold Simulation diagram showing the relationship between the average covert communication rate and the data. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] like Figure 1 As shown, the full-duplex UAV-assisted integrated sensing and covert communication system includes a full-duplex ground base station, a full-duplex aerial mobile UAV, an aerial radar-sensing target, and an eavesdropper with an uncertain location. The full-duplex ground base station and the ground eavesdropper are ground nodes, while the full-duplex aerial mobile UAV and the aerial radar-sensing target are aerial nodes. Both the full-duplex ground base station and the full-duplex aerial UAV are equipped with dual antennas capable of transmitting and receiving. The full-duplex ground base station transmits a superimposed signal combining sensing and communication, while the full-duplex aerial UAV receives the covert communication signal and simultaneously transmits artificial interference to the detection of the ground eavesdropper. The ground eavesdropper uses an average power detector to monitor the covert communication link. The full-duplex ground base station completes radar target detection by receiving the echo signal from the aerial radar-sensing target.

[0026] The specific channel modeling for each link is as follows: Legitimate covert communication link: Full-duplex ground base stations transmit superimposed signals of converged communication and sensing in the same frequency band, among which the covert communication signal is sent to full-duplex aerial mobile drones; Sensing echo link: The sensing signal illuminates the target sensed by the airborne radar, and is reflected by the target sensed by the airborne radar to the full-duplex ground base station for target detection; Artificial noise jamming link: To counter ground-based eavesdroppers and improve system stealth, full-duplex aerial mobile drones transmit artificial noise towards the eavesdropper while receiving covert communication signals. However, because the aerial mobile drone operates in full-duplex mode, the transmitted artificial noise generates irremovable internal residual self-interference at its own receiving antenna. Cross-interference links: The artificial noise transmitted by full-duplex aerial mobile drones propagates in space, inevitably causing spatial cross-interference with the radar echo received by full-duplex ground base stations; Signal leakage link: The high-power superimposed signal transmitted by a full-duplex terrestrial base station inevitably spreads towards the ground eavesdropper during spatial propagation. Ground eavesdroppers with uncertain locations attempt to use average power detectors to listen to and detect the transmission status of the concealed signal. The eavesdropping signal they actually receive is composed of the signal leakage link leakage signal and the artificial noise from the cross-interference link superimposed in space.

[0027] To ensure the robust operation of the aforementioned communication system and maximize the covert communication rate, this application provides a full-duplex UAV-assisted sensor-integrated covert communication optimization method, such as... Figure 2 As shown, the optimization method includes the following steps: S100. Establish air-to-ground and ground-to-ground channel models respectively, derive the channel characteristic parameters of the legitimate communication link, sensing link, and eavesdropping link, and calculate the communication signal-to-interference-plus-noise ratio (SIR) of the full-duplex airborne mobile UAV receiver, the sensing SIR of the full-duplex ground base station receiver, and the target covert communication rate, specifically: S110. For the legitimate communication link from a full-duplex ground base station to a full-duplex airborne mobile UAV, and the sensing link from a full-duplex ground base station to an airborne radar-sensing target, a hybrid line-of-sight / non-line-of-sight probability air-to-ground channel model is adopted. Only the large-scale average channel gain is used for optimization. The average channel power gain is:

[0028] Among them, the time slot index is , To accept node indexes, Represents a full-duplex aerial mobile drone. This represents targets detected by airborne radar. The elevation angle is the distance between the transmitting node and the receiving node. , This represents the node height. The line-of-sight probability function is based on the elevation angle, and the elevation angle between nodes is... Relatedly, the S-curve model is used for calculation, and its expression is: ;in, and These are constants related to the environment. The channel gain at a reference distance of 1m; The distance between nodes is the Euclidean distance. and These are the path loss indices for line-of-sight links and non-line-of-sight links, respectively.

[0029] S120. For the interference link between a full-duplex aerial mobile UAV and a ground eavesdropper, an air-to-ground composite channel model is adopted, where the instantaneous channel power gain is the product of the large-scale average gain and the small-scale Rayleigh fading: .

[0030] Its large-scale average gain Following a similar line-of-sight / non-line-of-sight probability model to equation (1), it is specifically expressed as:

[0031] in, The Euclidean distance between the drone and the eavesdropper. This represents the corresponding line-of-sight probability. is the small-scale complex Gaussian fading coefficient.

[0032] S130. For the ground-to-ground eavesdropping link from the full-duplex ground base station to the ground eavesdropper, a non-line-of-sight probability-dominated Rayleigh fading channel model is adopted, with the instantaneous channel power gain being... ,in This represents the large-scale average gain. The path loss index for ground-to-ground eavesdropping links. The Euclidean distance between a full-duplex ground base station and a ground eavesdropper. For small-scale fading factors that follow an exponential distribution, .

[0033] S140, Derivation of Communication Signal-to-Interference-plus-Noise Ratio and Achievable Rate for Full-Duplex Aerial Mobile UAVs: A full-duplex ground base station transmits a superimposed signal combining sensing and communication. The signal received by the full-duplex aerial mobile UAV includes not only the effective communication and sensing signals transmitted by the base station, but also residual self-interference signals caused by artificial noise transmitted synchronously in full-duplex mode. (In time slots) Instantaneous signal-to-interference-plus-noise ratio at a full-duplex aerial mobile drone The expression is deduced as follows:

[0034] in, and These refer to base station communication and sensing transmission power, respectively. The residual self-interference coefficient in full-duplex mode of the UAV. The self-interference channel gain follows an exponential distribution. ; Artificial noise power emitted by full-duplex aerial mobile drones , The Gaussian white noise power of the full-duplex aerial mobile drone receiver.

[0035] Based on this, considering that the randomness of the residual self-interference channel can lead to communication interruptions, the interruption probability is defined as the instantaneous rate being lower than the target rate. The probability of that, i.e.

[0036] To meet the maximum allowable interruption probability Constraints (i.e.) ), and deduce the system's performance under the worst artificial interference power. The target's covert communication rate is:

[0037] Derivation of the Signal-to-Interference-plus-Noise Ratio (SIR) of the Sensing Echo from a Full-Duplex Ground Base Station: The radar echo signal received by a full-duplex ground base station not only contains the superimposed signal reflected from the sensed target, but is also inevitably affected by artificial noise emitted by full-duplex aerial mobile drones. After the full-duplex ground base station processes the echo signal using a linear receiver, the SIR of its radar sensing echo is derived as follows:

[0038] in, To characterize the combined parameters of radar cross section and two-way path loss, The sensing link gain from the base station to the target. For the interference link gain from the drone to the base station, This represents the Gaussian white noise power at the base station's sensing end. To ensure effective radar target detection, the system must ensure that the sensing signal-to-interference-plus-noise ratio (SINR) is greater than a preset threshold. .

[0039] S200, based on the signal detection mechanism of ground eavesdroppers, constructs a binary hypothesis detection model, and, considering the uncertainty of the ground eavesdropper's location, solves the worst-case channel gain extremum through geometric relationships, constructing a robust constraint boundary to ensure system concealment, such as... Figure 4 As shown, specifically: S210. Based on the average power detection mechanism of ground eavesdroppers, define the assumption of uncovered communication. The assumption of covert communication Deriving the false alarm probability With the probability of missed alarms A binary hypothesis detection model is constructed, and the minimum detection error probability under the optimal detection threshold is solved. ; Among them, the detection error probability of ground eavesdroppers satisfies By analyzing the values ​​of the detection threshold τ[t] in different intervals, the minimum detection error probability under the optimal detection threshold is obtained: ; in, For the instantaneous channel power gain of the full-duplex aerial mobile UAV to ground eavesdropper link, The instantaneous channel power gain for the full-duplex ground base station to ground eavesdropper link; ; ; This represents the maximum value of artificial noise in the current time slot. This represents the minimum value of artificial noise in the current time slot.

[0040] S220, Minimum error probability The expectation is calculated over the joint distribution of the full-duplex ground base station to ground eavesdropper and the full-duplex aerial mobile drone to ground eavesdropper channels. This expectation is then solved using double integrals and the Frullani integral identity to obtain a closed-form solution with the minimum average detection error probability. in, The ratio of interference to leakage power. , For the large-scale average gain of the full-duplex aerial mobile UAV to ground eavesdropper link, For the large-scale average gain of the full-duplex ground base station to ground eavesdropper link, The sensing transmit power for a full-duplex ground base station; S230, Obtain the estimated center and uncertainty radius of the eavesdropper's location. ; S240. Calculate the worst-case extreme distances between the ground eavesdropper and the full-duplex ground base station, and between the ground eavesdropper and the full-duplex aerial drone, based on geometric relationships, and calculate the worst-case extreme channel gain based on these extreme distances; wherein the geometric relationships are expressed using the triangle inequality; specifically: The maximum horizontal distance between a ground-based eavesdropper and a full-duplex aerial mobile drone is calculated using the triangle inequality. , The full-duplex aerial mobile UAV's planar projection coordinates in time slot t, corresponding to the maximum 3D distance. With minimum elevation angle Substituting into the air-to-ground channel model, the minimum line-of-sight probability is obtained. Finally, the minimum channel gain for the worst-case scenario of a full-duplex aerial drone to a ground eavesdropper was calculated: Solving for the minimum horizontal distance between a ground eavesdropper and a full-duplex ground base station using the triangle inequality. Substituting into the ground-to-ground channel model, we obtain the maximum channel gain from the full-duplex ground base station to the ground eavesdropper in the worst-case scenario: ; in, For the fixed location of a full-duplex ground base station, This is the ground-to-ground link path loss index. S250, Calculate the minimum interference to leakage power ratio based on the channel gain extreme value. Specifically: Based on the radius of uncertainty of the location of ground eavesdroppers Derive the minimum disturbance to leakage power ratio in the worst-case scenario. ,in Minimum large-scale gain corresponding to the full-duplex aerial mobile UAV to ground eavesdropper link Maximum large-scale gain of the link between full-duplex terrestrial base station and ground eavesdropper ; S260. Construct robust constraints that ensure the average minimum detection error probability is not lower than a preset concealment threshold: ; Among them, ∈[0,1] is the preset concealment requirement threshold.

[0041] S300. With maximizing the average covert communication rate as the optimization objective, this paper constructs a multi-constraint coupled optimization problem by integrating the three-dimensional flight maneuver constraints of a full-duplex aerial mobile UAV, the transmit power constraints of a full-duplex ground base station, the performance constraints of airborne radar sensing targets, and the aforementioned robust covertness constraint boundary. Specifically:

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048] in, The total number of task time slots. For the purpose of concealing communication rates for the target. To sense the echo signal-to-interference-plus-noise ratio for base stations. The minimum threshold for perceived signal-to-interference-plus-noise ratio; The three-dimensional coordinates of a full-duplex aerial mobile UAV in time slot t. and These are the three-dimensional coordinates of a full-duplex aerial mobile UAV in the initial and final time slots, respectively. and These are the system-defined starting and ending coordinates of the full-duplex aerial mobile UAV mission. The maximum flight speed for a full-duplex aerial mobile drone, This represents the maximum total transmission power of the base station. and These represent the lower and upper bounds of the three-dimensional flight area of ​​a full-duplex aerial mobile UAV.

[0049] S400. Define the state space and action space, design a reward function that includes a penalty for constraint violation, transform the multi-constraint optimization problem into a Markov decision process, and use a deep reinforcement learning algorithm to train and obtain the optimal control strategy. The state space includes the position state, perception performance state, and stealth state of the full-duplex aerial mobile UAV, while the action space includes the power allocation parameters of the full-duplex ground base station and the flight control parameters of the full-duplex aerial mobile UAV. The reward function includes a communication rate gain term and a constraint violation penalty term; The deep reinforcement learning algorithm employs the soft actor-critic algorithm, which includes a double-Q network structure, reparameterization techniques, and a target network soft update mechanism.

[0050] S410. Transform the multi-constraint optimization problem into a Markov decision process, specifically: S411. Map the strongly coupled multi-constraint optimization problem to a Visio Markov Decision Process (MDP), where the state space of the Markov Decision Process is defined as... in, Let be the three-dimensional coordinates of the full-duplex aerial mobile UAV in the t-th time slot; For the first Radar sensing signal-to-interference-plus-noise ratio in each time slot; For the first Minimum average detection error probability per time slot; Action space ,in Configure a composite reward function including a penalty term for the three-dimensional velocity vector of a full-duplex aerial mobile UAV:

[0051] in, Positive rewards are given for the distance to the destination. For communication rate Incentive weights, and These are the penalty weight constants for perceived violations and concealed violations, respectively. It is a non-negative truncation operator; S412. Based on the maximum entropy reinforcement learning framework, the optimization objective is to maximize the expected reward and policy entropy. sum:

[0052] in, To balance the temperature parameter between expected return and action exploration policy entropy weights, the parameters of the double soft Q network are updated by minimizing the soft Bellman residual, and the actor network policy parameters are updated by minimizing the KL divergence and utilizing reparameterization techniques. .

[0053] S420. The optimal control strategy is obtained by training using a deep reinforcement learning algorithm, such as... Figure 3 As shown, specifically: S421. Network and Buffer Initialization: Initialize Actor Network Parameters Two soft Q network parameters Two target Q-network parameters and order , Initialize the experience replay buffer .

[0054] S422, State Observation and Action Interaction: In each training round, the environment is reset and the initial state is observed. Full-duplex aerial mobile drones are based on the current strategy network. Sampling action Interact with the environment and observe the rewards from the environment's feedback. Next state With end mark ; S423, Experience Storage and Data Sampling: Transferring Data Tuples Store to experience playback buffer When the amount of data in the buffer meets the update conditions, a batch of transferred data is randomly sampled from the buffer, and subsequent network parameter updates are performed. S424. Target Value Calculation: Sampling the Action at the Next Moment Calculating target value based on the maximum entropy framework :

[0055] in, As a discount factor, The temperature parameter is used to balance expected returns with the entropy weight of the action exploration strategy; S425, SoftQ Network Update: Update the two softQ networks by minimizing the mean squared error loss. The loss function is:

[0056] S426, Actor Network Update: Employs reparameterized motion sampling techniques. ,in Noise sampled from a standard Gaussian distribution. This is a reparameterized function. The actor network is updated by minimizing the KL divergence loss, with the loss function being:

[0057] S427, Target Network Soft Update: Soft update the target Q-network parameters using the Polyak averaging method.

[0058] in, The update rate is soft. Repeat the above training process until the reward function converges, ultimately outputting the optimal control policy network. At this point, a power allocation and flight decision-making method is needed to maximize the average covert communication rate while meeting stringent constraints such as perception and concealment. S500: The trained optimal control strategy is deployed to the control system, generating and executing full-duplex ground base station power allocation commands, full-duplex aerial mobile UAV flight trajectory commands, and jamming signal transmission commands in real time to achieve closed-loop control. The deployment and execution of the optimal control strategy includes: Within each control time slot, the optimal control strategy is input based on the current system state, and the full-duplex ground base station power allocation parameters and the full-duplex aerial mobile UAV flight speed vector parameters are output. Control commands are sent to full-duplex ground base stations and full-duplex aerial mobile drones via the control link; the full-duplex ground base stations adjust the power distribution of communication signals and sensing signals according to the commands. Full-duplex aerial mobile drones perform spatial maneuvers and emit artificial noise interference signals according to instructions.

[0059] This application also provides a full-duplex UAV-assisted sensor-integrated covert communication optimization system. The communication optimization system is used to execute the above-mentioned full-duplex UAV-assisted sensor-integrated covert communication optimization method. The communication optimization system includes a channel modeling module, a constraint construction module, an optimization problem construction module, an intelligent solution module, and a strategy execution module. The channel modeling module establishes air-to-ground and ground-to-ground channel models respectively, derives the channel characteristic parameters of legitimate communication links, sensing links and eavesdropping links, and calculates the communication signal-to-interference-plus-noise ratio of the full-duplex airborne mobile UAV receiver, the sensing signal-to-interference-plus-noise ratio of the full-duplex ground base station receiver, and the target covert communication rate. The constraint construction module constructs a binary hypothesis detection model based on the signal detection mechanism of ground eavesdroppers, and solves the worst-case channel gain extreme value through geometric relationships to address the uncertainty of the ground eavesdropper's location, thereby constructing a robust constraint boundary to ensure the system's concealment. The optimization problem construction module takes maximizing the average covert communication rate as the optimization objective and integrates the three-dimensional flight maneuver constraints of full-duplex aerial mobile UAVs, the transmit power constraints of full-duplex ground base stations, the performance constraints of aerial radar sensing targets, and the robust covert constraint boundary to construct a multi-constraint coupled optimization problem. The intelligent solution module defines the state space and action space, designs a reward function that includes a penalty term for constraint violation, transforms the multi-constraint optimization problem into a Markov decision process, and uses a deep reinforcement learning algorithm to train and obtain the optimal control strategy. The strategy execution module deploys the trained optimal control strategy to the communication system, and generates and executes full-duplex ground base station power allocation commands, full-duplex aerial mobile UAV flight trajectory commands, and interference signal transmission commands in real time to achieve closed-loop control.

[0060] To verify the performance of the method proposed in this invention, the following simulation experiments were conducted. The core parameters of the simulation environment are set as shown in the table below: Node 3D coordinate settings: Full-duplex ground base station: Initial position of full-duplex aerial mobile drone End position Airborne radar detects target Tom: Ground eavesdroppers estimate the center location: .

[0061] Simulation Results and Analysis (1) Three-dimensional maneuver trajectory optimization and constraint satisfaction performance: such as Figure 5 As shown, the three-dimensional flight trajectory of a full-duplex aerial mobile UAV, optimized based on the soft actor-critic algorithm of this invention, exhibits superior global search and dynamic decision-making capabilities compared to traditional baseline algorithms such as PPO and TD3. This is achieved while strictly meeting the stealth requirement threshold. and the perceived signal-to-interference-plus-noise ratio threshold Under stringent conditions, the soft actor-critic algorithm intelligently balances the requirements for target perception, the need for stealth from ground eavesdroppers, and physical flight energy consumption. The drone successfully completed its maneuver from start to finish with a smooth trajectory free from any physical or communication constraints, fully demonstrating the excellent decision-making ability of the soft actor-critic algorithm in handling multi-constraint, high-dimensional spatial problems.

[0062] (2) Maximizing the performance of covert communication rate: such as Figure 6 As shown in the time-step rate evolution diagram, the communication rate curve output by the soft actor-commentator algorithm of this invention is more stable than that of the PPO and TD3 algorithms, effectively avoiding communication interruptions caused by violations of concealment constraints. The average concealed communication rate of the system stably reaches [a certain value]. This resulted in a significant increase in throughput. Furthermore, such as... Figure 7 As shown, with the threshold of concealment requirements... As the restrictions gradually loosen, the average communication rate shows an upward trend; especially in strictly concealed areas, the soft actor-critic algorithm still achieves a higher average concealed communication rate than other algorithms by accurately matching the optimal power allocation strategy in the continuous action space, proving that it still performs well under extreme and constrained conditions.

[0063] (3) Anti-interference and system robustness: such as Figure 8 As shown, with the radius of uncertainty in the location of the ground eavesdropper... As the uncertainty increases, the average communication rate inevitably decreases due to the adoption of the worst-case extreme value conservative boundary derived in this invention. However, the soft actor-critic algorithm exhibits the gentlest rate decrease, maintaining high covert communication performance even in extremely high uncertainty scenarios. This directly verifies the effectiveness of the robust covertness constraint derived in this invention at the physical level; simultaneously, as... Figure 9 As shown, facing the inherent self-interference challenge under the full-duplex system, with the self-interference residual coefficient... As the noise level increases, the soft actor-critic algorithm can dynamically reduce the artificial noise power or compensate for the communication transmit power based on the instantaneous channel state. Through this game-theoretic mechanism, the system effectively suppresses the degradation of legitimate receiver performance caused by severe self-interference, maintaining its performance lead over the baseline algorithm in simulations.

[0064] (4) Sensing and communication coordination performance: such as Figure 10 As shown, with the perceived SINR threshold The rigidity of the system's overall transmit power is inevitably limited, leading to a compression of communication power and a reasonable decrease in average communication rate. This accurately reflects the inherent sensory resource competition characteristics of the soft actor-commentator system. However, under the same stringent high-sensitivity requirements, this invention utilizes the soft actor-commentator algorithm's micro-allocation capability of radio frequency resources to achieve a communication rate superior to the PPO and TD3 algorithms, realizing a dynamic and coordinated balance between radar target detection performance and covert data transmission.

[0065] The simulation results above demonstrate that by constructing an integrated sensing and covert communication system model for UAV base stations, deriving robust covertness constraints, and implementing multi-dimensional joint optimization based on the soft actor-critic algorithm, this invention effectively improves the covert communication rate of the system while ensuring sensing performance and covertness, thus verifying the effectiveness and superiority of the system and method proposed in this invention.

[0066] Secondly, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described full-duplex UAV-assisted sensor-integrated covert communication optimization method.

[0067] Specifically, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0068] In this embodiment, the computer-readable storage medium may include, but is not limited to: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0069] Preferably, the computer-readable storage medium is a non-volatile storage medium, such as a solid-state drive (SSD) or flash memory.

[0070] Additionally, this application provides an electronic device, which may be a server, desktop computer, laptop computer, or embedded image processing device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the aforementioned full-duplex UAV-assisted sensor-integrated covert communication optimization method.

[0071] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In this embodiment, the processor is mainly used to perform convolution operations, linear attention matrix operations, and regression calculations for fully connected layers.

[0072] Memory, as a computer-readable storage medium, is used to store software programs, computer-executable programs, and data. Memory can primarily include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a given function (such as an image processing driver, a quality assessment algorithm library, etc.); the data storage area may store data created based on terminal usage (such as a light field image dataset to be evaluated, model weight files, etc.). Furthermore, memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0073] It should be noted that the electronic device in this embodiment can be a standalone computing device or part of a cloud server cluster. When applied in the cloud, the electronic device receives light field images uploaded by the terminal via the network, completes quality assessment, and then feeds back the scoring results to the terminal.

[0074] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in this application, based on the technical solution and concept of this application, should be included within the scope of protection of this application.

Claims

1. A method for optimizing full-duplex unmanned aerial vehicle-assisted sensor-integrated covert communication, characterized in that, The optimization method is applied to a full-duplex unmanned aerial vehicle (UAV)-assisted integrated sensing and covert communication system, which includes a full-duplex ground base station, a full-duplex aerial mobile UAV, an aerial radar-sensing target, and a ground-based eavesdropper with an uncertain location; the method includes: Air-to-ground and ground-to-ground channel models were established separately. Channel characteristic parameters of legitimate communication links, sensing links, and eavesdropping links were derived. The communication signal-to-interference-plus-noise ratio (SIR) of the full-duplex airborne mobile UAV receiver, the sensing SIR of the full-duplex ground base station receiver, and the target covert communication rate were calculated. Based on the signal detection mechanism of ground eavesdroppers, a binary hypothesis detection model is constructed. Considering the uncertainty of the location of ground eavesdroppers, the worst-case channel gain extremum is solved through geometric relationships, and a robust constraint boundary is constructed to ensure the concealment of the system. With the goal of maximizing the average covert communication rate, a multi-constraint coupled optimization problem is constructed by integrating the three-dimensional flight maneuver constraints of full-duplex aerial mobile UAVs, the transmit power constraints of full-duplex ground base stations, the performance constraints of aerial radar sensing targets, and the robust covertness constraint boundary. Define the state space and action space, design a reward function that includes a penalty term for constraint violation, transform the multi-constraint optimization problem into a Markov decision process, and use a deep reinforcement learning algorithm to train and obtain the optimal control policy; The optimal control strategy, once trained, is deployed to the communication system to generate and execute full-duplex ground base station power allocation commands, full-duplex aerial mobile UAV flight trajectory commands, and jamming signal transmission commands in real time, thereby achieving closed-loop control.

2. The full-duplex UAV-assisted sensor-integrated covert communication optimization method according to claim 1, characterized in that, The air-to-ground channel model adopts a probabilistic line-of-sight channel model, and its large-scale average channel power gain is dynamically calculated based on the three-dimensional distance, elevation angle and environmental parameters between the full-duplex airborne mobile UAV and the full-duplex ground base station. The ground-to-ground channel model adopts the Rayleigh fading channel model, and its channel gain follows an exponential distribution.

3. The full-duplex UAV-assisted sensor-integrated covert communication optimization method according to claim 1, characterized in that, The robust constraint boundaries for constructing the system to ensure concealment include: Obtain the estimated center and uncertainty radius of the eavesdropper's location; The worst-case extreme distances between the ground eavesdropper and the full-duplex ground base station, and between the ground eavesdropper and the full-duplex aerial mobile drone are calculated based on geometric relationships, and the worst-case extreme channel gain is calculated based on the extreme distances. Calculate the minimum interference to leakage power ratio based on the channel gain extremum; Construct robust constraints that ensure the average minimum detection error probability is not lower than a preset concealment threshold.

4. The full-duplex UAV-assisted sensor-integrated covert communication optimization method according to claim 3, characterized in that, The geometric relationship is expressed using the triangle inequality.

5. The full-duplex UAV-assisted sensor-integrated covert communication optimization method according to claim 1, characterized in that, The deep reinforcement learning algorithm employs the soft actor-critic algorithm, and the decision optimization problem is modeled as a Markov decision process. The state space includes the position state, perception performance state, and stealth state of the full-duplex aerial mobile UAV. The action space includes the power allocation parameters of the full-duplex ground base station and the flight control parameters of the full-duplex aerial mobile UAV; The reward function includes a communication rate gain term and a constraint violation penalty term.

6. The full-duplex UAV-assisted sensor-integrated covert communication optimization method according to claim 5, characterized in that, The soft actor-critic algorithm includes a double-Q network structure, reparameterization techniques, and a target network soft update mechanism.

7. The full-duplex UAV-assisted sensor-integrated covert communication optimization method according to claim 1, characterized in that, The optimal control strategy for deployment execution includes: Within each control time slot, the optimal control strategy is input based on the current system state, and the full-duplex ground base station power allocation parameters and the full-duplex aerial mobile UAV flight speed vector parameters are output. Control commands are sent to full-duplex ground base stations and full-duplex aerial mobile drones via the control link; the full-duplex ground base stations adjust the power distribution of communication signals and sensing signals according to the commands. Full-duplex aerial mobile drones perform spatial maneuvers and emit artificial noise interference signals according to instructions.

8. A full-duplex UAV-assisted sensor-integrated covert communication optimization system, the system being used to execute the full-duplex UAV-assisted sensor-integrated covert communication optimization method according to any one of claims 1-7, characterized in that, The communication optimization system includes: The channel modeling module establishes air-to-ground and ground-to-ground channel models respectively, derives the channel characteristic parameters of legitimate communication links, sensing links, and eavesdropping links, and calculates the communication signal-to-interference-plus-noise ratio (SIR) of the full-duplex airborne mobile UAV receiver, the sensing SIR of the full-duplex ground base station receiver, and the target covert communication rate. The constraint construction module constructs a binary hypothesis detection model based on the signal detection mechanism of ground eavesdroppers. It also solves the worst-case channel gain extremum through geometric relationships to address the uncertainty of the ground eavesdropper's location, thus constructing a robust constraint boundary to ensure the system's concealment. The optimization problem construction module aims to maximize the average covert communication rate. It integrates the three-dimensional flight maneuver constraints of full-duplex aerial mobile UAVs, the transmit power constraints of full-duplex ground base stations, the performance constraints of aerial radar sensing targets, and the robust covert constraint boundary to construct a multi-constraint coupled optimization problem. The intelligent solution module defines the state space and action space, designs a reward function that includes a penalty term for constraint violation, transforms the multi-constraint optimization problem into a Markov decision process, and uses a deep reinforcement learning algorithm to train and obtain the optimal control strategy. The strategy execution module deploys the trained optimal control strategy to the communication system, and generates and executes full-duplex ground base station power allocation commands, full-duplex aerial mobile UAV flight trajectory commands, and jamming signal transmission commands in real time to achieve closed-loop control.

9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the full-duplex UAV-assisted sensor-integrated covert communication optimization method as described in any one of claims 1-7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the full-duplex UAV-assisted sensor-integrated covert communication optimization method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method, device and electronic device for determining UAV trajectory and beamforming vector

    CN117880817B

  • Physical layer secure transmission method and system in communication and sensing integrated unmanned aerial vehicle network

    CN118042454A

  • Hidden transmission strategy of NOMA-RIS-assisted communication and inductance integrated system

    CN118432672A

  • Trajectory and power design method and system for UAV covert relay communication network

    CN118869037B