DRL-driven low-altitude unmanned aerial vehicle high-energy-efficiency sensing integrated optimization method
By optimizing beamforming and flight trajectory of a low-altitude unmanned aerial vehicle (UAV) system using a DRL-driven dual-agent algorithm, the energy consumption and dynamic adaptability issues in the integration of sensing and communication in low-altitude UAVs were resolved. This achieved high-efficiency optimization of the integration of sensing and communication, improving the system's endurance and communication quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-03
AI Technical Summary
Existing low-altitude UAV integrated sensing technology has failed to effectively optimize energy consumption, resulting in high system operating costs and weak endurance. Furthermore, traditional methods have poor adaptability and real-time performance in dynamic environments, failing to meet the high energy efficiency requirements of multiple UAV missions.
A DRL-driven dual-agent deep deterministic policy gradient algorithm is adopted. Through the collaborative work of beam optimization agent and trajectory optimization agent, combined with intelligent reflector technology, the beamforming and flight trajectory of the base station and UAV are optimized. The reward functions for energy-efficient communication and rate are designed to maximize the system's energy efficiency.
It effectively optimizes energy consumption in dynamic environments, improves system endurance, adapts to multiple UAV mission scenarios, enhances communication link quality and sensing accuracy, and meets the demand for high-efficiency integrated sensing under multiple constraints.
Smart Images

Figure CN121789515A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-altitude unmanned aerial vehicle (UAV) technology, and in particular to a high-efficiency integrated sensor optimization method for low-altitude UAVs driven by DRL. Background Technology
[0002] With the rapid development of the Low Altitude Economy (LAE), unmanned aerial vehicles (UAVs) are increasingly being used in logistics, environmental monitoring, and agricultural plant protection, and scenarios where multiple UAVs perform dedicated tasks in parallel have become commonplace. To ensure safe and efficient low-altitude operations, it is necessary to use Sensor-Integrated Communication and Navigation (ISAC) technology to enable seamless communication and navigation services between base stations and authorized UAVs, while simultaneously conducting real-time sensing and monitoring of low-altitude targets (such as unauthorized aircraft).
[0003] Intelligent Reflectors (RIS), a key 6G technology, can flexibly optimize the wireless channel propagation environment by adjusting the phase offset of passive reflective units, significantly improving communication link quality and sensing accuracy, and providing a new path for enhancing the performance of integrated sensing and communication systems. However, in low-altitude scenarios, the transmission energy consumption and circuit energy consumption of base stations, as well as the propulsion energy consumption of UAVs (accounting for more than 90% of the total energy consumption of UAVs), constitute the main energy consumption overhead of the system. Traditional integrated sensing and communication designs often focus on maximizing communication rate and sensing performance, neglecting energy consumption constraints, resulting in high system operating costs and weak endurance, which restricts the large-scale application of LAEs.
[0004] Existing optimization methods (such as iterative algorithms and convex optimization) largely rely on accurate Channel State Information (CSI). However, the movement of UAVs and the dynamic nature of targets at low altitudes cause CSI to vary over time and be difficult to obtain accurately. Traditional methods suffer from poor adaptability and insufficient real-time performance. Deep Reinforcement Learning (DRL) has the advantage of iteratively optimizing long-term goals through trial and error in unknown dynamic environments and has shown potential in wireless communication resource allocation. However, it has not yet formed a RIS-assisted optimization scheme for energy-efficient communication and rate maximization, making it difficult to meet the high-efficiency sensing integration requirements under the constraints of multiple UAV tasks.
[0005] Existing traditional low-altitude UAV sensing and communication integration technologies do not consider energy consumption and adaptability to dynamic scenarios. Most solutions focus only on maximizing communication rate or sensing accuracy, failing to incorporate base station transmission circuit energy consumption and UAV propulsion energy consumption into joint optimization. This results in high system operating costs, weak UAV endurance, and difficulty in supporting long-term multi-mission scenarios. Furthermore, due to dynamic environmental changes, traditional convex optimization and alternating iteration methods rely on accurate channel state information (CSI) and fixed environment models, which cannot cope with the time-varying channel characteristics caused by UAV movement and dynamic changes in target trajectory at low altitudes, leading to poor robustness and real-time performance. Existing DRL-related solutions do not design specific reward functions and constraint handling mechanisms for energy-efficient sensing objectives, resulting in slow algorithm convergence and poor practical deployment.
[0006] Therefore, existing technologies need further improvement and refinement. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide a high-efficiency integrated sensing optimization method for low-altitude unmanned aerial vehicles driven by DRL.
[0008] To achieve the above objectives, the technical solution provided by this invention is as follows:
[0009] A high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL mainly includes the following steps:
[0010] Step S1: Establish the intelligent reflector-assisted UAV sensory integration system environment, and initialize the network parameters using the dual-agent deep deterministic policy gradient algorithm; the dual agents include a beam optimization agent and a trajectory optimization agent, and construct their respective Actor policy network, Critic value network and their corresponding target network, and initialize the experience replay buffer;
[0011] Step S2: Obtain the current state of the environment This includes predicted integrated channel state information from base stations to drones and targets, as well as the location status of multiple drones and detection targets;
[0012] Step S3: Based on the current state The current action is selected using a dual-agent cooperative strategy; the beam optimization agent generates beamforming actions between the base station and the intelligent reflector, the trajectory optimization agent generates flight direction actions of the UAV, and exploration noise is added to the actions to enhance robustness.
[0013] Step S4: Perform the action and calculate the reward function that includes system energy efficiency, collision avoidance constraints, and perception accuracy constraints. And observe the next state based on environmental feedback. ;
[0014] Step S5: Transform the tuples Store the data in the experience replay buffer; randomly sample mini-batch data from the buffer, update the Critic network parameters of the two agents using truncated double-Q learning and target policy smoothing mechanism, and delay updating the Actor network parameters and target network parameters.
[0015] Step S6: Change the new state Set to the current state, repeat steps S3 to S6 until the preset training termination condition is met or the network reaches convergence, and output the optimal beamforming and trajectory planning strategy.
[0016] As a preferred embodiment of the present invention, the network structure initialized in step S1 specifically involves: initializing an Actor network for the beam optimization agent and the trajectory optimization agent, respectively. , and two Critic networks , and , Simultaneously initialize the corresponding target network. , and , and , ;in and These are network parameters.
[0017] As a preferred embodiment of the present invention, the current environmental state in step S2 is specifically as follows:
[0018] The current environmental state is the predicted integrated channel state information for each time slot from the base station to all drone users and detection targets: And the location status of multiple drones and detection targets: The predicted integrated channel state information includes communication links between the base station and UAV users, communication links between the base station and the smart reflector, sensing links between the smart reflector and the sensing target, and communication links between the smart reflector and the UAV users. Therefore, time slots... The current state of the environment can be defined as:
[0019] .
[0020] As a preferred embodiment of the present invention, the action selection in step S3 specifically refers to: beam-optimized intelligent agent outputting an action. This includes base station communication precoding matrix, radar precoding matrix, and intelligent reflector phase shift matrix; the trajectory optimization agent outputs actions. Including all drone flight directions; actual actions performed. Gaussian noise is added to the model's output action. :
[0021] Beam-optimized agent output action :
[0022]
[0023] Trajectory optimization agent output action :
[0024]
[0025] Actual actions :
[0026]
[0027] In the formula The mean is 0 and the variance is Gaussian noise.
[0028] As a preferred embodiment of the present invention, collision avoidance constraints, perception accuracy constraints, system energy efficiency communication and rate, and reward function are included. The definition is as follows:
[0029] To avoid collisions between different drones, the following collision avoidance constraints must be implemented:
[0030]
[0031] To ensure target detection, a signal-to-noise ratio threshold needs to be detected. Require:
[0032]
[0033] System energy efficiency communication and rate are defined as follows:
[0034]
[0035] Under the above constraints, the reward function of the two agents is... Defined as:
[0036] (8)
[0037] In the above formula, when only considering the optimization objective of maximizing the total speed of all drones, the reward function is: When collision avoidance constraints cannot be met, the reward function is a penalty. ,in This is the penalty coefficient corresponding to the collision avoidance constraint, used to balance utility and cost; when the agent needs to meet the preset perception signal-to-noise ratio requirement, the average perception signal-to-noise ratio throughout the flight cycle is less than the minimum threshold. When the reward function is a penalty term .
[0038] Specifically, in order to guide the agent to maximize communication performance while strictly satisfying safety and perception constraints, this invention designs a piecewise reward function that includes positive utility and violation penalty. More specifically, the agent will achieve a communication and speed equal to the system's energy efficiency if and only if both the collision avoidance constraint (Condition 5) and the average sensing signal-to-noise ratio constraint (Condition 6) are satisfied simultaneously. Positive rewards are given when the agent makes an unwise decision that violates the collision avoidance constraint (i.e., condition 5 is not met). In such cases, the positive energy efficiency communication and rate rewards are cancelled and replaced with a fixed penalty. ,in This is a penalty coefficient corresponding to the collision avoidance constraint, forcing it to prioritize flight safety; furthermore, when the average perceived signal-to-noise ratio throughout the flight cycle is below a preset threshold... (That is, condition 6 is not met) Regardless of whether a collision occurs, an additional penalty proportional to the difference in signal-to-noise ratio will be added to the reward function. This reward mechanism effectively incentivizes agents to optimize their strategies, maximizing the system's overall energy efficiency and communication rate while ensuring that collision and perception constraints are met.
[0039] In the formula, condition (5) corresponds to the collision avoidance constraint, and condition (6) corresponds to the preset sensing signal-to-noise ratio requirement. , These correspond to the penalty coefficients for the two constraints, respectively. Energy efficiency, communication, and speed are the objective functions that need to be maximized.
[0040] As a preferred embodiment of the present invention, the update formula of the Critic network in step S5 is: calculate the target action. Introducing smoothing noise:
[0041]
[0042] Calculate the target Q value Take the minimum value from the outputs of the two target Critic networks:
[0043] (10)
[0044] By minimizing the mean square error loss function Update the Critic network:
[0045] (11)
[0046] In the formula For sampling batch size, This is the discount factor.
[0047] As a preferred embodiment of the present invention, the delayed update strategy of the Actor network in step S5 is: only update the Critic network. After that, update the Actor network parameters once. The goal is to maximize the Q value:
[0048] (12)
[0049] Simultaneously, soft updates are performed on the target network parameters:
[0050] (13)
[0051] In the formula This is the soft update coefficient.
[0052] Compared to existing technologies, the DRL-driven high-efficiency integrated sensing optimization method for low-altitude UAVs provided in this invention no longer relies on accurate Channel State Information (CSI). It also designs a dual-agent scheme to optimize the beamforming of multiple UAV trajectories, base stations, and intelligent reflectors, respectively. This avoids poor algorithm convergence due to high input correlation. Furthermore, by incorporating energy efficiency, it designs energy-efficient communication and data rates, effectively avoiding high energy consumption caused by focusing solely on communication and data rates. This method is more suitable for scenarios involving multiple UAVs performing tasks. Combined with intelligent reflector technology, it improves communication links, ensuring sustainable communication quality and better addressing the issue of poor channel gain in line-of-sight paths in urban scenarios.
[0053] Compared with existing technologies, the principles and advantages of this technical solution are as follows:
[0054] 1. The DRL-driven high-efficiency integrated sensing optimization method for low-altitude UAVs provided in this invention incorporates energy consumption into the model system, effectively avoiding energy waste, better planning multi-UAV trajectories, and efficiently executing tasks while saving energy. Furthermore, to address channel variations in dynamic scenarios, a deep reinforcement learning (DRL) approach is employed, designing constraints as corresponding reward functions. This effectively transforms the traditional optimization problem into a reinforcement learning task for training the network, ultimately enabling effective handling of time-varying scenarios.
[0055] 2. The DRL-driven high-efficiency integrated sensing optimization method for low-altitude UAVs provided by this invention, under the constraints of sensing accuracy, maximum base station transmit power, intelligent reflector amplitude and phase, and hybrid UAV start and end positions, provides an optimization scheme that maximizes the maximum energy-efficient communication and rate of legally authorized UAVs for the joint optimization of energy-efficient communication and rate of legal users, hybrid UAV trajectory, and waveforms of base stations and intelligent reflectors.
[0056] 3. The DRL-driven high-efficiency integrated sensing optimization method for low-altitude UAVs provided in this invention optimizes the communication link using a smart reflector, with the core objective of maximizing energy-efficient communication and data rate. Through DRL technology, it achieves joint optimization of base station beamforming, UAV trajectory planning, and smart reflector phase offset, without relying on accurate channel state information (CSI). This better adapts to dynamically changing environments while strictly meeting multiple constraints, including base station power limitations, UAV collision avoidance, sensing accuracy requirements, smart reflector amplitude and phase constraints, and fixed UAV take-off and end positions. Attached Figure Description
[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the services required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a flowchart illustrating the high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL provided by the present invention.
[0059] Figure 2 This is a schematic diagram of the link of the DRL-driven high-efficiency integrated sensing system (communication, sensing, trajectory) for low-altitude unmanned aerial vehicles provided by the present invention.
[0060] Figure 3 This is a schematic diagram of the TD3 algorithm framework based on dual agents provided by the present invention. Detailed Implementation
[0061] The present invention will be further described below with reference to specific embodiments:
[0062] Example 1
[0063] like Figures 1 to 3 As shown, this embodiment provides a high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL, which mainly includes the following steps:
[0064] Step S1: Establish the intelligent reflector-assisted UAV sensory integration system environment, and initialize the network parameters using the dual-agent deep deterministic policy gradient algorithm; the dual agents include a beam optimization agent and a trajectory optimization agent, and construct their respective Actor policy network, Critic value network and their corresponding target network, and initialize the experience replay buffer;
[0065] Step S2: Obtain the current state of the environment This includes predicted integrated channel state information from base stations to drones and targets, as well as the location status of multiple drones and detection targets;
[0066] Step S3: Based on the current state The current action is selected using a dual-agent cooperative strategy; the beam optimization agent generates beamforming actions between the base station and the intelligent reflector, the trajectory optimization agent generates flight direction actions of the UAV, and exploration noise is added to the actions to enhance robustness.
[0067] Step S4: Perform the action and calculate the reward function that includes system energy efficiency, collision avoidance constraints, and perception accuracy constraints. And observe the next state based on environmental feedback. ;
[0068] Step S5: Transform the tuples Store the data in the experience replay buffer; randomly sample mini-batch data from the buffer, update the Critic network parameters of the two agents using truncated double-Q learning and target policy smoothing mechanism, and delay updating the Actor network parameters and target network parameters.
[0069] Step S6: Change the new state Set to the current state, repeat steps S3 to S6 until the preset training termination condition is met or the network reaches convergence, and output the optimal beamforming and trajectory planning strategy.
[0070] As a preferred embodiment of the present invention, the network structure initialized in step S1 specifically involves: initializing an Actor network for the beam optimization agent and the trajectory optimization agent, respectively. , and two Critic networks , and , Simultaneously initialize the corresponding target network. , and , and , ;in and These are network parameters.
[0071] As a preferred embodiment of the present invention, the current environmental state in step S2 is specifically as follows:
[0072] The current environmental state is the predicted integrated channel state information for each time slot from the base station to all drone users and detection targets: And the location status of multiple drones and detection targets: The predicted integrated channel state information includes communication links between the base station and UAV users, communication links between the base station and the smart reflector, sensing links between the smart reflector and the sensing target, and communication links between the smart reflector and the UAV users. Therefore, time slots... The current state of the environment can be defined as:
[0073] .
[0074] As a preferred embodiment of the present invention, the action selection in step S3 specifically refers to: beam-optimized intelligent agent outputting an action. This includes base station communication precoding matrix, radar precoding matrix, and intelligent reflector phase shift matrix; the trajectory optimization agent outputs actions. Including all drone flight directions; actual actions performed. Gaussian noise is added to the model's output action. :
[0075] Beam-optimized agent output action :
[0076]
[0077] Trajectory optimization agent output action :
[0078]
[0079] Actual actions :
[0080]
[0081] In the formula The mean is 0 and the variance is Gaussian noise.
[0082] As a preferred embodiment of the present invention, collision avoidance constraints, perception accuracy constraints, system energy efficiency communication and rate, and reward function are included. The definition is as follows:
[0083] To avoid collisions between different drones, the following collision avoidance constraints must be implemented:
[0084]
[0085] To ensure target detection, a signal-to-noise ratio threshold needs to be detected. Require:
[0086]
[0087] System energy efficiency communication and rate are defined as follows:
[0088]
[0089] Under the above constraints, the reward function of the two agents is... Defined as:
[0090] (8)
[0091] In the above formula, when only considering the optimization objective of maximizing the total speed of all drones, the reward function is: When collision avoidance constraints cannot be met, the reward function is a penalty. ,in This is the penalty coefficient corresponding to the collision avoidance constraint, used to balance utility and cost; when the agent needs to meet the preset perception signal-to-noise ratio requirement, the average perception signal-to-noise ratio throughout the flight cycle is less than the minimum threshold. When the reward function is a penalty term .
[0092] Specifically, in order to guide the agent to maximize communication performance while strictly satisfying safety and perception constraints, this invention designs a piecewise reward function that includes positive utility and violation penalty. More specifically, the agent will achieve a communication and speed equal to the system's energy efficiency if and only if both the collision avoidance constraint (Condition 5) and the average sensing signal-to-noise ratio constraint (Condition 6) are satisfied simultaneously. Positive rewards are given when the agent makes an unwise decision that violates the collision avoidance constraint (i.e., condition 5 is not met). In such cases, the positive energy efficiency communication and rate rewards are cancelled and replaced with a fixed penalty. ,in This is a penalty coefficient corresponding to the collision avoidance constraint, forcing it to prioritize flight safety; furthermore, when the average perceived signal-to-noise ratio throughout the flight cycle is below a preset threshold... (That is, condition 6 is not met) Regardless of whether a collision occurs, an additional penalty proportional to the difference in signal-to-noise ratio will be added to the reward function. This reward mechanism effectively incentivizes agents to optimize their strategies, maximizing the system's overall energy efficiency and communication rate while ensuring that collision and perception constraints are met.
[0093] In the formula, condition (5) corresponds to the collision avoidance constraint, and condition (6) corresponds to the preset sensing signal-to-noise ratio requirement. , These correspond to the penalty coefficients for the two constraints, respectively. Energy efficiency, communication, and speed are the objective functions that need to be maximized.
[0094] As a preferred embodiment of the present invention, the update formula of the Critic network in step S5 is: calculate the target action. Introducing smoothing noise:
[0095]
[0096] Calculate the target Q value Take the minimum value from the outputs of the two target Critic networks:
[0097] (10)
[0098] By minimizing the mean square error loss function Update the Critic network:
[0099] (11)
[0100] In the formula For sampling batch size, This is the discount factor.
[0101] As a preferred embodiment of the present invention, the delayed update strategy of the Actor network in step S5 is: only update the Critic network. After that, update the Actor network parameters once. The goal is to maximize the Q value:
[0102] (12)
[0103] Simultaneously, soft updates are performed on the target network parameters:
[0104] (13)
[0105] In the formula This is the soft update coefficient.
[0106] Example 2
[0107] like Figures 1 to 3 As shown in the figure, this embodiment provides a high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL. The method mainly includes the following specific steps:
[0108] Step 1: Initialize the dual-agent network and experience replay buffer. At the beginning of system operation, construct the following: Figure 2The diagram shows a smart reflector-assisted UAV-ISAC system model. This invention designs two independent agents: 1. A beam optimization agent: responsible for optimizing the base station's communication / radar precoding matrix and the phase shift matrix of the smart reflector. 2. A trajectory optimization agent: responsible for optimizing the horizontal flight direction of multiple UAVs. A TD3 network structure is constructed for each agent, including: an Actor network (policy network). Used for fitting deterministic strategies. Critic Network (Value Network) , The TD3 algorithm utilizes two Critic networks to mitigate the Q-value overestimation problem. Target network: Its parameters are initialized the same as those of the main network. Experience replay buffer. Used to store historical interaction data Let its capacity be .
[0109] Step 2: Obtain the current environmental state information in each time slot. Intelligent agents observe environmental conditions The state-space design of this invention no longer uses accurate channel state information (CSI), but instead combines predicted CSI with location information:
[0110]
[0111] Step 3: Select the current action based on the dual-agent strategy and the current state. Two agents generate actions respectively: 1. Beam-optimized agent generates actions. (Corresponding to communication precoding, radar precoding, and intelligent reflector phase shift, respectively). 2. Trajectory optimization agent generates actions. (correspond (The flight angle of each drone). The TD3 agent will generate the motion direction of each drone. Specifically, sub-actions To maintain exploratory nature during training, truncated normal distribution noise is added to the output action:
[0112]
[0113] This allows the agent to try different strategies to discover more energy-efficient solutions.
[0114] Step 4: Execute actions and calculate reward function. The system executes actions. The environment transitions to the next state. The reward function designed in this invention The aim is to maximize energy efficiency (EE) while satisfying multiple constraints. Energy efficiency is defined as:
[0115]
[0116] in .
[0117] Step 5: Store experience and update network parameters (core algorithm step) This involves storing the tuples. Store in experience buffer When the buffer size meets the requirements, a batch of data is randomly sampled (Batch size). (Training)
[0118] Target policy smoothing and target value calculation: To reduce variance, noise is added when calculating the target action (TargetPolicy Smoothing).
[0119]
[0120] Calculating the target Q value using truncated double Q learning Taking the minimum of the two target Critic networks effectively suppresses the overestimation bias of the Q value:
[0121]
[0122] Updating the Critic Networks: Updating both Critic Networks by minimizing the Loss function using gradient descent.
[0123]
[0124] Delayed update of the Actor network: To ensure training stability, the Actor network is updated less frequently than the Critic network. Whenever the Critic network is updated... Update the Actor network once:
[0125]
[0126] Soft update of the target network: The parameters of the target network are slowly updated using a soft update method to maintain the stability of the training target.
[0127]
[0128]
[0129] Step Six: Iterate until convergence, repeating steps three through five. After each training episode, determine if the maximum number of training episodes has been reached. As training progresses, explore noise. The size gradually decreases. Finally, when the network converges, the trained beam-optimized agent and trajectory-optimized agent models are output for high-efficiency sensory integration mission execution by low-altitude UAVs.
[0130] The models used in both of the above embodiments are described below:
[0131] I. System Model
[0132] This invention proposes a method such as Figure 2 The illustrated intelligent reflective surface-assisted UAV-ISAC system, in which the base station is equipped with The root transceiver antenna is A single-antenna licensed UAV provides downlink communication and simultaneously detects a low-altitude target (e.g., an unlicensed UAV). The fixed-position smart reflector has... One reflective element is used to provide a non-line-of-sight link, ensuring sustainable communication quality.
[0133] (i) Trajectory model of UAV and perceived target
[0134] The drone flies at a fixed altitude to perform the mission, and we will consider the duration. Divided into Each time slot, using Indicates the duration of the time slot. It is short enough that we can assume that the positions of the target and the drone remain constant within each time slot.
[0135] Drone flight model: during a specific time period Inside, drones The trajectory is represented as a A sequence of length, i.e. ,in Composition, representing the horizontal dimension coordinates, This represents the vertical coordinate (i.e., flight altitude). Assume a drone. at a constant speed Flight, therefore the distance traveled in a single time interval is Therefore, the following constraints are derived:
[0136]
[0137] Furthermore, the initial and final positions of the multiple UAVs (i.e., flight mission constraints) are given by the following formula:
[0138]
[0139] in Indicates drone The initial position, Indicates drone The final position.
[0140] To avoid collisions between different drones, the following collision avoidance constraints must be implemented:
[0141]
[0142] 2) Target movement model: To simplify calculations, it is assumed that the target moves at a constant speed. The target moves, and its flight distance within a single time slot is... To more accurately simulate the movement characteristics of real targets, this study introduces a Gaussian-Markov process to capture the temporal correlation of the motion direction. Let... and These represent the target in the time slot. The mathematical model for the azimuth and elevation angles at time can be expressed as:
[0143]
[0144] in and It is the time correlation coefficient, which adjusts for the degree of time dependence.
[0145] Therefore, given and The target is in the time slot. The position +1 can be represented as:
[0146]
[0147] in It is a target detection Time slot horizontal coordinates It is a target detection Vertical coordinates of time slots.
[0148] 3) Unmanned Aerial Vehicle (UAV) Energy Consumption Model
[0149] At flight speed We can derive the propulsion energy consumption of a rotary-wing UAV as follows:
[0150]
[0151] constant and These represent the induced power and blade profile power in the hovering state, respectively. It is the tip speed of the rotor blades. This represents the average rotor induced velocity during hovering. Furthermore, It is the fuselage drag ratio. It's the rotor solidity, finally. It is air density. It is the rotor disk area.
[0152] (ii) Signal Model
[0153] In each time slot In this context, the base station's transmitted signal is a weighted sum of communication symbols and radar detection signals, expressed as:
[0154]
[0155] in It is a communication signal. , It is a radar detection signal. , It is a communication precoding matrix. It is the radar precoding matrix. and It is statistically independent. The power constraint is given below:
[0156]
[0157] 1) Communication model: such as Figure 2 As shown, Time-slot drones The received signal at the location consists of two parts: the line-of-sight path and the BS-RIS-UAV two-hop link, which can be represented as follows: ,in It is additive white Gaussian noise (AWGN). , , , These are the channel gain from the base station to the smart reflector, and the base station's first... Channel gain and smart reflector of each drone user to the first The channel gain and intelligent reflector passive beamforming matrix of each drone user. Therefore, drones The signal-to-interference-plus-noise ratio (SIR) at the communication point is:
[0158]
[0159] Accordingly, the sum rate of all drones can be calculated:
[0160]
[0161] 2) Perceptual models: such as Figure 2As shown, the echo signal reflected from the target also consists of two parts: one part comes from the directly reflected signal, and the other part is reflected to the integrated sensing base station by the intelligent reflective surface. Therefore, the echo signal from the target to the base station can be represented as follows: ,in , It refers to the composite channel gain from base station to target and from base station to smart reflector to target. It's AWGN. This is the noise power. The target's perceived signal-to-noise ratio is then given by the following formula:
[0162]
[0163] The goal of this invention is to maximize the performance of all drones. The expected total energy-efficient communication rate within the time slot, while simultaneously satisfying sensor signal-to-noise ratio requirements, flight mission constraints, collision avoidance, and maximum transmit power constraints. Base station beamforming, intelligent reflector beamforming, and each UAV are jointly optimized. The optimization problem can be formulated as:
[0164]
[0165] in It is a collision avoidance safety distance, a constraint. It ensures that the average perceived signal-to-noise ratio threshold is met. Requirements for perceiving targets and constraints The base station's transmission power was limited. , It is the amplitude and phase constraint of the beam of the intelligent reflector, constraint It refers to the start and end points of the drone's mission.
[0166] like Figure 3 As shown, this technical solution is decomposed into basic MDP elements, including agent, environment, state, action, and reward, as described below:
[0167] A. Base station and smart reflector beamforming
[0168] 1) Status The state of the first TD3 agent is the predicted integrated channel state information (CSI) for each time slot from the base station to all drone users and detection targets. And the location status of multiple drones and detection targets: Therefore, time slot The system state can be defined as:
[0169]
[0170] 2) Actions TD3 agents will generate , and As an action, to handle complex-valued inputs, all inputs are divided into real and imaginary parts, which are then combined and input into the network. Therefore, the agent's action can be defined as:
[0171]
[0172] 3) Rewards :
[0173]
[0174] First, considering only the optimization objective of maximizing the energy efficiency, communication, and speed of all drones, the reward is determined by... First, the agent needs to recognize its unwise decision—that is, the collision avoidance constraint cannot be met—and the corresponding reward should be a penalty. ,in This corresponds to the penalty coefficient for collision avoidance constraints, balancing utility and cost; third, the agent needs to meet a preset perception signal-to-noise ratio requirement, if the average perception signal-to-noise ratio over the entire flight cycle is less than a minimum threshold. Then a penalty item should be added. In this way, the agent can be encouraged to optimize its strategy in order to meet the average signal-to-noise ratio requirement of the perceived target in subsequent rounds.
[0175] B. Drone trajectory optimization
[0176] 1) Status The state of the first TD3 agent is the predicted integrated channel state information (CSI) for each time slot from the base station to all drone users and detection targets. And the location status of multiple drones and detection targets: Therefore, time slot The system state can be defined as:
[0177]
[0178] 2) Actions The TD3 agent will generate the movement directions of each drone. Specifically, sub-actions for drones The horizontal position of the next time slot becomes Each action element is a continuous variable optimized by the agent. Therefore, the agent's action can be defined as:
[0179]
[0180] 3) Rewards The reward for this agent is the same as the reward function defined for the previous agent, because the two agents have the same constraints and the same optimization objective, namely, maximizing... .
[0181]
[0182] The above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Therefore, any changes made in accordance with the shape and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A high-efficiency integrated sensing optimization method for low-altitude unmanned aerial vehicles driven by DRL, characterized in that, Includes the following steps: Step S1: Establish the intelligent reflector-assisted UAV sensory integration system environment, and initialize the network parameters using the dual-agent deep deterministic policy gradient algorithm; the dual agents include a beam optimization agent and a trajectory optimization agent, and construct their respective Actor policy network, Critic value network and their corresponding target network, and initialize the experience replay buffer; Step S2: Obtain the current state of the environment This includes predicted integrated channel state information from base stations to drones and targets, as well as the position status of multiple drones and detection targets; Step S3: Based on the current state The current action is selected using a dual-agent cooperative strategy; the beam optimization agent generates beamforming actions between the base station and the intelligent reflector, the trajectory optimization agent generates flight direction actions of the UAV, and exploration noise is added to the actions to enhance robustness. Step S4: Perform the above actions and calculate the reward function that includes system energy efficiency, collision avoidance constraints, and perception accuracy constraints. And observe the next state based on environmental feedback. ; Step S5: Transform the tuples Store the data in the experience replay buffer; randomly sample mini-batch data from the buffer, update the Critic network parameters of the two agents using truncated double-Q learning and target policy smoothing mechanism, and delay updating the Actor network parameters and target network parameters. Step S6: Change the new state Set to the current state, repeat steps S3 to S6 until the preset training termination condition is met or the network reaches convergence, and output the optimal beamforming and trajectory planning strategy.
2. The high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL according to claim 1, characterized in that, The network structure initialized in step S1 specifically involves: initializing the Actor network for the beam optimization agent and the trajectory optimization agent, respectively. , and two Critic networks , and , Simultaneously initialize the corresponding target network. , and , and , ;in and These are network parameters.
3. The high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL according to claim 1, characterized in that, The current state of the environment in step S2 is specifically as follows: The current environmental state is the predicted integrated channel state information for each time slot from the base station to all drone users and detection targets: And the location status of multiple drones and detection targets: The predicted integrated channel state information includes communication links between the base station and UAV users, communication links between the base station and the smart reflector, sensing links between the smart reflector and the sensing target, and communication links between the smart reflector and the UAV users. Therefore, time slots... The current state of the environment can be defined as: 。 4. The high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL according to claim 1, characterized in that, The action selection in step S3 specifically refers to: beam optimization agent outputting an action. This includes base station communication precoding matrix, radar precoding matrix, and intelligent reflector phase shift matrix; the trajectory optimization agent outputs actions. Including all drone flight directions; actual actions performed. Gaussian noise is added to the model's output action. : Beam-optimized agent output action : ; Trajectory optimization agent output action : ; Actual actions : ; In the formula The mean is 0 and the variance is Gaussian noise.
5. The high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL according to claim 1, characterized in that, The collision avoidance constraints, perception accuracy constraints, system energy efficiency communication and rate, and reward function in step S4 are mentioned above. The definition is as follows: To avoid collisions between different drones, the following collision avoidance constraints must be implemented: ; To ensure target detection, a signal-to-noise ratio threshold needs to be detected. Require: ; System energy efficiency communication and rate are defined as follows: ; Under the above constraints, the reward function of the two agents is... Defined as: (8); In the above formula, when only considering the optimization objective of maximizing the total speed of all drones, the reward function is: When collision avoidance constraints cannot be met, the reward function is a penalty. ,in This is the penalty coefficient corresponding to the collision avoidance constraint, used to balance utility and cost; when the agent needs to meet the preset perception signal-to-noise ratio requirement, the average perception signal-to-noise ratio throughout the flight cycle is less than the minimum threshold. When the reward function is a penalty term .
6. The high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL according to claim 1, characterized in that, The update formula for the Critic network in step S5 is: calculate the target action. Introducing smoothing noise: ; Calculate the target Q value Take the minimum value from the outputs of the two target Critic networks: (10); By minimizing the mean square error loss function Update the Critic network: (11); In the formula For sampling batch size, This is the discount factor.
7. The high-efficiency integrated sensing optimization method for low-altitude UAVs driven by DRL according to claim 1, characterized in that, The delayed update strategy for the Actor network in step S5 is: update only the Critic network. After that, update the Actor network parameters once. The goal is to maximize the Q value: (12); Simultaneously, soft updates are performed on the target network parameters: (13); In the formula This is the soft update coefficient.