MATD3-Based Multi-UAV Integrated Sensing and Communication Trajectory and Beamforming Optimization Method
By introducing a multi-agent reinforcement learning method with centralized training and distributed execution in multi-unmanned airport scenarios, combining K-means clustering and MATD3 algorithms, the drone trajectory and beamforming are optimized, and the multi-agent collaboration problem under dynamic user roaming is solved, and the performance of the communication and perception integrated system is improved.
Patent Information
- Application Number
- CN202510653967.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the multi-unmanned airport scenario, the existing drone-assisted communication and perception integrated system is difficult to deal with the multi-agent collaboration problems in dynamic user roaming and complex environments, resulting in limited system performance.
Multi-agent reinforcement learning (MARL) method based on centralized training and distributed execution is adopted, combined with K-means clustering and MATD3 algorithm, optimize drone trajectory and beamforming, consider user roaming and three-dimensional deployment, and formulate optimization problems to maximize communication speed and perceptual performance.
Under the constraints of transmission power and perceived beam pattern gain, the drone trajectory and beamforming are optimized, the system's communication rate and perceived performance are improved, and the system's efficiency and stability are improved.
Smart Images

Figure CN120186624B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communication technologies, and particularly relates to a method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3. Background Art
[0002] With the rapid development of the sixth-generation mobile communication, communication and sensing integration (ISAC) has gradually become a research hotspot. It integrates communication and sensing functions into the same system and shares hardware resources. This integration not only improves the efficiency of the system but also enables the communication and sensing functions to promote each other.
[0003] Applying unmanned aerial vehicles (UAVs) to the ISAC system has become an increasingly prominent research field because it has great potential to enhance communication and sensing capabilities. Compared with ground channels, UAVs provide important line-of-sight (LoS) links and flexible support for the deployment of the dual-functional ISAC system. Specifically, in terms of communication, the transmission of communication signals becomes more stable and can provide a relatively high data transmission rate. In terms of sensing, the high-altitude position of UAVs enables them to obtain better sensing performance. For example, radar signals are not blocked or scattered by ground obstacles, thus achieving high-precision and stable sensing.
[0004] Although in the past few years, UAV-assisted ISAC systems have been widely explored in terms of UAV deployment and resource allocation, research in dynamic user roaming and complex multi-UAV scenarios is still lacking. In multi-UAV scenarios, single-agent reinforcement learning methods are difficult to handle the cooperation problems among multiple agents. In previous UAV-assisted ISAC research, traditional methods such as convex optimization performed well in the limited static user environment of UAVs and users. However, when the environment changes dynamically due to user roaming or the system becomes complex with the increase of targets, the problem becomes an NP-hard problem, and it is infeasible to solve it with these methods.
[0005] Original reinforcement learning methods, such as Q-learning, provide solutions for tasks with a limited state-action space based on tables. And some advanced deep reinforcement learning (DRL) algorithms, such as deep Q-network (DQN), proximal policy optimization (PPO), soft actor-critic (SAC), deep deterministic policy gradient (DDPG), etc. have also been used in various tasks. In these scenarios, each UAV only needs to calculate and execute operations related to its local observation. In multi-UAV scenarios, single-agent reinforcement learning methods are difficult to handle the cooperation problems among multiple agents. Although existing MARL algorithms show superiority in agent cooperation, the local exploration of each agent has not been fully studied, and this deficiency limits the performance of UAV-assisted ISAC systems. Summary of the Invention
[0006] Objective of the Invention: The objective of the present invention is to provide a method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3. Compared with the existing UAV ISAC systems that assume static users or two-dimensional UAV trajectories, multi-agent reinforcement learning (MARL) based on centralized training and distributed execution (CTDE) is used in the multi-UAV scenario, considering a more realistic scenario of roaming users and three-dimensional deployment of UAVs. An optimization problem of joint flight trajectory and beamforming is formulated to maximize the long-term communication sum rate under the constraints of transmit power and ensuring the beam pattern gain of the sensing target.
[0007] Technical Solution: A method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3 of the present invention includes the following steps:
[0008] Step 1: Construct a communication and sensing integrated system supported by multi-UAVs, where the UAVs act as dual-functional antenna base stations to perform downlink communication with ground users and sense ground users simultaneously.
[0009] Step 2: In terms of communication, calculate the communication link gain and signal-to-interference-plus-noise ratio (SINR), and then derive the communication sum rate therefrom; in terms of sensing, calculate the beam pattern gain.
[0010] Step 3: Under the condition of ensuring sensing performance, establish an optimization problem with the objective of maximizing the total data rate of all communication users within the task time period.
[0011] Step 4: Use a hierarchical user association algorithm based on K-means to associate communication users, sensing users with UAVs.
[0012] Step 5: Use the MATD3 algorithm with centralized training and distributed execution (CTDE) to train the UAVs to complete the optimization of the trajectory and beamforming of multi-UAV communication and sensing integration.
[0013] Further, Step 1 is specifically: Construct a communication and sensing integrated system supported by multi-UAVs, including several UAVs acting as dual-functional antenna base stations, performing downlink communication with ground users, and sensing ground users simultaneously; assume that the UAVs and users are respectively equipped with linear uniformly spaced array antennas and a single receive antenna, and the users randomly roam in the service area; the positions of UAV , communication user , and sensing user at time slot are respectively represented as , , , wheret Indicates a time slot, .
[0014] Furthermore, in step 2, in terms of communication, the communication link gain and signal-to-interference-plus-noise ratio (SINR) are calculated, and the communication sum rate is derived therefrom as follows:
[0015] In terms of the communication model, a free space path loss model is adopted, and from the unmanned aerial vehicle (UAV) to the user the channel gain is expressed as:
[0016] ;
[0017] wherein, is the channel gain at a reference distance of 1 m, and are the transmit antenna gain and receive antenna gain of the communication user, respectively; is the array response vector towards the user and is expressed as:
[0018] ;
[0019] wherein, and are the antenna space and wavelength, respectively;
[0020] ;
[0021] wherein, is the pitch angle between the UAV and the communication user;
[0022] The UAVs are assigned to communication and sensing users, and each user can only be matched with one of the UAVs; Let indicate the association indicator between the UAV and the target user in the time slot . When it indicates that the user is associated with the UAV in the time slot ; Therefore, the system has a user association constraint:
[0023] ;
[0024] wherein, represents the user, ;
[0025] The UAV The transmitted signal is given by:
[0026] ;
[0027] in and Respectively indicate that Symbols of information signals and drones User The beamforming vector of
[0028] user The received signal is represented as:
[0029] ;
[0030] in, represents additive white Gaussian noise, represents the noise power, Indicates drone To the user associated with it The channel gain, Indicates other drones and users The channel gain, and Respectively represent drones Information signals sent to other communication users and drones The beamforming vectors of other communication users;
[0031] From the above formula, we can conclude that the communication user Not only will the user be interfered with by the same drone service area, but also by the user in other drone service areas. The received signal-to-interference-and-noise ratio SINR is:
[0032] ;
[0033] use represents bandwidth, and the achievable sum rate is expressed as:
[0034]
[0035] Furthermore, in step 2, in terms of perception, the beam pattern gain is calculated as follows:
[0036] For sensing signals, drones and target sensing users The bidirectional channel power gain between is given by:
[0037] ;
[0038] Among them, is the channel power at a reference distance of 1 m, and are the transmit antenna gain and receive antenna gain of the sensing user, respectively; represents the average value of the radar cross section of the target user, is the wavelength of the carrier; is the array response vector towards the user and is expressed as:
[0039] ;
[0040] ;
[0041] Among them, the unmanned aerial vehicle to its associated sensing user the pitch angle;
[0042] In this ISAC system, the communication signal is used for user sensing, and the transmit beam pattern gain of the unmanned aerial vehicle to its associated sensing user is used as a criterion to measure the sensing performance, that is, the power of the sensing signal to the target direction, and is expressed as:
[0043] ;
[0044] Among them, is the array response vector towards the user represents the beamforming vector of the user at the position of the unmanned aerial vehicle represents the association indicator between the unmanned aerial vehicle and the target user in the time slot
[0045] To meet the requirements of sensing performance, considering the minimum threshold and the echo distance for the beam pattern gain of different sensing users, the following requirements for sensing performance are as follows:
[0046]
[0047] Furthermore, step 3 is specifically as follows:
[0048] In each time slot , the ISAC signal is simultaneously sent to communication users and sensing users. For any UAV , its position satisfies the following constraint conditions:
[0049] ;
[0050] ;
[0051] ;
[0052] where and are the lower and upper bounds of the UAV's flight trajectory respectively, and the maximum flight speed of the UAV in each time slot is , and the distance between UAVs is greater than the safety distance ;
[0053] By jointly optimizing the UAV's trajectory , beamformer and user association indicator , under the constraints of beam pattern gain, maximum transmit power and UAV flight, maximize the communication sum rate. Therefore, the optimization problem is formulated as:
[0054] ;
[0055] ;
[0056] ;
[0057] ;
[0058] ;
[0059] ;
[0060] ;
[0061] ;
[0062] ;
[0063] Among them, is the communication rate of the time-slot user, and the optimization objective (17) is expressed as the sum of the communication rates provided by all UAVs for communication users; (18) is the user association constraint; (19) means that the total power allocated to the user does not exceed the maximum power of each UAV, is the maximum power of the UAV; (20) means to ensure that the sensing requirement of each sensing target in the time slot meets the minimum threshold , represents the UAV, , represents the sensing user, ; (21)-(25) represent the maximum flight distance constraint, anti-collision constraint and flight range constraint of the UAV.
[0064] Furthermore, step 4 is specifically as follows: Use the random walk model to simulate the change of the user's position. At any discrete time slot , the user moves randomly along the axis with respect to the normal distribution and , where is the maximum speed of the user's movement; Use the hierarchical user association algorithm based on K-means to associate the communication users, sensing users with the UAVs; The communication users are first clustered by the K-means algorithm according to the spatial distance, and each communication user cluster is assigned to the UAV closest to it; Then the UAV adjusts its movement trajectory and beamformer to provide the maximum communication rate for this user cluster; Based on the existing clusters, the sensing user association is performed, the beam pattern gain between the sensing user and the UAV is calculated respectively, and each sensing user is assigned to the UAV that provides the maximum beam pattern gain by comparison.
[0065] Furthermore, in step 5, the MATD3 algorithm specifically includes the following steps:
[0066] Step 5.1, Initialization: Initialize two Critic networks and one Actor network for each agent, and initialize the corresponding target networks;
[0067] Step 5.2, Interaction and experience storage: Each agent interacts with the environment and records the current state, action, reward and the next state;
[0068] Step 5.3, Update the Critic network: Sample a batch of data from the experience replay pool;
[0069] Step 5.4, Calculate the target value and update the two Critic networks;
[0070] Step 5.5. Delay the update of the Actor network: Update the parameters of the Actor network every several steps to maximize the value of its corresponding Critic network. value;
[0071] Step 5.6. Soft-update the target network: Update the parameters of the target Critic network and the target Actor network.
[0072] Step 5.7. Repeat Steps 5.2 - 5.6 until the agent learns to optimize the policy in the environment.
[0073] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method of the present invention.
[0074] The present invention also discloses a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.
[0075] The present invention also discloses a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.
[0076] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:
[0077] The present invention takes into account more practical user roaming scenarios and three-dimensional deployments of drones, and adopts a two-step method for the proposed optimization problem, which can decompose the complex optimization problem into two separate problems. First, cluster the users to determine the users served by the drones before providing communication and sensing services, and then solve the joint optimization of the flight trajectory and beamforming vector by using the MATD3 algorithm, which introduces a delayed update and a target policy smoothing mechanism on the traditional MADDPG algorithm to avoid instability caused by too fast policy updates and can perform better in a mixed collaborative and competitive environment.
[0078] Under the constraints of transmit power and ensuring the gain of the sensing target beam pattern, this invention formulates an optimization problem for joint flight trajectory and beamforming to maximize the long-term communication sum rate. To address the challenges brought by dynamic and high-dimensional characteristics, this optimization problem is formulated as a partially observable Markov decision process (POMDP) problem, and a two-step method for dynamic scenarios is proposed: (1) A hierarchical user association algorithm based on K-means is proposed to update user association regularly; (2) The multi-agent twin-delayed deep deterministic policy gradient (MATD3) method is used to jointly optimize the flight trajectory and beamforming vector. Reinforcement learning learns the optimal policy through the interaction between the agent and the environment. In the optimization of UAV trajectory and beamforming, reinforcement learning can dynamically adjust the behavior of the UAV to maximize the long-term benefit.
[0079] In a multi-UAV scenario, single-agent reinforcement learning methods are difficult to handle the cooperation problems between multiple agents. Therefore, this invention further introduces the multi-agent reinforcement learning (MARL) method. By treating each UAV as an independent agent, the computational burden is decentralized, and each UAV is allowed to execute actions based on local observations, thereby improving the efficiency and performance of the system. Description of the Drawings
[0080] Figure 1 Model of the UAV trajectory and beamforming optimization method for the communication and sensing integrated system based on MATD3.
[0081] Figure 2 Structural diagram of the centralized MATD3 algorithm.
[0082] Figure 3 Convergence performance of the MATD3 algorithm under different learning rates.
[0083] Figure 4 Performance comparison diagram of the MATD3 algorithm with other different algorithms.
[0084] Figure 5 Variation of the communication sum rate under different numbers of UAVs. Detailed Implementation Modes
[0085] To better understand the purpose, structure, and function of this invention, the following further describes in detail a UAV trajectory and beamforming optimization method for a communication and sensing integrated system based on the MATD3 algorithm of this invention with reference to the drawings.
[0086] UAV Trajectory and Beamforming Optimization Method for Communication Sensing Integrated System Based on MATD3 Algorithm. Aiming at the problems that existing research fails to consider user roaming and three-dimensional UAV deployment, a multi-UAV ISAC system model is established, and the achievable sum rate is used as the performance index of the system. This invention focuses on studying the UAV trajectory and beamforming optimization scheme that maximizes the achievable sum rate under the condition of meeting the minimum sensing performance requirements, and proposes the corresponding mixed-integer non-linear programming problem. For the problem of user roaming, this invention first uses the K-means clustering algorithm to cluster communication users and sensing users respectively, and then uses the MATD3 algorithm to jointly optimize the flight trajectory and beamforming vector to solve this problem.
[0087] Step 1: As Figure 1 shown, construct a communication sensing integrated system supported by multiple UAVs. N UAVs act as dual-functional antenna base stations, communicate with ground users in the downlink, and simultaneously sense ground users. There are a total of users. Assume that the UAVs and users are respectively equipped with linear uniformly spaced array antennas and a single receiving antenna, and the users randomly roam in the service area; the UAVs , communication users , sensing users at the time slot are respectively represented as , , , where t represents the time slot, .
[0088] Step 2: In terms of the communication model, this invention adopts the free space path loss model. The channel gain from the UAV to the user can be expressed as:
[0089]
[0090] where is the channel gain at the reference distance of 1m, and are the transmit antenna gain and receive antenna gain of the communication user respectively. is the array response vector towards the user , and it can be expressed as:
[0091]
[0092] where and are the antenna space and wavelength respectively;
[0093]
[0094] The drones are assigned to a certain number of communication and sensing users, and each user can only be matched with one of the drones. Let denote the association indicator between the drone and the target user at time slot . When , it means that user is associated with the drone at time slot . Therefore, the system has the following user association constraint:
[0095]
[0096] The transmitted signal of the drone is given by:
[0097]
[0098] where and represent the information signal with symbols and the beamforming vector of user at the drone , respectively.
[0099] The signal received by user can be expressed as
[0100]
[0101] where denotes the additive white Gaussian noise, and denotes the noise power.
[0102] It can be observed from the above equation that the communication user is not only interfered by the users within the service area of the same drone but also by the users within the service areas of other drones. Therefore, the received signal-to-interference-plus-noise ratio (SINR) of user is:
[0103]
[0104] Using to denote the bandwidth, the achievable sum rate can be expressed as:
[0105]
[0106] For the sensing signal, the drone and the target sensing user The two-way channel power gain between is given by the following formula
[0107]
[0108] where is the channel power at a reference distance of 1 m, and are the transmit antenna gain and receive antenna gain of the sensing user, respectively. represents the average radar cross section of the target user, is the wavelength of the carrier. is the array response vector towards user k, which can be expressed as:
[0109]
[0110]
[0111] In this ISAC system, communication signals are used for user sensing. Using an unmanned aerial vehicle to its associated sensing user The transmit beam pattern gain is used as a criterion to measure the sensing performance, that is, the power of the sensing signal to the target direction, which can be expressed as:
[0112]
[0113] To meet the requirements of sensing performance, the beam pattern gain for different sensing users needs to consider the minimum threshold and the echo distance . Therefore, the following requirements for sensing performance are:
[0114]
[0115] Step 3: In each time slot , the ISAC signal is simultaneously sent to the communication user and the sensing user. For any unmanned aerial vehicle , its position must satisfy the following constraint conditions:
[0116]
[0117] where and are the lower and upper bounds of the unmanned aerial vehicle's motion trajectory, respectively, and the maximum flight speed of the unmanned aerial vehicle in each time slot is , the distance between unmanned aerial vehicles must be greater than the safety distance .
[0118] By jointly optimizing the trajectory of the unmanned aerial vehicle , the beamformer and the user association indicator , maximize the communication sum rate under the constraints of beam pattern gain, maximum transmit power, and UAV flight constraints. Therefore, the optimization problem can be formulated as:
[0119] ;
[0120] ;
[0121] ;
[0122] ;
[0123] ;
[0124] ;
[0125] ;
[0126] ;
[0127] ;
[0128] Step 4: To consider the mobility of users, the present invention uses a random walk model to simulate the change of user positions. At any discrete time slot , the user moves randomly along the and axes with respect to the normal distribution , where is the maximum speed of user movement.
[0129] Use a hierarchical user association algorithm based on K-means to associate communication users and sensing users with the UAV. Since the LoS channel is adopted, users close to each other are likely to have similar channel state information. Communication users are first clustered by the K-means algorithm according to their spatial distances. Each communication user cluster is assigned to the UAV closest to it. Then the UAV can adjust its flight trajectory and beamformer to provide the maximum communication rate for this user cluster. Next is the association of sensing users based on the existing clusters. The specific process is to calculate the beam pattern gain between the sensing user and the UAV respectively, and assign each sensing user to the UAV that can provide the maximum beam pattern gain by comparison.
[0130] The hierarchical user association algorithm based on K-means is summarized as follows:
[0131]
[0132] Step 5: Considering that multiple drones are deployed in this system scenario, the present invention introduces the MATD3 method to jointly optimize the drone trajectories and beamforming vectors. MATD3 is an improved Actor-Critic (AC) algorithm, mainly used to solve the non-stationary multi-agent problems in continuous action spaces. It has decentralized Actor networks, a centralized Critic network, an experience replay buffer, and additional noise to ensure stable and efficient learning. TD3 is an optimization of the Deep Deterministic Policy Gradient (DDPG), and MATD3 is developed by extending it to multi-agent scenarios. Similar to the Multi-Agent Deep Deterministic Policy Gradient (MADDG), each agent includes an Actor module and a Critic module. The Actor module realizes the interaction between the agent and the environment and directly outputs deterministic actions; the Critic module evaluates the policy of the Actor module and guides the policy improvement. In a multi-agent scenario, the policies of each agent are updated iteratively, resulting in the environment being dynamically unstable for a specific agent. An agent cannot adapt to the dynamically unstable environment only by changing its own policy. To make it easier for agents to exhibit cooperative behaviors, the MATD3 model adopts a framework of centralized training and distributed execution. During the training process, the Critic needs to obtain the observation and action information of other agents. During the execution phase, the agent only selects actions based on its own observations.
[0133] The several main structures of the MATD3 algorithm are as follows:
[0134] (1) Environment setting:
[0135] The system state is , and the observation of each agent is , and the action is .
[0136] Each agent selects an action according to its policy and obtains an immediate reward according to the reward function .
[0137] (2) Objective: The objective of each agent is to maximize the expected cumulative return:
[0138]
[0139] where is the immediate reward of agent at time, and is the discount factor.
[0140] (3) Critic network: Each agent The Critic network has two value functions and is used to reduce the bias of value estimation. These two value functions are updated in a similar way to that in TD3:
[0141] First, calculate the target value, using the next actions of all agents :
[0142]
[0143] where is the target network, and is the action generated by the target Actor network.
[0144] The Critic network learns the value function by minimizing the following formula:
[0145]
[0146] where and are the parameters of the Critic network and the target Critic network respectively, is the policy taken by the agent is the discount factor, and
[0147] is the experience pool. The value of the Critic network is also used to update the Actor network. The gradient of the Actor policy can be updated by maximizing the following formula :
[0148]
[0149] where is the policy gradient of the agent
[0150] (5) Target policy smoothing: To reduce the estimation bias, when calculating the target value, noise is added to the action of each agent, which is the policy noise for enhancing agent exploration. This can avoid the policy from overfitting to certain specific actions, thus improving exploration and stability.
[0151] (6) Delayed update: To further improve stability, the Actor network is updated less frequently than the Critic network. For example, the Actor network is updated once for every two updates of the Critic network. In addition, the parameter update of the target network is also slower, following , where is the rate parameter for soft update.
[0152] Under this algorithm, each drone is regarded as an independent agent, and drones do not need to perform additional communication during training and can achieve cooperation by obtaining information from the environment and selecting their own actions.
[0153] In the present invention, the three components of the agent are described in detail as follows:
[0154] (1) State: The state space includes the signal-to-interference-plus-noise ratio of communication users, the beam pattern gain of sensing users, and the user association indicator, which can be expressed as .
[0155] (2) Action: The action space includes the flight position of the drone and the beamforming vector, which can be expressed as .
[0156] (3) Reward: According to the proposed optimization problem, the optimization goal is to maximize the communication sum rate of all communication users while meeting the beam pattern gain requirements of sensing users within the user cluster. The present invention designs the reward function as follows
[0157]
[0158]
[0159] where means that if the beam pattern gain cannot be guaranteed to meet the minimum sensing requirement, a penalty will be imposed.
[0160] The specific process of the MATD3 algorithm is as follows:
[0161] (1) Initialization: Initialize two Critic networks and one Actor network for each agent, and initialize the corresponding target networks.
[0162] (2) Interaction and experience storage: Each agent interacts with the environment and records the current state, action, reward, and next state.
[0163] (3) Update the Critic network: Sample a batch of data from the experience replay pool.
[0164] (4) Calculate the target value according to the above formula and update the two Critic networks.
[0165] (5) Delayed update of the Actor network: Update the parameters of the Actor network every several steps to maximize the value of its corresponding Critic network. value.
[0166] (6) Soft update of the target network: Update the parameters of the target Critic network and the target Actor network.
[0167] (7) Repeat the steps until the agent learns to optimize the policy in the environment.
[0168] The centralized MATD3 structure is in Figure 2 .
[0169] The simulation experiment design of the present invention is specifically described as follows: A square service area of 500 meters × 500 meters is simulated. Each drone carries four linear uniform array antennas and is associated with two communication users and one sensing user. The maximum transmit power is 20 dBm. The drones fly at a speed of 5 m / s at an altitude of 100 - 150 meters, and the maximum speed of user movement is 0.5 m / s.
[0170] Figure 3 shows the convergence performance of the MATD3 algorithm when changing the learning rate. As the number of iterations increases, it is expected that the average reward of the entire ISAC system during training will also increase. When the learning rate is too small, the convergence speed is slow. On the other hand, due to the continuous random movement of users, too large a learning rate will cause the content learned by the algorithm to become outdated, resulting in a performance reduction. In this experiment, when the learning rate is 5e - 4, the average reward converges quickly with the highest average reward.
[0171] Figure 4 shows a detailed comparison of the average rewards of three different algorithms. The MATD3 algorithm has the most significant effect, indicating that this algorithm can obtain the maximum communication sum rate while meeting the needs of the sensing user. Followed by the MADDPG algorithm, and finally the DDQN algorithm. It shows that the algorithm of the present invention can better realize the communication and sensing functions of the ISAC system.
[0172] Figure 5 shows the influence of the number of drones on the system communication sum rate. In the simulation, the number of drones is set to 3 - 7, the number of communication users is 6, and the number of sensing users is 3. It can be seen that the communication sum rate of the ISAC system increases with the increase in the number of drones.
Claims
1. A multi-UAV integrated communication and sensing trajectory and beamforming optimization method based on MATD3, characterized in that It includes the following steps: Step 1: Construct a communication and sensing integrated system supported by multiple UAVs. Use UAVs as dual-functional antenna base stations to conduct downlink communication with ground users and sense ground users simultaneously; Step 1 specifically includes: constructing a communication and sensing integrated system supported by multiple unmanned aerial vehicles (UAVs), which includes several UAVs acting as dual-functional antenna base stations to perform downlink communication with ground users and simultaneously sense ground users; assume that the UAVs and users are respectively equipped with linear uniformly spaced array antennas and a single receiving antenna, and the users randomly roam in the service area; the UAVs , communication users , and sensing users at the positions in time slots are respectively represented as , , , where t represents a time slot, ; Step 2: In terms of communication, calculate the communication link gain and signal-to-interference-plus-noise ratio (SINR), and derive the communication sum rate therefrom. In terms of sensing, calculate the beam pattern gain; In Step 2, the calculation of the beam pattern gain in terms of sensing is specifically as follows: For the sensing signal, the UAV and the target sensing user The two-way channel power gain between them is given by: ; Among them, is the channel power at a reference distance of 1 m, and are the transmit antenna gain and receive antenna gain of the sensing user, respectively; represents the average value of the radar cross section of the target user, is the wavelength of the carrier; is the array response vector towards the user which is expressed as: ; ; Among them, drone to the pitch angle of its associated sensing user ; Using communication signals for user sensing in an ISAC system, using drones to its associated sensing user The transmit beam pattern gain to the target direction of the sensing signal, which is used as a measure of sensing performance, i.e., the power of the sensing signal to the target direction, is expressed as: ; Among them, is the array response vector towards the user and represents the beamforming vector of the user at the UAV , and represents the association indicator between the UAV and the target user in the time slot . To meet the requirements of sensing performance, the beam pattern gain for different sensing users is considered with a minimum threshold and the echo distance , and the following requirements for sensing performance are as follows: ; Step 3: Under the condition of ensuring sensing performance, establish an optimization problem with the goal of maximizing the total data rate of all communication users during the mission time period; Step 3 is specifically as follows: In each time slot , the ISAC signal is simultaneously sent to communication users and sensing users. For any drone , its position satisfies the following constraint conditions: ; ; ; where and are the lower and upper bounds of the UAV's movement trajectory respectively, and the maximum flight speed of the UAV in each time slot is , and the distance between UAVs is greater than the safety distance ; By jointly optimizing the trajectory of the UAV , beamformer and user association indicator , under the constraints of beam pattern gain, maximum transmit power, and UAV flight, maximize the communication sum rate. Therefore, the optimization problem is formulated as: ; ; ; ; ; ; ; ; ; Among them, is the communication rate of the time-slot user, and the optimization objective (17) is expressed as the sum of the communication rates provided by all UAVs for communication users; (18) is the user association constraint; (19) means that the total power allocated to the users does not exceed the maximum power of each UAV, is the maximum power of the UAV; (20) means to ensure that the sensing requirement of each sensing target in the time slot meets the minimum threshold , represents the UAV, , represents the sensing user, ; (21)-(25) represent the maximum flight distance constraint, anti-collision constraint and flight range constraint of the UAV; Step 4: Use a hierarchical user association algorithm based on K-means to associate communication users, sensing users with UAVs; Step 5: Use the MATD3 algorithm of centralized training and distributed execution CTDE to train the UAVs to complete the optimization of the communication and sensing integrated trajectory and beamforming of multiple UAVs.
2. The multi-UAV integrated communication and sensing trajectory and beamforming optimization method based on MATD3 according to claim 1, characterized in that In Step 2, the calculation of the communication link gain and signal-to-interference-plus-noise ratio (SINR) and the derivation of the communication sum rate therefrom are specifically as follows: In terms of the communication model, the free space path loss model is adopted, from the UAV to the user The channel gain is expressed as: ; Among them, is the channel gain at a reference distance of 1 m, and are the transmit antenna gain and receive antenna gain of the communication user, respectively; is the array response vector towards the user and is expressed as: ; Among them, and are the antenna space and wavelength respectively; ; Among them, is the pitch angle between the drone and the communication user; The UAVs are assigned to communication and sensing users, and each user can only be matched with one of the UAVs; let denote the association indicator between the UAV and the target user . When , it means that user is associated with the UAV in time slot . Therefore, the system has user association constraints: ; Among them, denotes the user, ; Unmanned aerial vehicle The transmitted signal of which is given by the following formula: ; wherein and respectively represent an information signal with symbols and the beamforming vector of the user at the UAV ; User The received signal is expressed as: ; Among them, represents additive white Gaussian noise, represents the noise power, represents the unmanned aerial vehicle (UAV) to the user associated therewith channel gain, represents the channel gain between other UAVs and users ; and respectively represent the information signal sent by the UAV to other communication users and the beamforming vector of other communication users at the UAV ; It can be obtained from the above formula that communication users are not only interfered by users within the same UAV service area, but also interfered by users within other UAV service areas. Therefore, the received signal-to-interference-plus-noise ratio (SINR) of the user is: ; Use to represent the bandwidth, and the achievable sum rate is expressed as: 。 3. A method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3 according to claim 1, characterized in that, Step 4 specifically is: Use a random walk model to simulate the change of the user's location. At any discrete time slot , the user moves randomly along the and axes with respect to a normal distribution , where is the maximum speed of the user's movement; Use a hierarchical user association algorithm based on K-means to associate communication users, sensing users, and drones; The communication users are first clustered by the K-means algorithm according to the spatial distance, and each communication user cluster is assigned to the nearest drone; Then the drone adjusts its movement trajectory and beamformer to provide the maximum communication rate for this user cluster; Based on the existing cluster, perform sensing user association, calculate the beam pattern gain between the sensing user and the UAV respectively, and assign each sensing user to the UAV that provides the maximum beam pattern gain by comparison.
4. A method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3 according to claim 1, characterized in that, In Step 5, the MATD3 algorithm specifically includes the following steps: Step 5.1: Initialization: Initialize two Critic networks and one Actor network for each agent, and initialize the corresponding target networks; Step 5.2: Interaction and experience storage: Each agent interacts with the environment, records the current state, action, reward, and next state; Step 5.3: Update the Critic network: Sample a batch of data from the experience replay pool; Step 5.4, calculate the target value and update the two Critic networks; Step 5.
5. Delay the update of the Actor network: Update the parameters of the Actor network every several steps to maximize the value of its corresponding Critic network; Step 5.6: Soft update the target network: Update the parameters of the target Critic network and target Actor network; Step 5.7: Repeat Steps 5.2 - 5.6 until the agent learns to optimize the strategy in the environment.
5. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method described in Claim 1.
6. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method described in Claim 1 are implemented.
7. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method described in Claim 1 are implemented.
Citation Information
Patent Citations
Resource allocation method for unmanned aerial vehicle communication and sensing integrated system
CN117295090A
Unmanned aerial vehicle trajectory and beamforming design method for communication computing resource fusion network
CN117595905A
Cited By
Multi-target pointing intelligent anti-interference beam generation method
CN121966635A