MATD3-based multi-unmanned aerial vehicle sensing integrated trajectory and beam forming optimization method

Through the multi-agent reinforcement learning method using the MATD3 algorithm and K-means user association algorithm in multi-unmanned airport scenes, the problem of collaboration between multiple drones is solved, and efficient communication and perceptual performance optimization in dynamic environments is achieved.

CN120186624AActive Publication Date: 2025-06-20NANJING UNIV OF POSTS & TELECOMM +1

Patent Information

Application Number
CN202510653967.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

In the multi-unmanned airport scenario, single agent reinforcement learning methods are difficult to deal with the problem of collaboration between multiple agents, and traditional methods are difficult to effectively solve the NP-difficulty problem in dynamic user roaming and complex environments.

Method used

The MATD3-based multi-agent reinforcement learning (MARL) method is adopted, combined with centralized training and distributed execution (CTDE) framework, and the UAV is trained using the K-means hierarchical user association algorithm and the MATD3 algorithm to optimize the synesthesia integrated trajectory and beamforming of multi-drone.

Benefits of technology

Maximize long-term communication and speed while meeting perceptual performance and transmit power constraints, improve system efficiency and performance, and enable effective collaboration in dynamic and complex multi-unmanned airport scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186624A_ABST
    Figure CN120186624A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-unmanned aerial vehicle communication sensing integrated trajectory and beam forming optimization method based on MATD3. The method comprises the following steps: constructing a communication sensing integrated system supported by multiple unmanned aerial vehicles; in the aspect of communication, a communication link gain and a signal to interference plus noise ratio (SINR) are calculated, and communication and rate are derived therefrom; in the sensing aspect, the beam pattern gain is calculated; establishing an optimization problem by taking maximization of the total data rate of all communication users in the task time period as an optimization target; associating the communication user and the sensing user with the unmanned aerial vehicle by using a K-means-based hierarchical user association algorithm; and training the unmanned aerial vehicles by using an MATD3 algorithm of centralized training and distributed CTDE execution, and completing optimization of the multi-unmanned aerial vehicle sensing integrated trajectory and beam forming. According to the method, a more practical roaming user scene and three-dimensional deployment of the unmanned aerial vehicle are considered, and relatively high communication and rate are achieved while sensing requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication technologies, and particularly relates to a method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3. Background Art

[0002] With the rapid development of the sixth-generation mobile communication, communication and sensing integration (ISAC) has gradually become a research hotspot. It integrates communication and sensing functions into the same system and shares hardware resources. This integration not only improves the efficiency of the system but also enables the communication and sensing functions to promote each other.

[0003] Applying unmanned aerial vehicles (UAVs) to the ISAC system has become an increasingly prominent research field because of its great potential for enhancing communication and sensing capabilities. Compared with ground channels, UAVs provide important line-of-sight (LoS) links and flexible support for the deployment of the dual-functional ISAC system. Specifically, in terms of communication, the transmission of communication signals becomes more stable and can provide a relatively high data transmission rate. In terms of sensing, the high-altitude position of UAVs enables them to obtain better sensing performance. For example, radar signals are not blocked or scattered by ground obstacles, thus achieving high-precision and stable sensing.

[0004] Although in the past few years, UAV-assisted ISAC systems have been widely explored in terms of UAV deployment and resource allocation, research in dynamic user roaming and complex multi-UAV scenarios is still lacking. In multi-UAV scenarios, single-agent reinforcement learning methods are difficult to handle the cooperation problems among multiple agents. In previous UAV-assisted ISAC research, traditional methods such as convex optimization performed well in the limited static user environment of UAVs and users. However, when the environment changes dynamically due to user roaming or the system becomes complex with the increase of targets, the problem becomes an NP-hard problem, and it is infeasible to solve it with these methods.

[0005] Original reinforcement learning methods, such as Q-learning, provide solutions for tasks with a limited state-action space based on tables. And some advanced deep reinforcement learning (DRL) algorithms, such as deep Q-network (DQN), proximal policy optimization (PPO), soft actor-critic (SAC), deep deterministic policy gradient (DDPG), etc., have also been used in various tasks. In these scenarios, each UAV only needs to calculate and execute operations related to its local observation. In multi-UAV scenarios, single-agent reinforcement learning methods are difficult to handle the cooperation problems among multiple agents. Although existing MARL algorithms show superiority in agent cooperation, the local exploration of each agent has not been fully studied, and this deficiency limits the performance of UAV-assisted ISAC systems. Summary of the Invention

[0006] Objective of the Invention: The objective of the present invention is to provide a method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3. Compared with the existing UAV ISAC systems that assume static users or two-dimensional UAV trajectories, multi-agent reinforcement learning (MARL) based on centralized training and distributed execution (CTDE) is used in the multi-UAV scenario, considering a more realistic scenario of roaming users and three-dimensional deployment of UAVs. An optimization problem of joint flight trajectory and beamforming is formulated to maximize the long-term communication sum rate under the constraints of transmit power and ensuring the beam pattern gain of the sensing target.

[0007] Technical Solution: A method for optimizing the trajectory and beamforming of multi-UAV communication and sensing integration based on MATD3 of the present invention includes the following steps: Step 1: Construct a communication and sensing integrated system supported by multiple UAVs, where the UAVs act as dual-functional antenna base stations to conduct downlink communication with ground users and sense the ground users simultaneously; Step 2: In terms of communication, calculate the communication link gain and signal-to-interference-plus-noise ratio (SINR), and therefrom derive the communication sum rate; in terms of sensing, calculate the beam pattern gain; Step 3: Under the condition of ensuring sensing performance, establish an optimization problem with the objective of maximizing the total data rate of all communication users within the task time period; Step 4: Use a hierarchical user association algorithm based on K-means to associate communication users, sensing users with UAVs; Step 5: Use the MATD3 algorithm of centralized training and distributed execution (CTDE) to train the UAVs to complete the optimization of the trajectory and beamforming of multi-UAV communication and sensing integration.

[0008] Further, Step 1 is specifically as follows: Construct a communication and sensing integrated system supported by multiple UAVs, including several UAVs acting as dual-functional antenna base stations, conducting downlink communication with ground users and sensing ground users simultaneously; assume that the UAVs and users are respectively equipped with linear uniformly spaced array antennas and a single receive antenna, and the users roam randomly in the service area; the positions of UAV , communication user , sensing user at time slot are respectively represented as , , , where t represents the time slot, and .

[0009] Further, in step 2, in terms of communication, calculate the communication link gain and signal-to-interference-plus-noise ratio (SINR), and therefrom derive the communication sum rate specifically as follows: In terms of the communication model, adopt the free space path loss model, from the unmanned aerial vehicle (UAV) to the user The channel gain is expressed as: ; where is the channel gain at a reference distance of 1 m, and are the transmit antenna gain and receive antenna gain of the communication user respectively; is the array response vector towards the user and is expressed as: ; where and are the antenna space and wavelength respectively; ; where is the elevation angle between the UAV and the communication user; The UAVs are assigned to communication and sensing users, and each user can only be matched with one of the UAVs; Let represent the association indicator between the UAV and the target user in time slot . When , it means that the user is associated with the UAV in time slot ; Therefore, the system has the user association constraint: ; where represents the user, ; The transmitted signal of the UAV is given by the following formula: ; where and represent the information signal with symbols and the beamforming vector of the user at the UAV respectively; The user The received signal is expressed as: ; wherein, represents additive white Gaussian noise, represents the noise power, represents the unmanned aerial vehicle (UAV) to the user associated therewith channel gain, represents the channel gain between other UAVs and the user ; and respectively represent the information signal transmitted by the UAV to other communication users and the beamforming vector of other communication users at the UAV ; It can be obtained from the above formula that the communication user is not only interfered by the users within the service area of the same UAV, but also interfered by the users within the service areas of other UAVs. Therefore, the received signal-to-interference-plus-noise ratio (SINR) of the user is: ; Using to represent the bandwidth, the achievable sum rate is expressed as:

[0010] Furthermore, in step 2, in terms of sensing, the specific calculation of the beam pattern gain is as follows: For the sensing signal, the two-way channel power gain between the UAV and the target sensing user is given by the following formula: ; wherein, is the channel power at a reference distance of 1 m, and are the transmitting antenna gain and receiving antenna gain of the sensing user respectively; represents the average value of the radar cross section of the target user, is the wavelength of the carrier; is the array response vector towards the user and is expressed as: ; ; wherein, the UAV to its associated sensing user of the pitch angle; In this ISAC system, the communication signal is used for user sensing, and a drone is used to its associated sensing user The transmit beam pattern gain to is used as a criterion to measure the sensing performance, that is, the power of the sensing signal to the target direction, expressed as: ; where, is the array response vector towards user , represents the beamforming vector of the user at the drone , represents the association indicator between the drone and the target user at time slot ; To meet the requirements of sensing performance, the beam pattern gain for different sensing users considers the minimum threshold and the echo distance , and the following requirements for sensing performance are as follows:

[0011] Furthermore, step 3 is specifically as follows: At each time slot , the ISAC signal is simultaneously sent to the communication user and the sensing user. For any drone , its position satisfies the following constraint conditions: ; ; ; where and are the lower and upper bounds of the drone's movement trajectory respectively, and the maximum flight speed of the drone at each time slot is , and the distance between drones is greater than the safety distance ; By jointly optimizing the trajectory of the drone , the beamformer and the user association indicator , under the constraints of beam pattern gain, maximum transmit power, and drone flight, maximize the communication sum rate. Therefore, the optimization problem is formulated as: ; ; ; ; ; ; ; ; ; Among them, is the communication rate of the time-slot user, and the optimization objective (17) is expressed as the sum of the communication rates provided by all UAVs for communication users; (18) is the user association constraint; (19) means that the total power allocated to the users does not exceed the maximum power of each UAV, is the maximum power of the UAV; (20) means to ensure that the sensing requirement of each sensing target in the time slot meets the minimum threshold , represents the UAV, , represents the sensing user, ; (21)-(25) represent the maximum flight distance constraint, anti-collision constraint and flight range constraint of the UAV.

[0012] Furthermore, step 4 is specifically as follows: Use the random walk model to simulate the change of the user's position. At any discrete time slot , the user moves randomly along the axis with respect to the normal distribution and , where is the maximum speed of the user's movement; Use the hierarchical user association algorithm based on K-means to associate the communication users, sensing users with the UAVs; The communication users are first clustered by the K-means algorithm according to the spatial distance, and each communication user cluster is assigned to the UAV closest to it; Then the UAV adjusts its movement trajectory and beamformer to provide the maximum communication rate for this user cluster; Based on the existing clusters, the sensing user association is performed. The beam pattern gain between the sensing user and the UAV is calculated respectively, and each sensing user is assigned to the UAV that provides the maximum beam pattern gain by comparison.

[0013] Furthermore, in step 5, the MATD3 algorithm specifically includes the following steps: Step 5.1, Initialization: Initialize two Critic networks and one Actor network for each agent, and initialize the corresponding target networks; Step 5.2, Interaction and Experience Storage: Each agent interacts with the environment and records the current state, action, reward, and next state; Step 5.3, Update the Critic Networks: Sample a batch of data from the experience replay pool; Step 5.4, Calculate the target value and update the two Critic networks; Step 5.5, Delayed Update of the Actor Network: Update the parameters of the Actor network every several steps to maximize the value of its corresponding Critic network ; Step 5.6, Soft Update of the Target Networks: Update the parameters of the target Critic network and target Actor network; Step 5.7, Repeat steps 5.2 - 5.6 until the agent learns to optimize the policy in the environment.

[0014] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the method of the present invention.

[0015] The present invention also discloses a computer - readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.

[0016] The present invention also discloses a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.

[0017] Beneficial Effects: Compared with the prior art, the present invention has the following remarkable advantages: The present invention considers a more realistic user roaming scenario and the three - dimensional deployment of drones, and adopts a two - step method for the proposed optimization problem, which can decompose the complex optimization problem into two separate problems. First, cluster the users to determine the users served by the drones before providing communication and sensing services, and then solve the joint optimization of the flight trajectory and beamforming vector by using the MATD3 algorithm. This algorithm introduces a delayed update and a target policy smoothing mechanism on the basis of the traditional MADDPG algorithm to avoid instability caused by too fast policy updates and can perform better in a mixed collaborative and competitive environment.

[0018] Under the constraints of transmission power and ensuring the gain of the sensing target beam pattern, an optimization problem of joint flight trajectory and beamforming is formulated to maximize the long-term communication sum rate. To address the challenges brought by dynamic and high-dimensional characteristics, the optimization problem is formulated as a partially observable Markov decision process (POMDP) problem, and a two-step method for dynamic scenarios is proposed: (1) A hierarchical user association algorithm based on K-means is proposed to update user association regularly; (2) The multi-agent twin-delayed deep deterministic policy gradient (MATD3) method is used to jointly optimize the flight trajectory and beamforming vector. Reinforcement learning learns the optimal policy through the interaction between the agent and the environment. In the optimization of UAV trajectory and beamforming, reinforcement learning can dynamically adjust the behavior of the UAV to maximize the long-term benefit.

[0019] In a multi-UAV scenario, single-agent reinforcement learning methods are difficult to handle the cooperation problems among multiple agents. Therefore, this invention further introduces the multi-agent reinforcement learning (MARL) method. By treating each UAV as an independent agent, the computational burden is decentralized, and each UAV is allowed to execute actions based on local observations, thereby improving the efficiency and performance of the system. Brief Description of the Drawings

[0020] Figure 1 Model of the UAV trajectory and beamforming optimization method for the communication and sensing integrated system based on MATD3.

[0021] Figure 2 Structural diagram of the centralized MATD3 algorithm.

[0022] Figure 3 Convergence performance of the MATD3 algorithm under different learning rates.

[0023] Figure 4 Performance comparison diagram of the MATD3 algorithm with other different algorithms.

[0024] Figure 5 Variation of the communication sum rate under different numbers of UAVs. Detailed Description of the Invention

[0025] To better understand the purpose, structure, and function of the present invention, the following further describes in detail a method for optimizing the UAV trajectory and beamforming of a communication and sensing integrated system based on the MATD3 algorithm of the present invention with reference to the accompanying drawings.

[0026] UAV Trajectory and Beamforming Optimization Method for Communication Sensing Integrated System Based on MATD3 Algorithm. Aiming at the problems that existing research does not consider user roaming and three-dimensional deployment of UAVs, a multi-UAV ISAC system model is established, and the achievable sum rate is used as the performance index of the system. This invention focuses on studying the UAV trajectory and beamforming optimization scheme that maximizes the achievable sum rate under the condition of meeting the minimum sensing performance requirements, and proposes the corresponding mixed integer non-linear programming problem. Aiming at the user roaming problem, this invention first uses the K-means clustering algorithm to cluster the communication users and sensing users respectively, and then uses the MATD3 algorithm to jointly optimize the flight trajectory and beamforming vector to solve this problem.

[0027] Step 1: As Figure 1 shown, construct a communication sensing integrated system supported by multiple UAVs. N UAVs act as dual-functional antenna base stations, and communicate with ground users in the downlink, and at the same time sense ground users. There are a total of users. Assume that the UAVs and users are respectively equipped with linear uniformly arranged array antennas and a single receiving antenna, and the users randomly roam in the service area; the UAVs , communication users , sensing users at time slot are respectively represented as , , , where t represents the time slot, .

[0028] Step 2: In terms of the communication model, this invention adopts the free space path loss model. The channel gain from the UAV to the user can be expressed as: where is the channel gain at a reference distance of 1m, and are the transmit antenna gain and receive antenna gain of the communication user respectively. is the array response vector towards the user , and it can be expressed as: where and are the antenna space and wavelength respectively; The UAVs are assigned to a certain number of communication and sensing users, and each user can only be matched with one of the UAVs. Let denote the association indicator between the UAV and the target user in time slot . When , it means that user is associated with the UAV in time slot . Therefore, the system has the user association constraint: The transmitted signal of the UAV is given by: where and represent the information signal with symbols and the beamforming vector of user at the UAV , respectively.

[0029] The signal received by user can be expressed as where denotes the additive white Gaussian noise, and represents the noise power.

[0030] It can be observed from the above formula that the communication user is not only interfered by the users within the service area of the same UAV but also by the users within the service areas of other UAVs. Therefore, the received signal-to-interference-plus-noise ratio (SINR) of user is: Using to represent the bandwidth, the achievable sum rate can be expressed as: For the sensing signal, the two-way channel power gain between the UAV and the target sensing user is given by where is the channel power at a reference distance of 1 m, and and are the transmit antenna gain and receive antenna gain of the sensing user, respectively. represents the average value of the radar cross section of the target user, and is the wavelength of the carrier. is the array response vector towards user k, which can be expressed as: In this ISAC system, communication signals are used for user sensing. Using drones to their associated sensing users The transmit beam pattern gain as a measure of the sensing performance, i.e., the power of the sensing signal to the target direction, can be expressed as: To meet the requirements of sensing performance, the beam pattern gain for different sensing users needs to consider the minimum threshold and the echo distance . Therefore, the following requirements for sensing performance are: Step 3: In each time slot , the ISAC signal is simultaneously sent to communication users and sensing users. For any drone , its position must satisfy the following constraint conditions: where and are the lower and upper bounds of the drone's motion trajectory respectively, and the maximum flight speed of the drone in each time slot is , the distance between drones must be greater than the safety distance .

[0031] By jointly optimizing the drone's trajectory , the beamformer and the user association indicator , under the constraints of beam pattern gain, maximum transmit power, and drone flight, maximize the communication sum rate. Therefore, the optimization problem can be formulated as: ; ; ; ; ; ; ; ; ; Step 4: To account for the mobility of users, the present invention uses a random walk model to simulate the change of user location. At any discrete time slot , the user moves randomly along the and axes with respect to a normal distribution , where is the maximum speed of user movement.

[0032] The communication users and sensing users are associated with the UAVs using a hierarchical user association algorithm based on K-means. Since the LoS channel is adopted, users in close proximity are likely to have similar channel state information. The communication users are first clustered by the K-means algorithm according to their spatial distances. Each communication user cluster is assigned to the UAV closest to it. Then the UAV can adjust its flight trajectory and beamformer to provide the maximum communication rate for this user cluster. Next is the association of sensing users based on the existing clusters. The specific process is to calculate the beam pattern gain between the sensing users and the UAVs respectively, and each sensing user is assigned to the UAV that can provide the maximum beam pattern gain by comparison.

[0033] The hierarchical user association algorithm based on K-means is summarized as follows:

[0034] Step 5: Considering that multiple UAVs are deployed in this system scenario, the present invention introduces the MATD3 method to jointly optimize the UAV trajectories and beamforming vectors. MATD3 is an improved Actor-Critic (AC) algorithm mainly used to solve non-stationary multi-agent problems in continuous action spaces. It has decentralized Actor networks, a centralized Critic network, an experience replay buffer, and additional noise to ensure stable and efficient learning. TD3 is an optimization of the Deep Deterministic Policy Gradient (DDPG), and MATD3 is developed by extending it to multi-agent scenarios. Similar to the Multi-Agent Deep Deterministic Policy Gradient (MADDG), each agent includes an Actor module and a Critic module. The Actor module realizes the interaction between the agent and the environment and directly outputs deterministic actions; the Critic module evaluates the policy of the Actor module and guides the policy improvement. In a multi-agent scenario, the policies of each agent are updated iteratively, resulting in a dynamically unstable environment for a specific agent. An agent cannot adapt to a dynamically unstable environment only by changing its own policy. To make it easier for agents to exhibit cooperative behaviors, the MATD3 model adopts a framework of centralized training and distributed execution. During the training process, the Critic needs to obtain the observations and action information of other agents. During the execution phase, the agent selects actions only based on its own observations.

[0035] The several main structures of the MATD3 algorithm are as follows: (1) Environment setting: The system state is , and the observation of each agent is , and the action is .

[0036] Each agent selects an action according to its policy and obtains an immediate reward according to the reward function .

[0037] (2) Goal: The goal of each agent is to maximize the expected cumulative return: where is the immediate reward of agent at time , and is the discount factor.

[0038] (3) Critic network: The Critic network of each agent has two value functions and to reduce the bias of value estimation. The update methods of these two value functions are similar to those in TD3: First, calculate the target value, using the next actions of all agents: where is the target network, and is the action generated by the target Actor network .

[0039] The Critic network learns the value function by minimizing the following formula: where and are the parameters of the Critic network and the target Critic network respectively, is the policy taken by agent , is the discount factor, and is the experience pool.

[0040] ​​​​​​​(4) Actor Network: The decentralized Actor network is used to make actions based on local observations. The update of the Actor network is also carried out by maximizing the value. The gradient of the Actor policy can update its parameters by maximizing the following formula : where is the policy gradient of the agent .

[0041] (5) Target Policy Smoothing: To reduce the estimation bias, when calculating the target value, noise is added to the action of each agent , which is the policy noise for enhancing the exploration of the agent. This can avoid the policy from overfitting to certain specific actions, thereby improving the exploration and stability.

[0042] (6) Delayed Update: To further improve the stability, the update frequency of the Actor network is lower than that of the Critic network. For example, the Actor network is updated once every two updates of the Critic network. In addition, the parameter update of the target network is also slow, following , where is the rate parameter of soft update.

[0043] Under this algorithm, each drone is regarded as an independent agent. The drones do not need to perform additional communication during training and can achieve cooperation by obtaining information from the environment and choosing their own actions.

[0044] In the present invention, the three components of the agent are described in detail as follows: (1) State: The state space includes the signal-to-interference-plus-noise ratio of communication users, the beam pattern gain of sensing users, and the user association indicator, which can be expressed as .

[0045] (2) Action: The action space includes the flight position of the drone and the beamforming vector, which can be expressed as .

[0046] (3) Reward: According to the proposed optimization problem, the optimization goal is to maximize the communication sum rate of all communication users while satisfying the beam pattern gain requirements of sensing users within the user cluster. The present invention designs the reward function as follows where It means that if the beam pattern gain cannot meet the minimum sensing requirements, a penalty will be imposed.

[0047] The specific process of the MATD3 algorithm is as follows: (1) Initialization: Initialize two Critic networks and one Actor network for each agent, and initialize the corresponding target networks.

[0048] (2) Interaction and experience storage: Each agent interacts with the environment, recording the current state, action, reward, and next state.

[0049] (3) Update the Critic network: Sample a batch of data from the experience replay pool.

[0050] (4) Calculate the target value according to the above formula and update the two Critic networks.

[0051] (5) Delayed update of the Actor network: Update the parameters of the Actor network every several steps to maximize the value of its corresponding Critic network. value.

[0052] (6) Soft update of the target network: Update the parameters of the target Critic network and target Actor network.

[0053] (7) Repeat the steps until the agent learns to optimize the strategy in the environment.

[0054] The centralized MATD3 structure is in Figure 2 .

[0055] The simulation experiment design of the present invention is specifically described as follows: A square service area of 500 meters × 500 meters is simulated. Each UAV carries four linear uniform array antennas and is associated with two communication users and one sensing user. The maximum transmit power is 20 dBm. The UAVs fly at a speed of 5 m / s at an altitude of 100 - 150 meters, and the maximum speed of the user movement is 0.5 m / s.

[0056] Figure 3 It shows the convergence performance of the MATD3 algorithm when changing the learning rate. As the number of iterations increases, it is expected that the average reward of the entire ISAC system during training will also increase. When the learning rate is too small, the convergence speed is slow. On the other hand, due to the continuous random movement of the users, an overly large learning rate will cause the content learned by the algorithm to become outdated, resulting in a performance reduction. In this experiment, when the learning rate is 5e-4, the average reward converges quickly with the highest average reward.

[0057] Figure 4It shows a detailed comparison of the average rewards of three different algorithms. The MATD3 algorithm has the most significant effect, indicating that this algorithm can obtain the maximum communication sum rate while meeting the needs of sensing users. Followed by the MADDPG algorithm, and finally the DDQN algorithm. It shows that the algorithm of the present invention can better realize the communication and sensing functions of the ISAC system.

[0058] Figure 5 It shows the impact of the number of UAVs on the communication sum rate of the system. In the simulation, the number of UAVs is set to 3 - 7, the number of communication users is 6, and the number of sensing users is 3. It can be seen that the communication sum rate of the ISAC system increases with the increase in the number of UAVs.

Claims

1. A multi-UAV synaesthesia integrated trajectory and beamforming optimization method based on MATD3, characterized in that: The steps include: Step 1: Build a communication perception integrated system supported by multiple drones, with drones acting as dual-function antenna base stations to perform downlink communications with ground users and perceive ground users at the same time; Step 2: In terms of communication, calculate the communication link gain and signal-to-interference-and-noise ratio (SINR), and derive the communication sum rate from them; in terms of perception, calculate the beam pattern gain; Step 3: Under the condition of ensuring the perception performance, an optimization problem is established with the optimization goal of maximizing the total data rate of all communication users in the task time period; Step 4: Use the K-means-based hierarchical user association algorithm to associate the communication users with the perception users and the drones; Step 5: Use the MATD3 algorithm with centralized training and distributed execution of CTDE to train the drones and optimize the trajectory and beamforming of multi-drone synaesthesia integration.

2. The multi-UAV synaesthesia integrated trajectory and beamforming optimization method based on MATD3 according to claim 1 is characterized in that: Step 1 is as follows: Build a multi-UAV-supported communication perception integrated system, including several UAVs acting as dual-function antenna base stations, and The ground users perform downlink communication and sense ground users; suppose the drone and the user are equipped with Linear uniform array antennas and a single receiving antenna, users roam randomly in the service area; UAV , communication users , perceive users exist The positions of the time slots are represented as , , ,in, t Indicates time slot, .

3. The multi-UAV synaesthesia integrated trajectory and beamforming optimization method based on MATD3 according to claim 2 is characterized in that: In step 2, in terms of communication, the communication link gain and signal to interference and noise ratio SINR are calculated, and the communication sum rate is derived from them as follows: In terms of communication model, the free space path loss model is adopted. To User The channel gain is expressed as: ; in, is the channel gain at the reference distance of 1m, and are the transmitting antenna gain and receiving antenna gain of the communication user respectively; It is user-oriented An array response vector, which is expressed as: ; in, and are antenna space and wavelength respectively; ; in, is the pitch angle between the UAV and the communication user; Drones are assigned to communication and perception users, and each user can only be matched with one of the drones; Indicates time slot Drones and target users The association indicator between When the user In time slot With drones associated; therefore, the system has user association constraints: ; in, Represented as a user, ; Drones The transmitted signal is given by: ; in and Respectively indicate that Symbols of information signals and drones User The beamforming vector of user The received signal is represented as: ; in, represents additive white Gaussian noise, represents the noise power, Indicates drone To the user associated with it The channel gain, Indicates other drones and users The channel gain, and Respectively represent drones Information signals sent to other communication users and drones The beamforming vectors of other communication users; From the above formula, it can be concluded that the communication user Not only will the user be interfered with by the same drone service area, but also by the user in other drone service areas. The received signal-to-interference-and-noise ratio SINR is: ; use represents bandwidth, and the achievable sum rate is expressed as: 。 4. The multi-UAV synaesthesia integrated trajectory and beamforming optimization method based on MATD3 according to claim 2 is characterized in that: In step 2, in terms of perception, the beam pattern gain is calculated as follows: For sensing signals, drones and target sensing users The bidirectional channel power gain between is given by: ; in, is the channel power at the reference distance of 1m, and They are the transmit antenna gain and receive antenna gain of the perceived user, respectively; represents the average value of the radar cross section of the target user, is the wavelength of the carrier; It is user-oriented An array response vector, which is expressed as: ; ; in, Drones to its associated sensor user The pitch angle angle; In this ISAC system, communication signals are used for user sensing, using drones to its associated sensor user The transmit beam pattern gain is used as a measure of the perception performance, that is, the power of the perception signal to the target direction, expressed as: ; in, It is user-oriented An array response vector, Indicates drone The beamforming vector of the user at Indicates time slot Drones and indicators of association with the target user; In order to meet the requirements of perception performance, the minimum threshold is considered for the beam pattern gain of different sensing users. and echo distance , the following requirements are placed on perception performance: 。 5. The multi-UAV synaesthesia integrated trajectory and beamforming optimization method based on MATD3 according to claim 4 is characterized in that: Step 3 is as follows: In each time slot , the ISAC signal is sent to the communication user and the perception user at the same time. For any UAV , whose position satisfies the following constraints: ; ; ; in and are the lower and upper bounds of the UAV’s trajectory, and the maximum flight speed of the UAV in each time slot is , the distance between drones is greater than the safe distance ; By jointly optimizing the trajectory of the UAV , beamformer and user association indicator , under the constraints of beam pattern gain, maximum transmit power and UAV flight, the communication and rate are maximized. Therefore, the optimization problem is formulated as: ; ; ; ; ; ; ; ; ; in, for The communication rate of the time slot user, the optimization objective (17) is expressed as the sum of the communication rates provided by all UAVs to the communication users; (18) is the user association constraint; (19) indicates that the total power allocated to the user does not exceed the maximum power of each UAV, is the maximum power of the UAV; (20) indicates that each sensing target is guaranteed to be The sensing of the time slot requires to meet the minimum threshold , Indicates drone, , Indicates the perceived user, ; (21)-(25) represent the maximum flight distance constraint, collision avoidance constraint and flight range constraint of the UAV.

6. The multi-UAV synaesthesia integrated trajectory and beamforming optimization method based on MATD3 according to claim 1 is characterized in that: Step 4 is as follows: Use the random walk model to simulate the change of user location in any discrete time slot , users follow the normal distribution of and The axes are randomly moved, where is the maximum speed of user movement; a hierarchical user association algorithm based on K-means is used to associate communication users with perception users and drones; communication users are first clustered by the K-means algorithm according to spatial distance, and each communication user cluster is assigned to the drone closest to it; then the drone adjusts its motion trajectory and beamformer to provide the maximum communication rate for the user cluster; Based on the existing clusters, sensing users are associated, the beam pattern gains between sensing users and drones are calculated respectively, and each sensing user is assigned to the drone that provides the maximum beam pattern gain through comparison.

7. The multi-UAV synaesthesia integrated trajectory and beamforming optimization method based on MATD3 according to claim 1 is characterized in that: In step 5, the MATD3 algorithm specifically includes the following steps: Step 5.1, Initialization: Initialize two Critic networks and one Actor network for each agent, and initialize the corresponding target network; Step 5.2, Interaction and experience storage: Each agent interacts with the environment and records the current state, action, reward, and next state; Step 5.3, Update the Critic network: Sample a batch of data from the experience replay pool; Step 5.4: Calculate the target Value, and update the two Critic networks; Step 5.5, Delayed update of the Actor network: Every few steps, update the parameters of the Actor network to maximize the corresponding Critic network value; Step 5.6, soft update target network: update the parameters of the target Critic network and the target Actor network; Step 5.

7. Repeat steps 5.2-5.6 until the agent learns to optimize the strategy in the environment.

8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method of claim 1.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

Citation Information

Patent Citations

  • Resource allocation method for unmanned aerial vehicle communication and sensing integrated system

    CN117295090A

  • Unmanned aerial vehicle trajectory and beamforming design method for communication computing resource fusion network

    CN117595905A

  • Joint waveform design method and device in enabling and inductance integrated network of unmanned aerial vehicle

    CN118400005A

  • AoI optimization method and system of communication and perception system supported by multiple unmanned aerial vehicles

    CN118945693A

  • Rate optimization method of unmanned aerial vehicle information acquisition system based on sensing integration

    CN119233285A

Cited By

  • Multi-unmanned aerial vehicle converged communication perception calculation method

    CN121078457A