Multi-unmanned aerial vehicle communication perception integration method based on adaptive reward weight adjustment

Through adaptive reward weight adjustment and multi-agent near-end strategy optimization methods, the perceived fairness and information age optimization of the integrated drone communication and perception system in dynamic scenarios is solved, efficient collaborative communication and perception of multi-drone systems are realized, and system throughput and task execution efficiency are improved.

CN120264350APending Publication Date: 2025-07-04FUDAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510386879.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing integrated UAV communication and perception system fails to effectively coordinate the optimization of perception fairness and information age in dynamic scenarios, and has high computational complexity, making it difficult to achieve real-time and efficient communication and perception in multiple mobile ground nodes.

Method used

The integrated communication and perception method of multi-drone communication based on adaptive reward weight adjustment is adopted. The trajectory planning, collision avoidance and resource allocation strategies of the drone are trained through multi-agent near-end strategy optimization (MAPPO). Combined with the MAPPO method of adaptive reward weight adjustment, the UAV decision-making process is dynamically guided to achieve low-complexity real-time decision-making.

Benefits of technology

It improves the communication throughput of edge ground nodes, realizes collaborative service of multiple drones in dynamic environments, improves the real-time and accuracy of perceived tasks, and significantly improves system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264350A_ABST
    Figure CN120264350A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-unmanned aerial vehicle communication perception integration method based on adaptive reward weight adjustment. The method comprises the following steps: a plurality of unmanned aerial vehicles equipped with communication sensing integrated modules are matched to provide downlink communication service for large-scale mobile ground nodes under the condition of ensuring sensing fairness and information age; according to the method, multi-agent near-end strategy optimization is introduced to train trajectory planning in tasks of multiple unmanned aerial vehicles, collision, node association and resource allocation strategies are avoided, and the communication throughput of the system is maximized; for the problem of frequent and repeated training caused by ground node movement, an adaptive reward weight adjustment MAPPO method is introduced for dynamically guiding the decision process of each unmanned aerial vehicle, and compared with a reference scheme, the method can improve the communication throughput of edge ground nodes; experimental results prove that the method can enable multiple unmanned aerial vehicles to cooperate in a ground node mobile environment to obtain higher communication throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication and sensing integration, and specifically relates to a multi-UAV communication and sensing integration method based on adaptive reward weight adjustment. Background Art

[0002] Currently, the wireless communication technology centered on the fifth-generation mobile communication technology is facing huge transformation and practical application pressures. With the decreasing marginal benefits emerging from the technological iteration in the wireless communication field, the phenomenon has become increasingly significant. At the technical level, although technologies such as large-scale antenna arrays and millimeter-wave communications have improved the utilization rate of bandwidth resources in scientific research work, they have not given rise to disruptive application modes like mobile payment and the explosion of short videos in the 4G era in real applications. Under the above background, the research on the sixth-generation mobile communication technology takes the combination of the space-air-ground-sea integrated network and artificial intelligence technology as the breakthrough direction, and the scientific community tries to expand the boundary of communication technology through disruptive technologies such as communication and sensing integration.

[0003] In the technical architecture and application scenario planning of the 6G network system, as a core supporting element, UAVs play an indispensable role in heterogeneous network integration and function realization. Due to their flexible and mobile characteristics in three-dimensional space and the unique advantages of their air-to-ground transmission links, they are applied to the auxiliary wireless communication systems in disaster rescue and traffic management. China's low-altitude economy has entered a strategic opportunity period of large-scale industrialization. The communication and sensing integration system based on UAVs has obvious advantages compared with the traditional ground communication and sensing integration system. The communication and sensing integration enabled by UAVs can not only provide a wider coverage range but also optimize the performance of communication and sensing functions through flexible aerial deployment.

[0004] Although many studies have deeply explored the UAV communication and sensing integration system in different aspects, most studies have failed to fully consider the combined effects of sensing fairness and age of information in mobile ground nodes. Although some scholars have enhanced the management of the age of information using dynamic programming methods while maintaining the positioning accuracy of static ground nodes, this study only focuses on static ground nodes and does not cover the dynamic changes of mobile ground nodes. Existing research more focuses on the improvement of communication performance or the separate optimization of sensing tasks, while ignoring how to balance and optimize the relationship between sensing fairness and age of information in a dynamic environment. Therefore, existing research has not achieved the collaborative optimization of sensing fairness and age of information in mobile ground nodes, and further research and innovation are urgently needed.

[0005] Designing an efficient integrated communication and sensing system for unmanned aerial vehicles (UAVs), especially in dynamic scenarios, faces multiple challenges. First, considering the flexibility of UAVs during aerial operations, the channels between UAVs and ground nodes change frequently due to factors such as environmental variations, movement speeds, and obstacles. In a moving environment, especially when the movement trajectories of UAVs and the relative positions of ground nodes are constantly changing, the time-varying nature of the channels is exacerbated, posing significant challenges to the stability, reliability, and real-time performance of the communication and sensing systems. Problems such as signal attenuation, multipath effects, and interference make the system's performance vary greatly at different time points, imposing high requirements on the accuracy of sensing, the rate of communication, and delay control. Second, the timeliness and fairness requirements of sensing tasks are also issues that cannot be ignored in the design. In dynamic scenarios, the movement of ground nodes and environmental uncertainties make the requirements for real-time performance and high accuracy in sensing tasks even more stringent. Especially in the case of multiple ground nodes with a wide distribution, how to reasonably schedule system resources to ensure that sensing tasks can obtain effective information in real-time at different time points and achieve fairness among various ground nodes has become one of the core problems in the invention. Summary of the Invention

[0006] In view of the problems of sensing fairness and limited information age faced by UAV sensing tasks in the dynamic integrated communication and sensing scenario with wide distribution and high demand of ground nodes, as well as the shortcomings of high computational complexity and the need for retraining of the model due to environmental changes in the implementation of the UAV-assisted integrated communication and sensing solution, the present invention provides a multi-UAV integrated communication and sensing method based on adaptive reward weight adjustment. By introducing multi-agent proximal policy optimization to train the trajectory planning, collision avoidance, node association, and resource allocation strategies in the tasks of multiple UAVs, the communication throughput of the system is maximized; aiming at the problem of frequent repeated training caused by the movement of ground nodes, an adaptive reward weight adjustment multi-agent proximal policy optimization MAPPO method is introduced to dynamically guide the decision-making process of each UAV during training, obtaining a policy network that can map the observations of UAVs on the environment to decision-making actions, so as to achieve low-complexity instant decision-making in the execution stage. Compared with the baseline scheme, the present invention can improve the communication throughput of edge ground nodes; experimental results prove that the present invention enables multiple UAVs to cooperate in the ground node movement environment to obtain higher communication throughput.

[0007] The technical solution of the present invention is specifically introduced as follows.

[0008] The present invention provides a multi-UAV integrated communication and sensing method based on adaptive reward weight adjustment, including:

[0009] A group of multiple drones equipped with a communication and sensing integrated module provide downlink communication services for a group of ground nodes. At the same time, the drones send sensing signals to the ground nodes for ranging and observe and collect environmental data.

[0010] Based on the sensing data and environmental data, multiple drones use the trained MAPPO model to make real-time decisions;

[0011] The training process of the MAPPO model includes:

[0012] Step 1: Construct a multi-drone communication and sensing integrated system model, including γ drones and K ground nodes; each drone provides communication services for multiple ground nodes through orthogonal frequency division multiple access technology and realizes distance measurement by sending sensing signals, and constructs an optimization problem model for the throughput of the multi-drone communication and sensing integrated system;

[0013] Step 2: Model the communication and sensing tasks as a partially observable Markov decision process, define the joint state space, local observation space, action space and multi-objective reward function, obtain states, observations, decisions through training rounds, and calculate rewards;

[0014] Step 3: According to the designed adaptive reward weight adjustment method, dynamically adjust the weight coefficients of each reward item through curriculum error and standardize the weight coefficients;

[0015] Step 4: Based on the centralized training - decentralized execution framework, use the multi-agent proximal policy optimization algorithm to collaboratively optimize the policy, calculate the global advantage estimate, and execute gradient descent to update the policy network and value network;

[0016] Step 5: Repeat Steps 1 - 4 until all training rounds are completed, and experimentally verify the impact of the number of drones on the system performance to achieve the trajectory collaborative planning and communication resource scheduling of multiple drones.

[0017] Furthermore, in order to achieve the above objectives, it is necessary to establish a multi-drone communication and sensing integrated system model, and the model includes:

[0018] (1) Drone flight area constraints, heading angle change rate constraints, maximum speed and anti-collision distance constraints: The communication and sensing integrated task of the drone serves K ground nodes and lasts for N time slots, numbered n = 0,..., N - 1. The horizontal coordinates of the i-th drone at time slot n are q i [n] = (x i [n], y i [n]) T , and the drone flight area constraint is: x min ≤ x i [n] ≤ x max , ymin ≤y i [n]≤y max 。The drone controls its flight trajectory during the mission by controlling its speed and heading angle. Let v i [n]≤v max represent the speed of drone i at time slot n, and the heading angle θ i [n]∈(0,2π], where v max represents the maximum speed of the drone. The travel distance l i [n] of the i-th drone between two consecutive time slots is i [n]-q i [n-1]∥≤v max τ; To ensure the flight stability of the drone, a constraint on the rate of change of the heading angle is introduced: |θ i [n]-θ i [n-1]|≤Δθ max , where Δθ max represents the maximum instantaneous steering angle corresponding to the servo response limit of the drone flight control system. In the scenario of multi-drone collaborative work, a group safety isolation condition needs to be introduced. For drones i and i'; it satisfies ∥q i [n]-q i′ [n]∥≥d min , where d min represents the distance threshold set to avoid drone collisions.

[0019] (2) Total transmit power upper limit and unique ground node service constraint: In each time slot n, let P i [n] represent the total transmit power of drone i, which is constrained by the maximum power . The total power is allocated to the communication power and the sensing power . The two are divided by the power allocation coefficient ρ i [n], where the power allocated to communication is The power allocated to sensing is satisfies Then, by introducing a joint resource allocation strategy in the three-dimensional space domain and time-frequency domain, a multi-dimensional mapping relationship between the drone and the ground node is established to ensure that each ground node is served by only one drone in the same time slot. Specifically, the binary decision variable X i,k [n]∈{0,1} represents the association variable between drone i and the k-th ground node. The variable being 1 indicates an association between the two, otherwise it is 0. It satisfies the unique ground node service constraint: where γ represents the total number of drones.

[0020] (3) Detection variance threshold constraint based on signal-to-noise ratio for perceived quality, perceived fairness, and information age constraint: The signal-to-noise ratio of the single-static radar system equipped on the UAV is: where Γ TX and Γ RX represent the transmitting and receiving antenna gains respectively; λ represents the signal wavelength; σ k represents the radar cross-sectional area of the target; physical constants k, T0, B s [n], and Ψ represent the Boltzmann constant, the receiver equivalent noise temperature, the signal bandwidth, the noise factor, and the detection loss respectively. The device distance estimation performance is only affected by the environmental noise in the form of an additive Gaussian white random process. The Cramér-Rao lower bound is used to evaluate the performance of distance detection, and the signal-to-noise ratio ξ i,k [n] satisfies: where c is the speed of light, v S is the number of sampling samples, the carrier frequency of the i-th UAV, is the detection variance threshold, is the perceived bandwidth of the i-th UAV, β i [n] is its perceived bandwidth allocation coefficient. The perceived fairness index is inspired by the Jain fairness metric. The perceived fairness where ψ i,k [n] represents the latest effective perception time of the UAV i for the ground node k, and ψ i,k [n] is updated to the latest time when the perceived detection variance constraint is satisfied. Λ2 is the threshold to ensure perceived fairness. Then we evaluate the freshness of the perceived data through the information age. The information age of the perceived information of the ground node k obtained by the UAV i where ∈→0 + is a small constant used to prevent division by zero. By optimizing the information age, it is ensured that the UAV uses the latest perceived data as much as possible, thereby improving the execution efficiency and accuracy of the task.

[0021] (4) The main objective of the designed multi-UAV communication and sensing integrated system is to maximize the downlink communication service throughput by jointly optimizing the trajectory planning, ground node association, power allocation, and spectrum resource allocation of multiple UAVs while satisfying the given constraints, so as to support the simultaneous access and service of more ground nodes. Therefore, the system objective function is:

[0022]

[0023] where, represents the aggregated transmission rate obtained by the ground node k, R i,k[n] represents the transmission rate between the UAV i and the ground node k, and σ 2 is the noise power. In the formula, the interference term represents the co-channel interference caused by the spectrum reuse of the multi-UAV system. represents the communication sub-channel bandwidth obtained by a single ground node, and L i,k [n] = β0(d i,k [n]) -2 is the communication channel gain between the UAV i and the ground node k, β0 is the path loss factor. represents the distance between the UAV i and the ground node k, and h represents the cruising altitude of the UAV.

[0024] Furthermore, the multi-objective reward function described in step 2 includes the following items:

[0025] (1) Boundary control reward: Each UAV needs to fly within a specific task boundary to ensure the effectiveness of the task. Considering the characteristics of the rectangular task area, the boundary reward function is

[0026]

[0027] where: The central coordinates q of the task area c =(x c , y c ), the rectangular boundary size L x ×L y , L x =x max -x min , L y =y max -y min , and the boundary deviation where

[0028] By punishing the behavior of the UAV deviating from the task boundary, the effective execution of the task can be guaranteed, and the safety and efficiency of the UAV can be ensured during the task process.

[0029] (2) Collision avoidance reward between UAVs: In the multi-UAV cooperative flight task, avoiding collisions is an important goal to ensure the smooth completion of the task. To achieve this goal, we design the collision avoidance reward. where

[0030] N pairs =Υ(Υ - 1) / 2 is the number of UAV pairs, and d safe =2d minLet \(d_0\) be the safety distance threshold. The design purpose of this reward function is to encourage the UAV to maintain a sufficient safety distance during flight to avoid collisions.

[0031] (3) Distributed age-of-information reward: In the case of multiple UAVs, each UAV may be responsible for different nodes, so the reward needs to be distributed to each UAV. Therefore, an association variable \(X_{i}[n]\) is introduced to distinguish the nodes served by UAV \(i\). The distributed age-of-information reward for UAV \(i\) is: i,k [n] to distinguish the nodes served by UAV \(i\), and the distributed age-of-information reward for UAV \(i\) is:

[0032]

[0033] The responsibility division of each UAV is achieved through the association variable. This reward design ensures high scalability and will not become computationally complex due to the increase in the number of UAVs.

[0034] (4) Perceived fairness reward: In a multi-UAV system, each UAV has a different coverage area, so an independent fairness metric needs to be designed for each UAV to avoid some nodes being ignored. In addition, the perceived fairness reward needs to be balanced at the whole system level and the single UAV level. The perceived fairness reward is where represents the global perceived fairness. This reward combines the local fairness of UAV \(i\) with the global fairness to balance the individual and group optimization goals.

[0035] (5) Throughput reward and minimum throughput reward: To avoid numerical inconsistencies caused by scale differences, the system throughput is normalized. Define the normalized throughput contribution of UAV \(i\) as where \(L\) max is the maximum channel gain, which is determined by the minimum vertical distance \(h\) between the UAV and the ground node, corresponding to the free space path loss when the UAV hovers directly above the ground node. To ensure that each ground node in the system, especially those at the edge, can meet the basic communication quality requirements, for the set of associated nodes \(K_{i}[n]=\{k|\chi_{k}[n]=1\}\) of UAV \(i\), define its minimum throughput reward as i [n] = {k|χ i,k [n] = 1}, and define its minimum throughput reward as This reward guarantees the service quality of edge nodes and drives UAV \(i\) to preferentially improve the performance of edge nodes.

[0036] (6) Node association reward: Through an exponential penalty mechanism, it is ensured that only one UAV is associated with each ground node \(k\). The ground node association reward for UAV \(i\) is This design satisfies the following mathematical property: when and only when all associated nodes of UAV \(i\) are uniquely served, the cumulative penalty amount in the exponential term is zero, and the reward reaches the maximum value of 1. Each conflicting ground node introduces \((N link -1)2 Type of punishment, where N link is the number of UAVs actually associated with ground nodes, ensuring that multiple association conflicts cause superlinear attenuation; at the same time, ensuring non-active neutrality, that is, when the UAV is not associated with any node, the reward is always 1, avoiding ineffective punishment for non-active UAVs.

[0037] Furthermore, an adaptive reward weight adjustment method is designed to dynamically adjust the reward function during training to help the agent adapt to the training process and environmental changes. The adaptive reward weight adjustment method includes:

[0038] Calculate the population curriculum error at the m-th adjustment Quantify the deviation between each reward item and the target performance, by Dynamically update the reward weight coefficient at the m-th adjustment, where D j is the initial difficulty coefficient, E is the preset error upper limit, and κ is the adaptive coefficient. Then through Normalize the weights, and the weight coefficients adjusted centrally during the training phase are synchronized to all UAV policy networks to ensure behavioral consistency during decentralized execution.

[0039] Furthermore, due to the high-dimensional coupled decision space, group collaboration constraints, partial observability and non-stationarity in dynamic environments brought by the multi-UAV communication and sensing integrated system. In addition, the centralized solution is sensitive to the number of UAVs, increasing the number of UAVs will significantly increase the algorithm complexity, and the failure of the control center will lead to the paralysis of the entire network. Therefore, combined with the dynamic characteristics of the considered optimization environment including mobile ground nodes, we designed a multi-agent proximal policy optimization algorithm (Multi-Agent Proximal Policy Optimization, MAPPO).

[0040] The implementation of the MAPPO algorithm includes:

[0041] Step 41: Initialize the policy network of the i-th UAV and the value network model parameters;

[0042] Step 42: Each UAV samples an action according to the policy network and executes it, calculates the reward and updates the global state s[n];

[0043] Step 43: Calculate the population curriculum error Update and standardize the weight coefficients

[0044] Step 44: Calculate the global advantage estimation of the i-th UAV globally where is the time difference error and γ is the discount factor. Optimize the policy network objective function for the m-th adjustment of the UAV Perform the update of the policy network parameters.

[0045] Step 45: Minimize the loss function of the value network in the m-th adjustment Perform the update of the value network parameters;

[0046] Step 46: Repeat steps 41 to 45 until the maximum number of training rounds M is completed.

[0047] Furthermore, to ensure the feasibility of the actions output by the policy network, each UAV policy network generates actions in the following way: The network adopts a multi-layer perceptron structure, the LeakyReLU function is used for activation between layers, the Tanh function is used as the activation function for the output layer to output the mean of continuous actions, which is linearly mapped to the physically feasible range, and the straight-through estimator is used to map the associated variables to binary values. The heading angle and position update follow the discrete-time kinematic equations.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows: Under the proposed integrated communication and sensing method for multiple UAVs, the collaborative cooperation among multiple UAVs serves multiple mobile ground nodes and obtains communication and sensing coordination gains. The trajectory planning and node association strategy of multiple UAVs achieve adaptive adjustment to node movement. The minimum node throughput of the integrated communication and sensing system is improved. Multiple UAVs cooperate to perform tasks, and the total system throughput is greatly improved compared with the single-UAV scenario. At the same time, this method avoids the need to retrain the model due to environmental changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments.

[0050] Figure 1 It is the network structure diagram of the integrated communication and sensing system for multiple UAVs in Embodiment 1 of the present invention.

[0051] Figure 2 It is the training flow chart of the integrated communication and sensing system for multiple UAVs in Embodiment 1 of the present invention.

[0052] Figure 3 It is the graph of the total system throughput and the minimum ground node throughput of the system varying with the number of UAVs for the task in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of the present technology to understand the purpose, technical solutions, and advantages of the present invention. It should be clear, however, that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0054] As mentioned in the background art, in a dynamic scenario, the movement of ground nodes and the uncertainty of the environment make the requirements for real-time performance and high accuracy of sensing tasks more stringent. The present invention considers a dynamic scenario involving a wide distribution of multiple ground nodes, realizes the collaborative optimization of sensing fairness and age of information, and ultimately maximizes the system throughput.

[0055] Embodiment 1

[0056] In this example, in a communication-sensing integrated system assisted by unmanned aerial vehicles (UAVs), 3 UAVs serve 5 dynamic ground nodes and 25 static ground nodes. The horizontal coordinates of the starting positions of the three UAVs are q1[0] = (1000, 500) T , q2[0] = (750, 500) T , q3[0] = (1250, 500) T . The duration of the entire communication-sensing integrated task is 40 seconds, and the duration of each time slot is τ = 0.2 seconds. The flight altitude of all UAVs in the task scenario is fixed at h = 200 meters. In the multi-UAV communication-sensing integrated system, the node distribution range is wider, and the horizontal flight range of the UAVs to execute tasks is [x min , x max = [0, 2000] meters, [y min , y max = [0, 2000] meters. The maximum speed of the UAVs is v max = 50 m / s. The maximum change in the flight angle within each time slot is Δθ max = π / 6 radians. The maximum transmit power is 23 dBm, the total bandwidth B = 100 MHz. In the sensing system, the transmit antenna gain Γ TX = 13 dBi, the receive antenna gain Γ RX = 13 dBi. Without loss of generality, the radar cross-section σ k = 0.3 m 2 , the Boltzmann constant k = 1.38×10 -23 J / K, the thermal noise temperature T0 = 290 K, the noise factor F = 5 dB, and the detection loss Ψ = -6 dB. In the path loss modeling of the communication system, the path loss factor is set to β0 = 10 -5 , and the noise power is σ2 = 10 -10 Watts.

[0057] See Figure 1 , which is the network structure of the integrated communication and sensing system for multi-UAVs. This figure presents the architecture of the reinforcement learning decision-making system for multi-UAV cooperation in this example. Both the policy network (also known as the action network actor) and the value network are composed of four fully connected layers. The architectural structures of the two networks are similar, including a multi-layer perceptron design. The first fully connected layer contains 64 neurons, while each subsequent layer contains 128 neurons. The activation function in the network is selected as LeakyReLU. Both update the parameters and gradients bidirectionally through the optimizer to form a dynamic policy optimization closed-loop. The action network of UAV i receives the observation o i [n] as input, and after passing through the above action network structure, outputs the mean μ(o i [n]) and standard deviation σ(o i [n]) of the multivariate Gaussian distribution under the observation o i [n]. Then, UAV i samples from the above multivariate Gaussian distribution to obtain the action a i [n] at time slot n. The value network of UAV i receives the state s i [n] as input. Since in the centralized training framework, s[n] = s i [n], after passing through the above value network structure, it outputs the value evaluation Calculate the global advantage estimation when the i-th UAV takes the action a i [n] after the m-th adjustment Then, the optimizer of the action network of UAV i updates the network parameters according to the objective function Execute. For every 4 sets of data of the integrated communication and sensing tasks for multi-UAVs (corresponding to a batch size of 800), the neural network performs 15 parameter updates for training. The action a i [n] is sampled from the multivariate normal distribution based on the policy network, and the standard deviation σ(o i [n]) of each action distribution is directly fitted by the policy network.

[0058] See Figure 2 , the present invention provides a method for integrated communication and sensing of multi-UAVs. This method includes steps 1-5, and the following will specifically describe each step in detail in combination with specific embodiments.

[0059] Step 1: Construct a multi-UAV communication and sensing integrated system model, including γ UAVs and K ground nodes; each UAV provides communication services for multiple ground nodes through orthogonal frequency division multiple access technology and realizes distance measurement by sending sensing signals. Construct the throughput optimization problem model of the above multi-UAV communication and sensing integrated system, including the following constraints and optimization objectives:

[0060] (1) UAV flight area constraint, heading angle change rate constraint, maximum speed and anti-collision distance constraint

[0061] The communication and sensing integrated task of the UAV serves K ground nodes and lasts for N time slots, numbered n = 0, …, N - 1. The horizontal coordinate of the i-th UAV at time slot n is q i [n]=(x i [n],y i [n]) T , and the UAV flight area constraint is: x min ≤x i [n]≤x max ,y min ≤y i [n]≤y max , x min is the minimum allowable value of the UAV horizontal coordinate in the x-axis direction, x max is the maximum allowable value of the UAV horizontal coordinate in the x-axis direction; y min is the minimum allowable value of the UAV horizontal coordinate in the y-axis direction, y max is the maximum allowable value of the UAV horizontal coordinate in the y-axis direction.

[0062] The UAV controls its flight trajectory during the mission by controlling its speed and heading angle.

[0063] Let v i [n]≤v max represent the speed of the i-th UAV at time slot n, and the heading angle θ i [n]∈(0,2π], where v max represents the maximum speed of the UAV. The travel distance l i [n]=∥q i [n]-q i [n - 1]∥≤v max τ; To ensure the stability of the UAV flight, the heading angle change rate constraint is introduced: |θ i [n]-θ i [n - 1]|≤Δθ max , where Δθ maxDenote the maximum instantaneous steering angle corresponding to the servo response limit of the UAV flight control system. In the scenario of multi-UAV collaborative operation, it is necessary to introduce the group safety isolation condition. For UAV i and UAV i'; satisfy ∥q i [n] - q i′ [n]∥≥d min , where d min represents the distance threshold set to avoid UAV collisions.

[0064] (2) Total transmission power upper limit and unique ground node service constraint

[0065] In each time slot n, let P i [n] represent the total transmission power of UAV i, which is constrained by the maximum power . The total power is allocated to the communication power and the sensing power . The two are divided by the power allocation coefficient ρ i [n]. Among them, the power allocated to communication is The power allocated to sensing is Satisfy

[0066] By introducing a joint resource allocation strategy in the three-dimensional space domain and time-frequency domain, a multi-dimensional mapping relationship between the UAV and the ground node is established to ensure that the ground node is served by only one UAV in the same time slot. Specifically, define the binary decision variable χ i,k [n] ∈ {0, 1} to represent the association variable between UAV i and the k-th ground node. The variable being 1 indicates their association, otherwise 0. Satisfy the unique ground node service constraint: where γ represents the total number of UAVs.

[0067] (3) Detection variance threshold constraint based on signal-to-noise ratio for sensing quality, sensing fairness, and information age constraint The signal-to-noise ratio of the single-static radar system equipped on the UAV is:

[0068]

[0069] where Γ TX and Γ RX represent the transmitting and receiving antenna gains respectively; λ represents the signal wavelength; σ k represents the radar cross-sectional area of the target; the physical constants k, T0, B s [n], and Ψ represent the Boltzmann constant, the receiver equivalent noise temperature, the signal bandwidth, the noise factor, and the detection loss respectively.

[0070] The device distance estimation performance is only affected by the environmental noise in the form of an additive white Gaussian random process. The Cramér-Rao lower bound is used to evaluate the performance of distance detection, and the signal-to-noise ratio ξ is obtained. i,k [n] satisfies:

[0071] where c is the speed of light, v S is the number of sampling samples, the carrier frequency of the i-th drone, is the detection variance threshold, is the sensing bandwidth of the i-th drone, β i [n] is its sensing bandwidth allocation coefficient.

[0072] The sensing fairness metric is inspired by the Jain fairness metric. The sensing fairness

[0073] where ψ i,k [n] represents the latest valid sensing time of drone i for ground node k, ψ i,k [n] is updated to the latest time when the sensing detection variance constraint is satisfied, and Λ2 is the threshold to ensure sensing fairness.

[0074] Next, we evaluate the freshness of the sensing data through the age of information. The age of information of the sensing information of ground node k obtained by drone i:

[0075]

[0076] where ∈ → 0 + is a small constant used to prevent division by zero. By optimizing the age of information, it is ensured that the drones use the latest sensing data as much as possible, thereby improving the execution efficiency and accuracy of the tasks.

[0077] (4) The main objective of the designed integrated communication and sensing system for multiple drones is to maximize the downlink communication service throughput of the system by jointly optimizing the trajectory planning, ground node association, power allocation, and spectrum resource allocation of multiple drones while satisfying the given constraints, so as to support the simultaneous access and service of more ground nodes. Therefore, the system objective function is:

[0078]

[0079] where, represents the aggregated transmission rate obtained by ground node k, R i,k [n] represents the transmission rate between drone i and ground node k, σ 2 is the noise power, and the interference term in the formula represents the co-channel interference caused by spectrum reuse in the multi-drone system, Denote the communication sub-channel bandwidth obtained by a single ground node, L i,k [n] = β0(d i,k [n]) -2 is the communication channel gain between the UAV i and the ground node k, and β0 is the path loss factor. Denote the distance between the UAV i and the ground node k, and h represents the cruising altitude of the UAV.

[0080] Step 2: Model the communication sensing task as a partially observable Markov decision process, define the joint state space, local observation space, action space, and multi-objective reward function, and obtain states, observations, and decisions through training episodes, and calculate the rewards. The multi-objective reward function includes the following items:

[0081] (1) Boundary control reward

[0082] Each UAV needs to fly within the task boundary to ensure the effectiveness of the task. For the characteristics of the rectangular task area, the boundary reward function is

[0083]

[0084] where: The center coordinates q of the task area c =(x c , y c ), the rectangular boundary size L x ×L y , L x =x max -x min , L y =y max -y min , the boundary deviation where

[0085] By punishing the behavior of the UAV deviating from the task boundary, the effective execution of the task can be guaranteed, and the safety and efficiency of the UAV can be ensured during the task process.

[0086] (2) Collision avoidance reward between UAVs

[0087] In the multi-UAV cooperative flight task, collision avoidance is an important goal to ensure the smooth completion of the task. To achieve this goal, we design a collision avoidance reward. where N pairs =Υ(Υ - 1) / 2 is the number of UAV pairs, d safe =2d minLet \(d_0\) be the safety distance threshold. The design purpose of this reward function is to encourage the UAV to maintain a sufficient safety distance during flight and avoid collisions.

[0088] (3) Distributed age-of-information reward

[0089] In the case of multiple UAVs, each UAV may be responsible for different nodes, so the reward needs to be distributed to each UAV. Therefore, the correlation variable \(\chi_{i}[n]\) is introduced i,k to distinguish the nodes served by UAV \(i\). The distributed age-of-information reward for UAV \(i\) is:

[0090]

[0091] where \(\psi_{k}[n] =\) k \(\min\) i \(\psi\) i,k \([n]\) represents the comprehensive latest effective perception time of ground node \(k\); \(\epsilon\) is a small constant, \(\epsilon\rightarrow0\) + ;

[0092] The responsibility of each UAV is divided through the correlation variable. This reward design ensures high scalability and will not become computationally complex due to the increase in the number of UAVs.

[0093] (4) Perception fairness reward

[0094] In a multi-UAV system, each UAV has a different coverage area, so an independent fairness index needs to be designed for each UAV to avoid some nodes being ignored. In addition, the perception fairness reward needs to be balanced at the whole system level and the single UAV level. The perception fairness reward is where represents the global perception fairness. This reward combines the local fairness of UAV \(i\) with the global fairness to balance the individual and group optimization goals.

[0095] (5) Throughput reward and minimum throughput reward

[0096] To avoid numerical inconsistencies caused by scale differences, the system throughput is normalized. The normalized throughput contribution of UAV \(i\) is defined as where \(L\) max is the maximum channel gain, which is determined by the minimum vertical distance \(h\) between the UAV and the ground node, corresponding to the free space path loss when the UAV hovers directly above the ground node. To ensure that each ground node in the system, especially those at the edge, can meet the basic communication quality requirements, for the set of associated nodes \(K_{i}[n]=\{k|\chi_{i}[n] = 1\}\) of UAV \(i\), its minimum throughput reward is defined as i [n] = {k|χ i,k [n] = 1}, and its minimum throughput reward is defined as This reward guarantees the service quality of edge nodes and drives UAV i to preferentially improve the performance of edge nodes.

[0097] (6) Node association reward

[0098] Through the exponential penalty mechanism, it is ensured that only the ground node k is associated with one UAV. The ground node association reward of UAV i is This design satisfies the following mathematical properties: when and only when all associated nodes of UAV i are uniquely served, the cumulative penalty amount in the exponential term is zero, and the reward reaches the maximum value of 1. Each conflicting ground node introduces (N link -1) 2 type penalty, where N link is the actual number of associated UAVs of the ground node, ensuring that multiple association conflicts cause superlinear attenuation; at the same time, it ensures non-activity neutrality, that is, when the UAV is not associated with any node, the reward is always 1, avoiding ineffective penalties for inactive UAVs.

[0099] Step 3: According to the designed adaptive reward weight adjustment method, calculate the population curriculum error at the m-th adjustment:

[0100]

[0101] where r j,i [n] represents the j-th reward obtained by UAV i at time slot n; m′ represents the number of training rounds between the (m - 1)-th adjustment and the m-th adjustment;

[0102] Quantify the deviation between each reward item and the target performance through:

[0103]

[0104] Dynamically update the reward weight coefficient at the m-th adjustment, where D j is the initial difficulty coefficient, E is the preset error upper limit, and κ is the adaptive coefficient. Then through Normalize the weights and synchronize the weight coefficients adjusted centrally during the training phase to all UAV policy networks to ensure behavioral consistency during decentralized execution.

[0105] Step 4: Based on the centralized training - decentralized execution framework, use the multi-agent proximal policy optimization algorithm to optimize the policy network for multi-agent cooperation, calculate the global advantage estimation, and perform gradient descent to update the policy network and value network; specifically, the implementation of the multi-agent proximal policy optimization algorithm includes:

[0106] Step 41: Initialize the policy network and value network model parameters of the i-th UAV;

[0107] Step 42: Each drone samples an action according to the policy network:

[0108]

[0109] And executes it, calculating the reward:

[0110]

[0111] And updates the global state s[n];

[0112] Step 43: Calculate the population curriculum error Update and normalize the weight coefficients

[0113] Step 44: Calculate the global advantage estimate of the i-th drone globally:

[0114]

[0115] Where is the temporal difference error, and γ is the discount factor; is the value evaluation of the system state at the (n + 1)-th time slot by the value network of drone i after the m-th adjustment; is the value evaluation of the system state at the n-th time slot by the value network of drone i after the m-th adjustment.

[0116] Then optimize the policy network objective function of this drone in the m-th adjustment Execute the policy network parameter update;

[0117]

[0118] Where the probability ratio is the probability that the policy network of drone i outputs the action a i [n] according to the observation o i [n] in the iterative process of the current adjustment and the probability that the policy network of drone i outputs the action a i [n] according to the observation o i [n] after the m-th adjustment of the ratio; is the global advantage estimate of the i-th drone after the m-th adjustment; the clip function is used to clip the probability ratio to the range [1 - ε0, 1 + ε0], where ε0 is the clip coefficient; is the entropy regularization term; μ is the entropy regularization coefficient.

[0119] Step 45: Minimize the loss function of the value network in the m-th adjustment:

[0120]

[0121] Execute the value network parameter update; where represents the value evaluation of the system state at the nth time slot by the value network of the UAV i in the currently adjusted iteration process.

[0122] Step 46: Repeat Steps 41 to 45 until the maximum number of training rounds M is completed.

[0123] Step 5: Repeat Steps 1 - 4 until all training rounds are completed, and experimentally verify the impact of the number of UAVs on the system performance, so as to achieve the collaborative trajectory planning and communication resource scheduling of multiple UAVs.

[0124] Through the above method, this embodiment can achieve efficient trajectory planning for multiple UAVs in a ground node scenario with a wide distribution range, and can quickly respond to the dynamic movement of ground nodes. During the trajectory planning process, the collaborative cooperation and adaptive adjustment of the multi - UAV system are realized, achieving an impressive system throughput performance. Figure 3 It is a graph showing the changes in the total system throughput and the minimum ground node throughput of the communication - perception integrated system with the number of UAVs performing tasks and the progress of the training rounds. As the number of UAVs increases, the system throughput increases significantly, but the growth gradually slows down, showing a marginal effect. The minimum throughput of a single UAV is 5 Mbit, which increases to 26.3 Mbit (5.3 times) when there are 2 UAVs, and reaches 43.7 Mbit (8.7 times) when there are 3 UAVs. The total system throughput increases from 607.4 Mbit for a single UAV to 2104.6 Mbit (3.5 times) for 2 UAVs, and reaches 3428.5 Mbit (5.6 times) for 3 UAVs. Multiple UAVs improve efficiency through task assignment and spectrum reuse, but the throughput growth rate slows down when there are 4 UAVs, and the training time is extended to 2.5 times that of 3 UAVs. The training stability decreases as the number of UAVs increases. The single - UAV system tends to be stable after 2000 rounds, while the 3 - UAV system still fluctuates significantly after 8000 rounds, with a significant increase in complexity. The system action and state space grow as the Cartesian product, and longer training rounds and adaptive strategies are required to ensure stability. In summary, increasing the number of UAVs improves the throughput, but the marginal effect is obvious, and the training complexity and volatility increase. It is necessary to set according to the requirements of ground nodes, comprehensively considering throughput and complexity. Therefore, 3 UAVs are used as the task standard setting in this embodiment. Experiments show that in the considered dynamic scenario involving multiple ground nodes with a wide distribution, the present invention can achieve the collaborative optimization of sensing fairness and age of information, and finally obtain an impressive system throughput performance.

[0125] The above technical details and the step - by - step process of algorithm implementation are only used as demonstration examples to better illustrate the method proposed by the present invention, and should not be construed as a limitation of the present invention. Other researchers in the field can perform deformation and recombination within the scope of the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A multi-UAV communication and sensing integration method based on adaptive reward weight adjustment, characterized in that The method includes: Multiple drones equipped with communication and sensing integrated modules provide downlink communication services for a group of ground nodes. At the same time, the drones send sensing signals to the ground nodes for ranging and observe and collect environmental data. Based on the sensing data and environmental data, multiple drones use the trained multi-agent proximal policy optimization (MAPPO) model for real-time decision-making. Among them: The training process of the multi-agent proximal policy optimization (MAPPO) model includes: Step 1: Construct a multi-drone communication and sensing integrated system model: including Υ drones and K ground nodes; each drone provides communication services for multiple ground nodes through orthogonal frequency division multiple access technology, and realizes distance measurement by sending sensing signals, and constructs an optimization problem for the throughput of the multi-drone communication and sensing integrated system. Step 2: Model the communication and sensing tasks as a partially observable Markov decision process, define the joint state space, local observation space, action space and multi-objective reward function, obtain states, observations, decisions through training episodes, and calculate rewards. Step 3: According to the adaptive reward weight adjustment method, dynamically adjust the weight coefficients of each reward term through curriculum error and standardize the weight coefficients. Step 4: Based on the centralized training - decentralized execution framework, use the multi-agent proximal policy optimization algorithm to collaboratively optimize the policy network of multi-agent cooperation, calculate the global advantage estimation, and perform gradient descent to update the policy network and value network. Step 5: Repeat Steps 1 - 4 until all training episodes end, and experimentally verify the impact of the number of drones on the system performance to achieve the collaborative trajectory planning and communication resource scheduling of multiple drones.

2. The integrated method for communication and sensing of multiple UAVs according to claim 1, characterized in that In Step 1, the constructed multi-drone communication and sensing integrated system model satisfies the following constraints: (1) Drone flight area constraint, heading angle change rate constraint, maximum speed and anti-collision distance constraint; (2) Total transmit power upper limit and unique service constraint for ground nodes; (3) Detection variance threshold constraint based on signal-to-noise ratio for sensing quality, sensing fairness and age of information constraint.

3. The integrated communication and sensing method for multiple UAVs according to claim 2, wherein, In constraint condition (2), In each time slot n, let P i [n] denote the total transmission power of UAV i, which is subject to the maximum power constraint. The total power is allocated to the communication power and the sensing power which are divided by the power allocation coefficient ρ i [n], satisfying By introducing a joint resource allocation strategy in the three-dimensional space domain and time-frequency domain, a multi-dimensional mapping relationship between the UAV and the ground nodes is established to ensure that each ground node is served by only one UAV in the same time slot. Specifically, a binary decision variable χ i,k [n] represents the association variable between the i-th UAV and the k-th ground node, x i,k [n] ∈ {0, 1}, where the variable being 1 indicates their association, otherwise 0, satisfying the unique service constraint for the ground nodes: where γ represents the total number of drones.

4. The integrated method for communication and sensing of multiple UAVs according to claim 2, wherein In the constraint condition (3), let the signal-to-noise ratio of the monostatic radar system equipped on the UAV be ξ i,k [n], the device distance estimation performance is affected by the environmental noise in the form of an additive white Gaussian random process. The Cramér-Rao lower bound is used to evaluate the performance of distance detection, and the signal-to-noise ratio ξ i,k [n] satisfies: Where: c is the speed of light, v S is the number of sampling samples, the carrier frequency of the i-th drone, is the detection variance threshold, is the sensing bandwidth of the i-th drone, β i [n] is its sensing bandwidth allocation coefficient; The sensing fairness metric is inspired by the Jain fairness measure. The sensing fairness is as follows: where ψ i,k [n] represents the latest valid sensing time of UAV i for ground node k, and ψ i,k [n] is updated to the latest time when the sensing detection variance constraint is satisfied. Λ2 is the threshold to ensure sensing fairness, and χ i,k [n] represents the association variable between UAV i and the k-th ground node, and χ i,k [n] ∈ {0, 1}; The freshness of the sensing data is evaluated through the age of information. The age of information of the sensing information of ground node k obtained by drone i: where ∈ → 0 + , where ∈ is a small constant used to prevent division by zero, and Λ3 is a threshold to ensure the age of information.

5. The integrated method for communication and sensing of multiple unmanned aerial vehicles according to claim 2, wherein While satisfying the given constraints, the multi-drone communication and sensing integrated system model maximizes the downlink communication service throughput of the system through the collaborative optimization of the trajectory planning, ground node association, power allocation, and spectrum resource allocation of multiple drones, so as to support the simultaneous access and service of more ground nodes. The objective function of the multi-drone communication and sensing integrated system model is: where θ i [n] represents the heading angle of the i-th UAV at time slot n; l i [n] represents the travel distance of the i-th UAV between two consecutive time slots; χ i,k [n] represents the association variable between the drone i and the k-th ground node; ρ i [n] represents the power distribution coefficient; β i [n] is the sensing bandwidth allocation coefficient of UAV i; τ represents the slot duration; R k [n] represents the aggregated transmission rate obtained by ground node k; Among them, R i,k [n] represents the transmission rate between the UAV i and the ground node k, and σ 2 is the noise power. In the formula, the interference term represents the co-channel interference caused by spectrum reuse in the multi-UAV system, represents the communication sub-channel bandwidth obtained by a single ground node, and L i,k [n] is the communication channel gain between the UAV i and the ground node k, and β0 is the path loss factor.

6. The integrated communication and sensing method for multiple UAVs according to claim 1, wherein In Step 2, the multi-objective reward function includes: boundary control reward, collision avoidance reward between drones, distributed age of information reward, sensing fairness reward, throughput reward and minimum throughput reward, node association reward.

7. The integrated method for communication and sensing of multiple unmanned aerial vehicles according to claim 6, wherein The distributed age of information reward for drone i is: x i,k [n] represents the association variable between the drone i and the k-th ground node; ψ k [n] = min i ψ i,k [n] represents the comprehensive latest effective sensing time of the ground node k; ∈ is a small constant, ∈ → 0 + ; The sensing fairness reward is: where, ψ i,k [n] represents the latest effective sensing time of UAV i for ground node k, F global [n] represents the global sensing fairness, The normalized throughput reward for drone i is: where L max is the maximum channel gain, determined by the cruising altitude h of the UAV, corresponding to the free space path loss when the UAV hovers above the ground node; R i,k [n] represents the transmission rate between the UAV i and the ground node k, P max represents the maximum power of the UAV, σ 2 is the noise power; The set of associated nodes K for the UAV i i [n] = {k|χ i,k [n] = 1}, and the minimum throughput reward is defined as:

8. The integrated method for communication and sensing of multiple unmanned aerial vehicles according to claim 1, wherein, In step 3, the adaptive reward weight adjustment method includes: First, calculate the population course error at the m-th adjustment where r j,i [n] represents the j-th reward obtained by the UAV i at time slot n; m ′ represents the training rounds between the (m - 1)-th adjustment and the m-th adjustment; Quantify the deviation between each reward item and the target performance, and dynamically update the reward weight coefficient for the m-th adjustment Among them, D j is the initial difficulty coefficient, E is the preset error upper limit, and κ is the adaptive coefficient, is the reward weight coefficient for the (m - 1)-th adjustment; Then normalize the weights: Finally, the weight coefficients adjusted during the training phase are synchronized to all UAV policy networks to ensure consistent behavior during decentralized execution.

9. The integrated communication and sensing method for multiple UAVs according to claim 8, wherein, In step 4, the implementation of the MAPPO algorithm includes: Step 41: Initialize the policy network and value network model parameters of the i-th drone and value network model parameters; Step 42: Each UAV samples an action according to the policy network and executes it: Where: θ i [n] represents the heading angle of the i-th UAV at time slot n, l i [n] represents the travel distance between two consecutive time slots of the i-th UAV, ρ i [n] represents the power allocation coefficient, β i [n] is the sensing bandwidth allocation coefficient of UAV i, x i,k [n] represents the association variable between UAV i and the k-th ground node Calculating rewards And update the global state s[n]; Step 43: Calculate the group course error Update and standardize the weight coefficients Step 44: Calculate the global advantage estimate of the i-th UAV globally: where is the time difference error, γ is the discount factor; is the value evaluation of the value network of UAV i for the system state at the (n + 1)-th time slot after the m-th adjustment; is the value evaluation of the value network of UAV i for the system state at the n-th time slot after the m-th adjustment; Optimize the objective function of the policy network for the drone in the m-th adjustment to perform the update of the policy network parameters: Among them, the probability ratio is the probability that the policy network of UAV i outputs action a i [n] according to the observation o i [n] in the current iteration of adjustment; and the probability that the policy network of UAV i outputs action a i [n] according to the observation o i [n] after the m-th adjustment; The ratio; is the global advantage estimate of the i-th UAV after the m-th adjustment; the clip function is used to clip the probability ratio to the range [1 - ∈0, 1 + ∈0], where ∈0 is the clip coefficient; is the entropy regularization term; μ is the entropy regularization coefficient; Step 45: Minimize the loss function of the value network in the m-th adjustment Perform the update of the value network parameters: Among them, represents the value evaluation of the system state at the nth time slot by the value network of UAV i during the current adjustment iteration process; Step 46: Repeat steps 41 to 45 until the maximum number of training episodes M is completed.

10. The integrated method for multi-UAV communication and sensing according to claim 1, wherein The policy network of the UAV generates actions in the following way: The network adopts a multi-layer perceptron structure, with the LeakyReLU function used for activation between layers, and the Tanh function used as the activation function for the output layer to output the mean of continuous actions, which is linearly mapped to the physically feasible range. The straight-through estimator is used to map the associated variables to binary values, and the heading angle and position update follow the discrete-time kinematic equations.

Citation Information

Cited By

  • Unmanned aerial vehicle multi-altitude network cooperation method and system based on low-altitude airspace operation risk

    CN121433289A