Joint resource allocation and trajectory design method and system based on deep reinforcement learning

By optimizing the flight trajectory and power allocation of UAVs through deep reinforcement learning models, the complexity of resource allocation and trajectory design in UAV-assisted multi-cell ISAC systems is solved, achieving efficient communication and perception services, reducing energy consumption and improving user fairness.

CN122120842APending Publication Date: 2026-05-29GUANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGXI NORMAL UNIV
Filing Date
2026-01-20
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In UAV-assisted multi-cell ISAC systems, resource allocation and trajectory design are highly coupled and have high optimization complexity, making it difficult to meet the requirements of communication and sensing service quality. At the same time, deployment flexibility and cost are high.

Method used

A joint resource allocation and trajectory design method based on deep reinforcement learning is adopted. The flight trajectory and power allocation of the UAV are optimized by using a deep reinforcement learning model. User scheduling and power allocation are combined for unified modeling and joint optimization. The DDPG framework is used to process the continuous action space.

Benefits of technology

This technology enables UAVs to adaptively adjust their flight paths in dynamic environments, meeting the requirements of communication service quality and perception accuracy while minimizing the overall system transmission power, thereby improving the overall system throughput and user fairness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120842A_ABST
    Figure CN122120842A_ABST
Patent Text Reader

Abstract

The application provides a joint resource allocation and trajectory design method and system based on deep reinforcement learning, which is suitable for a multi-cell integrated sensing and communication system assisted by a UAV. The method comprises the following steps: initializing a starting position of the UAV, communication and sensing system parameters, and a deep reinforcement learning model; scheduling users of multi-cell base stations under the condition of a given UAV position; allocating transmission power of each communication node under the current user scheduling result to meet the communication performance constraint and the sensing accuracy constraint; outputting a motion control strategy of the UAV by the deep reinforcement learning model based on the power allocation result and system state information, updating the flight trajectory of the UAV; and judging whether the joint optimization process converges or not, and outputting the optimization result if the joint optimization process converges, and then continuing the iterative optimization. The application alleviates the coupling problem between resource allocation and trajectory design in a multi-cell ISAC scenario, and improves the adaptive ability and comprehensive performance of the system in a dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field] This invention relates to the field of communication technology, and in particular to a method and system for joint resource allocation and trajectory design based on deep reinforcement learning. [Background Technology] The low-altitude economy refers to economic activities that rely on low-altitude airspace (typically below 1000 meters) and integrate unmanned aerial vehicles (UAVs), communications, sensing, and artificial intelligence. In recent years, it has become a significant driver of digital and intelligent services. Unlike traditional high-altitude aviation, the low-altitude economy targets short-range, low-speed, and diverse applications. Typical application scenarios include urban air mobility, emergency rescue, precision agriculture, environmental monitoring, and smart logistics. For example, in emergency rescue, responders need to maintain reliable communication with the command center while simultaneously acquiring terrain and location information through sensing to support timely decision-making. However, traditional architectures implement communication and sensing separately. This separate design suffers from low resource utilization, high system complexity, and significant information exchange latency, making it difficult to meet the needs of integrated services. To address these limitations, Integrated Sensing and Communication (ISAC) has emerged as a promising technology for achieving high-reliability communication and high-precision sensing. With the rapid development of the low-altitude economy, more and more low-altitude operational scenarios are emerging. These scenarios present an urgent need to simultaneously meet the requirements of integrated communication and sensing. In terms of communication, this includes data transmission between UAVs, ground base stations (TBS), and control centers. In terms of perception, it involves environmental perception, obstacle detection, and target tracking.

[0001] Unlike traditional systems that focus solely on communication or sensing capabilities, ISAC systems emphasize the synergistic integration of both. To jointly evaluate communication and sensing performance, performance metrics such as the Cramer-Rao lower bound (CRLB) and signal-to-interference-plus-noise ratio (SINR) are typically used. The limitations of single-cell ISACs have spurred research into multi-cell ISAC systems, where multiple fixed terrestrial base stations (TBS) work collaboratively to improve coverage and overall system performance. Recent advancements in multi-cell ISACs demonstrate their ability to support user communication while simultaneously detecting moving targets. Furthermore, multi-cell ISACs can facilitate collaborative sensing and information exchange across multiple base stations.

[0002] Despite the potential for improved Quality of Service (QoS), the TBS collaborative service ISAC itself still faces challenges in deployment flexibility and high costs. Specifically, fixed TBSs are difficult to configure for temporary or emergency scenarios, requiring rapid expansion of coverage, while deploying new TBSs incurs significant time and financial costs. As a key player in the low-altitude economy, drones possess unique advantages such as high mobility, flexible deployment, and wide coverage, offering a unique solution to overcome the limitations of fixed TBSs. Unlike fixed TBSs, drones can be dynamically deployed as "flying base stations," reaching cell edge areas and adjusting altitude and location to optimize communication coverage and sensing performance. Therefore, researching drone-assisted multi-cell ISAC systems is of great significance. By leveraging the mobility of drones and the stability of TBSs, such systems can achieve complementary advantages, providing continuous, high-quality integrated communication and sensing services for the low-altitude economy.

[0003] Despite the aforementioned advantages, implementing UAV-assisted multi-cell ISAC systems still faces several challenges: (a) Given the stringent wireless communication and computational resources required for UAV-assisted multi-cell ISAC systems, minimizing the power consumption of both the UAV and the TBS is crucial. This necessitates the development of efficient resource allocation algorithms to ensure that users' communication and sensing quality of service (QoS) can be guaranteed with moderate overhead. (b) Unlike traditional single-base station or multi-base station TBS networks, UAV-assisted multi-base station ISAC systems require joint consideration of air-to-ground BS cooperation, communication and sensing QoS guarantees, and UAV trajectory planning. These factors result in tightly coupled and highly non-convex decision variables, which are difficult to handle using traditional convex optimization methods alone. [Summary of the Invention] To overcome the above problems, this invention proposes a joint resource allocation and trajectory design method and system based on deep reinforcement learning that can effectively solve the above problems.

[0004] The present invention provides a technical solution to the above-mentioned technical problems: a joint resource allocation and trajectory design method based on deep reinforcement learning, applicable to UAV-assisted multi-cell sensor integration (ISAC) systems, comprising the following steps: Step 1, Initialization Steps: Initialize the UAV's starting position, communication and perception system parameters, and deep reinforcement learning model; Step 2, User Scheduling Step: Given the location of the UAV, schedule the communication users of the base station according to the multi-cell communication environment to determine the set of users participating in the current time slot service; Step 3, Power Allocation Step: Based on the current user selection results, the transmission power of each communication node is allocated to meet the communication performance constraints and sensing accuracy constraints. Step 4, Trajectory Update Step: Based on the power allocation results and system state information, the deep reinforcement learning model outputs the motion control strategy of the UAV and updates the UAV flight trajectory accordingly. Step 5, Convergence Judgment Step: Determine whether the reward curve of the joint optimization process has converged. If it has, output the optimized UAV flight trajectory, user scheduling scheme and power allocation results; otherwise, return to Step 2 and continue iterative optimization.

[0005] Assuming the drone base station is at a fixed altitude over the target area Flight, the total duration of each flight mission is For ease of optimization and analysis, continuous time periods will be used. Uniformly discretized There are 1 time slot, and the length of each time slot is 1. By selecting a sufficiently large It can make It is small enough to ensure that the displacement of the UAV in each time slot is negligible compared to its air-to-ground (A2G) communication distance; this approximation allows the UAV's ground channel to be considered quasi-static in each time slot and assumes that the latency requirements of the sensing task are met. make Represents the user scheduling matrix, where Indicates in time slot user Scheduled Represents the power allocation matrix. Let the horizontal trajectory matrix of the UAV be represented; in this case, the optimization problem of the present invention can be expressed as: P1: (1) st (1a) ; (1b) ; (1c) ; (1d) ; (1e) ; (1f) . here, Let represent a vector of all ones with the corresponding dimension, and (1a) Ensure that the transmit power of each base station is non-negative in each time slot. (1b) and (1c) constrain the number of communication users scheduled in each time slot; L represents the number of base stations. In addition, considering user fairness, (1d) limits the maximum number of time slots scheduled to each user. The constraints in (1e) ensure that the SINR requirements of communication users are met. (1f) Ensure that the CRLB of the target location estimation does not exceed the threshold. .

[0006] The user scheduling and power allocation problem can be viewed as the inner part of a two-layer algorithm. For any given drone starting point, the user scheduling problem can be formulated as: P2: (2) st (1a)-(1d) Each base station is assumed to be randomly distributed within its coverage area. I. Candidate communication users; Lemma 1 is proposed below; Lemma 1: In a UAV-assisted multi-cell ISAC system, when the sensing constraints are relatively relaxed, the total transmission power consumption decreases monotonically with the SCG of each communication user. prove: For the UAV-assisted multi-cell ISAC system under consideration, when the CRLB constraint is relatively relaxed, the power consumption is mainly determined by the communication requirements. In this case, the base station's transmit power monotonically decreases with the increase of SCG, as shown in the following proof: Assume it exists One communication user ,as well as There are 1 base stations, of which ( ) is TBS; Indicates from base station To communication users SCG, Transmission power is denoted as ; in each time period, Therefore, the constraints in (1f) can be equivalently restated using the following matrix representation: .

[0007] The constraints can then be restated as the following matrix:

[0008] here, ,and Represents an invertible square matrix composed of SCG coefficients; the total transmitted power is relative to The partial derivatives are:

[0009] The inverse function can be expressed as , Indicates and The associated adjoint matrix, Corresponding to its determinant; it can then be restated as:

[0010] in Indicates that it is located at The row and number Column elements, It is worth noting that, The diagonal entries correspond to the SCG of the communication user, while the off-diagonal elements are composed of user interference and The product is obtained by using; Given a specific structure, we obtain:

[0011] in Indicates the corresponding base station SCG serving users; positive item and Indicates the user's position in the first month. The product of channel gain and interference in each time slot; it is worth noting that for the base station For users of the service, direct SCG Typically, the cross-channel gain is much greater than that of other base stations; therefore, we have This means Similarly, using a similar methodology, we can also prove... ; Therefore, we conclude that: ; According to Lemma 1, the system transmission power continuously decreases as the communication user SCG increases. Based on this, users can schedule according to their SCG. Based on (P1), to make it more manageable, the power allocation problem can be reformulated in vector form: (3) here, ; Due to the quadratic terms in the constraints, problem (P3) remains nonconvex; given the separability of the objective function and constraints, problem (P3) can be decomposed into... Each independent sub-problem, in each time slot This corresponds to one question: (4) This allows for the independent resolution of the power allocation problem for each time slot, yielding the optimal solution for the power allocation problem for a given UAV trajectory point within each time slot; each sub-problem corresponds to one time slot. Joint power control in the process; to solve the non-convex power control problem for each time slot n, we adopt the SDR method, as detailed below: First, define as well as The objective function in (P3.1) can be expressed as: ; Based on this, the constraint can be redefined as: , (5) express There are vectors, where the i-th element is 1 and all other elements are 0. ; Furthermore, the constraints can be rewritten as: (6) here This represents the trace operator; therefore, constraint (19d) is equivalent to: (7) Subsequently, problem (P3.1) was equivalently restated as the following problem; (8) Due to the rank-one constraint, problem (P3.2) remains non-convex. To address this, the SDR technique is applied to relax the rank-one constraint, thus reformulating the problem into a semi-positive definite programming (SDP) form. The solution to an SDP is typically a covariance matrix, not a rank-one feasible vector. To obtain feasible power allocations from the relaxed SDR solution, a Gaussian randomization procedure is used to extract rank-one power vectors that satisfy the original non-convex constraint. To improve robustness and avoid suboptimal local solutions, the randomization process is repeated multiple times to obtain a set of feasible power vectors. Among all candidates, the algorithm selects the scheme that minimizes the total transmit power, thereby achieving a near-optimal feasible solution close to the lower bound of the SDR.

[0012] Step 4, the drone trajectory optimization problem, can be viewed as the outer layer of a two-layer algorithm. At the external level, UAV trajectory design is reformulated as a Markov decision process (MDP) and adopts the DDPG framework to optimize UAV trajectories by efficiently processing the continuous action space. The non-convex UAV trajectory optimization problem is reformulated as an MDP, defined as a quintuple (S, A, P, R, γ), with each component detailed below: 1) State Space Time slot The state at a given point encapsulates the information needed for trajectory decision-making: , (9) in Indicates the horizontal coordinates of the drone. Indicates the current time slot number. Indicates the location of the perceived target; 2) Action Space Continuous motion spatial control for drone movement: , (10) in Represents the direction of flight. Indicates the distance traveled; 3) Transition probability State transitions are deterministic and governed by the following conditions: (11) The perceived target moves according to its prescribed trajectory; 4) Reward Function Given the location of the UAV, calculate the power allocation using SDR. The reward is defined as: , (12) Minimize power while satisfying the SINR and CRLB constraints of the power-optimal subproblem; 5) Discount Factor Discount factor Weight future rewards in the calculation of cumulative returns; The MDP process transforms the complex trajectory optimization problem into a sequential decision framework, which can be effectively solved using deep reinforcement learning algorithms. The DDPG framework combines the advantages of policy-based and value-based reinforcement learning methods, enabling efficient handling of continuous action spaces. It consists of two main neural networks: an Actor network that determines the policy and a Critic network that evaluates the policy. Mapping states to actions, and the Critic network Approximate action-value function; two networks are trained simultaneously to optimize the policy; DDPG uses a target network. and To improve training stability, an experience replay buffer is provided to improve sampling efficiency; During training, the agent selects actions based on the following conditions: (13) in Representing exploration noise; the Critic network updates by minimizing the loss: (14) in The minimum value of the experience pool is used; the Actor network is updated based on the gradient of the expected Q value. , (15) The parameter update methods for the two networks are as follows: (16) in and The learning rate is used; the target network is updated via soft updates. (17) (18) in Control the update rate; In step 5, the result of the inner algorithm is fed back to the outer algorithm as a reward. The agent learns through exploration, and when the change in the reward value of the learned model is less than a preset threshold, the joint optimization process is determined to have converged.

[0013] This invention also provides a joint resource allocation and trajectory design system based on deep reinforcement learning, applicable to UAV-assisted multi-cell ISAC systems, comprising: The initialization module is used to initialize the UAV's starting position, communication and perception system parameters, and deep reinforcement learning model. The user scheduling module is used to schedule communication users of multiple cell base stations given the location of the UAV. The power allocation module is used to allocate transmit power to each communication node under the current user scheduling result; The trajectory update module is used to output the UAV motion control strategy based on the power allocation results and update the UAV flight trajectory; The convergence judgment module is used to determine whether the joint optimization process has converged. If it has, the optimization result is output; otherwise, the process is returned to the user scheduling module to continue execution.

[0014] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention addresses the challenges of unmanned aerial vehicle (UAV)-assisted multi-cell sensing integrated systems by proposing a joint resource allocation and trajectory design method based on deep reinforcement learning. By unifying and jointly optimizing user scheduling, power allocation, and UAV flight trajectory design, it effectively solves the problems of high coupling and high optimization complexity in multi-cell ISAC scenarios. More specifically, this invention utilizes a deep reinforcement learning model to adaptively decide on the UAV's flight trajectory, enabling the UAV to adjust its flight path in real time according to dynamic changes in the communication and sensing environment. This minimizes the overall system transmit power while meeting communication service quality and sensing accuracy requirements. Furthermore, by jointly optimizing user scheduling and power allocation strategies, this invention balances multi-cell communication performance and sensing performance, effectively improving the overall system throughput and user fairness. Simulation results demonstrate that this invention exhibits fast convergence speed and good optimization performance in multi-cell ISAC scenarios, making it suitable for UAV-assisted communication and sensing systems in complex dynamic environments. [Attached Image Description] Figure 1 This is a model diagram of an unmanned aerial vehicle-assisted multi-unit ISAC system; Figure 2 This is the reward convergence training graph of the DDPG-JRO algorithm; Figure 3 This is a hyperparameter sensitivity analysis diagram from DDPG training; Figure 4 This is a graph showing the relationship between total power consumption and the SINR threshold; Figure 5 This is a graph showing the relationship between total power consumption and the CRLB threshold; Figure 6 This is the drone trajectory diagram under power minimization. Figure 7 This is a diagram showing the impact of the maximum scheduling time slot on the total transmit power consumption; Figure 8 This is a graph showing the relationship between total power consumption and the number of users; Figure 9 It is a power consumption diagram of different base stations in N time slots; Figure 10 This is a flowchart of the joint resource allocation and trajectory design method based on deep reinforcement learning of the present invention.

Detailed Implementation Methods

[0015] It should be noted that in the embodiments of the present invention, all directional indications (such as up, down, left, right, front, back, etc.) are limited to relative positions on the specified view, rather than absolute positions.

[0016] Furthermore, in this invention, descriptions involving "first," "second," etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0017] System Model This invention considers unmanned aerial vehicle-assisted multi-cell ISAC systems, such as... Figure 1 As shown, the system consists of a heterogeneous set of base stations and multiple users. The TBS in each cell serves communication users within its respective coverage area. It is worth noting that users located at the edges of multiple cells typically experience poor channel conditions due to longer propagation distances and severe inter-cell interference. Furthermore, due to the mobility of target sensing users, fixed TBS struggles to guarantee the required sensing performance, especially in temporary or emergency situations. To ensure communication and sensing QoS requirements, drones are deployed as aerial base stations to transmit communication signals to these edge users, while simultaneously collaborating with the TBS to provide services to mobile sensing users.

[0018] Assuming the drone base station is at a fixed altitude over the target area Flight, the total duration of each flight mission is For ease of optimization and analysis, continuous time periods will be used. Uniformly discretized There are 1 time slot, and the length of each time slot is 1. By selecting a sufficiently large It can make It is small enough to ensure that the displacement of the UAV within each time slot is negligible compared to its air-to-ground (A2G) communication distance. This approximation allows the UAV's ground-to-ground channel to be considered quasi-static within each time slot, and assumes that the latency requirements of the sensing task are met.

[0019] communication model Assuming ground base station With users The channel coefficients between them are: (1) here This represents the small-scale fading coefficient that follows a Rayleigh distribution. This represents the large-scale path loss, measured in decibels (dB). Therefore, the corresponding channel power gain can be expressed as: .

[0020] Unlike TBS (Traffic-Based Service), drones, acting as airborne base stations, operate at relatively high altitudes, making their communication links with ground users less susceptible to disruption. Therefore, the A2G (Area-to-Growth) channel between the drone and the ground user is primarily characterized by line-of-sight (LoS) link features, and its channel state information depends mainly on the distance between the drone and the user. Consequently, the A2G channel follows a free-space path loss model.

[0021] Assuming the Doppler effect caused by the drone's motion and the fact that ground users can be perfectly compensated, let Indicates drone In the time slot The horizontal position of the drone. With ground users In the time slot The SCG can be represented as: (2) In the above expressions, Indicates the reference distance SCG at the location, and Indicates drone With ground users In the time slot The distance. Let Indicates base station In the time slot Data symbols transmitted to its terrestrial communication users. Indicates base station In the time slot The transmit power. According to these definitions, terrestrial communication users... The received signal can be represented as: , (3) here, Indicates terrestrial communication users Received complex Gaussian noise. Therefore, terrestrial communication users The signal-to-interference-plus-noise ratio (SINR) can be expressed as: (4) make Indicates the base station in the time slot The transmission power, This indicates the transpose operation.

[0022] Perceptual Model The perception model primarily describes the cooperative distance measurement and localization process between the UAV and the TBS for target sensing users. It characterizes sensing accuracy by combining channel models, noise characteristics, and the geometric configuration of the observations. To achieve reliable system optimization, accurate evaluation of parameter estimation performance is necessary. This invention employs the Cramero Lower Bound (CRLB) for target localization as a sensing performance metric, particularly emphasizing the impact of the geometric configuration between the UAV and the ground-based detection target on estimation accuracy. CRLB provides a theoretical lower bound on the variance of any unbiased estimator. By limiting CRLB to below a specified threshold, the system ensures that the localization error of the moving target remains within acceptable limits. This facilitates joint optimization of sensing performance and communication quality.

[0023] The system includes ( One ground base station and one drone serve as an aerial base station, working together to provide services. Multiple communication users simultaneously perform the location and sensing of mobile target users. The system has... One transmitter and Each base station is equipped with a receiver. Each base station also has a transmitter. and receiving end It operates in full-duplex mode, enabling simultaneous communication and sensing. The target is assumed to be located on a two-dimensional plane. Positioning is achieved by jointly transmitting detection signals from the UAV and detection satellites and receiving the corresponding echo signals. A time-division multiplexing mode is used, assuming that the signals within each time slot satisfy both ergodicity and statistical independence, for example: (5) (6) set up This indicates the time delay of radar detection between different TBSs. This represents the conjugate operator. Further, it is assumed that the radar detection period is long enough to ensure accurate estimation. Therefore, the base station... In the time slot The received signal is: (7) Noise item Modeled as having zero mean and autocorrelation function Circularly symmetric complex Gaussian noise. Coefficients This reflects the impact of the radar cross section (RCS) of the sensed target, as well as the transmission ground base station. With receiving ground base station The radar propagation path loss between them. When and At that time, we assume and They are statistically independent. Let... Indicates the transmitting radar With receiving radar The radar path loss between them is as follows: (8) For the transmitting end With the receiving end The propagation delay between them can be expressed by combining the drone's altitude information, and its formula is: (9) (10) , (11) here, Represents the speed of light. Indicates the transmitting end The altitude, where if the transmitter is a ground base station, then If the transmitter is a drone, then The CRLB metric we employ provides a theoretical lower bound for the mean squared error (MSE) of any unbiased estimator and is widely used in wireless sensing systems. Based on the theory of the Fisher information matrix, the target position estimation matrix of CRLB can be expressed as: (12) (13) in , Indicates the radar transmitter transmitting a signal The effective bandwidth, ,in express bandwidth, express Frequency domain transformation.

[0024] By summing the elements of the CRLB matrix, we obtain the lower bound of the total MSE of the target location estimate, denoted as . .here, and These represent the estimated measurement targets. and MSE of coordinates. According to the literature, the CRLB matrix... The trace can be further represented as: , (14) here, , , , , .

[0025] The CRLB expression in equation (14) can be used to calculate the MSE of the maximum likelihood estimate. Furthermore, since this invention assumes that the RCS of the target is known, the system mainly relies on the signal propagation delay to determine the target's position. Considering the prior information about the target's position, this invention uses the trace of the CRLB matrix as a sensing performance index. Therefore, optimizing the CRLB is a feasible method to improve the accuracy of real-time sensing estimation.

[0026] Problem Construction make Represents the user scheduling matrix, where Indicates in time slot user Scheduled Represents the power allocation matrix. Let represent the horizontal trajectory matrix of the UAV. The resulting power minimization problem with SINR and CRLB constraints is expressed as: P1: (15) st (1a) ; (1b) ; (1c) ; (1d) ; (1e) ; (1f) . here, Let represent a vector of all ones with the corresponding dimension, and (1a) Ensures that the transmit power of each base station is non-negative in each time slot. (1b) and (1c) constrain the number of communication users scheduled in each time slot. L represents the number of base stations. In addition, considering user fairness, (1d) limits the maximum number of time slots scheduled to each user. The constraints in (17e) ensure that the SINR requirements of communication users are met, and (1f) ensures that the CRLB of the target location estimation does not exceed the threshold. .

[0027] Optimization Algorithm A. User scheduling based on service channel gain As mentioned above, we consider by A system consisting of one TBS and one drone base station. Each base station is assumed to be randomly distributed within its coverage area. There are several candidate communication users. Lemma 1 is presented below.

[0028] Lemma 1: In a UAV-assisted multi-cell ISAC system, when the sensing constraints are relatively relaxed, the total transmission power consumption decreases monotonically with the SCG of each communication user.

[0029] prove: For the UAV-assisted multi-cell ISAC system under consideration, when the CRLB constraint is relatively relaxed, power consumption is mainly determined by communication requirements. In this case, the base station's transmit power monotonically decreases with increasing SCG, as shown in the following proof.

[0030] Assume it exists One communication user ,as well as There are 1 base stations, of which ( ) is TBS. Indicates from base station To communication users SCG, Transmission power is denoted as In each time period, Therefore, the constraints in (17f) can be equivalently restated using the following matrix representation: (16) The constraints can then be restated as the following matrix: (17) here, ,and This represents an invertible square matrix composed of SCG coefficients. The total transmitted power is relative to... The partial derivatives are: (18) The inverse function can be expressed as , Indicates and The associated adjoint matrix, This corresponds to its determinant. It can then be restated as: (19) in Indicates that it is located at The row and number Column elements, It is worth noting that, The diagonal entries correspond to the communication user's SCG, while the off-diagonal elements are determined by user interference and... The product is obtained by using... Given a specific structure, we obtain: (20) in Indicates the corresponding base station SCG serves users. (Positive item) and Indicates the user's position in the first month. The product of channel gain and interference in each time slot. It is worth noting that for base stations... For users of the service, direct SCG This is typically much greater than the cross-channel gain of other base stations. Therefore, we have This means Similarly, using a similar methodology, we can also prove... .

[0031] Therefore, we conclude that: . (twenty one) According to Lemma 1, the system transmission power continuously decreases as the communication user SCG increases. Based on this, users can be scheduled according to their SCG. Assume... Indicates from base station To communication users The SCG. The detailed user scheduling process can be described as follows.

[0032] In each time slot In the middle, each Candidate communication users according to their Sort from lowest to lowest to form a Ranking list Then, from arrive Inside, when the time slot The number of user scheduling times is less than At that time, for each from the top of the list Starter users Calculate. If , Then the user In the time slot Scheduled, and set Otherwise, if ,user will directly from Remove from the middle. Repeat the above steps until... Communication users Each time slot is scheduled.

[0033] B. Optimal transmit power allocation Based on (P1), to make it more manageable, the power allocation problem can be reformulated in vector form: (twenty two) here, .

[0034] Due to the quadratic terms in the constraints, problem (P2) remains nonconvex. Given the separability of the objective function and constraints, problem (P2) can be decomposed into... Each independent sub-problem, in each time slot This corresponds to one question: (twenty three) This decomposition greatly simplifies our method because we can solve the power allocation problem for each time slot independently. We can then obtain the optimal solution to the power allocation problem for a given UAV trajectory point in each time slot.

[0035] Each subproblem corresponds to a time slot. Joint power control in the process. To address the non-convex power control problem for each time slot n (P2.1), we employ the SDR method, detailed below.

[0036] First, define as well as The objective function in (P2.1) can be expressed as: .

[0037] Based on this, the constraint can be redefined as: , (twenty four) express There are vectors, where the i-th element is 1 and all other elements are 0. .

[0038] Furthermore, constraint (19d) can be rewritten as: (25) here This represents the trace operator. Therefore, the constraint is equivalent to: (26) Subsequently, problem (P2.1) is equivalently restated as the following problem.

[0039] (27) Due to the rank-one constraint, problem (P2.2) remains non-convex. To address this issue, the SDR technique is applied to relax the rank-one constraint, thereby reformulating the problem into a semidefinite programming (SDP) form. The solution to an SDP is typically a covariance matrix, rather than a rank-one feasible vector.

[0040] To obtain feasible power allocation from the relaxed SDR solution, a Gaussian randomization procedure is used to extract the rank-one power vector that satisfies the original non-convex constraints. After solving the SDR problem, the covariance matrix obtained by eigenvalue decomposition is first decomposed... , Includes feature vectors, It is a diagonal matrix of eigenvalues. This decomposition makes it possible to... Multiplying with a randomly generated Gaussian vector makes it possible to construct random candidate power vectors. These random vectors preserve the statistical structure of the SDR solution and serve as potential first-order approximations. Since these original candidates may not satisfy the SINR and CRLB constraints, each candidate condition is... The scaling factor ensures that all constraints are feasible. This scaling guarantees that the resulting vectors are effective for both communication and sensing requirements. To improve robustness and avoid suboptimal local solutions, the randomization process is repeated multiple times to obtain a set of feasible power vectors. Among all candidates, the algorithm selects the scheme that minimizes the total transmit power, thus achieving a near-optimal feasible solution close to the lower bound of the SDR. The complete process, including the SDR solution, Gaussian randomization, constraint enforcement, and its integration with UAV trajectory optimization, is summarized in Algorithm 1.

[0041] C. Joint resource optimization based on DDPG In this section, we will describe in detail the proposed DDPG-JRO algorithm, which is designed with a two-layer structure. Specifically, at the outer layer, the UAV trajectory design is reformulated as a Markov Decision Process (MDP) and uses the DDPG framework to optimize the UAV trajectory by efficiently handling the continuous action space. In the inner layer, user scheduling is determined according to the SCG principle, while optimal power allocation is obtained through SDR and Gaussian randomization. The inner layer solution is then incorporated into the reward function of the outer layer, thus combining UAV trajectory optimization with user scheduling and power allocation to minimize the total transmit power across all N time slots.

[0042] Markov decision process description In this section, we reformulate the non-convex UAV trajectory optimization problem as an MDP, which is defined as a quintuple (S, A, P, R, γ), with each component detailed below.

[0043] 1) State Space Time slot The state at a given point encapsulates the information needed for trajectory decision-making: , (28) in Indicates the horizontal coordinates of the drone. Indicates the current time slot number. It indicates the location of the perceived target.

[0044] 2) Action Space Continuous motion spatial control for drone movement: , (29) in Represents the direction of flight. Indicates the distance traveled.

[0045] 3) Transition probability State transitions are deterministic and governed by the following conditions: (30) The perceived target moves according to its prescribed trajectory.

[0046] 4) Reward Function Given the location of the UAV, calculate the power allocation using SDR. The reward is defined as: (31) Minimize power while satisfying the SINR and CRLB constraints of the power-optimal subproblem.

[0047] 5) Discount Factor Discount factor Weighted future rewards are added to the cumulative return calculation.

[0048] The MDP process transforms the complex trajectory optimization problem into a sequential decision framework, which can be effectively solved using deep reinforcement learning algorithms. The DDPG framework combines the advantages of policy-based and value-based reinforcement learning methods, enabling efficient handling of continuous action spaces. It consists of two main neural networks: an Actor network that determines the policy, and a Critic network that evaluates the policy. Mapping states to actions, and the Critic network Approximate action-value function. Two networks are trained simultaneously to optimize the policy. DDPG uses the target network. and To improve training stability, an experience replay buffer is provided to improve sampling efficiency.

[0049] During training, the agent selects actions based on the following conditions: (32) in This represents the exploration of noise. The Critic network updates by minimizing the loss function: (33) in This represents the minimum value in the experience pool. The Actor network updates based on the gradient of the expected Q-value: (34) The parameter update methods for the two networks are as follows: (35) in and The learning rate is used. The target network is updated via soft updates: (36) (37) The proposed DDPG-JRO algorithm is designed as a two-layer framework. In the outer layer, the Actor-Critic network, the target network, and the replay buffer are initialized, and the UAV trajectory is optimized within the DDPG framework. During each training session, the outer agent selects the UAV action based on the reward feedback from the previous time slot. In the inner layer, given the current UAV position, user scheduling is determined according to the SCG principle, and the power allocation problem is solved through SDR and Gaussian randomization to minimize the total transmit power of the time slot while satisfying SINR and CRLB constraints. The reward for the outer layer trajectory optimization is then calculated using the inner layer solution. State transitions, including UAV position, selected action, and obtained reward, are stored in the replay buffer. Soft target updates of the Actor-Critic network ensure stable training. After all time slots in each round, the total transmit power is compared with the current best power; if improvement is achieved, the best trajectory and power allocation are updated. The proposed DDPG-JRO algorithm achieves dynamic coordination between trajectory, scheduling, and power allocation, enabling the UAV to adaptively reduce energy consumption under communication and sensing constraints. The overall process of the algorithm is summarized as Algorithm 2.

[0050] D. Overall Algorithm Description Table 1 Algorithm 1 ; Table 2 Algorithm 2 ; Simulation results This section presents numerical results to evaluate the effectiveness of the proposed DDPG-JRO algorithm in UAV-assisted multi-cell ISAC systems.

[0051] Figure 2The reward curves of the proposed DDPG-JRO algorithm during training are shown. As training progresses, the reward steadily increases and eventually stabilizes. Due to the inherent exploration in DDPG, random noise is introduced during policy updates, resulting in performance fluctuations during iterations. To highlight the overall performance trend, the solid line represents the smoothed moving average of the reward, and the shaded area represents the standard deviation. These results validate the effectiveness of DDPG in achieving performance optimization in complex environments.

[0052] Figure 3 The impact of different learning rates and discount factors on the convergence performance of the DDPG-JRO algorithm is demonstrated. Here, the learning rate is determined by α (i.e., α = αa = αc). If γ is large, the unpredictability of future channel states can significantly affect the estimation of the Critic value, potentially impacting the learning process. Furthermore, the system operates under strict instantaneous quality of service constraints; setting γ = 0.1 encourages the intelligent agent to prioritize these immediate constraints and optimize real-time energy efficiency, ensuring robust performance for each time slot. Therefore, using a smaller discount factor allows the agent to obtain higher reward values ​​within the proposed framework. Additionally, a reasonably chosen learning rate accelerates reward growth, improves fusion behavior, and thus enhances overall training efficiency. In contrast, excessively large or small parameter values ​​may lead to slow convergence or decreased reward performance.

[0053] Figure 4 The relationship between total base station power consumption and the SINR threshold was analyzed. Clearly, as the SINR threshold increases, each user requires higher transmit power to meet this constraint. This not only directly increases transmission power but also exacerbates interference to other users, forcing them to further increase their power. Crucially, the proposed DDPG-JRO algorithm effectively mitigates this escalation by leveraging the synergistic effect between UAV maneuverability and dynamic control. Unlike static allocation schemes, DDPG agent learning dynamically adjusts the UAV trajectory to improve line-of-sight channel conditions for weaker users. By optimizing position to enhance desired signal strength, the algorithm reduces the need for excessively high transmit power, demonstrating its ability to find energy-saving strategies in continuous operational spaces under high-interference scenarios.

[0054] Figure 5The relationship between total base station power consumption and the CRLB threshold is illustrated. When the CRLB threshold is relaxed from 5 to 10, power consumption decreases significantly, indicating a strong negative correlation between sensing accuracy requirements and power consumption. More importantly, the figure highlights the adaptability of the proposed DDPG-JRO algorithm under strict sensing constraints. In environments with low CRLB thresholds, the algorithm does not simply rely on increasing sensor signal power; instead, it optimizes the UAV's flight path to maintain a favorable geometric configuration relative to the target. By actively guiding the UAV to a higher-precision detection location, the algorithm significantly reduces its dependence on high-power signal transmission, thereby keeping total power consumption within acceptable limits while meeting stringent accuracy requirements.

[0055] The drone's flight trajectory is as follows Figure 6 As shown, the sub-graphs correspond to the target moving in the upper left, upper right, and upper right directions, respectively. It can be observed that the proposed algorithm initially guides the drone towards users at the cell edge, effectively reducing path loss by shortening the communication distance. Subsequently, under the constraint of sensing requirements, the trajectory is dynamically adjusted to simultaneously meet communication and sensing needs.

[0056] Figure 7 The impact of the maximum number of schedulable cycles on the power consumption of a UAV-assisted multi-cell ISAC system was examined. The results show that at lower ρ1 values, the total transmission power of the base station remains at a high level, indicating improved fairness among users. However, as the ρ1 value increases, the total transmission power decreases accordingly. This behavior pattern can be attributed to a fundamental trade-off between total transmission power allocation and fairness optimization in wireless communication systems.

[0057] Figure 8 The changes in total transmit power with varying numbers of users are illustrated, with each curve corresponding to different SINR and CRLB constraints. As the number of users increases, the total power fluctuates slightly, but remains within a reasonable range. These fluctuations primarily stem from the randomness of the candidate user distribution and the dynamic trajectory adjustments of the UAV BS. The results demonstrate that the proposed system can effectively select the optimal user while maintaining fairness, thus achieving efficient power control. Overall, the system exhibits flexibility in handling multi-user scenarios, possessing controllable total power consumption and robust performance, highlighting its applicability in multi-user environments.

[0058] Figure 9 The power consumption of the three base stations in each time slot is shown. In the current scenario, SINR=1 and CRLB=1 are set. After satisfying the constraints, the drone base station will track the movement of the target user after the 15th time slot, causing its power consumption to increase significantly. In contrast, the power fluctuation of TBS is relatively stable and gradually increases.

[0059] Figure 10The flowchart of the joint resource allocation and trajectory design method based on deep reinforcement learning of this invention includes: Step 1: Initialize the UAV's starting position, communication and perception system parameters, and deep reinforcement learning model.

[0060] Step 2: Given the location of the UAV, perform communication user scheduling for the base station.

[0061] Step 3: Perform power allocation under the current user selection.

[0062] Step 4: Update the drone's flight trajectory based on the power allocation results.

[0063] Step 5: Determine whether the algorithm has converged.

[0064] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention addresses the challenges of unmanned aerial vehicle (UAV)-assisted multi-cell sensing integrated systems by proposing a joint resource allocation and trajectory design method based on deep reinforcement learning. By unifying and jointly optimizing user scheduling, power allocation, and UAV flight trajectory design, it effectively solves the problems of high coupling and high optimization complexity in multi-cell ISAC scenarios. More specifically, this invention utilizes a deep reinforcement learning model to adaptively decide on the UAV's flight trajectory, enabling the UAV to adjust its flight path in real time according to dynamic changes in the communication and sensing environment. This minimizes the overall system transmit power while meeting communication service quality and sensing accuracy requirements. Furthermore, by jointly optimizing user scheduling and power allocation strategies, this invention balances multi-cell communication performance and sensing performance, effectively improving the overall system throughput and user fairness. Simulation results demonstrate that this invention exhibits fast convergence speed and good optimization performance in multi-cell ISAC scenarios, making it suitable for UAV-assisted communication and sensing systems in complex dynamic environments.

[0065] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any modifications, equivalent substitutions and improvements made within the concept of the present invention should be included within the patent protection scope of the present invention.

Claims

1. A joint resource allocation and trajectory design method based on deep reinforcement learning, applicable to UAV-assisted multi-cell ISAC systems, characterized in that... Includes the following steps: Step 1, Initialization Steps: Initialize the UAV's starting position, communication and perception system parameters, and deep reinforcement learning model; Step 2, User Scheduling Step: Given the location of the UAV, perform communication user scheduling for the base station; Step 3, Power Allocation Steps: Perform power allocation under the current user selection; Step 4, Trajectory Update Step: Update the UAV flight trajectory based on the power allocation results; Step 5, Convergence Judgment Step: Determine whether the reward curve of the joint optimization process has converged. If so, output the optimization result. Otherwise, return to step 2 and continue iterating.

2. The joint resource allocation and trajectory design method based on deep reinforcement learning according to claim 1, characterized in that, Assuming the drone base station is at a fixed altitude over the target area Flight, the total duration of each flight mission is For ease of optimization and analysis, continuous time periods will be used. Uniformly discretized There are 1 time slot, and the length of each time slot is 1. By selecting a sufficiently large It can make It is small enough to ensure that the displacement of the UAV in each time slot is negligible compared to its air-to-ground (A2G) communication distance; this approximation allows the UAV's ground channel to be considered quasi-static in each time slot and assumes that the latency requirements of the sensing task are met. make Represents the user scheduling matrix, where Indicates in time slot user Scheduled Represents the power allocation matrix. Let the horizontal trajectory matrix of the UAV be represented; in this case, the optimization problem of the present invention can be expressed as: P1: (1) st (1a) ; (1b) ; (1c) ; (1d) ; (1e) ; (1f) ; here, Let represent a vector of all ones with the corresponding dimension, and (1a) Ensure that the transmit power of each base station is non-negative in each time slot. (1b) and (1c) constrain the number of communication users scheduled in each time slot; L represents the number of base stations. In addition, considering user fairness, (1d) limits the maximum number of time slots scheduled to each user. The constraints in (1e) ensure that the SINR requirements of communication users are met. (1f) Ensure that the CRLB of the target location estimation does not exceed the threshold. .

3. The joint resource allocation and trajectory design method based on deep reinforcement learning according to claim 2, characterized in that, The user scheduling and power allocation problem can be viewed as the inner part of a two-layer algorithm, characterized by: For any given drone starting point, the user scheduling problem can be formulated as: P2: (2) st (1a)-(1d) Each base station is assumed to be randomly distributed within its coverage area. I. Candidate communication users; Lemma 1 is proposed below; Lemma 1: In a UAV-assisted multi-cell ISAC system, when the sensing constraints are relatively relaxed, the total transmission power consumption decreases monotonically with the SCG (Serving Channel Gain) of each communication user. prove: For the UAV-assisted multi-cell ISAC system under consideration, when the CRLB (Cramer-Rao lower bound) constraint is relatively relaxed, the power consumption is mainly determined by the communication requirements. In this case, the base station's transmit power monotonically decreases with the increase of SCG, as shown in the following proof: Assume it exists One communication user ,as well as There are 1 base stations, of which ( ) is TBS; Indicates from base station To communication users SCG, Transmission power is denoted as ; in each time period, Therefore, the constraints in (1f) can be equivalently restated using the following matrix representation: , The constraints can then be restated as the following matrix: ; here, ,and Represents an invertible square matrix composed of SCG coefficients; the total transmitted power is relative to The partial derivatives are: ; The inverse function can be expressed as , Indicates and The associated adjoint matrix, Corresponding to its determinant; it can then be restated as: , in Indicates that it is located at The row and number Column elements, It is worth noting that, The diagonal entries correspond to the SCG of the communication user, while the off-diagonal elements are composed of user interference and The product is obtained by using; Given a specific structure, we obtain: , in Indicates the corresponding base station SCG serving users; positive item and Indicates the user's position in the first month. The product of channel gain and interference in each time slot; it is worth noting that for the base station For users of the service, direct SCG Typically, the cross-channel gain is much greater than that of other base stations; therefore, we have This means Similarly, using a similar methodology, we can also prove... ; Therefore, we conclude that: ; According to Lemma 1, the system transmission power continuously decreases as the communication user SCG increases. Based on this, users can schedule according to their SCG. Based on (P1), to make it more manageable, the power allocation problem can be reformulated in vector form: (3) here, ; Due to the quadratic terms in the constraints, problem (P3) remains nonconvex; given the separability of the objective function and constraints, problem (P3) can be decomposed into... Each independent sub-problem, in each time slot This corresponds to one question: (4) This allows for the independent resolution of the power allocation problem for each time slot, yielding the optimal solution for the power allocation problem for a given UAV trajectory point within each time slot; each sub-problem corresponds to one time slot. Joint power control in the process; to solve the non-convex power control problem for each time slot n, we adopt the SDR method, as detailed below: First, define as well as The objective function in (P3.1) can be expressed as: ; Based on this, the constraint can be redefined as: , (5) express There are vectors, where the i-th element is 1 and all other elements are 0. ; Furthermore, the constraints can be rewritten as: (6) here This represents the trace operator; therefore, constraint (19d) is equivalent to: (7) Subsequently, problem (P3.1) was equivalently restated as the following problem; (8) Due to the rank-one constraint, problem (P3.2) remains non-convex. To address this, the SDR technique is applied to relax the rank-one constraint, thus reformulating the problem into a semi-positive definite programming (SDP) form. The solution to an SDP is typically a covariance matrix, not a rank-one feasible vector. To obtain feasible power allocations from the relaxed SDR solution, a Gaussian randomization procedure is used to extract rank-one power vectors that satisfy the original non-convex constraint. To improve robustness and avoid suboptimal local solutions, the randomization process is repeated multiple times to obtain a set of feasible power vectors. Among all candidates, the algorithm selects the scheme that minimizes the total transmit power, thereby achieving a near-optimal feasible solution close to the lower bound of the SDR.

4. The joint resource allocation and trajectory design method based on deep reinforcement learning according to claim 3, characterized in that, Step 4, the drone trajectory optimization problem, can be viewed as the outer layer of a two-layer algorithm, characterized by: At the external level, UAV trajectory design is reformulated as a Markov decision process (MDP) and adopts the DDPG framework to optimize UAV trajectories by efficiently processing the continuous action space. The non-convex UAV trajectory optimization problem is reformulated as an MDP, defined as a quintuple (S, A, P, R, γ), with each component detailed below: 1) State Space Time slot The state at a given point encapsulates the information needed for trajectory decision-making: , (9) in Indicates the horizontal coordinates of the drone. Indicates the current time slot number. Indicates the location of the perceived target; 2) Action Space Continuous motion spatial control for drone movement: , (10) in Represents the direction of flight. Indicates the distance traveled; 3) Transition probability State transitions are deterministic and governed by the following conditions: (11) The perceived target moves according to its prescribed trajectory; 4) Reward Function Given the location of the UAV, calculate the power allocation using SDR. The reward is defined as: (12) Minimize power while satisfying the SINR and CRLB constraints of the power-optimal subproblem; 5) Discount Factor Discount factor Weight future rewards in the calculation of cumulative returns; The MDP process transforms the complex trajectory optimization problem into a sequential decision framework, which can be effectively solved using deep reinforcement learning algorithms. The DDPG framework combines the advantages of policy-based and value-based reinforcement learning methods, enabling efficient handling of continuous action spaces. It consists of two main neural networks: an Actor network that determines the policy and a Critic network that evaluates the policy. Mapping states to actions, and the Critic network Approximate action-value function; two networks are trained simultaneously to optimize the policy; DDPG uses a target network. and To improve training stability, an experience replay buffer is provided to improve sampling efficiency; During training, the agent selects actions based on the following conditions: (13) in Representing exploration noise; the Critic network updates by minimizing the loss: (14) in The minimum value of the experience pool is used; the Actor network is updated based on the gradient of the expected Q value. (15) The parameter update methods for the two networks are as follows: (16) in and The learning rate is used; the target network is updated via soft updates. (17) (18) in Control the update rate.

5. The joint resource allocation and trajectory design method based on deep reinforcement learning according to claim 4, characterized in that, In step 5, the result of the inner algorithm is fed back to the outer algorithm as a reward. The agent learns through exploration, and when the change in the reward value of the learned model is less than a preset threshold, the joint optimization process is determined to have converged.

6. A joint resource allocation and trajectory design system based on deep reinforcement learning, applicable to UAV-assisted multi-cell ISAC systems, characterized in that... include: The initialization module is used to initialize the UAV's starting position, communication and perception system parameters, and deep reinforcement learning model. The user scheduling module is used to schedule communication users of multiple cell base stations given the location of the UAV. The power allocation module is used to allocate transmit power to each communication node under the current user scheduling result; The trajectory update module is used to output the UAV motion control strategy based on the power allocation results and update the UAV flight trajectory; The convergence judgment module is used to determine whether the joint optimization process has converged. If it has, the optimization result is output; otherwise, the process is returned to the user scheduling module to continue execution.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program configured to, when invoked by a processor, implement the steps of the joint resource allocation and trajectory design method based on deep reinforcement learning as described in any one of claims 1-5.