A multi-moving target tracking method and system for integrated unmanned aerial vehicle
By introducing a dual-timescale structure and an optimized network into the UAV system, the problems of tracking fairness and low resource utilization efficiency in multi-target tracking of UAVs are solved, and efficient multi-target tracking and data transmission are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-05-19
- Publication Date
- 2026-06-16
Smart Images

Figure CN122218684A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of wireless communication and radar sensing, specifically to a multi-moving target tracking method and system for an integrated sensing and sensing UAV. Background Technology
[0002] Low-altitude intelligent networks (LAWNs) are becoming the core infrastructure of next-generation intelligent systems, supporting transformative applications such as smart transportation, emergency rescue, and logistics delivery. These applications rely on continuous and accurate tracking of ground entities and efficient transmission of sensing data. However, the unpredictable mobility of targets, coupled with intermittent communication and sensing links in complex environments, poses severe challenges to system design.
[0003] Unmanned aerial vehicles (UAVs) equipped with Integrated Sensing and Communication (ISAC) capabilities, leveraging their inherent maneuverability, can dynamically track targets and improve perception accuracy and data transmission timeliness through air-to-ground line-of-sight links, making them a promising solution. However, existing research on multi-target tracking for ISAC UAVs still has the following shortcomings: First, at the performance index level, existing research typically uses "total performance" or "single time-slot index" as optimization objectives. This optimization strategy, centered on total performance, is inherently greedy, tending to allocate resources to easily trackable targets while ignoring perception fairness issues. This results in some targets remaining in the estimation blind zone for extended periods, weakening global situational awareness. Furthermore, target tracking is essentially a continuous state evolution process. Single time-slot indexes cannot effectively characterize the long-term cumulative effect of state estimation errors. In safety-critical scenarios such as autonomous driving or emergency rescue, long-term inaccuracies in certain target state information can lead to decision-making errors and serious safety hazards. Second, at the perception scheduling level, most existing mechanisms employ uniform and fixed perception intervals, failing to consider the time-varying nature of wireless channels and target dynamics. This fixed strategy lacks the ability to flexibly adjust perception timing according to real-time tracking needs. For example, it cannot prioritize the perception of targets with high accumulated errors, leading to a surge in tracking errors for some targets while wasting perception resources for other targets. Third, at the time scale level, existing research generally jointly optimizes UAV trajectories and beamforming within a single time scale framework. Due to the wide coverage area of UAVs, frequent trajectory adjustments are both unnecessary and costly, while ISAC signals typically require only short dwell times, and antenna arrays support rapid beamforming adjustments. Forcibly constraining these variables, which have fundamental differences in time granularity, to the same time scale will lead to unnecessary decision-making overhead, limit adjustment flexibility, and thus restrict the full utilization of communication and perception performance.
[0004] Therefore, existing technologies suffer from poor tracking fairness, rigid perception scheduling, and a single time scale, leading to long-term inaccuracies in multiple objectives and low resource utilization efficiency. Summary of the Invention
[0005] To address the aforementioned issues, this invention proposes a multi-moving target tracking method and system for integrated sensing and communication unmanned aerial vehicles (UAVs). This method can flexibly match the time granularity requirements of different decision variables while ensuring tracking accuracy and fairness, thereby achieving cross-timescale collaborative optimization of UAV trajectory, sensing scheduling, and communication sensing beamforming.
[0006] According to some embodiments, the present invention adopts the following technical solution: A multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle (UAV) divides the total service time into multiple large time slots, each of which further contains multiple smaller time slots. The method iterates through each large time slot, including: Acquire the state of the integrated sensor drone under dual time scales in the current large time slot, including the drone's flight state at the beginning of the large time slot, the predicted state and state age value of all moving targets, and the remaining power and perception completion mark at the beginning of the small time slot; At the upper-level large time-slot scale, the flight status of the UAV and the predicted status of each moving target are input into the flexible action-evaluation network to plan the flight speed and direction of the UAV within the large time slot in order to update its flight trajectory. At the lower-level hourly slot scale, the state age value, remaining power, and perception completion identifier of each moving target are input into the multi-head discrete near-end strategy optimization network to jointly optimize perception scheduling and determine the perception target and power coefficient of the current hourly slot UAV. Based on the perception scheduling results, the perception beamforming vector and communication beamforming vector of the UAV are solved to perform perception of moving targets and data communication with the ground data fusion station. After receiving the sensing data of the current large time slot, the ground data fusion station uses the extended Kalman filter method with multi-step prediction and backtracking correction to update the state of each moving target and the state age value of each target, and then enters the next large time slot iteration.
[0007] According to some embodiments, the present invention adopts the following technical solution: A multi-moving target tracking system for a sensor-integrated unmanned aerial vehicle (UAV) divides the total service time into multiple large time slots, each large time slot further contains multiple small time slots, and iterates through each large time slot, including: The state acquisition module is configured to acquire the state of the integrated sensor drone under dual time scales in the current large time slot, including the drone's flight state at the beginning of the large time slot, the predicted state and state age value of all moving targets, and the remaining power and perception completion indicator at the beginning of the small time slot. The trajectory update module is configured to: input the UAV flight state and the predicted state of each moving target into the flexible action-evaluation network at the upper-level large time slot scale, plan the UAV's flight speed and direction within the large time slot, and update its flight trajectory. The perception scheduling module is configured to: input the state age value, remaining power and perception completion identifier of each moving target into the multi-head discrete near-end strategy optimization network at the lower hourly slot scale, perform joint optimization of perception scheduling, and decide the perception target and power coefficient of the current hourly slot UAV; The beamforming module is configured to: solve the UAV's perception beamforming vector and communication beamforming vector based on the perception scheduling results, in order to perform perception of moving targets and data communication with the ground data fusion station; The state update module is configured as follows: after receiving the sensing data of the current large time slot, the ground data fusion station uses the extended Kalman filter method with multi-step prediction and backtracking correction to update the state of each moving target and the state age value of each target, and then enters the next large time slot iteration.
[0008] According to some embodiments, the present invention adopts the following technical solution: A computer program product includes a computer program that, when executed by a processor, implements the aforementioned method for tracking multiple moving targets in a sensor-integrated unmanned aerial vehicle.
[0009] According to some embodiments, the present invention adopts the following technical solution: A non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the multi-moving target tracking method of a sensor-integrated unmanned aerial vehicle.
[0010] According to some embodiments, the present invention adopts the following technical solution: An electronic device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the multi-moving target tracking method of the integrated sensor drone.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle (UAV). By dividing the total service time into a dual-timescale structure consisting of large time slots and small time slots, and acquiring the UAV's flight status, target prediction status, and status age value at the beginning of the large time slot, and acquiring the remaining power and perception completion flag at the beginning of the small time slot, this method achieves decoupling of trajectory decision-making and signal-level decision-making at the time granularity. This allows the UAV to flexibly perform perception scheduling and beamforming updates at the small time slot level without frequently adjusting its trajectory, effectively reducing decision-making overhead and improving system response flexibility.
[0012] The upper layer of this invention employs a flexible motion-evaluation network to plan flight trajectories based on the UAV's flight status and the target's predicted status. Its maximum entropy regularization mechanism encourages the full exploration of diverse trajectory strategies and prevents strategies from converging prematurely to suboptimal paths. The lower layer employs a multi-head discrete near-end strategy optimization network to jointly optimize the perceived targets and power coefficients based on the state age value, remaining power, and perception completion identifier of each target. This allows for the dynamic allocation of perception resources according to real-time tracking requirements, prioritizing targets with higher state ages and larger estimation errors, thereby suppressing the long-term accumulation of tracking errors.
[0013] This invention solves for the sensing beamforming vector used for target state sensing and the communication beamforming vector used for data backhaul based on the sensing scheduling results, so that the same UAV antenna array can perform sensing and communication functions simultaneously or in time-division under shared spectrum resources. This ensures both the signal-to-noise ratio of the sensing echo to improve measurement accuracy and the efficient transmission of sensing data to the ground data fusion station. The ground data fusion station of this invention uses an extended Kalman filter method with multi-step prediction and backtracking correction to update the state and state age value of each moving target. This mechanism uses the received sensing data to backtrack and correct the state at historical moments, and then obtains the current and future target states through multi-step prediction. This effectively compensates for the time misalignment between data acquisition and processing, and ensures the tracking continuity and state estimation accuracy under non-uniform sensing intervals. At the same time, the updated state age value serves as the input for the next iteration, forming a closed-loop long-term tracking optimization, overcoming the limitation that traditional single-time-slot indicators cannot characterize the long-term cumulative effect of tracking errors.
[0014] Therefore, while ensuring the continuous tracking accuracy of multiple moving targets, the present invention achieves a balanced allocation of sensing resources among the targets, and is particularly suitable for complex scenarios in low-altitude intelligent networks where targets are highly mobile and communication links are dynamically changing. Attached Figure Description
[0015] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0016] Figure 1 This is the architecture diagram of the ISAC UAV multi-moving target tracking system proposed in Example 1.
[0017] Figure 2 This is the execution flowchart of the TH-SP algorithm proposed in Example 1.
[0018] Figure 3 This is a convergence performance graph of the TH-SP algorithm proposed in Example 1.
[0019] Figure 4 This is a diagram showing the trajectory planning results of the ISAC UAV obtained by the algorithm proposed in Example 1.
[0020] Figure 5 This is the AoS evolution diagram of each target obtained by the algorithm proposed in Example 1.
[0021] Figure 6 This is a comparison chart of the convergence and multidimensional performance of the algorithm proposed in Example 1 and the benchmark algorithm. Detailed Implementation
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0024] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0025] Terminology Explanation: ISAC: Integrated Communication and Sensing; DFS: Ground Data Fusion Station; PCRB: Posterior Cramerlow Bound; AoS: Status Age; EKF: Extended Kalman Filter; SAC: Flexible Motion Evaluation Algorithm; PPO: Proximity Policy Optimization Algorithm; TH-SP: A framework for evaluating and optimizing hierarchical flexible motions with dual time scales and proximal strategies; JFI: Jain Fairness Index; SCA: Continuous Convex Approximation; SDR: Semidefinite relaxation.
[0026] Example 1 One embodiment of the present invention provides a multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle (UAV), which divides the total service time into multiple large time slots, each large time slot further comprising multiple small time slots, and iterates over each large time slot, including: Step S1: Obtain the state of the integrated sensor drone under dual time scales in the current large time slot, including the drone's flight state at the beginning of the large time slot, the predicted state and state age value of all moving targets, and the remaining power and perception completion indicator at the beginning of the small time slot. Step S2: At the upper-level large time-slot scale, the UAV flight status and the predicted status of each moving target are input into the flexible action-evaluation network to plan the UAV's flight speed and direction within the large time slot in order to update its flight trajectory; Step S3: At the lower-level hourly slot scale, input the state age value, remaining power and perception completion flag of each moving target into the multi-head discrete near-end strategy optimization network to jointly optimize perception scheduling and decide the perception target and power coefficient of the current hourly slot UAV. Step S4: Based on the perception scheduling results, solve for the UAV's perception beamforming vector and communication beamforming vector to perform perception of moving targets and data communication with the ground data fusion station; Step S5: After receiving the sensing data of the current large time slot, the ground data fusion station uses the extended Kalman filter method with multi-step prediction and backtracking correction to update the state of each moving target and the state age value of each target, and then enters the next large time slot iteration.
[0027] As an example, in response to the application requirements of continuous tracking and data transmission of multiple moving targets in low-altitude intelligent networks, this invention proposes a dual-timescale sensor-integrated UAV multi-moving target tracking framework. It utilizes an ISAC UAV to continuously sense and track multiple ground moving targets and transmits the acquired sensed data to a ground data fusion station (DFS) to achieve real-time updates and data sharing of the global target status.
[0028] The main challenges that this framework aims to address include: 1) how to define an evaluation index that can characterize the evolution of long-term tracking accuracy and overcome the limitations of traditional single-slot indexes; 2) how to achieve accurate target state prediction and backtracking correction under non-uniform sensing intervals; and 3) how to handle mixed-integer optimization variables coupled across time scales and achieve efficient and stable policy learning in solving NP-hard problems.
[0029] To address the above challenges, the specific solution in this embodiment is as follows: (1) Construct the architecture of ISAC UAV multi-moving target tracking system.
[0030] like Figure 1 As shown, the system architecture includes: a UAV with ISAC functionality, a fixed-location ground data fusion station (DFS), and... K There are [number] ground moving targets, and the set of moving targets is [number]. , Indicates the first A moving target. ISAC UAV equipment includes... Uniform planar array (UPA) of antennas, number of antennas The spacing between adjacent antennas is half a wavelength.
[0031] To avoid interference between communication and sensing functions, the UAV flexibly switches between communication and sensing modes at a micro-timeslot granularity, sharing the same spectrum resources throughout the mission cycle. T Inside, the ISAC drone continuously tracked... K The system identifies a moving target and transmits the acquired sensing data to the DFS to achieve data aggregation and target status updates.
[0032] (2) A dual-timescale unmanned aerial vehicle (UAV) ISAC (TTU-ISAC) operation mechanism is proposed.
[0033] To coordinate the time granularity difference between low-frequency UAV trajectory adjustment (i.e., step S2) and high-frequency communication sensing beamforming update (step S3), this embodiment designs a TTU-ISAC dual-layer time slot structure.
[0034] Specifically, the total service time Divided into Each large time slot further contains [number] large time slots, and each large time slot further contains [number] large time slots. There are several hourly gaps of equal length, with each hourly gap lasting for a period of... In total · The flight direction and speed of the UAV are updated at the beginning of each large time slot, while the target sensing scheduling and communication sensing beamforming are updated in each small time slot.
[0035] In each large time slot Inside, the ISAC drone was the first to... Target perception is performed in one perception hour slot, and then in the remaining hour slot... Each communication hour slot will transmit the sense data to the DFS, where After receiving data, DFS updates the target state information. This process occurs during... N It executes in a loop within a large time slot until the task is completed.
[0036] Due to limitations in the size, weight, and power of UAVs, the antenna array size is limited. To avoid severe performance degradation caused by sensing too many targets simultaneously, the UAV senses 1 to 2 targets within each sensing time slot. K s There are several targets. To ensure reliable tracking and prevent over-sensing, each target is sensed exactly once within each large time slot; therefore, a binary scheduling indicator variable for the sensing task is defined. When the target moves k In the The value is 1 when the hour gap is sensed, and 0 otherwise.
[0037] Ground data fusion station DFS fixed at location At a fixed altitude, the ISAC drone... H Flights at various locations, in each large time slot n Adjust speed and direction at the start of the time slot, and reach the position at the end of the large time slot. Due to the short duration of the hourly slots, it is assumed that both the UAV and the moving target are in a quasi-stationary state within a single hourly slot; moving target k In the i The state vector of an hourly slot is defined as , Let k be the position of the moving target k in the i-th hour slot. Let be the first derivative of the position, and let represent the velocity information of the moving target k in the i-th hour slot.
[0038] (3) Construct an ISAC UAV communication sensing signal model.
[0039] Sensing signal model: Within a sensing time slot i, the UAV transmits a sensing beam toward the moving target to be sensed. The sensing beam serves as a carrier to transmit a composite sensing signal vector. It can be expressed as a weighted superposition of the sensing signals of each target, and can be represented by the formula: (1) in, For the corresponding sensing beamforming vector, To sense signals, To perceive the binary scheduling indicator variable of the task, when the target is moved k The value is 1 when the i-th hour slot is sensed, and 0 otherwise.
[0040] After the sensing beam reaches the target, it is reflected back to the UAV. The UAV receives the echo signal, which contains Doppler frequency shift, round-trip time delay, and noise components. The sensing channel between the UAV and the target adopts a probabilistic path loss model. The channel contains line-of-sight (LoS) and non-line-of-sight (NLoS) components. The probability of the LoS component is related to the elevation angle between the UAV and the target. In subsequent communication time slots, the DFS receives the sensing echo data sent by the UAV, performs matched filtering on it, and jointly estimates the time delay, Doppler, and angle parameters to obtain target measurement information, which is used for the backtracking estimation of the target state in the subsequent MPRC-EKF algorithm.
[0041] The effective signal-to-noise ratio for UAV target perception is expressed as: (2) in, For signal processing gain, For the sensing channel matrix, The variance is Gaussian noise. To sense beamforming vectors, For the trace operation of a matrix, the superscript H represents the conjugate transpose of the matrix or vector.
[0042] During communication time slots i Internally, the transmitted communication signal vector Communication beamforming vector and communication signals The product constitutes The communication channel also uses a probabilistic path loss model to obtain the communication signal-to-noise ratio. and transmission rate for: (3) (4) in, For communication channel matrix, For communication noise variance, For bandwidth.
[0043] The transmit power of each sensing and communication hour slot is constrained by the maximum power budget of the hour slot. To ensure the complete transmission of sensing data, a constraint needs to be imposed: the total transmission volume in the communication phase must not be less than the amount of data generated in the sensing phase, where the amount of data is proportional to the number of sensing hour slots, the signal sampling rate, and the number of quantization bits.
[0044] (4) Construct an extended Kalman filter (MPRC-EKF) target tracking model.
[0045] The Extended Kalman Filter (MPRC-EKF) framework proposed in this embodiment integrates multi-step prediction and backtracking correction mechanisms to address the time misalignment between sensing data acquisition and DFS processing. The specific process is as follows: At the beginning of each hourly slot, a state evolution model is used to predict the state of all targets. The target state evolution adopts a uniform motion model, and the state transition equation is: The covariance prediction equation is ,in, Here is the state transition matrix, and the transition noise is... The noise follows a zero-mean Gaussian distribution, and the covariance matrix of the transferred noise is: .
[0046] Assuming in the ( n -1) the first large time slot In a short interval, the ISAC drone targeted the target. k Perform sensing and obtain sensory echoes; in the ( n At the end of the -1)th large time slot, DFS receives the th Slot-level sensing echoes and predicted states temporarily stored in the drone Covariance Matrix DFS performs matched filtering on the sensed echo to obtain the measurement data vector. ,in (5) For two-way delay, For Doppler frequency shift, The Doppler shift is the trigonometric function value of the vertical angle between the UAV and the target. The value is the trigonometric function of the horizontal angle.
[0047] Moving target k In the i hourly interval measurement data vector With the target state The relationship between them is determined by a nonlinear observation model. Control, among which, It is a nonlinear measurement function. To measure the noise, the noise covariance matrix is: Its elements and perceived signal-to-noise ratio They are inversely proportional, which can be expressed by the formula: (6) Due to the time lag between data acquisition and processing, DFS performs backtracking correction: Using historical prediction status Covariance Matrix First, the nonlinear observation function Performing a first-order Taylor expansion at the historical prediction state yields the Jacobian matrix. Then, based on the standard EKF procedure, the Kalman gain is calculated: (7) Next, update the historical state vector and error covariance matrix to their posterior form: (8) (9) Subsequently, DFS estimates the target state at the start time of the large time slot n based on the posterior state using multi-step prediction. It is used to guide UAV flight direction and speed decisions at the beginning of a large time slot.
[0048] Large time slot n Within each hourly slot, the target state is iteratively predicted through a state evolution model, which is used for dynamic sensing beamforming and target scheduling. The UAV performs communication and sensing operations and temporarily stores sensing data. During the communication hourly slot, the data is transmitted back to the DFS to trigger subsequent backtracking corrections, forming a continuous closed-loop tracking process.
[0049] (5) Derive the posterior lower bound of Cramerol PCRB, and define the AoS index and AoS gain.
[0050] This embodiment uses PCRB to evaluate target tracking accuracy. PCRB integrates information from the state evolution model and the nonlinear observation model, providing a basic lower bound for the mean squared error (MSE). PCRB is defined by the inverse of the Fisher information matrix (FIM), and the overall FIM consists of the FIM of the measurement information contribution terms. FIM and prior state prediction information contribution It consists of two parts: (10) in, The Jacobian matrix represents the partial derivative of the measured data vector with respect to the target state vector. The covariance matrix representing the measurement noise, Represents the state transition matrix. for The covariance matrix in the EKF algorithm for hourly slot target k. The superscript T represents the transpose of the matrix, and the superscript (-1) represents the inversion of the matrix.
[0051] Accordingly, the target k In the hour gap i The PCRB matrix is defined as the inverse of the FIM: .
[0052] To evaluate long-term tracking performance, this embodiment introduces the concept of State Age (AoS), which is similar to Information Age (AoI). AoS is used to characterize the temporal evolution of target state uncertainty during the tracking process.
[0053] Specifically, AoS utilizes the dynamic evolution of PCRB for quantification. Without sensing updates, state estimation relies solely on prediction, and PCRB monotonically increases due to the accumulation of prediction noise, meaning the deviation between the estimated state and the true target state gradually widens. Conversely, when the target is sensed, measurement information reduces PCRB and suppresses uncertainty through FIM, thereby refreshing AoS. The PCRB matrix... The evolution equation is: (11) in, For the initialized EKF covariance matrix, The Jacobian matrix represents the partial derivative of the measured data vector with respect to the target state vector. The covariance matrix representing the measurement noise, Represents the state transition matrix. for The covariance matrix in the EKF algorithm for hourly slot target k. State transition noise, To sense the binary scheduling indicator variable for the task, Represents the PCRB matrix of the previous hour slot. The superscript T indicates the transpose of the matrix, and the superscript (-1) indicates the inversion of the matrix.
[0054] When the target is perceived ( The PCRB is jointly updated by measurement information and prior target state information; otherwise, the PCRB is recursively predicted step by step through the state transition equation.
[0055] The state age (AoS) is defined below as the trace of the PCRB matrix: Since DFS only updates the target state after receiving sensing data at the end of a large time slot, AoS is only refreshed at these moments; to quantize the large time slot... n Inner hour gap i Perceive target k The contribution of AoS to subsequent tracking performance is defined as the AoS gain. Multiply the AoS reduction caused by the sensing operation by the remaining duration of the effect: (12) in, For the duration of the hourly slot, the AoS reduction amount For this sensing of large time slots before and after nThe difference in the final PCRB trace; To the hourly gap i Starting from the modified PCRB, a large time slot was obtained through multi-step prediction. n The PCRB at the end.
[0056] (6) Construct the overall optimization problem.
[0057] This embodiment aims to maximize the weighted sum of the AoS gains of all targets and the Jain Fairness Index (JFI), which quantifies the degree of balance in the distribution of AoS gains among all targets; the optimization problem spans two time scales: ISAC UAV trajectory optimization at the large time slot level, and target scheduling and communication-aware beamforming determined at the small time slot level.
[0058] The optimization problem is structured as follows: (13) in, N This is the total number of large time slots. n It is a large timeslot index. K It is the total number of perceived targets. k It is a perception target index. Representing the n Within a large time slot, if in a small time slot i For the target k The resulting AoS gain is achieved through sensing. It is a weighting constant; , These represent the positions reached by the ISAC UAV at the end of time slot n-1 and n, respectively. This indicates the maximum flight speed of the UAV. It is the duration of the small gap. It refers to the number of small time slots within each large time slot; To sense the binary scheduling indicator variable for the task, Indicates a large time slot n The set of indices for all sensing hour slots within the timeframe. It is the maximum number of sensed targets within each hourly slot; For communication beamforming matrix, For the goal k The sensing beamforming matrix, This represents the average transmit power budget value for large time slots. The set of indexes representing communication hour slots, This represents the maximum power budget value for each hourly slot. It is the first i The transmission rate of a time slot, It is the first n The amount of sensing data that needs to be transmitted in each large time slot.
[0059] The specific objective function of this optimization problem is to maximize... N The cumulative value of the sum of AoS gains over each large time slot and the weighted sum of the Jain Fairness Index (JFI).
[0060] Constraint C1 limits the drone's flight speed to no more than the maximum speed, where The maximum flight speed is specified. Constraints C2-C4 define the target sensing scheduling strategy, including the binary nature of the scheduling indicator variable, that each target is sensed exactly once in each large time slot, and that the number of targets sensed in each small time slot is between 1 and... K s Between; Constraint C5 imposes an average transmit power budget constraint at the large time slot level, Constraint C6 imposes a maximum power budget constraint for each hour slot; Constraint C7 ensures that all sensing data is completely transmitted to the DFS during the communication phase.
[0061] (7) A solution framework for hierarchical flexible action-evaluation and proximal policy optimization (TH-SP) with dual time scales is proposed, and its execution flowchart is as follows. Figure 2 As shown, this embodiment decouples the optimization problem into three sub-problems: 1) Joint optimization of trajectory and target scheduling across time scales; 2) Sensing beamforming design; 3) Communication beamforming design.
[0062] First, the upper-layer Flexible Action-Evaluation (SAC) agent determines the UAV's flight speed and direction within a large time slot, a decision that remains constant throughout the smaller time slots. At the smaller time slot level, the lower-layer Multi-Head Discrete PPO algorithm optimizes the agent, performing target scheduling and power allocation.
[0063] Based on the above decisions, the SCA-SDR optimizer derives the sensing beamforming vector. SCA is a successive convex approximation method, and SDR is a semidefinite relaxation method. The remaining time and power resources are used for communication beamforming based on the fast water-filling algorithm. Finally, the objective function value is fed back as a reward to the upper and lower layer agents, driving the policy to evolve in the optimal direction.
[0064] A. Upper-level Markov Decision Process and SAC Algorithm Each large time slot n Initially, the ISAC UAV acquires an upper-level observation state. , represented as It includes the current large time slot index, the UAV target position in the previous large time slot, the current state age value of all targets, the predicted state vector of all targets, and the small time slot index of each target when it was last perceived.
[0065] At the large time-slot level, the actions that need to be decided are: This includes the flight speed of the drone in that large time slot. and flight direction The ISAC drones execute the decisions and actions within this large time slot and receive rewards. .
[0066] To capture the lasting impact of top-level decisions, top-level rewards... Defined as the sum of the lower-level rewards of all perceived smaller time slots within this large time slot: (14) in, For large time slots n The set of perceptual hour slot indexes in the middle, To sense hour gaps i The reward for the large time slot is obtained after the internal small time slot decision is completed and executed, i.e., at the end of the large time slot.
[0067] This embodiment solves the sequential decision problem in a continuous action space using the upper-level SAC algorithm. The SAC architecture includes an actor network. Two evaluation commentator networks and and two corresponding target critic networks and Within each large time slot, the agent observes the state and executes trajectory actions. After the large time slot ends, it receives a convergence reward and transitions to the next state. The transition tuple is stored in the experience replay buffer. .
[0068] During training Mid-sample mini-batch data is used to update the commentator network by minimizing the flexible Bellman residuals. (15) in, It is the damage function of the critic network. Network parameters representing the current critic network, j It is a critics' online index, superscript h This represents a network of upper-level commentators. Indicates from the experience replay pool Randomly sampled state-action pairs To calculate the expectation, and These represent the samples obtained from the playback pool. n The actual state of each large time slot and the actions taken at the higher level. This refers to the current network of critics. The target value updated for the commentator network. For the first n A large time slot reward, As a discount factor, This represents the TargetQ network. These are its parameters. Represents the state of the next major time slot. To adopt the latest actor network in the next state The new actions obtained below. These are the temperature parameters of the SAC algorithm. Representative actor's online output action The logarithmic probability.
[0069] The actor network updates its policy parameters by minimizing the expected KL divergence: (16) in, Let the loss function be the actor network. For the network parameters of the upper actor network, Indicates sampling from the playback pool And according to the current actor network Received The expected value is calculated; meanwhile, the target commentator network synchronizes the parameters through soft updates.
[0070] B. Lower-level Markov decision process and multi-head discrete strategy optimization algorithm Large time slot n At the start of each sensing hour slot within the timeframe, the ISAC drone acquires a lower-level state. It includes: current hourly slot index, number of undetected targets, remaining power budget, number of remaining hourly slots, current location of the UAV, state age value of all targets, predicted state of all targets, hourly slot index of each target last detected, and binary variables indicating whether each target has been detected in the current large hourly slot.
[0071] Based on the lower-level state, the drone obtains the lower-level actions. This includes the number of sensing targets that need to be sensed during that hourly slot. Target selection indicator vector Normalized sensing power coefficient The coefficient is from Selected from discrete levels, the maximum power available for that sensing time slot is determined. , This indicates the remaining available power within the large time slot up to the current small time slot.
[0072] To align the local decisions of the perceived time slots with the global objective of the optimization problem in (13), this embodiment employs reward shaping based on the potential function. The time slot reward consists of two parts: the AoS gain and the JFI increment. (17) Leveraging the scaling summation property, the cumulative local reward during the sensing phase within a large time slot precisely matches the single large time slot component in the global objective. Furthermore, to prevent aggressive sensing strategies from exhausting shared resources and violating communication constraints, a communication penalty term is introduced. Penalties are imposed when communication needs cannot be met. Otherwise, it is 0; the final compound reward is .
[0073] The lower-level MDP process uses the Proximal Policy Optimization (PPO) algorithm for decision-making, obtains hourly slot actions, and ensures the stability of policy updates through a pruning mechanism; the algorithm architecture includes an actor network. A network of critics To cope with multidimensional discrete action space In this embodiment, an action branching architecture is introduced into the actor network. The actor network is transformed into a structure where the output of a shared latent feature extractor is connected to three parallel decision heads. The first two decision heads output the number of perceived targets, respectively. and power coefficient The probability distribution is determined, and then specific actions are obtained through sampling. The third decision head is responsible for target selection and output. K The probability of selecting an objective.
[0074] Furthermore, to avoid redundant selection of already perceived targets, this embodiment integrates an action masking mechanism in the third decision head, which will then be applied to smaller time slots within the larger time slot. i Previously perceived targets are masked by imposing a significant penalty on their original probability values. Then, an effective target probability distribution is obtained by using a softmax function, from which samples are then taken without replacement. The goal is to obtain the index of the target that is perceived in that hour slot.
[0075] Stable updates are achieved by minimizing the total loss function using stochastic gradient descent. The total loss function is... It consists of three parts: the pruning agent strategy loss, the commentator mean squared error loss, and the entropy reward. (18) in, The operation represents the calculation of expectation on a finite sample. For lower-level actors' networks, Its network parameters. For the network parameters of the lower-level commentator network, For the first i The state of a time gap, and There are two weighting coefficients. As an entropy reward, it is used to encourage full exploration of the actor network. To reduce agent loss, it ensures that the magnitude of each update is controllable by limiting the range of the ratio between the old and new strategies. The ratio of the new strategy to the old strategy. and These represent the current actor network and the old actor network, respectively. This is the generalized advantage estimate (GAE). The cropping threshold, This is for the cropping operation. The mean squared error loss is used to optimize the estimation of the state value function. Representing the current network of critics, , Represents the old network of critics.
[0076] C. Perceptual Beamforming Design For each sensing hour slot i The goal of sensing beamforming is to achieve effective power budget To maximize the fairness of the AoS gain of the target sensed in this time slot and the AoS gain of all targets; let... Given the target set of the current time slot scheduling, the sensing beamforming subproblem can be written as: (19) in, For the goal k The sensing beamforming matrix, Representing the n Within a large time slot, if in a small time slot i For the target k The resulting AoS gain is achieved through sensing. It is a weighting constant. K It is the total number of perceived targets. k It is a perception target index. It is a known constant, representing the time from the beginning of this large time slot to the end of the smaller time slot. i The sum of the AoS gains of the previously perceived targets, b It is also a known constant, representing the time from the start of this large time slot to the start of the smaller time slot. i The sum of the squares of the AoS gain of the previously perceived target. c It is also a known constant, representing the time from the start of this large time slot to the start of the smaller time slot. iThe fairness exponent of the AoS gain of the previously perceived target is multiplied by .
[0077] According to the telescope summation property, if the fairness index of the AoS gain of all targets is 0 at the beginning of the large time slot, then the summation of the objective function values of all sensing small time slots in the large time slot is the objective function value in formula (13). By adopting the objective function in (19), the global target of the large time slot can be "sunk down" and the small time slot decision can be kept consistent with the global target in (13).
[0078] This embodiment solves the problem through the following steps: First, using positive semidefinite relaxation (SDR), the outer product matrix of the beamforming vectors is... By using the non-convex power constraint as an optimization variable and relaxing the rank-one constraint, the non-convex power constraint is transformed into a linear constraint. Secondly, auxiliary variables are introduced. This decouples the objective function containing complex fairness fractions into equivalent constraints and slack variables. For legacy non-convex fractional constraints, slack variables are introduced. By using variable substitution to decompose the problem, and by constructing a convex lower bound using algebraic identities and a first-order Taylor expansion, the non-convex terms such as bilinear and quadratic terms are completely convex using the SCA method. After the above transformation, the sensing beamforming subproblem represented by formula (19) is reconstructed into a standard SDP problem: (20) in, This is the maximum power budget for the hour slot sensing beamforming. Indicate target k Perceived beamforming matrix It needs to be a positive semi-definite matrix. This is the target set for the current time slot scheduling.
[0079] The SDP problem is solved iteratively. In each iteration, the CVX convex optimization toolbox configured with the MOSEK solver is invoked to solve the SDP problem, and the local reference point is continuously updated until convergence. Finally, if the obtained optimal covariance matrix does not satisfy the rank-one condition, the Gaussian randomization method is further applied to extract the approximately optimal sensing beamforming vector that satisfies the original power constraints. .
[0080] D. Communication beamforming design After the perception phase is completed, the remaining Communication hourly slots and residual power budget Used for communication transmission; since the ISAC UAV and DFS form a point-to-point channel, the optimal beamforming strategy under single-slot peak power constraints is Maximum Ratio Transmission (MRT), that is, the beamforming vector... Aligning the channel direction can be expressed by the formula: , For the first i Communication transmission power per hour slot This is the communication channel vector between UAC and DFS. This is for the modulus extraction operation. After alignment, communication beamforming simplifies to a power allocation problem: minimizing the total transmit power while satisfying peak power and total rate constraints, expressed by the formula: (twenty one) in, For the first n A set of communication hourly slot indexes for large time slots. This is the maximum power budget value within the hourly slot. The amount of sensing data that needs to be transmitted. The duration of the hourly gap. For available bandwidth, These are the normalized channel coefficients.
[0081] This embodiment employs the Fast Water Filling and Power Limiting (FWF-PP) algorithm to efficiently solve the power allocation problem. Specifically, the algorithm dynamically calculates the optimal water filling level based on the current channel state information, prioritizes allocating corresponding power according to the channel conditions, and limits and truncates allocation results that exceed the maximum transmit power limit of a single node / single link. Subsequently, the excess power is rapidly redistributed in the remaining links that have not reached the saturation limit, thereby solving the optimal power allocation scheme under the current conditions with extremely low computational complexity.
[0082] After obtaining this optimal power allocation, further verification is needed to determine whether the total allocated power satisfies the residual power budget constraint within the micro-time slot: (twenty two) in, This represents the optimal transmit power value. This represents the remaining power budget after the perception phase is completed within the large time slot, and also the maximum available power for the communication phase. If this constraint is met, the communication transmission is considered successful; otherwise, the communication is considered a failure, and the corresponding failure penalty signal is fed back to the deep reinforcement learning framework to drive the agent to update its subsequent action policy.
[0083] Finally, based on the above description, the complete process of the TH-SP algorithm is as follows: First, initialize the upper-layer SAC network (including actor network, dual critic network, target critic network, and entropy temperature parameter) and the lower-layer policy optimization network (including actor network and critic network).
[0084] In each training round, iterate through the following steps. N Each large time slot acquires the predicted state of all targets at the start of each large time slot. The upper-layer SAC actor network decides on flight speed and direction actions based on the current upper-layer state to update the UAV target position and initialize the remaining power budget.
[0085] Subsequently, in each sensing hourly slot within this large time slot, the lower-level actor network makes sensing scheduling decisions based on the lower-level state, namely the sensing target and power coefficient. The SCA-SDR optimizer then solves the sensing beamforming and executes the sensing operation, while simultaneously updating the remaining power and sensing reward.
[0086] After the sensing phase ends, the FWF-PP algorithm uses the remaining communication hour slots and residual power to complete the communication beamforming, and feeds back the communication penalty based on whether the transmission was successful.
[0087] Based on this, calculate the lower-level composite reward and the upper-level convergence reward, and store the corresponding transfer tuples into the buffer. and DFS updates the target state via MPRC-EKF.
[0088] Finally from The SAC network parameters are updated by sampling small batches of data. Update the parameters of the lower-level PPO network and clear it when sufficient data has been accumulated. .
[0089] The above process is iterated repeatedly in multiple training rounds until the policy converges.
[0090] After training, the upper and lower layer actor networks are deployed to the onboard computing platform of the ISAC UAV. Flight trajectory decision and perception scheduling decision are performed in each large time slot and small time slot, and online inference is performed in combination with the beamforming optimizer.
[0091] This embodiment demonstrates the simulation verification of the proposed dual-timescale integrated sensory UAV multi-moving target tracking framework. The simulation platform used in the experiment is a Python environment, and the deep learning framework used is PyTorch. The simulation parameters are set as follows: each large time slot contains... hourly gap, hourly gap duration s, carrier frequency GHz, bandwidth MHz, the number of UAV antennas is 4, the gain of the UAV and DFS antennas is 10dBi and 0dBi respectively; the maximum transmit power budget is and The deep reinforcement learning framework was trained for 15,000 rounds, with a discount factor. Both SAC and the lower-level policy optimization network use a fully connected network with two hidden layers and 512 neurons in each layer.
[0092] The convergence curves of the TH-SP algorithm and the comparison algorithm (fixed upper-level trajectory, PPO algorithm in the lower layer; fixed lower-level strategy, SAC algorithm in the upper layer) are as follows: Figure 3 As shown in the figure, it can be observed that all algorithms can converge stably, but TH-SP achieved the highest convergence reward.
[0093] The ISAC UAV's flight trajectory planning results and target tracking performance for a service duration of 30 seconds and a target quantity of 6 are as follows: Figure 4 As shown in the figure, the actual trajectory of the target and the estimated trajectory obtained by the proposed algorithm are basically consistent, and the ISAC UAV can follow the target's movement direction to improve the perception accuracy.
[0094] The AoS evolution curves of each target at the DFS end are as follows: Figure 5 As shown in the figure, the curve exhibits a typical sawtooth shape: AoS increases due to the accumulation of process noise and decreases after the end of each large time slot and the uploading of sensing data; AoS remains within the range of [0.3, 3] for most of the time period, indicating that the lower bound of the target state's Cramero is effectively constrained, ensuring accurate state perception.
[0095] The multidimensional performance comparison between the TH-SP algorithm proposed in this embodiment and two benchmark algorithms is as follows: Figure 6 As shown in the figure, the TH-SP algorithm achieved the highest total AoS gain and JFI value. This is attributed to the entropy regularization strategy of SAC, which enables the exploration of more diverse trajectories and avoids greedily favoring easily accessible targets. The performance of both the fixed trajectory benchmark and the fixed lower-level policy benchmark is significantly lower than that of TH-SP, verifying that both UAV trajectory and perception scheduling require adaptive optimization.
[0096] In summary, the tracking method provided in this embodiment has the following characteristics or advantages: 1) The Age of State (AoS) index and its gain definition are proposed. Based on the time evolution characteristics of the posterior Cramer-Rao lower bound (PCRB), it effectively captures the long-term cumulative effect of state estimation error in multi-target tracking and overcomes the limitation that traditional single-slot indexes cannot reflect long-term tracking performance. Furthermore, by combining the Jain fairness index, a joint optimization objective that takes into account both tracking accuracy and fairness among multiple targets is constructed.
[0097] 2) Design a dual-timescale UAV ISAC (TTU-ISAC) working mechanism, dividing the total service time into two levels: a large time slot and a small time slot. At the large time slot level, the flight trajectory of the UAV is planned and adjusted, while at the small time slot level, the target scheduling and communication sensing beamforming are rapidly optimized. This mechanism effectively matches the essential difference in time granularity between trajectory decision and signal-level decision, releasing performance potential while reducing decision-making overhead.
[0098] 3) A multi-step prediction and backtracking correction extended Kalman filter (MPRC-EKF) mechanism is proposed. To address the time misalignment between data acquisition and processing, the mechanism backtracks and corrects historical states while combining multi-step prediction to estimate the current target state, ensuring tracking continuity and accuracy under non-uniform sensing intervals.
[0099] 4) A dual-timescale hierarchical flexible actor-critic and near-end policy optimization (TH-SP) solution framework is proposed. The upper layer uses the flexible actor-critic (SAC) algorithm to optimize UAV trajectories in large time slots, and prevents the policy from converging to a suboptimal solution prematurely through maximum entropy regularization. The lower layer uses a multi-head discrete near-end policy optimization algorithm, which efficiently handles the scheduling of sensing targets and power allocation at the small time slot level through action branching architecture and masking mechanism. Under this framework, sensing beamforming is solved by successive convex approximation (SCA) and semidefinite relaxation (SDR) methods, and communication beamforming is obtained by fast water injection algorithm.
[0100] Example 2 One embodiment of the present invention provides a multi-moving target tracking system for an integrated sensor-controlled unmanned aerial vehicle (UAV), which divides the total service time into multiple large time slots, each large time slot further comprising multiple small time slots, and iterates over each large time slot, including: The state acquisition module is configured to acquire the state of the integrated sensor drone under dual time scales in the current large time slot, including the drone's flight state at the beginning of the large time slot, the predicted state and state age value of all moving targets, and the remaining power and perception completion indicator at the beginning of the small time slot. The trajectory update module is configured to: input the UAV flight state and the predicted state of each moving target into the flexible action-evaluation network at the upper-level large time slot scale, plan the UAV's flight speed and direction within the large time slot, and update its flight trajectory. The perception scheduling module is configured to: input the state age value, remaining power and perception completion identifier of each moving target into the multi-head discrete near-end strategy optimization network at the lower hourly slot scale, perform joint optimization of perception scheduling, and decide the perception target and power coefficient of the current hourly slot UAV; The beamforming module is configured to: solve the UAV's perception beamforming vector and communication beamforming vector based on the perception scheduling results, in order to perform perception of moving targets and data communication with the ground data fusion station; The state update module is configured as follows: after receiving the sensing data of the current large time slot, the ground data fusion station uses the extended Kalman filter method with multi-step prediction and backtracking correction to update the state of each moving target and the state age value of each target, and then enters the next large time slot iteration.
[0101] Example 3 One embodiment of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method for tracking multiple moving targets in a sensor-integrated unmanned aerial vehicle.
[0102] Example 4 In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the multi-moving target tracking method of a sensor-integrated unmanned aerial vehicle.
[0103] Example 5 One embodiment of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the multi-moving target tracking method of the integrated sensor drone.
[0104] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0106] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for tracking multiple moving targets using a sensor-integrated unmanned aerial vehicle (UAV), characterized in that, The total service time is divided into multiple large time slots, each of which further contains multiple smaller time slots. Each large time slot is iterated upon, including: Acquire the state of the integrated sensor drone under dual time scales in the current large time slot, including the drone's flight state at the beginning of the large time slot, the predicted state and state age value of all moving targets, and the remaining power and perception completion mark at the beginning of the small time slot; At the upper-level large time-slot scale, the flight status of the UAV and the predicted status of each moving target are input into the flexible action-evaluation network to plan the flight speed and direction of the UAV within the large time slot in order to update its flight trajectory. At the lower-level hourly slot scale, the state age value, remaining power, and perception completion identifier of each moving target are input into the multi-head discrete near-end strategy optimization network to jointly optimize perception scheduling and determine the perception target and power coefficient of the current hourly slot UAV. Based on the perception scheduling results, the perception beamforming vector and communication beamforming vector of the UAV are solved to perform perception of moving targets and data communication with the ground data fusion station. After receiving the sensing data of the current large time slot, the ground data fusion station uses the extended Kalman filter method with multi-step prediction and backtracking correction to update the state of each moving target and the state age value of each target, and then enters the next large time slot iteration.
2. The multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle as described in claim 1, characterized in that, In the dual time scale, the UAV performs flight trajectory planning and updating once at the beginning of each large time slot, and performs perception scheduling and beamforming updating once within each small time slot; Each large time slot contains a preset number of sensing hour slots and communication hour slots, where the sensing hour slots are used for target sensing and the communication hour slots are used to transmit the sensing data to the ground data fusion station.
3. The multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle as described in claim 1, characterized in that, The state age value is the trace based on the posterior Craméraud lower bound matrix; The objective function of the joint optimization is to maximize the weighted sum of the state age gains of each moving objective and the Jain fairness index; The state age gain is defined as the amount of state age reduction caused by the sensing operation multiplied by the remaining duration of the effect.
4. The multi-moving target tracking method for an integrated sensor-guided UAV as described in claim 1, characterized in that, The flexible motion-evaluation network employs a maximum entropy regularization strategy. Its actor network outputs motion probability distributions to encourage trajectory exploration, and its critic network is a dual-critic structure consisting of two evaluation networks and two target networks. The network parameters are updated by minimizing the flexible Bellman residual.
5. The multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle as described in claim 1, characterized in that, The multi-head discrete proximal policy optimization network includes a shared feature extractor and three parallel decision heads: the first decision head outputs the probability distribution of the number of perceived targets, the second decision head outputs the probability distribution of the power coefficient, and the third decision head outputs the selection probability of each target. The third decision head integrates an action masking mechanism to set the selection probability of the target that has been detected in the current large time slot to zero, so as to prevent repeated scheduling.
6. The multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle as described in claim 1, characterized in that, Solving for the sensing beamforming vector of an unmanned aerial vehicle (UAV) involves: introducing relaxation variables to transform the original problem into a semi-positive definite programming form; making the non-convex constraints convex through successive convex approximations; iteratively solving the semi-positive definite programming problem; and applying Gaussian randomization to the obtained covariance matrix to extract the approximate optimal sensing beamforming vector under rank-one constraints.
7. The multi-moving target tracking method for a sensor-integrated unmanned aerial vehicle as described in claim 1, characterized in that, Solving for the communication beamforming vector of the UAV specifically includes: aligning the beamforming vector with the channel direction between the UAV and the ground data fusion station using maximum ratio transmission; employing a fast water injection and power limiting algorithm to dynamically calculate the water injection level based on the channel state, limiting and truncating the allocation results that exceed the single-link power limit, and quickly redistributing the remaining power in the unsaturated link.
8. A multi-moving target tracking system for an integrated sensor-guided unmanned aerial vehicle (UAV), characterized in that, The total service time is divided into multiple large time slots, each of which further contains multiple smaller time slots. Each large time slot is iterated upon, including: The state acquisition module is configured to acquire the state of the integrated sensor drone under dual time scales in the current large time slot, including the drone's flight state at the beginning of the large time slot, the predicted state and state age value of all moving targets, and the remaining power and perception completion indicator at the beginning of the small time slot. The trajectory update module is configured to: input the UAV flight state and the predicted state of each moving target into the flexible action-evaluation network at the upper-level large time slot scale, plan the UAV's flight speed and direction within the large time slot, and update its flight trajectory. The perception scheduling module is configured to: input the state age value, remaining power and perception completion identifier of each moving target into the multi-head discrete near-end strategy optimization network at the lower hourly slot scale, perform joint optimization of perception scheduling, and decide the perception target and power coefficient of the current hourly slot UAV; The beamforming module is configured to: solve the UAV's perception beamforming vector and communication beamforming vector based on the perception scheduling results, in order to perform perception of moving targets and data communication with the ground data fusion station; The state update module is configured as follows: after receiving the sensing data of the current large time slot, the ground data fusion station uses the extended Kalman filter method with multi-step prediction and backtracking correction to update the state of each moving target and the state age value of each target, and then enters the next large time slot iteration.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement a multi-moving target tracking method for an integrated sensor-controlled unmanned aerial vehicle as described in any one of claims 1-7.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to perform a multi-moving target tracking method for an integrated sensor-controlled unmanned aerial vehicle as described in any one of claims 1-7.