A prediction error robust inter-sensory algorithm cooperative multi-unmanned aerial vehicle scheduling method
Patent Information
- Application Number
- CN202610756561.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-22
AI Technical Summary
这种预测状态与真实物理位置之间的客观偏差,会直接导致无人机发射波束未对准,从而引起通信链路质量的下降
Smart Images

Figure CN122802913A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, specifically to the field of wireless communication technology, and more specifically, to a method for scheduling multiple unmanned aerial vehicles (UAVs) in a cooperative manner with robust prediction error. Background Technology
[0002] Multi-UAV systems, with their superior mobility and rapid deployment capabilities, are becoming a transformative solution for providing ubiquitous network services in limited scenarios where ground communication infrastructure is weak, lacking, or damaged.
[0003] To provide continuous and high-quality communication services to ground users, drones require precise beam alignment and dynamic trajectory adjustments, which depend on acquiring real-time, accurate user location information. In this context, integrated sensing and communication (ISAC) technology has become a key enabling factor. This technology integrates communication and radar functions on a single platform, enabling drones to perform communication data transmission and active user positioning (i.e., acquiring high-precision user location information) with a single device.
[0004] However, the topology of multi-UAV networks is highly dynamic. If the system's resource allocation strategy relies solely on the instantaneously acquired user locations, unavoidable signal processing and computational delays will cause the user location information to become outdated by the time scheduling is executed, making it difficult to guarantee the continuity of service quality. Therefore, existing communication networks generally adopt a two-stage proactive scheduling framework of "prediction first, optimization later." This framework first predicts future user location information based on historical user location data, and then uses this predicted information to make UAV scheduling decisions in advance. This proactive mechanism effectively compensates for processing delays and ensures the stability and seamless connection of communication links.
[0005] Given the importance of accurate user location information for multi-UAV scheduling decisions, existing multi-UAV networks can acquire user locations through two methods: active sensing using ISAC and prediction. However, there is a fundamental trade-off between the two. Specifically, high-frequency active sensing based on ISAC can acquire high-precision real-time user locations, effectively ensuring accurate beam alignment. However, frequent sensing inevitably consumes time-frequency resources that could be used for data transmission. Conversely, while prediction-based active scheduling frameworks save sensing overhead and compensate for processing latency, they largely rely on the idealized assumption of "completely accurate prediction." In real-world dynamic networks, due to the randomness of user movement and the limitations of location prediction models, location prediction errors are unavoidable. This objective deviation between the predicted state and the actual physical location directly leads to misalignment of the UAV's transmitted beam, thereby causing a degradation in communication link quality.
[0006] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solution of the present invention, and does not imply that the relevant information is necessarily prior art. The relevant information was submitted and disclosed together with the present invention, and should not be considered prior art unless there is evidence that the relevant information was disclosed before the filing date of the present invention. Summary of the Invention
[0007] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a method for scheduling multiple unmanned aerial vehicles (UAVs).
[0008] The objective of this invention is achieved through the following technical solution:
[0009] According to a first aspect of the present invention, a multi-UAV scheduling method is provided. This method uses multiple mobile UAVs as access points to provide wireless communication services to one or more users within a coverage area. The method includes: constructing a time-division sensing integrated frame structure for joint sensing and prediction, wherein the time for UAVs to perform tasks is divided into multiple periods of joint sensing and prediction, each period containing multiple frames, which are divided into sensing frames and prediction frames, and the proportion of sensing frames can be dynamically adjusted through a ratio variable; acquiring user locations, wherein, based on sensing integration technology, the actual user location is actively sensed and obtained in some time slots of the sensing frame, and the predicted user location is obtained based on the user's historical location in all time slots of the prediction frame; estimating the UAV's transmit beam angle deviation based on the error statistics of the historical predicted user location, and constructing a communication rate model for location prediction error perception between the UAV and the user to characterize the quantitative mapping relationship from deviation to communication rate attenuation; and, based on the communication rate model, jointly optimizing the ratio variable, the deployment location of each UAV, the association indication matrix of the served user, and the bandwidth allocation matrix to obtain a multi-UAV scheduling decision, with the objective of maximizing the total communication rate when location prediction error exists. This scheme can achieve at least the following beneficial technical effects: by setting up a time-division sensing integrated frame structure, the proportion of sensing frames can be dynamically adjusted efficiently and flexibly by a ratio variable; by establishing a quantitative mapping from position prediction error to communication rate attenuation, scheduling decisions can explicitly consider the impact of prediction error, thereby improving the system's communication performance under dynamic environments and prediction uncertainties; with the goal of maximizing the total communication rate when position prediction error exists, by jointly optimizing the ratio variable, deployment location, user association, and bandwidth allocation, the scheme achieves the coordinated and efficient utilization of multi-UAV resources, achieving the optimal trade-off between sensing accuracy and communication capacity.
[0010] Optionally, each of the multiple frames includes multiple time slots, wherein a single fixed time slot is designated in all sensing frames for transmitting signals to perform user location sensing, and the remaining time slots in the sensing frames are used for wireless communication. This scheme can achieve at least the following beneficial technical effects: by performing location sensing through a single time slot in the fixed sensing frame, the sensing scheduling mechanism is simplified and sensing overhead is reduced; the remaining time slots in the sensing frame are used for communication, improving the utilization rate of time and frequency resources, and maximizing communication data transmission time while ensuring location calibration capabilities.
[0011] Optionally, joint optimization is achieved through a multi-agent deep reinforcement learning algorithm, wherein: the multi-UAV scheduling problem is modeled as a Markov decision process, with each UAV agent participating in the decision-making; the local observation state of each UAV agent is defined as including the UAV position in the previous time slot, the positions of all users obtained in the current time slot, and the channel state information of the previous time slot; the local actions of each UAV agent are defined as including the individual expected ratio, the deployment position of the UAV, the association indication matrix, and the bandwidth allocation matrix; during scheduling, each UAV agent independently predicts its local actions based on its own local observation state, and optimizes its individual expected ratio, deployment position, association indication matrix, and bandwidth allocation matrix separately, but uses the ratio variable allocated in the current period to replace the individual expected ratio when executing local actions; when updating parameters during the reinforcement learning process of the UAV agents, the parameters of each UAV agent are updated by backpropagation with the goal of maximizing the total communication rate, and the ratio variable for the next period is selected based on the individual expected ratio of the most recently predicted local actions of each UAV agent. This scheme can achieve at least the following beneficial technical effects: by incorporating the ratio variable into the distributed action space, each UAV agent can autonomously explore the optimal perception frequency during training; by determining a unified ratio variable through global optimization, the flexibility of distributed training and the global consistency of frame structure synchronization can be balanced; and by using a reward-driven approach that maximizes the total communication rate, the trained UAV agents can make robust scheduling decisions in the event of prediction errors.
[0012] Optionally, during the scheduling process, each UAV agent collects its own real samples and places them into an experience replay pool. Real samples include the UAV agent's local observation state in the current time slot t, the local action performed in the current time slot t, the reward obtained after performing the local action, and the next local observation state entered after performing the local action. This scheme can achieve at least the following beneficial technical effects: by standardizing the composition of samples in the experience replay pool (state-action-reward-next state), it provides a standard data input format for reinforcement learning algorithms; by enabling each UAV to independently collect its own samples, it supports distributed experience replay, improving the parallelism and efficiency of training data collection.
[0013] Optionally, the reinforcement learning algorithm is configured as follows: the joint action value function, which represents the expected total reward of multiple UAV agents performing joint actions in a global state, is decomposed into a weighted sum of the local value functions of each UAV agent. The weights of the weighted sum are dynamically generated by a multi-head self-attention mechanism based on the importance and reliability of the current local observation state of each UAV agent. A conditional generative diffusion model is introduced as a scene generator, which takes real samples from the experience replay pool as input and gradually reconstructs simulated samples containing position prediction error features by performing forward noise addition and backward noise reduction on the real samples. The simulated samples are mixed with real samples and stored in the experience replay pool to enhance the diversity of training data. Based on the real and simulated samples in the experience replay pool, the policy network in the multiple UAV agents is trained by maximizing the cumulative reward. The trained policy network is used to predict local actions. This scheme can achieve at least the following beneficial technical effects: it solves the multi-agent credit allocation problem through value decomposition networks, enabling each UAV to make independent decisions based on local observations while coordinating to optimize the global objective; it dynamically adjusts the contribution weights of each agent through a multi-head self-attention mechanism, enhancing the ability to select state features under prediction error environments; and it synthesizes simulated samples containing prediction error features through a diffusion model, expanding the diversity of training data and improving the generalization ability and robustness of the policy network to prediction uncertainty.
[0014] Optionally, the predicted user location is determined by a Transformer-based user location prediction model, which includes: a data input module for receiving a historical trajectory sequence composed of the user's historical real locations acquired from sensing frames; a linear embedding module for mapping the location coordinate vector of each time slot in the historical trajectory sequence to a higher-dimensional feature space to obtain a location feature matrix; a location encoding module for element-wise superimposing learnable location codes in each location feature matrix and assigning a time identifier to each time step to obtain location feature sequence data after superimposed location encoding; a multi-head self-attention feature extraction module for feeding the location feature sequence data after superimposed location encoding into a multi-layer Transformer encoder, extracting trajectory features in parallel in different representation subspaces through a multi-head self-attention mechanism to obtain a location feature vector that fuses local and long-term spatiotemporal correlations; and a location prediction module for processing the location feature vector through residual connections, layer normalization, a feedforward neural network, and a coordinate regression network with linear mapping to output the predicted user location. This scheme can achieve at least the following beneficial technical effects: by capturing the long-distance spatiotemporal dependence of user movement trajectory through the Transformer architecture, it overcomes the limitations of traditional algorithms in handling highly dynamic and nonlinear trajectories; by extracting local state correlations and long-term movement trends in parallel through a multi-head self-attention mechanism, it improves the accuracy of position prediction; and by preserving temporal causal relationships through learnable position encoding, it enables the model to implicitly recover key motion attributes such as velocity and acceleration.
[0015] Optionally, the communication rate model is constructed as follows: A signal-to-interference-plus-noise ratio (SIR) model is established between the UAV and the user, where the UAV transmit beam angle deviation caused by the variance of the user position prediction error is explicitly embedded into the effective beamforming gain term, causing the value of this gain term to decrease as the deviation increases; the SIR determined by the SIR model is substituted into the Shannon capacity formula to obtain the communication rate model. This scheme can achieve at least the following beneficial technical effects: by explicitly embedding the variance of the user position prediction error into the SIR model, the impact of prediction error on communication performance is quantitatively introduced; by designing a mechanism for the effective beamforming gain to decrease as the deviation increases, the communication rate model more accurately reflects the actual achievable rate under prediction uncertainty, providing a more accurate performance evaluation basis for subsequent joint optimization.
[0016] Optionally, the communication rate model can be expressed as:
[0017]
[0018] in, This represents the communication rate for prediction error perception between UAV k and user n in time slot t. This represents the communication bandwidth between drone k and user n. Let represent the signal-to-interference-plus-noise ratio (SIR) including prediction error between UAV k and user n in time slot t. This scheme can achieve at least the following beneficial technical effects: it quantitatively analyzes the direct impact of unavoidable prediction errors on user communication performance, and provides mathematical expressions for the actual SIR including prediction error and the prediction error-aware communication rate.
[0019] Optionally, the signal-to-interference-plus-noise ratio (SIR) is determined by the following SIR model:
[0020]
[0021] in, This represents the transmit power of UAV k to user n in time slot t. This is the reference channel power located 1 meter from the UAV's transmitting antenna. It is the carrier wavelength. This indicates the position of UAV k in time slot t. This indicates the position of user n in time slot t. This represents the probability that there is a line-of-sight link between drone k and user n. This represents the probability of a non-line-of-sight link. It refers to the number of drone transmitting antennas. Indicates the beamwidth factor. It is the angle between drone k and user n. The variance representing the user location prediction error. Indicates the non-line-of-sight attenuation factor. This represents the transmit power of UAV i to user j in time slot t. This represents the channel vector between UAV k and user n in time slot t under the probabilistic channel model. This represents the beamforming vector between UAV i and user j. , This represents the antenna array response vector function. This represents the actual physical angle between drone i and user j in time slot t. This represents the transmit beam angle deviation of UAV i relative to user j caused by the predicted position error. This represents the power of additive white Gaussian noise. This scheme achieves at least the following beneficial technical effects: by using a signal-to-interference-plus-noise ratio (SINR) model that includes prediction error, it breaks through the ideal assumption of "completely accurate prediction" in traditional scheduling mechanisms, providing a performance analysis basis for real-world communication scenarios with prediction errors; by comprehensively considering the impact of line-of-sight and non-line-of-sight links through a probabilistic channel model, it improves the accuracy of channel estimation; by using the functional relationship between array antenna gain and prediction error variance, it accurately quantifies the gain loss caused by beam misalignment; and by separating and modeling interfering links and desired links, it accurately assesses the impact of co-channel interference on communication quality, providing refined channel state information for UAV deployment and correlation decisions.
[0022] According to a second aspect of the present invention, a wireless communication system is provided, comprising a plurality of unmanned aerial vehicles (UAVs) configured to: collaboratively execute a multi-UAV scheduling method as described in the first aspect to obtain a multi-UAV scheduling decision including a ratio variable, deployment locations of each UAV, a user association indication matrix, and a bandwidth allocation matrix; according to the multi-UAV scheduling decision, each UAV flies to the deployment location corresponding to the decision; sets a sensing frame and a prediction frame according to the ratio variable in the decision, and performs user location sensing in a specified time slot of the sensing frame; and according to the bandwidth allocated by the bandwidth allocation matrix in the decision, each UAV provides wireless communication services to the user specified by the user association indication matrix in the decision. Attached Figure Description
[0023] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0024] Figure 1 This is a flowchart illustrating a multi-UAV scheduling method according to an embodiment of the present invention;
[0025] Figure 2 This is a schematic diagram illustrating the principle of the time-division synesthesia integrated frame structure for joint sensing and prediction according to an embodiment of the present invention.
[0026] Figure 3This diagram illustrates the performance comparison between the location prediction model used in the multi-UAV scheduling method according to an embodiment of the present invention and existing models under different user mobility stochasticities.
[0027] Figure 4 This is a schematic diagram comparing the multi-UAV scheduling method according to an embodiment of the present invention with other existing scheduling schemes in terms of total user communication data rate. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.
[0029] As mentioned in the background section, in real-world dynamic networks, prediction errors are unavoidable due to the randomness of user movement and the limitations of prediction algorithms. This objective deviation between the predicted state and the actual physical location directly leads to misalignment of the UAV's transmission beam, resulting in a decrease in communication link quality. This invention addresses the fundamental trade-off between active sensing and prediction in ISAC (Interactive Sensing and Prediction), as well as the dependence of existing prediction scheduling frameworks on the ideal assumption of completely accurate predictions. It proposes a multi-UAV scheduling method, achieved through the following technical means: First, a time-division integrated sensing frame structure for joint sensing and prediction is constructed. The time for UAV task execution is divided into multiple cycles, each containing two types of frame structures: sensing frames and prediction frames. The proportion of sensing frames is dynamically adjusted through a ratio variable, achieving an adaptive balance between sensing accuracy and communication capacity. Second, a communication rate model for location prediction error perception is established. Addressing the deficiency of existing technologies that ignore prediction errors, this invention estimates the UAV's transmission beam angle deviation based on historical error statistics between predicted and actual user locations. This deviation is explicitly embedded into the effective beamforming gain term, constructing a quantitative mapping relationship from transmission beam angle deviation to communication rate attenuation. This model breaks through the ideal assumption of completely accurate predictions, enabling communication rate assessments to accurately reflect the actual achievable performance even with prediction errors, providing accurate performance data for subsequent scheduling optimization. Third, it jointly optimizes sensing, prediction, and UAV location and bandwidth allocation scheduling strategies. Based on a pre-defined communication rate model, it aims to maximize the total communication rate even with location prediction errors, jointly optimizing ratio variables, UAV deployment locations, the association indication matrix of served users, and the bandwidth allocation matrix. By explicitly incorporating prediction errors into the optimization objective, scheduling decisions collaboratively adapt to prediction uncertainties across multiple dimensions, including sensing resource allocation, UAV deployment, user association, and bandwidth utilization, thereby achieving robust and efficient communication in dynamic environments.
[0030] According to one embodiment of the present invention, a multi-UAV scheduling method is provided, which uses multiple mobile UAVs as access points to provide wireless communication services to one or more users within a coverage area. See also Figure 1 The multi-UAV scheduling method includes steps S1, S2, S3, and S4. To better understand this invention, each step will be described in detail below with reference to specific embodiments.
[0031] Step S1: Construct a time-division integrated sensing frame structure for joint perception and prediction, wherein the time for the UAV to perform the task is divided into multiple cycles of joint perception and prediction, each cycle contains multiple frames, the multiple frames are divided into perception frames and prediction frames, and the proportion of perception frames can be dynamically adjusted by a ratio variable.
[0032] According to one embodiment of the present invention, considering the conflict between active sensing and communication resources, a time-division integrated sensing frame structure for joint sensing and prediction is proposed, such as... Figure 2 As shown. This embodiment divides the time for the UAV to perform its mission into multiple joint perception and prediction cycles, where each cycle contains multiple frames, and each frame includes multiple time slots. For example, each cycle contains... The system consists of L frames, each containing L time slots. Each frame within a period can be functionally divided into sensing frames and prediction frames. To balance sensing accuracy and communication capacity, the system introduces a ratio variable between joint sensing and prediction frames to control the proportion of sensing frames within a period. For example, for a joint sensing and prediction period, the previous... Each frame is a perception frame. This represents a ratio variable. This design breaks away from the limitations of traditional fixed sensing cycles, allowing the ratio variable to be increased when the user's position changes drastically in the environment. The value is increased to improve the active sensing frequency, thereby obtaining rich real-time calibration data and preventing beam misalignment caused by the accumulation of prediction errors; when the user trajectory is stable and easily predictable, the ratio variable can be reduced. The value of is adjusted to maximize the allocation of time-frequency resources to communication data transmission. This mechanism dynamically controls the trade-off between proactive sensing computation and predictive computation based on historical data, significantly improving the overall spectral efficiency and communication throughput of the system while ensuring the accuracy of location calibration.
[0033] According to one embodiment of the present invention, a single fixed time slot is designated in all sensing frames for transmitting signals to achieve user location sensing, and the remaining time slots in the sensing frames are used for wireless communication. A specific time slot in the sensing frame, such as the nth time slot, is designated as the active sensing time slot. Taking n=1 as an example, in the first time slot of the sensing frame, the UAV actively transmits signals to achieve user location sensing and obtain sensing data. This sensing data can not only be used to obtain the actual user location but also serve as a calibration point for user location prediction, thereby reducing the error of subsequent user location prediction. It should be noted that since the duration of a single frame is extremely short (e.g., 5G NR specifies a frame as 10ms), the change in user location between two frames is minimal; therefore, only a single time slot needs to be set for sensing in a sensing frame. Furthermore, considering that sensing would occupy communication resources that could otherwise be used for data transmission, continuous and frequent active sensing is unnecessary. Therefore, in the prediction frame, all time slots are used for communication data transmission, and the system relies entirely on the location deduced by the prediction model for UAV communication scheduling.
[0034] Step S2: Obtain user location. Based on the integrated sensing technology, the real user location is actively sensed and obtained in some time slots of the sensing frame, and the predicted user location is obtained in all time slots of the prediction frame based on the user's historical location.
[0035] According to one embodiment of the present invention, user location perception can be implemented based on ISAC technology. For predicting user location, it can be determined by a time-series model, such as an LSTM model, based on historical user locations. Preferably, user location prediction can also be determined by a Transformer-based user location prediction model, which includes: a data input module, a linear embedding module, a location encoding module, a multi-head self-attention feature extraction module, and a location prediction module. Wherein:
[0036] The data input module receives historical trajectory sequences composed of the user's historical real-time locations acquired from the perception frames. Within these perception frames, the UAV actively senses the target user's past... The actual user location coordinates for each time slot are constructed, with a length of [length missing]. The historical trajectory sequence.
[0037] A linear embedding module is used to map the position coordinate vector of each time slot in the historical trajectory sequence to a higher-dimensional feature space, obtaining a position feature matrix. Since the original user position coordinate vector features have a low dimension (2D or 3D), direct processing is difficult to extract high-order features. Therefore, a linear embedding layer can be used to first map the position coordinate vector of each time slot in the historical trajectory sequence to a higher-dimensional feature space, resulting in a position feature matrix. In the high-dimensional feature space, a higher-dimensional position feature matrix is generated.
[0038] The location encoding module is used to element-wise superimpose learnable location codes into each location feature matrix, assigning a time identifier to each time step to obtain the location feature sequence data after superimposed location encoding. Since the Transformer architecture itself abandons the sequential processing logic of traditional recurrent neural networks (RNNs), in order to preserve the strict temporal causality of trajectory sequences in parallel computing, the model superimposes learnable location codes element-wise into the location feature matrix. That is, the learnable location codes are added element-wise to the location feature matrix to obtain the location feature sequence data after superimposed location encoding. This location encoding assigns a unique time identifier to each time step, enabling the model to implicitly recover and distinguish key time-series attributes such as velocity and acceleration in user mobility data without a recurrent structure.
[0039] The multi-head self-attention feature extraction module feeds the positional feature sequence data after superimposed position encoding into a multi-layered Transformer encoder. Through a multi-head self-attention mechanism, trajectory features are extracted in parallel across different representation subspaces, resulting in a positional feature vector that fuses local and long-term spatiotemporal correlations. The positional feature sequence data after superimposed position encoding is fed into a multi-layered Transformer encoder. Inside the encoder, the core feature extraction is performed by the multi-head self-attention mechanism. Illustratively, the model linearly maps the input sequence data into a query matrix (Q), a key matrix (K), and a value matrix (V), respectively. The self-attention weight matrix is obtained by calculating the dot product of the query matrix (Q) and the key matrix (K) and performing Softmax normalization. This self-attention weight matrix is used to perform weighted aggregation on the value matrix (V), generating an attention output that incorporates global contextual information. This multi-head mechanism allows the model to capture trajectory features simultaneously in different representation subspaces. For example, some attention heads focus on the local state associations of adjacent time slots to accurately identify short-term drastic trajectory fluctuations (such as a user suddenly turning or braking); while other attention heads span longer time steps to capture long-term movement direction trends (such as straight driving on a main road).
[0040] The location prediction module processes the location feature vector through residual connections, layer normalization, a feedforward neural network, and a coordinate regression network with linear mapping, outputting the predicted user location. The location feature vector, after being processed by a multi-head self-attention mechanism, undergoes residual connections and layer normalization to prevent gradient vanishing and accelerate convergence. It is then fed into a feedforward neural network composed of fully connected layers for nonlinear spatial transformation, yielding processed high-dimensional hidden features. Finally, the processed high-dimensional hidden features are compressed back into a spatial vector through a linear mapping layer (coordinate regression network), thus outputting the predicted user location. During the user location prediction model training phase, user trajectory sequences are acquired as training samples, mean squared error is used as the loss function, and the parameters of the user location prediction model are continuously optimized through backpropagation until the prediction model converges.
[0041] This embodiment employs a Transformer-based prediction model to predict user location, abandoning the word vector embedding and classification output layers traditionally used by Transformers for processing discrete text. At the feature input end, a linear embedding module for physical coordinates is designed to map low-dimensional position vectors to a high-dimensional feature space. At the output end, a coordinate regression network replaces the traditional Softmax classifier, recompressing the high-dimensional hidden layer features into continuous location space coordinates. This enables the Transformer model to accurately adapt to highly dynamic and continuous user spatial displacement changes, providing the ability to process continuous spatiotemporal trajectory data in wireless communication networks.
[0042] Step S3: Estimate the transmit beam angle deviation of the UAV based on the historical error statistics of user location prediction, and construct a communication rate model for position prediction error perception between the UAV and the user to characterize the quantitative mapping relationship from the deviation to the communication rate attenuation.
[0043] According to one embodiment of the present invention, the communication rate model is constructed in the following manner: a signal-to-interference-plus-noise ratio (SIR) model between the UAV and the user is established, wherein the UAV transmit beam angle deviation caused by the variance of the user position prediction error is explicitly embedded into the effective beamforming gain term, such that the value of the gain term decreases as the deviation increases; the SIR determined by the SIR model is substituted into the Shannon capacity formula to obtain the communication rate model.
[0044] For example, suppose there are K drones serving N users. Due to errors in the output of the user location prediction model, the actual signal-to-interference-plus-noise ratio (SIR / N) between drone k and user n, including the prediction error, is given as follows:
[0045]
[0046] in, This represents the transmit power of UAV k to user n in time slot t. It is the reference channel power at 1 meter. It is the carrier wavelength. This indicates the position of UAV k in time slot t. This indicates the position of user n in time slot t. This represents the probability that there is a line-of-sight link between drone k and user n. This represents the probability of a non-line-of-sight link. It refers to the number of drone transmitting antennas. Represents the beamwidth factor, where It is the angle between drone k and user n. This represents the variance of the user location prediction error, which is related to the root mean square error of the user location prediction model output. Indicates the non-line-of-sight attenuation factor. This represents the transmit power of UAV i to user j in time slot t. The channel vector between UAV k and user n in time slot t under the probabilistic channel model is given by the following equation:
[0047]
[0048] The line-of-sight channel part , The antenna array response vector is given by the following formula:
[0049]
[0050] in Indicates the spacing between antenna elements. The sine of the angle between time slot t, UAV k, and user n is represented. The non-line-of-sight channel portion is modeled as the product of large-scale path loss and small-scale fading. ,in It is a complex Gaussian random vector with zero mean and unit covariance. The beamforming vector between UAV i and user j is given by the following equation:
[0051]
[0052] in, This represents the antenna array response vector function. This represents the actual physical angle between drone i and user j in time slot t. This represents the deviation of the transmit beam angle of UAV i from user j caused by the predicted position error. This represents the power of additive white Gaussian noise.
[0053] This signal-to-interference-plus-noise ratio (SIR) model breaks through the idealized assumption of "completely accurate prediction results" in existing scheduling mechanisms. It establishes a mathematical mapping relationship from algorithm prediction error to antenna beam alignment deviation, and then to effective SIR attenuation, enabling the system to quantitatively assess the impact of unavoidable prediction errors on the actual communication rate.
[0054] Finally, based on the above signal-to-interference-plus-noise ratio (SIR) model, the communication rate for error perception between UAV k and user n is predicted. It can be represented as:
[0055]
[0056] in This represents the communication bandwidth between drone k and user n. This represents the signal-to-interference-plus-noise ratio (SINR) between time slot t, UAV k, and user n, including prediction errors.
[0057] Step S4: Based on the communication rate model, with the objective of maximizing the total communication rate when there is a location prediction error, jointly optimize the ratio variable, the deployment location of each UAV, the association indication matrix of the users served, and the bandwidth allocation matrix to obtain the multi-UAV scheduling decision.
[0058] According to one embodiment of the present invention, in order to maximize the communication rate perceived by the user for the total prediction error, this embodiment models the UAV scheduling problem with prediction error as a Markov decision process, the illustrative definition of its core elements is as follows:
[0059] State Space: The state space represents the basic environmental information that enables UAV agents to make decisions. For each UAV agent, its local observation state vector consists of system state parameters that are crucial for collaborative decision-making, including: the UAV's position in the previous time slot, the positions of all users obtained in the current time slot, and the channel state information from the previous time slot.
[0060] Action Space: The action space defines the control dimensions that a UAV agent can execute in each state. Each UAV agent's action vector encompasses four continuous and discrete control variables: 1) Individual Expectation Ratio (equivalent to the ratio variable of individual expectations), used to control the ratio of actively perceived frames to predicted frames; 2) UAV deployment location (three-dimensional spatial coordinates) to achieve dynamic trajectory deployment; 3) UAV-user association indication matrix, determining the target users served by the UAV in its current time slot; and 4) Bandwidth allocation matrix, used to orthogonally allocate spectrum resources to each associated user. For example, the association indication matrix... , where each element , This indicates that user n is served by drone k. Bandwidth allocation matrix. , where each element This represents the bandwidth allocated to user n for drone k.
[0061] Reward Function: The reward function drives the cooperative optimization direction of multiple UAVs. This application sets maximizing the total communication rate of all UAVs as the joint optimization objective. Joint optimization is achieved through a multi-agent deep reinforcement learning algorithm.
[0062] According to one embodiment of the present invention, during scheduling, each UAV agent independently predicts its local actions based on its own local observation state, and optimizes its individual expected ratio, deployment location, association indication matrix, and bandwidth allocation matrix separately. However, when executing local actions, the ratio variable allocated in the current cycle is used instead of the individual expected ratio. Furthermore, the joint action value function, representing the expected total benefit of multiple UAV agents performing joint actions in the global state, is decomposed into a weighted sum of the local value functions of each UAV agent. Illustratively, a robust diffusion-driven value decomposition network algorithm is proposed to overcome the extreme sensitivity of conventional multi-agent reinforcement learning to prediction errors. This algorithm collaboratively integrates a conditional generative diffusion model and an attention-enhanced value decomposition network, achieving efficient multi-UAV collaborative scheduling in dynamic networks with prediction uncertainty. First, in the value decomposition network architecture, the system follows a centralized training and decentralized execution framework, effectively solving the credit allocation problem in multi-agent reinforcement learning. Its core approach is to decompose the joint action value function, representing the expected total benefit of the entire multi-UAV team performing joint action a in the global state s, into a value decomposition network. It is decomposed into a weighted sum of the local value functions of each UAV intelligent agent. The calculation formula is:
[0063]
[0064] in, This represents the adaptive learning weights dynamically generated through a multi-head self-attention mechanism, used to adjust the contributions of each agent based on the importance and reliability of the current state information. For the parameters of each independent Q network, This represents the local action value function of the k-th drone agent. This represents the local observation state of the k-th UAV agent. This represents the local action performed by the k-th drone agent. This mechanism enables the drone to selectively focus on the most reliable state features, enhancing its coordinated generalization ability in the face of prediction errors.
[0065] Furthermore, during the reinforcement learning process of the UAV agents, when updating parameters, the goal is to maximize the total communication rate. The parameters of each UAV agent are updated via backpropagation, and the ratio variable for the next cycle is selected based on the individual expected ratio of each agent's most recently predicted local actions. Illustratively, in the distributed training phase, each UAV agent uses the individual expected ratio as an exploration variable; in the global parameter update phase, through gradient aggregation and the goal of maximizing the total communication rate, the ratio variable for unified execution in the next cycle is determined from the most recent individual expected ratios of each agent, achieving a balance between distributed exploration and global consistency.
[0066] According to one embodiment of the present invention, during the scheduling process, each UAV agent collects its own real samples and places them into an experience replay pool. The real samples include the UAV agent's local observation state in the current time slot t, the local action performed in the current time slot t, the reward obtained after performing the local action, and the next local observation state entered after performing the local action. To address the problem of insufficient training data diversity caused by prediction uncertainty, a conditional diffusion model is introduced as a high-fidelity scene generator. The diffusion model includes two processes: forward noise addition and backward noise reduction. The forward process uses a cosine annealing strategy to process the initial state transition samples through Z iterations. (That is, the real state-action transition pairs collected from the actual interactions between the multi-agent system and the environment, serving as the model's true prior data) are gradually injected with Gaussian noise. This is done while keeping the data from the previous step known. Under the given conditions, generate noisy data for the current step. The forward Markov transition conditional probability distribution is expressed as:
[0067]
[0068] in, Represents the Gaussian distribution operator, These are noise scheduling parameters used to control the variance scale of the Gaussian noise injected in the current step. It is a smoothing scaling factor applied to the mean of the previous sample to prevent the data variance from exploding after multiple iterations; The identity matrix represents the added noise as isotropic and independently identically distributed across all feature dimensions.
[0069] The reverse denoising process, on the other hand, preserves the true state of the current system. Under certain constraints, simulated samples containing various prediction error characteristics are reconstructed stepwise from pure Gaussian noise. Illustratively, given the current... Noisy data of the step and environmental conditions Generate the data from the previous step. The inverse Markov transition conditional probability distribution is expressed as:
[0070]
[0071] in, The representative is parameterized as The denoised conditional probability distribution predicted by the neural network model; This indicates that the neural network is based on the current noisy samples. System status and the number of diffusion steps The mean of the Gaussian distribution jointly inferred, This represents the variance corresponding to the reverse step, and its core lies in accurately predicting the noise components that should be filtered out through the network. It is an identity matrix.
[0072] According to one embodiment of the present invention, a wireless communication system is provided, comprising multiple unmanned aerial vehicles (UAVs) configured to: collaboratively execute the multi-UAV scheduling method of the aforementioned embodiment to obtain a multi-UAV scheduling decision including a ratio variable, deployment locations of each UAV, a user association indication matrix, and a bandwidth allocation matrix; according to the multi-UAV scheduling decision, each UAV flies to the deployment location corresponding to the decision; sets a perception frame and a prediction frame according to the ratio variable in the decision, and performs user location perception in a specified time slot of the perception frame; according to the bandwidth allocated by the bandwidth allocation matrix in the decision, each UAV provides wireless communication services to the user specified by the user association indication matrix in the decision. Illustratively, firstly, in the interactive collection phase, each UAV agent interacts with the dynamic environment based on the current strategy to collect real "state-action-reward" transition pairs. For example, during the scheduling process, each UAV agent collects its own real samples and places them into an experience replay pool. The real samples include the local observation state of the UAV agent in the current time slot t, the local action performed in the current time slot t, the reward obtained after performing the local action, and the next local observation state entered after performing the local action. During the data augmentation phase, to mitigate data distortion caused by prediction errors, the diffusion model uses real samples formed from actual interaction data as prior input. Through forward noise addition and backward denoising processes, it synthesizes a large number of simulated samples containing various prediction uncertainty features. Subsequently, in the policy update phase, real data and the synthetic data generated by diffusion are mixed in a preset ratio and stored in the experience replay pool. During centralized training, a multi-head self-attention mechanism is used to dynamically evaluate the reliability of the local state features of each agent, rationally decompose the joint reward that maximizes the total communication rate, and then update the parameters of the local value network (Q-network) of each UAV through gradient descent. Finally, through continuous iterative learning until the parameters of each UAV agent converge, a multi-UAV scheduling decision that maximizes the total communication rate even with location prediction errors is provided. Through these decisions, the multi-UAV system can ultimately maximize the total communication data rate for ground users in a highly dynamic environment with unavoidable prediction errors.
[0073] To verify the effectiveness of the method of the present invention, the inventors also conducted comparative experiments.
[0074] The following presents simulation results of resource scheduling for a swarm of sensor-integrated drones using the method of this invention. In the simulation, 100 users are randomly initialized and distributed in the region. In a two-dimensional space, the user's initial movement speed The initial speed is randomly selected from {0 km / h, 5 km / h, 80 km / h}, simulating the movement of stationary users, pedestrians, and drivers respectively, with the initial movement direction of the users as follows. exist Random initialization is performed. Furthermore, a Gaussian Markov user random movement model is used to simulate user movement trajectories, and the movement speed of the nth user in time slot t is... and direction of movement All updates are based on their values in the previous time slot, and their mathematical expressions are as follows:
[0075]
[0076]
[0077] in, and For the range of values within The memory constants between these two values are used to control the degree of randomness (or memory level) in the changes of user speed and direction over time. When At this time, the system has strong memory, and users tend to engage in highly deterministic linear motion (high trajectory predictability); when At that time, the system has no memory, and the user's movement degenerates into completely random Brownian motion (the trajectory is extremely unpredictable). and These represent the asymptotic standard deviations of the user's speed and direction of movement, respectively, and are used to measure the fluctuation range of the user's dynamic changes. and It follows a standard normal distribution. Independent random Gaussian noise variables are used to introduce unpredictable micro-random perturbations into the movement state of each time slot. The simulation parameters are summarized in Table 1 below.
[0078] Table 1
[0079]
[0080] First, to verify the prediction performance of the Transformer-based user location prediction model (Transformer-MP), Figure 3The performance of the proposed algorithm compared to GRU (Gated Recurrent Unit) and CNN-LSTM (Convolutional Neural Network-based Long Short-Term Memory neural network) algorithms under different user movement stochasticities is demonstrated. According to equations (9) and (10), it can be seen that as the stochasticity parameter increases, the user's movement pattern exhibits greater regularity, thus improving the prediction accuracy of all algorithms, i.e., reducing the root mean square error (RMSE) of the predicted user position. Notably, the proposed Transformer-MP model consistently maintains excellent performance at different stochastic levels, reducing the RMSE by up to 30.6% and 46.9% compared to GRU and CNN-LSTM algorithms, respectively.
[0081] Figure 4 The performance of different scheduling schemes in terms of total user communication data rate was further compared. The Greedy algorithm is a heuristic optimization algorithm in which the drone selects a time slot based on probability. Joint actions that maximize instantaneous estimated reward, while... The probabilistic exploration of random actions focuses primarily on maximizing immediate rewards. In the Multi-Agent Deep Q-Network (MADQN), each drone acts as an independent agent, equipped with a deep Q-network to learn the optimal policy. These agents aim to maximize global rewards through independent learning. Simulation results show that as the number of drones increases from 3 to 8, the proposed robust diffusion-driven value decomposition network (RD-VDN) algorithm consistently outperforms MADQN and the Greedy algorithm in terms of communication speed. When the number of drones is 5, the RD-VDN algorithm achieves performance improvements of 13.7% and 20.7% compared to MADQN and Greedy algorithms, respectively.
[0082] In summary, the solutions provided by some embodiments of the present invention can achieve at least one or more of the following beneficial effects:
[0083] (1) A robust scheduling method is proposed for a multi-UAV system integrating sensing and communication. This method dynamically balances the time-frequency resources of sensing and communication by jointly using the sensing and prediction frame structure. By introducing the Transformer prediction model and the diffusion-driven value decomposition network algorithm, it alleviates the sensitivity of traditional prediction scheduling to data errors and significantly improves the communication data rate of UAVs under dynamic environment and prediction uncertainty conditions.
[0084] (2) A time-division ISAC (TD-ISAC) frame structure was designed, which divides the system frames into two categories: sensing frames and prediction frames. By optimizing the ratio between sensing frames and prediction frames, the system can dynamically control the trade-off between active sensing and prediction based on historical data, thereby maximizing communication performance while ensuring the accuracy of the predicted location.
[0085] (3) The direct impact of unavoidable prediction errors on user communication performance was quantitatively analyzed, and mathematical expressions for the actual signal-to-interference-plus-noise ratio (SINR) and the communication rate perceived by the prediction error were given. This formula constructs a quantitative mapping model from "prediction location error" to "actual communication performance degradation". In actual dynamic networks, prediction algorithms inevitably have biases. This formula physically equates this user location prediction bias to the beam misalignment angle of the UAV antenna array. Furthermore, beam misalignment will lead to a significant decrease in the effective signal gain received by the target user. This formula rigorously derives the expression for the SINR including prediction errors, enabling the system to accurately quantify the communication performance loss caused by inaccurate predictions.
[0086] (4) To address the limitation that existing prediction algorithms are usually designed in isolation from the underlying communication process and are difficult to guide actual network scheduling, a mobility prediction algorithm based on the Transformer architecture is integrated into the multi-UAV ISAC system. At the feature input end, the model is closely integrated with the joint perception and prediction frame structure designed by the system, and specifically extracts the high-precision historical position obtained by the active perception time slot of ISAC as the discrete sequence input; at the feature extraction end, with the help of multi-head self-attention mechanism and position encoding, the shortcomings of traditional algorithms in capturing long-distance spatiotemporal dependence when dealing with highly dynamic and nonlinear user trajectories are effectively overcome, and high-precision extrapolation of user coordinates in future time slots is achieved.
[0087] (5) To solve complex non-convex joint optimization problems, a conditional generative diffusion model is combined with a value decomposition network algorithm based on an attention mechanism. The diffusion model acts as a scene generator, synthesizing enhanced training data containing prediction error features through forward noise addition and backward denoising processes. By iteratively training on the augmented uncertainty dataset, the multi-UAV scheduling strategy is endowed with robustness and generalization ability when facing real prediction errors.
[0088] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.
[0089] This invention can be a system, method, electronic device, computing device, computer program product and / or computer-readable medium.
[0090] Computer program products mainly refer to software products that implement various aspects of the present invention through computer programs, or hardware products that carry software that implements various aspects of the present invention.
[0091] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can include, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.
[0092] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A multi-UAV scheduling method, wherein the method uses multiple mobile UAVs as access points to provide wireless communication services to one or more users within a coverage area, characterized in that, The method includes: A time-division integrated sensing frame structure for joint perception and prediction is constructed, wherein the time for the UAV to perform a task is divided into multiple cycles of joint perception and prediction, each cycle contains multiple frames, the multiple frames are divided into perception frames and prediction frames, and the proportion of perception frames can be dynamically adjusted by a ratio variable. To obtain the user's location, based on the integrated sensing technology, the real user's location is actively sensed in some time slots of the sensing frame, and the predicted user's location is obtained based on the user's historical location in all time slots of the prediction frame. Based on the statistical estimation of historical user location prediction errors, the transmit beam angle deviation of the UAV is estimated, and a communication rate model for position prediction error perception between the UAV and the user is constructed to characterize the quantitative mapping relationship from the deviation to the communication rate attenuation. Based on the communication rate model, with the objective of maximizing the total communication rate when there is a location prediction error, the multi-UAV scheduling decision is obtained by jointly optimizing the ratio variable, the deployment location of each UAV, the association indication matrix of the users served, and the bandwidth allocation matrix.
2. The method according to claim 1, characterized in that, Each of the multiple frames includes multiple time slots, wherein a single fixed time slot is designated in all sensing frames for transmitting signals for user location sensing, and the remaining time slots in the sensing frames are used for wireless communication.
3. The method according to claim 1, characterized in that, The joint optimization is achieved through a multi-agent deep reinforcement learning algorithm, wherein: The problem of multi-drone scheduling is modeled as a Markov decision process, in which the individual drone agents of the multiple drones participate in the decision-making. The local observation state of each UAV agent is set to include the UAV position in the previous time slot, the positions of all users obtained in the current time slot, and the channel state information of the previous time slot; The local actions of each drone agent are defined by including the individual expected ratio, the drone's deployment location, the association indication matrix, and the bandwidth allocation matrix; During scheduling, each UAV agent independently predicts local actions based on its own local observation state, and optimizes its own individual expected ratio, deployment location, association indication matrix and bandwidth allocation matrix separately. However, when executing local actions, the ratio variable allocated in the current period is used instead of the individual expected ratio. When updating parameters during the reinforcement learning process of UAV agents, the parameters of each UAV agent are updated by backpropagation with the goal of maximizing the total communication rate, and the ratio variable for the next cycle is selected based on the individual expected ratio of the most recently predicted local actions of each UAV agent.
4. The method according to claim 3, characterized in that, During the scheduling process, each UAV agent collects its own real samples and puts them into the experience replay pool. The real samples include the local observation state of the UAV agent in the current time slot t, the local action performed in the current time slot t, the reward obtained after performing the local action, and the next local observation state entered after performing the local action.
5. The method according to claim 4, characterized in that, The reinforcement learning algorithm is configured as follows: The joint action value function, which is the expected total benefit of multiple UAV agents performing joint actions in the global state, is decomposed into a weighted sum of the local value functions of each UAV agent. The weights of the weighted sum are dynamically generated by a multi-head self-attention mechanism based on the importance and reliability of the current local observation state of each UAV agent. A conditional generative diffusion model is introduced as a scene generator. It takes real samples in the experience replay pool as input, and gradually reconstructs simulated samples containing location prediction error features by performing forward noise addition and reverse noise removal on the real samples. The simulated samples are mixed with real samples and stored in the experience replay pool to enhance the diversity of training data. Based on real and simulated samples in the experience replay pool, a policy network in a multi-UAV agent is trained by maximizing cumulative rewards. The trained policy network is then used to predict local actions.
6. The method according to claim 5, characterized in that, The predicted user location is determined by a Transformer-based user location prediction model, which includes: The data input module is used to receive the historical trajectory sequence composed of the user's historical real location obtained from the sensing frame; The linear embedding module is used to map the position coordinate vector of each time slot in the historical trajectory sequence to a higher-dimensional feature space to obtain a position feature matrix. The location encoding module is used to superimpose learnable location codes element by element in each location feature matrix, assign a time identifier to each time step, and obtain location feature sequence data after superimposing location codes. The multi-head self-attention feature extraction module is used to feed the position feature sequence data after superimposed position encoding into the multi-layer Transformer encoder. Through the multi-head self-attention mechanism, trajectory features are extracted in parallel in different representation subspaces to obtain a position feature vector that integrates local and long-term spatiotemporal correlations. The location prediction module is used to process the location feature vector through residual connections, layer normalization, feedforward neural networks and linear mapping coordinate regression networks, and output the predicted user location.
7. The method according to any one of claims 1-6, characterized in that, The communication rate model is constructed in the following way: A signal-to-interference-plus-noise ratio (SIR) model is established between the UAV and the user, wherein the UAV transmit beam angle deviation caused by the variance of the user position prediction error is explicitly embedded into the effective beamforming gain term, and the value of the gain term decreases as the deviation increases. Substituting the signal-to-interference-plus-noise ratio (SIR) determined by the SIR model into the Shannon capacity formula yields the communication rate model.
8. The method according to any one of claims 1-6, characterized in that, The communication rate model is expressed as follows: in, This represents the communication rate for prediction error perception between UAV k and user n in time slot t. This represents the communication bandwidth between drone k and user n. This represents the signal-to-interference-plus-noise ratio (SIR) between time slot t, UAV k, and user n, including prediction errors.
9. The method according to claim 8, characterized in that, The signal-to-interference-plus-noise ratio (SINR) is determined by the following SINR model: in, This represents the transmit power of UAV k to user n in time slot t. This is the reference channel power located 1 meter from the UAV's transmitting antenna. It is the carrier wavelength. This indicates the position of UAV k in time slot t. This indicates the position of user n in time slot t. This represents the probability that there is a line-of-sight link between drone k and user n. This represents the probability of a non-line-of-sight link. It refers to the number of drone transmitting antennas. Indicates the beamwidth factor. It is the angle between drone k and user n. The variance representing the user location prediction error. Indicates the non-line-of-sight attenuation factor. This represents the transmit power of UAV i to user j in time slot t. This represents the channel vector between UAV k and user n in time slot t under the probabilistic channel model. This represents the beamforming vector between UAV i and user j. , This represents the antenna array response vector function. This represents the actual physical angle between drone i and user j in time slot t. This represents the transmit beam angle deviation of UAV i relative to user j caused by the predicted position error. This represents the power of additive white Gaussian noise.
10. A wireless communication system comprising a plurality of unmanned aerial vehicles (UAVs), the plurality of UAVs being configured to: The multi-UAV scheduling method as described in any one of claims 1-9 is collaboratively executed to obtain a multi-UAV scheduling decision that includes a ratio variable, the deployment location of each UAV, a user association indication matrix, and a bandwidth allocation matrix. Based on the multi-drone scheduling decision, each drone flies to the deployment location corresponding to the decision. The perception frame and prediction frame are set according to the ratio variable in the decision-making process, and the user's location is perceived in the specified time slot of the perception frame. According to the bandwidth allocated by the bandwidth allocation matrix in the decision-making process, each UAV provides wireless communication services to the users specified by the user association indication matrix in the decision-making process.