A cross-attention driven time-frequency non-stationary channel prediction method for unmanned aerial vehicle communication

CN122179042BActive Publication Date: 2026-09-15NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610629172.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-09-15
Estimated Expiration
2046-05-09

AI Technical Summary

Technical Problem

然而,现有信道预测模型在适配低空复杂场景时仍存在明显不足

Benefits of technology

[0082] First, the cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication of this invention, based on a non-stationary channel modeling method, comprehensively considers dynamic factors such as transceiver antenna gain, equivalent Doppler phase perturbation, and fuselage attitude rotation, thereby obtaining a more accurate non-stationary channel representation. The resulting dataset provides highly realistic input for model training, thus significantly improving the generalization and robustness of the prediction model in complex low-altitude communication scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122179042B_ABST
    Figure CN122179042B_ABST
Patent Text Reader

Abstract

The application discloses a cross-attention driven time-frequency non-stationary channel prediction method for unmanned aerial vehicle communication, comprising the following steps: based on a cluster delay line model, introducing dynamic factors such as antenna attitude rotation and Doppler phase disturbance to optimize a low-altitude unmanned aerial vehicle multiple-input multiple-output communication system channel model, collecting and discretizing a time-varying channel transfer function to obtain frequency domain channel state information; then improving a Transform double-domain prediction model, embedding a non-stationary attention mechanism in an encoder to restore channel non-stationary dependence, simultaneously introducing a frequency domain-time delay domain cross-attention mechanism to fuse double-domain complementary information, and constructing a non-stationary channel prediction model; finally, jointly inputting the frequency domain channel state information and time delay domain channel state information obtained by inverse Fourier transform of the frequency domain channel state information into the model to realize accurate prediction of the time delay domain channel state information at a future moment. The application improves the non-stationary channel feature capturing capability and is suitable for a low-altitude unmanned aerial vehicle communication scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of low-altitude network channel prediction technology, and specifically to a cross-attention driven time-frequency non-stationary channel prediction method for UAV communication. Background Technology

[0002] With the accelerated evolution of integrated air-space-ground communication systems, the low-altitude economy is becoming an important component of future intelligent networks. The demand for UAVs in emergency communication, IoT relay, and aerial base station scenarios continues to grow. Therefore, building a reliable communication link prediction mechanism to support the autonomous perception and intelligent decision-making capabilities of UAVs has become a key requirement for low-altitude network construction. Low-altitude UAV communication is characterized by high dynamism, three-dimensional maneuverability, and broadband transmission, resulting in strong time-frequency non-stationarity in the wireless propagation mechanism. This non-stationarity is the core challenge of low-altitude channel modeling and a significant bottleneck limiting the achievement of high-level autonomous behavior by UAVs.

[0003] Against this backdrop, channel prediction technology has received increasing attention due to the need for predicting non-stationary channels in UAV communication scenarios. Its goal is to infer the evolution trend of future channel state information based on historical and current channel state information, thereby ensuring communication continuity and stability in low-altitude, highly dynamic propagation environments. However, low-altitude channels exhibit both rapidly changing time evolution characteristics and a broadband non-stationary structure in the frequency domain, making it difficult for existing prediction methods based on local stationary assumptions to maintain sufficient accuracy and robustness in real-world UAV scenarios.

[0004] Accurate acquisition of low-altitude non-stationary channel data is a prerequisite for channel prediction, but the inherent non-stationarity of low-altitude channels increases the difficulty of accurate channel characterization. On the one hand, the non-stationarity stems from the dynamic propagation environment, causing rapid changes in channel statistical properties over time, including dynamic motion characteristics, three-dimensional attitude features, scattering effects, and rapid changes in the scattering environment. On the other hand, in ultra-wideband systems, the propagation mechanism exhibited by the channel at different frequencies is inconsistent, resulting in varying degrees of frequency domain non-stationarity. To simplify the acquisition of low-altitude channel data, some researchers have treated the low-altitude channel environment as generalized stationary uniform scattering, meaning that the channel statistical properties remain stable within the observation window; however, this assumption only holds true in low-dynamic environments. Existing models have made progress in describing the multi-domain characteristics of channels, but they cannot be directly applied to channel prediction tasks with extremely high accuracy requirements, nor can they provide physically interpretable training sample data for model training.

[0005] Deep learning-based channel prediction frameworks using historical channel data are widely adopted to support adaptive channel modulation and coding and resource allocation. Existing research typically utilizes models such as CNNs, RNNs / LSTMs, and Transformers to infer future channel states by learning the temporal dependencies of historical channel state information. However, most existing methods assume that the channel is approximately stationary in local time. Therefore, in low-altitude non-stationary channel scenarios, their prediction accuracy usually decreases significantly, making it difficult to meet the needs of practical communication systems.

[0006] In summary, accurate channel prediction is crucial for supporting autonomous communication, sensing, and mission execution capabilities in UAV networks geared towards the low-altitude economy. However, existing channel prediction models still have significant shortcomings when adapting to complex low-altitude scenarios. First, the difficulty in obtaining high-precision non-stationary channel data limits model performance. Most existing studies rely on simplified stationary channel models or simulation data based on the generalized stationary uniform scattering assumption, leading to deviations between training data and the statistical characteristics of the real low-altitude environment. This makes it difficult for prediction models to accurately capture the rapid evolution of channels in complex scenarios. Second, there is a lack of a unified prediction framework that can characterize the time-frequency non-stationary characteristics of channels. Traditional prediction frameworks assume that channel statistical characteristics remain stationary within a local time window. However, in low-altitude scenarios, complex UAV trajectory switching and abrupt changes in scatterer distribution cause rapid changes in channel statistical characteristics. Ignoring this non-stationary characteristic significantly reduces the model's prediction accuracy. Finally, the lack of a cross-domain feature interaction and completion mechanism limits the improvement of prediction performance. Existing deep learning methods generally process channel state information in the frequency domain and channel state information in the time delay domain separately, ignoring the time-frequency evolution law between the two. This weakens the model's ability to understand global channel patterns and makes it difficult to accurately capture the joint time-frequency dynamic characteristics of low-altitude channels.

[0007] Therefore, accurately characterizing the non-stationary evolution mechanism of low-altitude channels in complex time-frequency dual-domain environments and constructing a predictive model that can effectively utilize cross-domain channel state information has become an important scientific problem that urgently needs to be solved in building UAV networks oriented towards the low-altitude economy. Summary of the Invention

[0008] The purpose of this invention is to provide a cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication, which integrates non-stationary attention mechanism and cross-attention mechanism. The former is used to dynamically capture the time-varying non-stationary characteristics of low-altitude channels, while the latter realizes the interactive completion of time delay domain and frequency domain information, thereby obtaining channel prediction results that are more consistent with actual propagation characteristics and significantly improving the accuracy and generalization ability of prediction.

[0009] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows:

[0010] A cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication, the method comprising the following steps:

[0011] S1, a cluster delay line model is used to model the multi-input multi-output communication system serving low-altitude UAV communication, introducing dynamic factors including antenna attitude rotation and Doppler phase perturbation to optimize the channel model; within a time length of , bandwidth is Within a certain range, the time-varying channel transfer function between any antenna pair is collected, discretized, and rearranged into a matrix form to obtain the frequency domain channel state information between the antenna pairs.

[0012] S2 improves the Transformer-based dual-domain prediction model to obtain a non-stationary channel prediction model. Specifically, on the one hand, a non-stationary attention mechanism is introduced into the encoder. By re-injecting explicit compensation terms into the statistical information of the original un-stationarized sequence, the non-stationary dependencies in the original channel sequence are recovered, resulting in the attention of the original un-stationarized sequence. Then, a multilayer perceptron is used to learn the residual time-varying trend and amplitude shift information from the stationary sequence, adaptively estimating the scaling scalar and translation vector. Finally, the attention of the original un-stationarized sequence, the scaling scalar, and the translation vector are combined to reconstruct the non-stationary components weakened by the stationarization step in the original sequence, thereby adaptively capturing the dynamics of channel features changing over time. On the other hand, a frequency-delay domain cross-attention mechanism is introduced. The attention weights of frequency domain features are dynamically guided and optimized using delay domain features as auxiliary information. The complementary attention information of the delay domain and frequency domain are fused, enabling the non-stationary channel prediction model to fully utilize the implicit time-varying information in the delay domain while focusing on significant features in the frequency domain.

[0013] S3 takes the frequency domain channel state information from the past time period as input, transforms it into time delay domain channel state information through inverse discrete Fourier transform, and inputs the frequency domain channel state information and time delay domain channel state information together into the non-stationary channel prediction model to predict the time delay domain channel state information at future time.

[0014] Furthermore, in step S1, a cluster delay line model is used to model the multiple-input multiple-output communication system serving low-altitude UAV communication. Dynamic factors, including antenna attitude rotation and Doppler phase perturbation, are introduced. The process of optimizing the channel model includes the following steps:

[0015] A cluster delay line model is used to model a multiple-input multiple-output (MIMO) communication system serving low-altitude unmanned aerial vehicle (UAV) communication. Dynamic factors, including antenna attitude rotation and Doppler phase perturbation, are introduced into the channel modeling. The time-varying channel transfer function expression between the p-th transmit antenna and the q-th receive antenna pair of the optimized MIMO communication system is as follows:

[0016] ;

[0017] Where N(t) is the number of time-varying non-line-of-sight paths, M is the number of subpaths in the nth non-line-of-sight path, and n=1 corresponds to the LOS case. and These are path delay and phase, respectively. Indicates the departure angle is And the angle of arrival is Antenna gain coefficient under the following conditions; It represents the path amplitude coefficient of the m-th sub-path in the n-th propagation path between the p-th transmitting antenna and the q-th receiving antenna at time t;

[0018] Considering the phase shift caused by the motion of the transmitter and receiver and the antenna spacing, the path phase is modeled as follows:

[0019] ;

[0020] in, To represent a uniformly distributed random variable with a random initial phase on the path, and These are the path Doppler phase and the antenna phase, respectively.

[0021] ;

[0022] ;

[0023] In the formula, k represents the wave number. and These are the departure and arrival direction vectors for non-line-of-sight paths, respectively. and These are the position vectors of the p-th transmitting antenna and the q-th receiving antenna, respectively. This represents the velocity vector of the transmitting antenna. This represents the velocity vector of the receiving antenna. For integration variables; and These represent the attitude rotation matrices of the transmitter and receiver, respectively.

[0024] Furthermore, in step S1, during a time period of... , bandwidth is The process of acquiring the time-varying channel transfer function between any antenna pair within a certain range, discretizing it, and rearranging it into matrix form to obtain the frequency domain channel state information between the antenna pairs includes the following steps:

[0025] Based on the established channel model, the CTF between any antenna pair is discretized to obtain the discretized time-varying channel transfer function:

[0026] ;

[0027] in, Indicates the center frequency point. and These are the time sampling interval and the frequency sampling interval, respectively, and the number of time sampling points. Number of frequency sampling points i and j represent the time sampling point index and the frequency sampling point index, respectively;

[0028] The discretized time-varying channel transfer function is rearranged into matrix form to obtain the frequency domain channel state information between antenna pairs. Its elements Used to characterize the transmission characteristics of the channel at time sampling point i and frequency sampling point j. Represents a matrix space with dimensions T rows and K columns;

[0029] The obtained frequency domain channel state information is vectorized along the time dimension. Transform it into a real-valued tensor of a neural network. , This represents a matrix space with dimensions of T rows and 2K columns.

[0030] For the obtained neural network real-valued tensor Position encoding is performed, resulting in the encoder input as follows:

[0031] ;

[0032] In the formula, It is the neural network input tensor after adding position encoding. This is a position information matrix. For a token at position i, position encoding is performed along the time dimension using sine and cosine functions, with the frequency dimension of 2K used as the feature size.

[0033] ;

[0034] ;

[0035] In the formula, i and j represent the time sampling point index and the frequency sampling point index, respectively. It is the real part position code value corresponding to the i-th time step and the j-th frequency component. It is the imaginary part position code value corresponding to the i-th time step and the j-th frequency component.

[0036] Step S2 further includes:

[0037] Suppose that the future window to be predicted contains frequency domain channel state information of F consecutive time slices, where the initial channel state information is denoted as . The historical observation window contains frequency domain channel state information for P consecutive time slices, where the initial channel state information is denoted as... The non-stationary channel prediction model can be represented by the following mapping relationship:

[0038] ;

[0039] in This is a two-domain prediction model based on Transformer. This represents the time-delay domain channel state information for the predicted future moment, where i represents the index of the time sampling point;

[0040] Normalized mean square error (NMSE) is used to measure the deviation between the prediction results and the actual channel. The optimization objective for training the non-stationary channel prediction model is set as minimizing NMSE.

[0041] ;

[0042] in, Denotes the Frobenius norm of a matrix. Indicates the desired operation; and Let represent the predicted channel state matrix and the actual channel state matrix at the nth future time, respectively.

[0043] Step S2 further includes:

[0044] The frequency domain channel state information after stabilization is represented as follows:

[0045] ;

[0046] in, It is a vector of dimension P consisting entirely of 1s. Let P be a real vector space with P rows. Then the query matrix, key matrix, and value matrix of a stationary sequence are represented as follows:

[0047] ;

[0048] ;

[0049] ;

[0050] in, yes Average along the time dimension, Indicates the number of rows. The real vector space; the magnitude of the self-attention is:

[0051] ;

[0052] The attention required to obtain the original, unstable sequence is:

[0053] ;

[0054] in, , Representing the set of real numbers, we simplify it to:

[0055] ;

[0056] In the formula, Indicates to The frequency domain channel state tensor that has undergone stabilization processing; , and These represent the input query matrix, input key matrix, and input value matrix of the model after stabilization, respectively. This represents the stabilized query matrix. This indicates the result after stabilization using the weight matrix. The key matrix obtained by mapping, This indicates the result after stabilization using the weight matrix. The value matrix obtained by mapping; This represents the variance of the frequency domain channel characteristics after stabilization. The mean vector representing the channel state information along the frequency dimension;

[0057] definition For scaling scalars Represents the positive real part of the set of real numbers. It is a translation vector. Indicates the number of rows. The real vector space is obtained by training a multilayer perceptron:

[0058] ;

[0059] ;

[0060] Self-attention within the source domain is represented as:

[0061] ;

[0062] In the formula, and Let these represent the frequency domain variance and the frequency domain mean, respectively. This represents the value matrix after stabilization.

[0063] Step S3 further includes:

[0064] Performing an inverse discrete Fourier transform on the frequency domain channel state information yields the time delay domain channel state information. And perform position encoding operation to obtain the time delay domain channel state information. ;

[0065] Frequency domain channel state information As the source domain, the time-delay domain channel state information after position coding processing For the target domain, effective attention information from the target domain is added to the source domain, enabling the source domain to obtain richer attention information. Specifically, in the non-stationary channel prediction model, the total attention between the source and target domains is represented as the sum of the self-attention of the source domain and the cross-attention from the target domain to the source domain. In this context, self-attention within the source domain is used to capture the local and global dependencies of historical channel state information in the frequency domain, while cross-attention from the target domain to the source domain is used to effectively supplement the multipath structure and abrupt change information contained in the time delay domain into the frequency domain, as expressed in:

[0066] ;

[0067] After fusing non-stationary self-attention and cross-domain cross-attention, the sequence representation with contextual information obtained after updating the token is as follows:

[0068] ;

[0069] The deep features of the currently updated token are extracted and fused through a fully connected network to output sequence features. :

[0070] ;

[0071] sequence features The input is fed into the decoder to predict the channel state information in the time delay domain for future time moments.

[0072] Furthermore, sequence features The process of inputting the information into the decoder to predict the channel state information in the time delay domain for future moments includes the following steps:

[0073] The last output of the encoder Using historical channel state information fragments as the known part, and simultaneously predicting future information... Zero-padding is performed at each time step to obtain the input of the decoder. The sequence updated using the non-stationary attention mechanism is represented as follows:

[0074] ;

[0075] Mask() is a masking operation used to prevent the unknown, predictable future sequence from interfering with the extraction of attention information from the known sequence. It is represented as: LN() represents the layer normalization operation, and softmax() represents the output operation of the attention mechanism; Scalar for scaling the source domain. Represents the mean offset correction vector of the source domain; , and These represent the input query matrix, input key matrix, and input value matrix of the decoder, respectively.

[0076] The predicted sequence is further updated by combining the encoder output:

[0077] ;

[0078] in, , , , , and The weight matrices represent the key matrix, query matrix, and value matrix, respectively; the predicted channel data is represented as:

[0079] ;

[0080] Extract a length of Future channel prediction sequence The final destationary output is represented as .

[0081] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0082] First, the cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication of this invention, based on a non-stationary channel modeling method, comprehensively considers dynamic factors such as transceiver antenna gain, equivalent Doppler phase perturbation, and fuselage attitude rotation, thereby obtaining a more accurate non-stationary channel representation. The resulting dataset provides highly realistic input for model training, thus significantly improving the generalization and robustness of the prediction model in complex low-altitude communication scenarios.

[0083] Second, the cross-attention driven time-frequency non-stationary channel prediction method for UAV communication of this invention addresses the non-stationarity problem caused by rapid time-varying channels in low-altitude communication by introducing a non-stationary attention mechanism. Within the Transformer framework, through sequence stabilization and non-stationary factor capture operations, the non-stationary components in the original channel sequence are adaptively reconstructed. The model can more accurately capture the temporal evolution of channel statistical characteristics, thereby significantly improving the prediction accuracy of rapidly changing low-altitude channels.

[0084] Third, the cross-attention driven time-frequency non-stationary channel prediction method for UAV communication of the present invention introduces a cross-domain attention mechanism to overcome the limitations of single frequency domain feature modeling. It uses the time-delay domain channel state as a supplementary information source and performs joint modeling with frequency domain features, thereby achieving collaborative attention among multi-domain features and improving the global consistency and accuracy of channel prediction. Attached Figure Description

[0085] Figure 1 This is a flowchart of the cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication according to the present invention;

[0086] Figure 2 This is a simulation diagram showing the performance comparison of channel coefficient prediction under non-stationary abrupt change conditions of the present invention.

[0087] Figure 3 The diagram shows the simulation results of the channel prediction error of the receiving UAV under the conditions of speed of 10-80km / h and bandwidth of 1-2.5GHz; where (a) corresponds to the speed condition of 10-80km / h and (b) corresponds to the bandwidth condition of 1-2.5GHz. Detailed Implementation

[0088] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0089] A cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication, the method comprising the following steps:

[0090] S1, a cluster delay line model is used to model the multi-input multi-output communication system serving low-altitude UAV communication, introducing dynamic factors including antenna attitude rotation and Doppler phase perturbation to optimize the channel model; within a time length of , bandwidth is Within a certain range, the time-varying channel transfer function between any antenna pair is collected, discretized, and rearranged into a matrix form to obtain the frequency domain channel state information between the antenna pairs.

[0091] S2 improves the Transformer-based dual-domain prediction model to obtain a non-stationary channel prediction model. Specifically, on the one hand, a non-stationary attention mechanism is introduced into the encoder. By re-injecting explicit compensation terms into the statistical information of the original un-stationarized sequence, the non-stationary dependencies in the original channel sequence are recovered, resulting in the attention of the original un-stationarized sequence. Then, a multilayer perceptron is used to learn the residual time-varying trend and amplitude shift information from the stationary sequence, adaptively estimating the scaling scalar and translation vector. Finally, the attention of the original un-stationarized sequence, the scaling scalar, and the translation vector are combined to reconstruct the non-stationary components weakened by the stationarization step in the original sequence, thereby adaptively capturing the dynamics of channel features changing over time. On the other hand, a frequency-delay domain cross-attention mechanism is introduced. The attention weights of frequency domain features are dynamically guided and optimized using delay domain features as auxiliary information. The complementary attention information of the delay domain and frequency domain are fused, enabling the non-stationary channel prediction model to fully utilize the implicit time-varying information in the delay domain while focusing on significant features in the frequency domain.

[0092] S3 takes the frequency domain channel state information from the past time period as input, transforms it into time delay domain channel state information through inverse discrete Fourier transform, and inputs the frequency domain channel state information and time delay domain channel state information together into the non-stationary channel prediction model to predict the time delay domain channel state information at future time.

[0093] Table 1 lists the main variable names and corresponding variable definitions involved in the non-stationary channel prediction model of this invention.

[0094] Table 1: Variable Comparison Table for the Non-Stationary Channel Prediction Model of the Present Invention

[0095]

[0096] For a multi-input multi-output (MIMO) communication system serving the high bandwidth and low latency requirements of low-altitude unmanned aerial vehicle (UAV) communication, considering the sparse distribution characteristics of scatterers in a broadband channel, a cluster delay line model is used to describe the multipath propagation mechanism of the channel, accurately characterizing the amplitude, phase, and time-frequency evolution of signals along each path. To simultaneously capture the time and frequency domain characteristics of the channel, the current channel state of the link between the p-th receiving antenna and the q-th transmitting antenna can be represented as a time-varying channel transfer function (CTF). Where p = 1, 2, ..., P, q = 1, 2, ..., Q. In a time period of... , bandwidth is By acquiring CTF data within a certain range and discretizing it, we can obtain...

[0097] ;

[0098] in, Indicates the center frequency point. and Given the time and frequency sampling intervals respectively, the number of time sampling points is... Number of frequency sampling points .

[0099] The discretized CTF is rearranged into a matrix form to obtain the frequency domain channel state information between specific antenna pairs. Its elements It fully characterizes the transmission characteristics of the channel at time sampling point i and frequency sampling point j.

[0100] The core objective of channel prediction is to use the temporal dependence and evolution patterns of historical channel data to accurately infer the future channel state. Specifically, it can be divided into two parts: prediction objective and optimization objective.

[0101] Suppose that the future window to be predicted contains frequency domain channel state information of F consecutive time slices, where the initial channel state information is denoted as . The historical observation window contains frequency domain channel state information for P consecutive time slices, where the initial channel state information is denoted as... The prediction model proposed in this invention can then be represented as a mapping relationship:

[0102] ;

[0103] in This is a two-domain prediction model based on Transformer. This represents the time-delay domain channel state information for the predicted future time.

[0104] Furthermore, to optimize the prediction performance, normalized mean square error (RMSE) is used to measure the deviation between the prediction results and the actual channel. The optimization objective of model training is to minimize the RMSE.

[0105] ;

[0106] in, Denotes the Frobenius norm of a matrix. This represents the desired operation. The objective is to ensure that the model's prediction accuracy remains consistent and comparable across different channel scenarios.

[0107] To address the problems of nonstationary modeling and time-frequency dual-domain prediction, this invention designs a Transformer-based dual-domain parallel prediction framework, the core process of which is as follows: Figure 1 As shown, the proposed model takes frequency-domain channel state information from past time periods as input, transforms it to time-delay domain channel state information using IDFT, and then jointly inputs the frequency-domain and time-delay domain channel state information into the proposed dual-domain Transformer architecture to predict the channel state information at future time moments. To characterize the non-stationary nature of real-world channels, the proposed model introduces a non-stationary attention mechanism in the encoder to adaptively capture the dynamic changes of channel features over time. Furthermore, to enhance the model's expressive power in frequency-domain feature extraction, a frequency-delay cross attention mechanism is introduced to fuse complementary attention information from the time-delay and frequency domains, thereby improving overall prediction accuracy and robustness. This framework combines physically driven feature modeling with data-driven attention learning, ensuring both the physical interpretability and prediction accuracy of the model.

[0108] The principles of each step of this invention will be explained in detail below.

[0109] (I) Non-stationary channel extraction and embedding

[0110] Existing channel prediction models often focus on fitting temporal features, neglecting the impact of channel physical characteristics, such as antenna attitude and Doppler effect, on non-stationary features. Furthermore, they fail to effectively handle the serialization problem of three-dimensional channel state information, resulting in limited model generalization capabilities. This section considers actual communication influencing factors within the channel modeling framework and explicitly models the non-stationary features of the channel. By optimizing CTF modeling, channel state information serialization processing, and position encoding, it constructs high-fidelity input features adapted to the Transformer.

[0111] To accurately characterize the non-stationary physical properties of the channel, dynamic factors such as antenna attitude rotation and Doppler phase perturbation are introduced into the channel modeling. The optimized CTF expression between the p-th transmit antenna and the q-th receive antenna pair in the multiple-input multiple-output system is as follows:

[0112] ;

[0113] Where N(t) is the time-varying number of non-line-of-sight paths. Note that in this model, if a line-of-sight (LoS) case exists, it corresponds to a propagation path with n=1. M is the number of sub-paths in the nth non-line-of-sight path. and These are path delay and phase, respectively. Indicates the departure angle is And the angle of arrival is The antenna gain coefficient is calculated based on the time-varying effects of antenna directivity and attitude changes on signal amplitude and phase. It can be represented as

[0114] ;

[0115] in, and These represent the components of the radiation pattern of the transmitting or receiving antenna in the vertical and horizontal planes, respectively. The proposed model incorporates an attitude rotation matrix. To describe the three-dimensional time-varying attitude of the transmitter and receiver, it can be represented as:

[0116] ;

[0117] Considering the phase shift caused by the motion of the transmitter and receiver and the antenna spacing, the path phase is modeled as...

[0118] ;

[0119] in, To represent a uniformly distributed random variable with a random initial phase on the path, and These are the path Doppler phase and the antenna phase, respectively.

[0120] ;

[0121] ;

[0122] In the formula, k represents the wave number. and These are the departure and arrival direction vectors for non-line-of-sight paths, respectively. and These are the position vectors of the q-th transmitting antenna and the p-th receiving antenna, respectively.

[0123] Multi-antenna systems increase the complexity of channel prediction; therefore, a method of parallel processing of single antenna pairs is adopted to improve the real-time performance of channel prediction and thus improve prediction efficiency. Based on the established channel model, the CTF between any antenna pair is discretized using formula (1) to form a matrix, which represents the channel state information. To meet the sequence processing requirements of the transformer, the obtained channel state information is vectorized along the time dimension to obtain... This is then transformed into a real-valued tensor that a neural network can process. .

[0124] The Transformer itself ignores the temporal order of the input sequence. To enable the Transformer to better capture contextual dependencies, positional encoding is performed on the input sequence. The encoder's input is...

[0125] ;

[0126] in, It is a location information matrix. For a token at position i, the position is encoded along the time dimension using sine and cosine functions.

[0127] ;

[0128] ;

[0129] In this design, the 2K dimension of sequence frequency is used as the feature size.

[0130] (II) Non-stationary attention mechanisms and encoders

[0131] In real-world wireless propagation environments, channel statistical characteristics evolve dynamically with changes in time, space, and frequency, exhibiting significant non-stationary features. However, most existing deep learning channel prediction models are based on the stationarity assumption, which assumes that channel statistical characteristics remain constant within a local time or frequency range. This assumption holds approximately in low-dynamic scenarios, but in high-frequency communication environments such as high-speed movement, complex scattering, or millimeter waves, the prediction models struggle to accurately capture rapid changes in channel characteristics, limiting their prediction accuracy and generalization performance in real-world communication systems.

[0132] This invention designs a non-stationary attention mechanism within the Transformer framework to explicitly model the non-stationarity of channel features evolving over time. This module adaptively adjusts the model's attention level to channel features at different times through both non-stationary and cross-attention mechanisms, thereby effectively enhancing the model's representational power and prediction accuracy in non-stationary channel scenarios.

[0133] The self-attention of a traditional Transformer can be represented as:

[0134] ;

[0135] in, For a length of P and a dimension of Query, key and value matrix, This represents the attention distribution generator.

[0136] To improve the performance of the attention mechanism in the Transformer core, the input sequence of the encoder module is adjusted. Element-wise stabilization yields a stabilized sequence. This operation can be represented as

[0137] ;

[0138] in, They can be calculated as follows:

[0139] ;

[0140] ;

[0141] Stationarizing the input sequence can effectively eliminate scale differences in the input channel state information and reduce the inconsistency of dynamic range between different subcarriers or time slices. This makes it easier for the Transformer to capture the relative change patterns of the sequence and stabilize gradient updates, which is beneficial for learning time-dependent relationships. However, global normalization can lead to over-stationarization, where non-stationary features such as power drift, delay spread evolution, and multipath structure changes inherent in the original channel sequence are compressed or even eliminated. This makes it difficult for the model to perceive the statistical changes of the real low-altitude channel over time, thus affecting prediction accuracy. Therefore, a non-stationary attention mechanism is introduced to capture the implicit non-stationary dependencies in the original data.

[0142] First, the channel state information after stabilization can be represented as follows:

[0143] ;

[0144] in, It is an all-one vector. Therefore, the query matrix of a stationary sequence can be represented as...

[0145] ;

[0146] in, yes The average over the time dimension, and the above formula for This also applies. In this case, the magnitude of self-attention depends on... ,Right now

[0147] ;

[0148] The attention value of the original, unstable sequence can be obtained as follows:

[0149] ;

[0150] in, , Then the above expression can be simplified to

[0151] ;

[0152] This can be understood as follows: the left side of the equation corresponds to the attention generated by the original sequence containing non-stationary components, while the right side, based on the stabilized query matrix and key matrix, re-injects statistical information from the original sequence through explicit compensation terms. This allows for the recovery of non-stationary dependencies in the original channel sequence.

[0153] Subsequently, the definition For scaling scalars These are translation vectors, obtained through training using a multilayer perceptron (MLP), i.e.

[0154] ;

[0155] ;

[0156] In this context, the MLP is used to learn the residual time-varying trend and amplitude shift information from the stationary sequence. Its role is to adaptively estimate the scaling scalar and translation vector, thereby reconstructing the non-stationary components in the original sequence that were weakened by the stationarization step. Therefore, the internally derived self-attention can be further expressed as…

[0157] .

[0158] (III) Frequency-Delay Cross-Attention Mechanism and Decoder

[0159] Channel state information in the time delay domain characterizes the multipath structure and time-varying properties of the channel, and it has a Fourier transform relationship with channel state information in the frequency domain. Compared to modeling using only frequency domain features, channel state information in the time delay domain contains more intuitive propagation path information and its dynamic characteristics evolving over time, providing richer correlation information for the frequency domain. Therefore, this section first maps the channel state information in the frequency domain to the time delay domain using Fourier transform to uncover the multipath feature distribution of the channel across different time delay components. Subsequently, a frequency-time delay cross-attention mechanism is introduced, using time delay domain features as auxiliary information to dynamically guide and optimize the attention weights of frequency domain features. This mechanism enables deep interaction and complementarity between cross-domain features, allowing the model to focus on significant features in the frequency domain while fully utilizing the implicit time-varying information in the time delay domain, thereby obtaining more accurate and stable channel prediction results.

[0160] First, perform IDFT transformation on the frequency domain channel state information. The channel state information in the time delay domain is obtained:

[0161] ;

[0162] Subsequently, the same position encoding operation was performed to obtain... Frequency domain channel state information As the source domain, the time-delay domain channel state information The goal is to supplement the source domain with effective attention information from the target domain, enabling the source domain to acquire richer attention information and achieve accurate channel prediction. The overall attention between the source and target domains in this invention's prediction model consists of two parts: first, self-attention within the source domain, used to capture the local and global dependencies of historical channel state information in the frequency domain; second, cross-attention from the target domain to the source domain, used to effectively supplement the multipath structure and abrupt change information contained in the time delay domain to the frequency domain. The total attention between the source and target domains can be expressed as the sum of the self-attention of the source domain and the cross-attention from the target domain to the source domain.

[0163] ;

[0164] The cross-attention from the target domain to the source domain can be represented as:

[0165] ;

[0166] After fusing non-stationary self-attention and cross-domain cross-attention, the sequence with contextual information obtained after updating the token can be represented as follows:

[0167] ;

[0168] Subsequently, the deep features of the currently updated token are further extracted and fused through a fully connected network (FCN), which can be represented as follows:

[0169] .

[0170] Thus, this invention, by introducing non-stationary attention, reconstructs non-stationary dependencies through learning scale scaling scalars and translation vectors, and simultaneously captures temporal dependencies and dual-domain associations through a combination of self-attention and cross-domain attention. The output of the non-stationary encoder is a deep update sequence rich in non-stationary features, providing structured input for subsequent accurate channel prediction.

[0171] In obtaining the sequence features output by the encoder Then, the channel state information for future time periods is input into the decoder. The decoder outputs the encoder after padding with zeros for the future time periods. As input, the final destationary output can be expressed as .

[0172] The decoder's input is the last bit of the encoder's output. Using historical channel state information fragments as the known part, and simultaneously predicting future information... Zero-padding is performed at each time step to ensure that the model does not encounter any future information in its initial state. Therefore, the input to the decoder can be represented as... Similarly, the sequence updated using the non-stationary attention mechanism is represented as follows:

[0173] ;

[0174] Mask() is a masking operation used to prevent the unknown, predictable future sequence from interfering with the extraction of attention information from the known sequence. This operation can be represented as...

[0175] ;

[0176] The predicted sequence is further updated by combining the encoder output.

[0177] ;

[0178] in, , , , , These represent the weight matrices of the key matrix, query matrix, and value matrix, respectively. Ultimately, the predicted channel data can be represented as...

[0179] ;

[0180] Based on the above results, a length of [length missing] can be extracted from it. Future channel prediction sequence The final destationary output can be expressed as

[0181] .

[0182] (iv) Numerical simulation and analysis

[0183] Figure 2 This paper presents a comparison of the prediction results of different models in the time domain under typical low-altitude non-stationary channels. The intervals marked in light gray in the figure correspond to non-stationary abrupt changes in channel statistical characteristics, such as environmental changes or scatterer evolution. It can be observed that in non-stationary abrupt change regions, both the model without non-stationary attention and LLM4P exhibit significant offsets and response lags. In contrast, the model incorporating the non-stationary attention mechanism maintains a faster response speed and higher fitting accuracy near the abrupt change points, and its prediction curve is more consistent with the actual channel trend. Within typical abrupt change intervals, the models incorporating non-stationary attention significantly reduce prediction errors, with the normalized root mean square error being significantly better than the baseline model LLM4CP.

[0184] To verify the adaptability of the proposed model to channel time-frequency nonstationarity, Figure 3 Figure (a) shows the channel prediction error performance under the condition of UAV speed of 10-80 km / h at the receiving end. Figure 3 Figure (b) shows a simulation of the channel prediction error performance under a bandwidth of 1-2.5 GHz. Figure 3 As shown in (a), within the speed range of 10-80 km / h, the normalized root mean square error (MMSE) gradually increases with increasing speed. This is because the increased speed leads to increased time-domain non-stationarity of the channel, making it more difficult to capture channel behavior and thus increasing the channel prediction error. However, the model of this invention consistently maintains a low MMSE, indicating that it has good adaptability to changes in UAV speed and maintains good prediction performance even at high speeds and under conditions of strong time-domain non-stationarity. Figure 3In (b) of the model, the normalized root mean square error (RMSE) curves of several models increase with increasing bandwidth. This is attributed to the increased frequency domain non-stationarity of the channel due to the increased bandwidth, making it difficult to accurately capture the frequency domain non-stationary information in the time series, thus leading to increased prediction errors. The model of this invention maintains a low RMSE and a slow rate of increase within the 1-2.5 GHz bandwidth range, indicating that the model has strong adaptability to frequency domain non-stationarity and can better cope with communication scenarios involving signals of various bandwidths. The simulation results above demonstrate that the model can adapt well to changes in the time-frequency non-stationarity of the channel and has strong prediction capabilities for time-frequency non-stationary channels.

[0185] This invention proposes a dual-domain non-stationary channel prediction framework for UAV communication scenarios, aiming to overcome the bottleneck of limited prediction performance of existing models under strong time-frequency non-stationary environments. This invention generates training data that better conforms to the physical propagation mechanism through an improved non-stationary channel modeling method, thereby effectively improving the model's adaptability to real-world low-altitude dynamic channels. Based on the Transformer architecture, a non-stationary attention mechanism is introduced to explicitly capture and reconstruct the non-stationary evolution characteristics of the channel in both the time and frequency domains, significantly improving the model's attention allocation accuracy for key time-varying features. A dual-domain cross-attention module is used to achieve deep coupling between frequency domain and time-delay domain features, enabling the model to simultaneously utilize the frequency-varying structure of the broadband frequency domain and the multipath evolution information of the time-delay domain, thus more comprehensively and precisely characterizing the time-frequency coupling laws and evolution patterns of low-altitude channels. Simulation results show that the proposed model exhibits superior performance compared to related model-based and deep learning-based channel prediction models under non-stationary channel conditions. This research provides key forward-looking channel knowledge support for UAVs, enabling them to have greater communication autonomy and intelligent decision-making capabilities in low-altitude economic networks, and providing an effective technical path for building future-oriented AI-driven UAV networks.

[0186] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0187] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication, characterized in that, The method includes the following steps: S1, a cluster delay line model is used to model the multi-input multi-output communication system serving low-altitude UAV communication, introducing dynamic factors including antenna attitude rotation and Doppler phase perturbation to optimize the channel model; within a time length of , bandwidth is Within a certain range, the time-varying channel transfer function between any antenna pair is collected, discretized, and rearranged into a matrix form to obtain the frequency domain channel state information between the antenna pairs. S2 improves the Transformer-based dual-domain prediction model to obtain a non-stationary channel prediction model. Specifically, on the one hand, a non-stationary attention mechanism is introduced into the encoder. By re-injecting explicit compensation terms into the statistical information of the original un-stationarized sequence, the non-stationary dependencies in the original channel sequence are recovered, resulting in the attention of the original un-stationarized sequence. Then, a multilayer perceptron is used to learn the residual time-varying trend and amplitude shift information from the stationary sequence, adaptively estimating the scaling scalar and translation vector. Finally, the attention of the original un-stationarized sequence, the scaling scalar, and the translation vector are combined to reconstruct the non-stationary components weakened by the stationarization step in the original sequence, thereby adaptively capturing the dynamics of channel features changing over time. On the other hand, a frequency-delay domain cross-attention mechanism is introduced. The attention weights of frequency domain features are dynamically guided and optimized using delay domain features as auxiliary information. The complementary attention information of the delay domain and frequency domain are fused, enabling the non-stationary channel prediction model to fully utilize the implicit time-varying information in the delay domain while focusing on significant features in the frequency domain. S3 takes the frequency domain channel state information in the past time period as input, transforms it into the time delay domain channel state information through the inverse discrete Fourier transform, and inputs the frequency domain channel state information and the time delay domain channel state information together into the non-stationary channel prediction model to predict the time delay domain channel state information at future time. Step S3 further includes: Performing an inverse discrete Fourier transform on the frequency domain channel state information yields the time delay domain channel state information. And perform position encoding operation to obtain the time delay domain channel state information. ; This represents the frequency domain channel state information between antenna pairs. Represents a matrix space with dimensions T rows and K columns; Frequency domain channel state information As the source domain, the time-delay domain channel state information after position coding processing For the target domain, effective attention information from the target domain is added to the source domain, enabling the source domain to obtain richer attention information. Specifically, in the non-stationary channel prediction model, the total attention between the source and target domains is represented as the self-attention of the source domain. The sum of cross-attention from the target domain to the source domain Among them, self-attention within the source domain Cross-attention from the target domain to the source domain is used to capture local and global dependencies of historical channel state information in the frequency domain. This is used to effectively supplement the multipath structure and abrupt change information contained in the time delay domain into the frequency domain, and is represented as: ; After fusing non-stationary self-attention and cross-domain cross-attention, the sequence representation with contextual information obtained after updating the token is as follows: ; In the formula, This represents the feature dimensions of the query matrix, key matrix, and value matrix. Indicates to The frequency domain channel state tensor that has undergone stabilization processing; and Let these represent the input key matrix and input value matrix before stabilization, respectively. This represents the stabilized query matrix. This indicates the result after stabilization using the weight matrix. The key matrix obtained by mapping, This indicates the result after stabilization using the weight matrix. The value matrix obtained by mapping; The deep features of the currently updated token are extracted and fused through a fully connected network to output sequence features. : ; sequence features The input is fed into the decoder to predict the channel state information in the time delay domain for future time moments.

2. The cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication according to claim 1, characterized in that, In step S1, a cluster delay line model is used to model the multiple-input multiple-output communication system serving low-altitude UAV communication. Dynamic factors, including antenna attitude rotation and Doppler phase perturbation, are introduced. The process of optimizing the channel model includes the following steps: A cluster delay line model is used to model a multiple-input multiple-output (MIMO) communication system serving low-altitude unmanned aerial vehicle (UAV) communication. Dynamic factors, including antenna attitude rotation and Doppler phase perturbation, are introduced into the channel modeling. The time-varying channel transfer function expression between the p-th transmit antenna and the q-th receive antenna pair of the optimized MIMO communication system is as follows: ; Where N(t) is the number of time-varying non-line-of-sight paths, M is the number of subpaths in the nth non-line-of-sight path, and n=1 corresponds to the LOS case. and These are path delay and phase, respectively. Indicates the departure angle is And the angle of arrival is Antenna gain coefficient under the following conditions; It represents the path amplitude coefficient of the m-th sub-path in the n-th propagation path between the p-th transmitting antenna and the q-th receiving antenna at time t; Considering the phase shift caused by the motion of the transmitter and receiver and the antenna spacing, the path phase is modeled as follows: ; in, To represent a uniformly distributed random variable with a random initial phase on the path, and These are the path Doppler phase and the antenna phase, respectively. ; ; In the formula, k represents the wave number. and These are the departure and arrival direction vectors for non-line-of-sight paths, respectively. and These are the position vectors of the p-th transmitting antenna and the q-th receiving antenna, respectively. This represents the velocity vector of the transmitting antenna. This represents the velocity vector of the receiving antenna. For integration variables; and These represent the attitude rotation matrices of the transmitter and receiver, respectively.

3. The cross-attention driven time-frequency non-stationary channel prediction method for UAV communication according to claim 1, characterized in that, In step S1, during a time period of , bandwidth is The process of acquiring the time-varying channel transfer function between any antenna pair within a certain range, discretizing it, and rearranging it into matrix form to obtain the frequency domain channel state information between the antenna pairs includes the following steps: Based on the established channel model, the CTF between any antenna pair is discretized to obtain the discretized time-varying channel transfer function: ; in, Indicates the center frequency point. and These are the time sampling interval and the frequency sampling interval, respectively, and the number of time sampling points. Number of frequency sampling points i and j represent the time sampling point index and the frequency sampling point index, respectively; The discretized time-varying channel transfer function is rearranged into matrix form to obtain the frequency domain channel state information between antenna pairs. Its elements Used to characterize the transmission characteristics of the channel at time sampling point i and frequency sampling point j. Represents a matrix space with dimensions T rows and K columns; The obtained frequency domain channel state information is vectorized along the time dimension. Transform it into a real-valued tensor of a neural network. , This represents a matrix space with dimensions of T rows and 2K columns.

4. The cross-attention driven time-frequency non-stationary channel prediction method for UAV communication according to claim 1, characterized in that, For the obtained neural network real-valued tensor Position encoding is performed, resulting in the encoder input as follows: ; In the formula, It is the neural network input tensor after adding position encoding. This is a position information matrix. For a token at position i, position encoding is performed along the time dimension using sine and cosine functions, with the frequency dimension of 2K used as the feature size. ; ; In the formula, i and j represent the time sampling point index and the frequency sampling point index, respectively. It is the real part position code value corresponding to the i-th time step and the j-th frequency component. It is the imaginary part position code value corresponding to the i-th time step and the j-th frequency component.

5. The cross-attention driven time-frequency non-stationary channel prediction method for UAV communication according to claim 1, characterized in that, Step S2 further includes: Suppose that the future window to be predicted contains frequency domain channel state information of F consecutive time slices, where the initial channel state information is denoted as . The historical observation window contains frequency domain channel state information for P consecutive time slices, where the initial channel state information is denoted as... The non-stationary channel prediction model can be represented by the following mapping relationship: ; in This is a two-domain prediction model based on Transformer. This represents the time-delay domain channel state information for the predicted future moment, where i represents the index of the time sampling point; Normalized mean square error (NMSE) is used to measure the deviation between the prediction result and the actual channel. The optimization objective for training the non-stationary channel prediction model is set as minimizing NMSE. ; in, Denotes the Frobenius norm of a matrix. Indicates the desired operation; and Let represent the predicted channel state matrix and the actual channel state matrix at the nth future time, respectively.

6. The cross-attention driven time-frequency non-stationary channel prediction method for UAV communication according to claim 5, characterized in that, Step S2 further includes: The frequency domain channel state information after stabilization is represented as follows: ; in, It is a vector of dimension P consisting entirely of 1s. Let P be a real vector space with P rows. Then the query matrix, key matrix, and value matrix of a stationary sequence are represented as follows: ; ; ; in, yes Average along the time dimension, Indicates the number of rows. The real vector space; the magnitude of the self-attention is: ; The attention required to obtain the original, unstable sequence is: ; in, , Representing the set of real numbers, we simplify it to: ; In the formula, Indicates to The frequency domain channel state tensor that has undergone stabilization processing; , and Let these represent the input query matrix, input key matrix, and input value matrix before they are stabilized. This represents the stabilized query matrix. This indicates the result after stabilization using the weight matrix. The key matrix obtained by mapping, This indicates the result after stabilization using the weight matrix. The value matrix obtained by mapping; This represents the variance of the frequency domain channel characteristics after stabilization. The mean vector representing the channel state information along the frequency dimension; definition For scaling scalars Represents the positive real part of the set of real numbers. It is a translation vector. Indicates the number of rows. The real vector space is obtained by training a multilayer perceptron: ; ; Self-attention within the source domain is represented as: ; In the formula, and Let these represent the frequency domain variance and the frequency domain mean, respectively. This represents the value matrix after stabilization.

7. The cross-attention-driven time-frequency non-stationary channel prediction method for UAV communication according to claim 1, characterized in that, sequence features The process of inputting the information into the decoder to predict the channel state information in the time delay domain for future moments includes the following steps: The last output of the encoder Using historical channel state information fragments as the known part, and simultaneously predicting future information... Zero-padding is performed at each time step to obtain the input of the decoder. The sequence updated using the non-stationary attention mechanism is represented as follows: ; Mask() is a masking operation used to prevent the unknown, predictable future sequence from interfering with the extraction of attention information from the known sequence. It is represented as: LN() represents the layer normalization operation, and softmax() represents the output operation of the attention mechanism; Scalar for scaling the source domain. Represents the mean offset correction vector of the source domain; , and These represent the input query matrix, input key matrix, and input value matrix of the decoder, respectively. The predicted sequence is further updated by combining the encoder output: ; in, , , , , and The weight matrices represent the key matrix, query matrix, and value matrix, respectively; the predicted channel data is represented as: ; Extract a length of Future channel prediction sequence The final destationary output is represented as .

Citation Information

Patent Citations

  • Dynamic evolution and periodic structure double-current cross attention fused radio frequency signal classification method and system

    CN120067913A

  • MSGSE-NSCF-based urban solid waste incineration NOx emission prediction method

    CN121148533A