Unmanned aerial vehicle sensing method and system based on OTFS signal, storage medium and product
Through the prediction-assisted estimation network PAEnet, the problem of low accuracy in OTFS sensing technology in high mobility scenarios and low accuracy in noise environments is solved, and high accuracy and efficient delay and Doppler parameter estimation of drone sensing is achieved, especially in impulse noise environments.
Patent Information
- Application Number
- CN202510410265.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
In the high mobility scenarios, existing OTFS sensing technology has problems such as inaccurate parameter settings, errors caused by differences between rectangular filters and ideal filters, high dimensionality of DD domain equivalent channel matrix leads to computational complexity, and model-based detection methods cannot describe the characteristics of impulse noise distribution, affecting the accuracy and efficiency of signal perception and analysis.
The pre-trained prediction-assisted estimation network PAEnet is adopted, including the target parameter extraction network, predictive subnet and estimation network. By obtaining OTFS data in real time, extracting and predicting delays and Doppler parameters, and performing compensation splicing and estimation, it uses deep learning technology to optimize signal processing, especially maintaining robustness in impulse noise environments.
The accuracy and efficiency of drone sensing are improved, especially in pulse distortion scenarios, accurate prediction and compensation of delay and Doppler parameters are achieved, and the accuracy and robustness of signal perception are improved.
Smart Images

Figure CN120301745A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of OTFS channel estimation, and in particular, to a method, a system, a computer-readable storage medium, and a computer program product for deep learning-assisted UAV perception based on OTFS signals. Background Art
[0002] In recent years, scholars have been deeply exploring the sixth-generation (6G) technology to achieve reliable communication and deep interconnection. For 6G, sensing capabilities have received more attention and are considered crucial for capturing additional context information, which can improve communication efficiency and broaden application conditions.
[0003] Exploration and research have been carried out on unmanned aerial vehicles (UAVs) and communication and sensing systems, and many remarkable results have been achieved. J. Mu, R. Zhang, and Y. Cui et al. proposed the concept of integrating UAVs with integrated sensing and communication (ISAC) in "UAV Meets Integrated Sensing and Communication: Challenges and Future Directions", and then discussed the integration among base stations, UAVs, and UAV swarms. X. Jing, F. Liu, and C. Masouros et al. optimized the UAV trajectory in "ISAC from the Sky: UAV Trajectory Design for Joint Communication and Target Localization", and proposed a multi-stage trajectory design method for communication and target localization, achieving excellent target estimation results. However, due to the influence of atmospheric turbulence, UAVs often experience unexpected translations and rotations, which are not taken into account. Among them, environmental conditions such as rainfall and snowfall can cause signal attenuation and scattering, while fog can exacerbate signal propagation path loss. In addition, lightning and thunder during thunderstorms may cause pulse noise interference. In this case, passive detection of unknown UAV communication signals may also generate incomplete information. In the face of these special situations, there have been some research results in the literature. W. Yuan, C. Liu, and F. Liu et al. proposed a method for using previous data and implementing jitter UAV predictive beamforming in "Learning-Based Predictive Beamforming for UAV Communication". In addition, C. Liu, W. Yuan, and D. W. K. Ng et al. utilized the temporal characteristics of the UAV trajectory for precise localization prediction in "Location-Aware Predictive Beamforming for UAV Communication: A Deep Learning Approach" to assist beam alignment.
[0004] To implement the C&S function of an unmanned aerial vehicle, it is natural to consider what kind of waveform to adopt. Due to continuous movement at high speeds, drones are vulnerable to dynamic channels and carrier frequency offset (CFO). In orthogonal frequency division multiplexing (OFDM) modulation technology, frequency offset will destroy the orthogonality of subcarriers in the frequency domain, resulting in inter-carrier interference (ICI). To solve the signal transmission problem, especially in the highly mobile aerial environment, we considered a new modulation scheme, namely orthogonal time frequency space (OTFS). OTFS has advantages in high-speed scenarios by modulating and demodulating signals in the delay-Doppler (DD) domain. Through domain transformation, OTFS presents a sparse channel in the DD domain, rather than a time-varying frequency-selective fading channel in the time-frequency (TF) domain. The DD domain representation enables OTFS to cancel frequency offset, which is beneficial for highly mobile communication and sensing. In addition, OTFS can also achieve direct interaction between the transmitted signal and the DD domain channel, such as delay and Doppler parameters, which is exactly consistent with the purpose of sensing. The channel targets in the DD domain are easier to identify and locate, which is closely related to physical reality and can provide reliable sensing information. In addition, the range and speed parameters can be directly converted from the delay and Doppler frequency shift in the DD domain channel, thus simplifying the sensing process. In short, the sensing technology based on OTFS provides a simpler physical model suitable for high-mobility scenarios.
[0005] Although the OTFS sensing technology has its advantages, it will also be subject to some limitations during implementation. Inaccurate parameter settings or differences between the rectangular filter adopted and the ideal filter may lead to errors, which may affect the overall efficacy of OTFS in signal perception and analysis. Another challenge comes from the high dimensionality of the DD domain equivalent channel matrix, resulting in computationally intensive matrix inversion in linear detection methods. In addition, the discrete representation of delay and Doppler parameters in the algorithm may not be able to fully capture the continuity of these phenomena, which is also an inherent limitation. This is the so-called fractional delay Doppler phenomenon, which brings complexity and potential errors to the OTFS system. Model-based detection methods can handle continuous noise, but cannot describe the distribution characteristics of impulse noise. In addition, the computational complexity of model-based OTFS detection and fractional delay and Doppler estimation poses challenges to the time efficiency and computational power in practical applications. Therefore, an alternative model-based sensing method is needed.
[0006] Through an iterative training process and gradient descent algorithm based on a large amount of data, deep learning (DL) can update model weight parameters and give accurate results in classification and regression tasks. DL is gradually being widely applied to communication systems, especially in the design of transmitters and receivers, and is expected to improve performance in complex C&S scenarios. "Adaptive Threshold Deep Learning for Channel Estimation" proposed by X. Zhang, W. Yuan, M. Wen, etc. considered OTFS channel estimation and proposed a deep residual shrinkage network that can implicitly learn residual noise to recover channel information. Y.K. Enku, B. Bai, F. Wan, C.U. Guyo, I.N. Tiba, C. Zhang, and S. Li designed an OTFS system detector based on a two-dimensional convolutional neural network (2D-CNN) in "Signal Detection for OTFS Systems Based on Two-Dimensional Convolutional Neural Networks" and achieved a lower time complexity. Q. Li, Y. Gong, Y. Liu, Z. Xu, etc. proposed a data-driven OTFS receiver based on a deep neural network (DNN) in "Exploring the Performance of DNN-Based OTFS Receiver Algorithms", which can jointly consider multiple tasks of the OTFS receiver. However, considering the interference of a noisy environment where the signal is strongly affected and the occurrence of impulse noise, the aforementioned model-based OTFS channel processing methods still have problems of insufficient accuracy and unstable model robustness. Summary of the Invention
[0007] The purpose of the embodiments of the present disclosure is to provide a drone sensing method, system, computer-readable storage medium, and computer program product according to OTFS signals, so as to solve the aforementioned problems existing in the prior art.
[0008] To achieve the above purpose, the technical solutions adopted in the embodiments of the present disclosure are as follows:
[0009] On the one hand, the embodiments of the present disclosure propose a drone sensing method according to OTFS signals, which adopts a pre-trained prediction-assisted estimation network PAEnet. The prediction-assisted estimation network PAEnet includes: a target parameter extraction network, a prediction sub-network, and an estimation network; the method includes:
[0010] Obtain target OTFS data in real time, and use a pre-constructed target parameter extraction network to extract the target delay and Doppler parameter information of all OTFS frames in a preset time period τ-1 before the current time t. Among them, the target parameter extraction network is trained based on a first ESnet with a residual network as the backbone;
[0011] Based on all the target delay and Doppler parameter information of the extracted preset time period τ-1, use the pre-trained prediction sub-network to predict the delay and Doppler parameters at time t; wherein, the prediction sub-network includes: trained based on the PR-subnet;
[0012] Compensate and splice the predicted delay and Doppler parameters at time t with the OTFS frame at time t;
[0013] Input the compensated and spliced data into the pre-trained estimation network to obtain the delay and Doppler parameters of the OTFS frame at time t, wherein, the estimation network includes: trained based on the second ESnet of the residual network.
[0014] Exemplarily, the PR-subnet network includes: two LSTM layers and one fully connected layer.
[0015] Exemplarily, after real-time obtaining the OTFS data, the method further includes:
[0016] Adopt a data preprocessing estimator based on two-dimensional correlation to preprocess all OTFS frames of the current time t and the previous preset time period τ-1 to obtain corresponding compressed frames. The data preprocessing formula is as follows;
[0017]
[0018] Among them, V[k, l] represents the data after preprocessing, and Y dd is the signal before correlation, that is, the set of received OTFS frames the data of the corresponding frame in the data, represents the data set of all OTFS frames of the time series from time t-τ+1 to time t. t represents the current time, τ∈R represents the length of the time series data set, and t-τ+1 represents the start time; * represents the conjugate operation, and X dd is the input delay Doppler frame, k represents the Doppler index, l represents the delay index, n and m are the time index and frequency index respectively, and the output matrix of the 2D pulse compression correlation is expressed as is regarded as the set of compressed frames;
[0019] After obtaining the compressed frames, divide all the compressed frame data into the real part and the imaginary part to be processed separately by the neural network channels. The expression is as follows:
[0020]
[0021] Among them, represents the result of the compression and separation of the real part and the imaginary part of the input D signal frame at time t.
[0022] Exemplarily, compensating and splicing the predicted delay and Doppler parameters at time t with the OTFS frame at time t includes:
[0023] Mapping the predicted OTFS frame delay and Doppler parameters at time t to the DD domain plane to obtain corresponding coordinate data;
[0024] Fusing the real and imaginary parts of the predicted delay and Doppler parameters at time t with the imaginary and real parts of the preprocessed compressed frame at time t respectively to obtain a fused frame.
[0025] Exemplarily, the first ESnet and the second ESnet respectively include: a pre-training layer, a fine-tuning layer, and a fully connected layer;
[0026] The pre-training layer includes a first residual block layer1 and a second residual block layer2, and the fine-tuning layer includes a third residual block layer3 and a fourth residual block layer4;
[0027] The first, second, third, and fourth residual blocks respectively include, in sequence: a convolutional layer CONV, a batch normalization layer BN and an activation function ReLU, a convolutional layer CONV, and a batch normalization layer BN;
[0028] The expression of the remaining connectivity is as follows:
[0029]
[0030] Where is the output of each residual block, including the original input x of the corresponding residual block and the optimized output of the corresponding residual block
[0031] Exemplarily, an attention mechanism is added before the pre-training layer, and the attention mechanism equation is as follows:
[0032] M s (x) = σ(conv([avg(x), max(x)]))
[0033] Where M s (·) represents the function of the spatial attention mechanism, σ represents the sigmoid non-linear activation function, conv(·) is the convolution operation, and avg(·) and max(·) respectively take the mean and maximum values of the input;
[0034] Connect the two feature sets of the mean and the maximum value, and use the sigmoid activation function to calculate the weight coefficient, scale the parameters to the range and learn the non-linearity of the features,
[0035]
[0036] Among them, σ(x) represents the output value after applying the Sigmoid function, and this value is between 0 and 1, where e -x represents the negative exponential power of the base e of the natural logarithm, e -x is an exponential decay term. The spatial attention mechanism calculates the importance weights of each position in the input data through convolution.
[0037] Exemplarily, the training process of the PAEnet network includes:
[0038] During the offline training process, the empirical MSE is used as the cost function, and the empirical MSE loss function is used together with the gradient descent and backpropagation algorithms to iteratively refine the weight parameters of the network. Finally, the trained PAEnet network is obtained through multiple iterations.
[0039] On the other hand, an embodiment of the present disclosure proposes a drone perception system based on OTFS signals. The system includes: a pre-trained prediction-assisted estimation network PAEnet and a fusion module: The prediction-assisted estimation network PAEnet includes: a target parameter extraction network, a prediction sub-network, and an estimation network;
[0040] The target parameter extraction network is used to extract the target delay and Doppler parameter information of all OTFS frames in the preset time period τ - 1 before the current time t according to the real-time acquired target OTFS data. Among them, the target parameter extraction network is obtained by training a first ESnet based on a residual network as the backbone;
[0041] The prediction sub-network is used to predict the delay and Doppler parameters at time t according to the extracted target delay and Doppler parameter information of all preset time periods τ - 1; Among them, the prediction sub-network includes: being obtained by training based on PR-subnet;
[0042] The compensation splicing module is used to perform compensation splicing on the predicted delay and Doppler parameters at time t and the OTFS frame at time t;
[0043] The estimation network is used to estimate the data after compensation splicing to obtain the delay and Doppler parameters of the OTFS frame at time t. Among them, the estimation network is obtained by training a second ESnet based on a residual network.
[0044] On the other hand, an embodiment of the present disclosure proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0045] On the other hand, an embodiment of the present disclosure proposes a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0046] The beneficial effects of the embodiments of the present disclosure are as follows:
[0047] The method for unmanned aerial vehicle (UAV) perception according to the OTFS signal in the embodiments of the present disclosure, especially in a scenario with pulse distortion, adopts an auxiliary estimation network PAEnet to accurately predict the delay and Doppler parameters. As a key information compensation, it improves the current sensing accuracy.
[0048] The first ESnet of the auxiliary estimation network PAEnet extracts the previous delay and Doppler parameters, and uses the PR - subnet to generate contemporary predictions. Combining the predictions and the current observations, another ESnet performs the final delay and Doppler estimation. Compared with the prior art, the deep learning DL adopted in the embodiments of the present disclosure has advantages in the denoising process and can maintain robustness in the presence of pulse noise. Description of the Drawings
[0049] Figure 1 is a schematic flow chart of a method for UAV perception according to the OTFS signal proposed in the embodiments of the present disclosure;
[0050] Figure 2 is a schematic diagram of an orthogonal time - frequency - space (OTFS) modulation structure;
[0051] Figure 3 is a schematic diagram of a parameter sensing model structure when the target UAV is in a simple scenario;
[0052] Figure 4 is a schematic diagram of the comprehensive framework structure of the compensation mechanism adopted by the method for UAV perception according to the OTFS signal in the embodiments of the present disclosure;
[0053] Figure 5 is a schematic diagram of the network structure of the deep learning adopted by the method for UAV perception according to the OTFS signal in the embodiments of the present disclosure;
[0054] Figure 6 is a schematic diagram before and after the pre - processing of OTFS data in the method for UAV perception according to the OTFS signal in the embodiments of the present disclosure. Among them, Fig. (a) is a schematic diagram before the pre - processing of the OTFS signal frame; Fig. (b) is a schematic diagram after the pre - processing of the OTFS signal frame:
[0055] Figure 7 is a schematic diagram of the first ESnet structure in the method for UAV perception according to the OTFS signal in the embodiments of the present disclosure:
[0056] Figure 8 is a schematic diagram of the structure of the prediction sub - network in the method for UAV perception according to the OTFS signal in the embodiments of the present disclosure:
[0057] Figure 9It is a schematic diagram for comprehensively simulating and comparing the prediction-assisted estimation network compensation result of the embodiments of the present disclosure with the prior art. Specific Embodiments
[0058] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer and more understandable, the following further elaborates on the embodiments of the present disclosure in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present disclosure and are not used to limit the embodiments of the present disclosure.
[0059] Orthogonal time-frequency-space (OTFS) modulation technology can provide reliable communication and accurate sensing in high-mobility scenarios and has become a potential solution for various unmanned aerial vehicle (UAV)-related applications. In the embodiments of the present disclosure, the UAV sensing problem based on OTFS communication waveforms is considered. Due to random gusts and variable weather conditions, the sensing signal may be suddenly interfered with. To effectively address this challenge, a deep learning (DL)-based framework is proposed in the embodiments of the present disclosure to compensate for impulse interference. This framework uses empirical information to generate real-time predictions and then combines the predictions with current unreliable observation data to achieve accurate sensing of UAVs. Specifically, the embodiments of the present disclosure develop a prediction-assisted estimation network (PAEnet, Prediction-Assisted Estimation Network) to implement the proposed framework. The core component of PAEnet - the estimation network (ESnet) - can directly extract delays and Doppler from the transmitted OTFS frames, thereby reducing the complexity of the sensing process.
[0060] First, OTFS is introduced. For an OTFS system, it modulates and demodulates signals in the delay-Doppler domain. Domain conversion is often used among DD, time-frequency, and time domains. As Figure 2 shown, the OTFS modulation structure and operation process are briefly introduced. Specifically as follows:
[0061] Consider a delay and Doppler grid of size M×N, where M and N are the number of subcarriers and time slots of the OTFS signal, respectively. Here, the delay and Doppler plane formulas are introduced:
[0062] ∧={(lΔτ,kΔν),l=0,1,2,...,M-1,k=0,1,2,...,N-1} (1)
[0063] where Δτ and Δv are the normalized time delay interval and Doppler interval, respectively. k represents the Doppler index, and l represents the delay index. In addition, the time-frequency plane is given here:
[0064]
[0065] Among them, Δf and T are the subcarrier spacing and duration respectively. The delay-Doppler (DD) domain and the time-frequency (TF) domain are related.
[0066]
[0067] Among them, τ and v represent the actual time delay and Doppler frequency. They are derived by multiplying the index values l, k by the corresponding resolutions respectively.
[0068] At the transmitter side, the symbol x[k, l] in the DD domain is mapped to the time-frequency (TF) domain through the inverse symplectic finite Fourier transform (ISFFT) to obtain the time-frequency domain signal X tf [n, m]:
[0069]
[0070] Among them, n and m are the time index and frequency index respectively. The transmitted time-frequency domain signal X tf [n, m] is transformed into the time domain signal s(t) through the Heisenberg transform, where e is the base of the natural logarithm and j represents the imaginary unit.
[0071]
[0072] Among them, ɡ tx (t) is a pulse shaping filter, t = 0, 1, …… T - 1, and t represents the time instant. Then, the time domain signal s(t) will pass through a time-varying multipath channel, whose characteristics can be described by the delay-Doppler domain channel spread function h(τ, v), and h(τ, v) is as shown in Equation (6):
[0073]
[0074] Among them, p represents the number of propagation paths, ɡ i , τ i , v i represent the complex path gain of the i-th path, the delay of the i-th path, and the Doppler shift respectively. τ ∈ [0, τ max , v ∈ [-v max , v max are the delays and Doppler shifts within the maximum delay and maximum Doppler shift of all the channel paths, and δ(·) represents the Dirac delta function. After passing through the above channel, the signal is captured by the receiver, and the received signal is r(t), as shown in Equation (7):
[0075] r(t) = ∫∫h(τ, v)s(t - τ)e j2πv(t-τ) dτdv + n(t) (7) Among them, n(t) is zero-mean additive white Gaussian noise, and s(t) represents the transmitted signal. The received signal is converted to the TF domain through the Wigner transform to obtain Yt,f (t, f), as shown in Equation (8):
[0076]
[0077] where A{ɡ rx , r}(t, f) is the cross ambiguity function between the pulse shaping filter ɡ rx (t) and the received signal; ɡ rx (t) represents the receiver filter, * represents conjugate, f represents the frequency variable, r(t) represents the received signal in the time domain; t’ is the integration variable, representing the time point currently being considered, and the integration needs to be performed over the entire time axis. Finally, through the symmetric finite Fourier transform (SFFT), the signal is converted back to the DD domain for further analysis to obtain y[k, l], as shown in Equation (9):
[0078]
[0079] where Y[n, m] is the signal after discrete sampling of Y t,f (t, f).
[0080] Further discuss the relationship between the received DD domain sample matrix Y dd and the symbol matrix X dd . The TF domain channel matrix H tf ∈ C N×M can be expressed as Equation (10):
[0081] H t,f [n, m] = ∫∫h(τ, v)e j2πvnT e -j2πΔfτ dτdv (10)
[0082] where h(τ, v) is the DD domain channel defined in Equation (6). For ɡ tx (t) and ɡ rx (t) satisfying the biorthogonality condition, the received TF domain sample matrix can be expressed as Equation (11):
[0083] Y tf [n, m] = H tf [n, m]X tf [n, m] + Z tf [n, m] (11)
[0084] where Z tf ∈ C N×M is the TF domain additive noise sample matrix at the base station BS. Applying the SFFT to both sides of Equation (11) gives a DD domain equation, as shown in Equation (12):
[0085]
[0086] where Z dd ∈ C N×M is the DD domain additive noise sample matrix at the base station BS, H dd is the DD domain channel matrix, X dd is the transmit symbol matrix, k’, l’, k - k’, etc. are all index values of the lattice points, as shown in Equation (13):
[0087]
[0088] For the convenience of parameter estimation, the OTFS system model in Equation (12) can be rewritten as the following Equation (14)
[0089]
[0090] where y ∈ C MN×1 is the vectorized form of Y dd The vector z ∈ C MN×1 is the vectorized form of Z dd The components follow k v = [k v1 ,,, k vP T ∈ R P×1 , l τ = [l τ1 ,,, l τP T ∈ R P×1 and collect the normalized Doppler shift, the normalized delay offset, and the channel coefficients respectively. The matrix Φ(k v , l τ ) ∈ C MN×P is the measurement matrix and can be expressed as Equation (15):
[0091] Φ(k v , l r ) = [φ(k v1 , l τ1 ),..., φ(k vP P, l rP )] (15)
[0092] where P represents the number of paths / targets in the scenario. The parameter estimation problem is to estimate the target k v , l τ and v , l τ and
[0093] To clearly illustrate the framework proposed in the embodiments of the present disclosure, a dynamic model of UAV parameter perception is described by taking a UAV as an example. It is considered that the UAV and the base station subsystem BSs have the same height, following the assumptions widely used in the literature, such as in Reference [Wu, W. Yuan, and L. Bai, "On the interplay between sensing and communications for UAV trajectory design," IEEE IoT J., vol. 10, no. 23, pp. 20383–20395, 2023.] and Reference [J. Wu, W. Yuan, and L. Hanzo, "When UAVs meet ISAC: Realtime trajectory design for secure communications," IEEE Trans. Veh. Technol., vol. 72, no. 12, pp. 16766 - 16771, 2023.]. Therefore, the coordinates of the UAV can be simplified to a two-dimensional case. It should be noted that for some operations, the extension to a 3D scenario is straightforward and will not be elaborated here.
[0094] As Figure 3 shown, a dynamic model of parameter perception is introduced in a simple scenario of a target UAV. Two base stations BS, BS - 1 and BS - 2 are separated by a distance d, transmitting OTFS signals and respectively processing the echo signals reflected by the UAV. It is assumed that the origin O is located at BS - 1, and the initial i-th target position of the target UAV is (x i,0 , y i,0 ), and the distances between the target and the BSs at time i are l i1,0 , l i2,0 which can be expressed by the Euclidean distance as the following formula (16):
[0095]
[0096] The relative motion angle between the target UAV and the BS receiving the echo signal can be expressed by the calculated distance as the following formula (17):
[0097]
[0098] where, v y i,0 , v x i,0 respectively represent the velocities of the UAV in the y-axis and x-axis directions at the position (x i,0 , y i,0 ), and θ i represents the UAV at the target position (x i,0, y i,0 ) The included angle between the connection line with the base station BS-2 and the horizontal plane.
[0099] When the initial velocity of the unmanned aerial vehicle is v i,0 , the acceleration a i The accelerations in the x-axis and y-axis directions are respectively a x i,0 , a y i,0 The unmanned aerial vehicle flies from the position (x i,0 , y i,0 ) to the position (x i,1 , y i,1 ) after a time interval T’, and the kinematic equation of the unmanned aerial vehicle is calculated, as shown in formula (18):
[0100]
[0101] Where, v x i,1 , v y i,1 respectively represent the velocities of the unmanned aerial vehicle in the x-axis and y-axis directions at the position (x i,1 , y i,1 ), and a x i,1 represents the acceleration of the unmanned aerial vehicle in the x-axis direction at the position (x i,1 , y i,1 ).
[0102] Using the above parameters of the unmanned aerial vehicle, the actual time delay τ i,0 and the Doppler frequency v of formula (3) can be calculated by the following formula (19):
[0103]
[0104] Where, f c is the carrier frequency, and θ is the included angle between the movement of the unmanned aerial vehicle and the BS. The time delay τ i,0 and the Doppler frequency v can be used as sample data for training and predicting the auxiliary estimation network PAEnet.
[0105] As Figure 1 shown, on the one hand, an embodiment of the present disclosure proposes a method for unmanned aerial vehicle perception based on OTFS signals, which adopts a pre-trained prediction auxiliary estimation network PAEnet. The auxiliary estimation network PAEnet includes: a target parameter extraction network, a prediction sub-network, and an estimation network; the method includes:
[0106] Step S100: Obtain the target OTFS data in real time, and use the pre-constructed target parameter extraction network to extract the target delay and Doppler parameter information of all OTFS frames in the preset time period τ-1 before the current time t. Among them, the target parameter extraction network is obtained by training the first ESnet based on the residual network as the backbone.
[0107] The purpose of the embodiments of the present disclosure is to achieve sensing based on OTFS in the case where the transmission signal of a moving target encounters impulse noise, and high-precision sensing effects can also be obtained. The moving target can be a drone, a vehicle, etc. The embodiments of the present disclosure take a drone as an example for specific illustration. In the method of the embodiments of the present disclosure, step S100 focuses on extracting the delay and Doppler indices corresponding to different environments of the moving target, which directly reflect the state of the moving target (drone). On the basis of existing technical research work, the embodiments of the present disclosure consider that the signal is strongly affected by a noisy environment, and introduce a horizontal compensation framework to eliminate the influence of impulse noise interference on the sensing signal. The PAEnet framework of the embodiments of the present disclosure utilizes the temporal features of the target (drone), which helps to accurately predict the delay and Doppler parameters. These predictions can be used as key information compensation to improve the current sensing accuracy, especially in scenarios characterized by pulse distortion. The embodiments of the present disclosure use the pre-constructed ESnet1 to extract the delay and Doppler parameters of the target in the received OTFS frames for subsequent estimation. The pre-constructed ESnet2 can improve the sensing efficiency of OTFS and estimate the component delay and Doppler.
[0108] The compensation mechanism and implementation framework principle of the method for deep learning-assisted drone perception according to the OTFS signal in the embodiments of the present disclosure are as follows:
[0109] For the impulse noise of the drone sensing signal, the embodiments of the present disclosure propose a compensation mechanism that uses empirical information to optimize the current sensing process. First, the basic framework of the compensation mechanism is introduced, and then the implementation of the specific PAEnet is proposed.
[0110] A. Compensation framework. Equation (12) describes the received DD-domain OTFS frame, specifically as follows in formula (20):
[0111]
[0112] Among them, the component of the DD-domain additive noise Z dd follows In addition, another noise component is affected by specific environmental factors and is characterized by a high degree of randomness, such as thunder and lightning. Similar to the reference [G. Wang, H. Zhan, G. Zhang, X. Gui, and D. Xu, "Adaptive compensation method of position estimation harmonic error for emf-based observer in sensorless ipmsm drives," IEEE Trans. Power Electron., vol. 29, no. 6, pp. 3055–3064, 2013.], the DD-domain impulse noise under special conditions is modeled as ∈ dd [k, l], whose elements follow The channel input-output relationship becomes:
[0113]
[0114] where p(t) represents the arrival probability of the impulse noise according to the time variable. ∈ dd The appearance of [k, l] depends on the changes in environmental factors, making it have variability and randomness properties. Due to its unpredictability in terms of time and frequency, traditional model-based estimators are difficult to provide a robust solution. To eliminate the ∈ dd destructive impact of [k, l], a compensation mechanism is proposed and will be discussed as follows.
[0115] As Figure 4 shown, assume that the sensing signal of the target UAV encounters impulse noise perturbation at time t, resulting in a damaged message in the receiving part, denoted as with insufficient information contained therein, making it challenging to directly estimate the channel target parameters. The estimated channel is denoted as Due to the transient nature of the impulse noise, the signals of the previous frame are not affected, resulting in compensating with the previous data sequence containing this information. Therefore, the complete data set can be described by the following formula (22):
[0116]
[0117] where, represents a data set of the time series from time t - τ + 1 to time t. t represents the current time, τ ∈ R represents the length of the time series data set, and t - τ + 1 represents the start time. By analyzing the previous signals, an attempt is made to extract valuable information related to the current state for compensation. It should be noted that the prediction of the current parameters is obtained as follows in formula (23):
[0118]
[0119] Among them, G(·) represents the information extraction and prediction function, represents the compensation parameter for prediction. The prediction can be decomposed into the exact Y t and the prediction error W p in two parts. The error is mainly caused by the estimator and affects the reliability of the compensation information. Similar to the prediction, can also be decomposed as the following formula (24):
[0120]
[0121] where, W z and W ∈ are the noises caused by normal noise and impulse noise respectively. Thus, two versions of the Y t information are obtained, and which respectively retain the information not affected by impulse noise and prediction error. W p and W z+ W τ The magnitudes of respectively determine the reliability of the prediction and the observation, which are the main components to be considered. Therefore, the final estimation requires an appropriate data fusion and utilization process. Assume that F(·) represents the function of the estimator and C(·) represents the data fusion function, represents the fused data, and the final estimation is proposed:
[0122]
[0123] where, the G(·) extraction and prediction function is equivalent to the first ESnet network and the PR-subnet network, the F(·) estimator function is equivalent to the second ESnet network, and the C(·) data fusion function is equivalent to splicing. This part of the content introduces the principle of PAEnet in the embodiments of the present disclosure. By optimizing the weight parameter λ caused by the input reliability, λ is pre-optimally set. The goal of the embodiments of the present disclosure is to minimize the noise effect and collect valuable channel information Y t for accurate estimation of time-delay Doppler and channel gain. Due to the integration of the prediction model based on previous data and contemporary observations, the estimation inherently exceeds the analysis that only relies on the interrupted current observation.
[0124] B. Network implementation. The above compensation framework outlines a method for compensating the unreliable information of the current state of the UAV based on previous parameters. The compensation framework includes parameter extraction, prediction, data fusion, and estimation. It can be used in vehicle networks, UAV sensing, and various application scenarios, especially for noise interference signals.
[0125] In the embodiments of the present disclosure, the above compensation framework is instantiated by using a corresponding neural network module. As Figure 5 shown, Y t-τ+1 etc. represent the data sets of the OTFS frames received at corresponding times. Based on the predicted Unmanned Aerial Vehicle Sensing Estimation Neural Network PAEnet, in order to clarify the PAEnet network structure, the preprocessing of the input data is introduced first, and then each subnet and network estimation of PAEnet are introduced in detail.
[0126] A. Input data preprocessing. The initial data used is the received OTFS frame as shown in Equation (26):
[0127]
[0128] where, represents the OTFS frame at time t. For example, there is interference, is the input data sequence with a time series length τ∈R. It should be noted that the signal frame passes through a time-varying multipath channel as shown in Equation (6). As Figure 6 (a) shows, the received frame is interfered with and appears messy, which makes it difficult for the neural network to detect the target. In order to improve the quality of the input data, a data preprocessing estimator based on two-dimensional correlation is adopted as shown in the following Equation (27):
[0129]
[0130] where, Y dd is the signal before correlation, that is, the data of the corresponding frame in the data, without channel estimation or equalization processing, * refers to the conjugate operation, and X dd is the input delay-Doppler frame. The output matrix of 2D pulse compression (correlation) is denoted as This matrix represents the correlation values at different time-frequency points and can be regarded as a set of compressed frames. As Figure 6 (b) shows, the noise is compressed along the delay and Doppler axes, and the target becomes prominent, which can be regarded as the input of the subsequent PAEnet.
[0131] In addition, the complex plane data is divided into real and imaginary neural network channels. The input target parameter extraction network is that the data can be described as shown in the following Equation (28):
[0132]
[0133] where, represents the result of preprocessing the OTFS frames at several moments from time t-τ + 1 to time t, represents the result of preprocessing the OTFS frame at time t, It represents the set after all compressed frame data are respectively divided into the real part and the imaginary part for neural network channel processing. It represents the result of input signal frame compression and separation into real and imaginary parts at time t of the target parameter extraction network.
[0134] The goal of the embodiments of the present disclosure is to accurately determine specific target parameters, namely delay and Doppler parameter information, by analyzing the delay and Doppler planes of the OTFS signal frames received in the DD domain. The embodiments of the present disclosure formulate this task as an extraction and prediction problem of delay and Doppler indices in the DD domain, which is accomplished through the proposed PAEnet architecture. Consistent with this framework, the current and previously received OTFS frames are utilized. First, data preprocessing is performed to denoise the input data to enhance its applicability to network processing. In front of PAEnet, an ESnet1 based on residual learning is introduced to extract delay and Doppler indices from previous data frames containing channel information. These indices serve as mappings of the target physical state information and thus exhibit temporal relationships due to the motion dynamics of the target. Subsequently, an LSTM-based PR-subnet is employed to delve into these temporal relationships to predict subsequent frame parameters. By compensating and splicing the observed frame at the current moment with the prediction information generated by the PR-subnet, the final estimates of the delay and Doppler indices of the OTFS frame and ESnet2 are obtained. The structure of the estimation network ESnet2 is similar to that of ESnet1 because they are both involved in the parameter extraction task.
[0135] The neural network-based compensation framework utilizes a large dataset, effectively leveraging the information embedded in the previous signals, thereby improving the accuracy and efficiency of UAV sensing.
[0136] B. The first ESnet in the PAEnet structure: Formulate the target information extraction problem as a regression task, and the output is the delay and Doppler indices. The first ESnet network structure is as Figure 7 shown.
[0137] As an example of the first ESnet, the first ESnet includes: a pre-training layer, a fine-tuning layer, and a fully connected layer;
[0138] The pre-training layer includes a first residual block layer1 and a second residual block layer2, and the fine-tuning layer includes a third residual block layer3 and a fourth residual block layer4.
[0139] The first, second, third, and fourth residual blocks respectively and sequentially include: a convolutional layer CONV, a batch normalization layer BN, and an activation function ReLU, a convolutional layer CONV, and a batch normalization layer BN.
[0140] To effectively utilize the input information and suppress the noise components, a residual network is adopted as the backbone of the proposed ESnet, and the target parameter extraction network can be trained with the data obtained by Equation 19. The residual network introduces residual connections between specific layers, facilitating information propagation and solving the problems of gradient vanishing and explosion. In addition, the connections preserve the original features, thus improving the denoising effect. By stacking multiple residual blocks, the residual network can increase its depth while maintaining a high performance level. The structure of the residual block is as shown in Figure 7 shown, which consists of a convolutional layer, a batch normalization layer, and some activation functions. In the figure, X is equivalent to x in the text, and z is equivalent to x in the text. The expression of the residual connectivity is as shown in Equation (29):
[0141]
[0142] where is the output of the current residual block, containing the original input x of the previous residual block and the optimized output of the previous residual block. Here, the process of the residual connection is directly implemented through addition.
[0143] The first ESnet proposed in the embodiments of the present disclosure is a relatively deep network, which requires a large amount of data and time during the training process. However, the true accessibility of the training data is usually limited. Therefore, appropriately fixing and fine-tuning specific convolutional layers provides a promising strategy. The convolutional layers are sorted at the top and participate in general information extraction tasks, such as edge detection. These layers are always executed in various scenarios and can be used without parameter re-optimization. On the contrary, the convolutional layers towards the back undertake different features in different extraction tasks, making them suitable candidates for fine-tuning. Therefore, Figure 7 layer1 and layer2 in are set as pre-trained layers, where the parameters are fixed by using a pre-trained model. At the same time, layer3 and layer4 are set as fine-tuning layers, where parameter updates are required. This structure will reduce the tasks of network training, save the training data, and reduce the training duration, especially for various signal-to-noise ratio (SNR) ratio scenarios.
[0144] It should be noted that the noise components are suppressed by the above-related algorithms, and the input frames are more prominent in the case of only a few scattered target points, showing sparse features. To help the network focus on the targets, a spatial attention mechanism is added at the front of the structure. The attention mechanism equation is introduced as Equation (30):
[0145] M s (x) = σ(conv([avg(x), max(x)])) (30)
[0146] where, M s(·) represents the function of the spatial attention mechanism, σ represents the sigmoid non-linear activation function, conv(·) is the convolution operation, and avg(·) and max(·) respectively take the mean and maximum values of the input. After obtaining the mean and maximum values, the convolution method is used to extract features. Then, these two feature sets are concatenated, and the sigmoid activation function is used to calculate the weight coefficients, scaling the parameters to the range (0, 1) and learning the non-linearity of the features:
[0147]
[0148] Among them, σ(x) represents the output value after applying the Sigmoid function, and this value is between 0 and 1. e -x represents the negative exponential power of the base e of the natural logarithm, e -x is an exponential decay term. The spatial attention mechanism calculates the importance weights of each position in the input data through convolution. Based on the weights, the target parameter extraction network can capture key information more effectively.
[0149] After traversing the residual layer, in addition to promoting the final regression optimization, the target parameter extraction network finally forms a fully connected layer responsible for compressing and reducing the data dimension. Subsequently, a smoother dimension transformation is performed to ensure output stability. The preprocessed data set of each data is input into the first ESnet to extract the target delay and Doppler parameter information of the OTFS frame in the τ - 1 time period. The output of the target parameter extraction network includes multiple groups of two-dimensional time delay and Doppler index coordinates, denoted as the following formula (32):
[0150]
[0151] Among them, represents the parameters at time t - 1, that is, the time delay and Doppler index extracted by the target parameter extraction network. E(·) is the total function of the first ESnet, is the extracted delay Doppler index value of the specific time t - 1 of target i, is the data at time t - 1 after preprocessing. The embodiments of the present disclosure can extract the delay and Doppler data of multiple targets in the OTFS frame.
[0152] Step S200, according to all the target delay and Doppler parameter information of the extracted preset time period τ - 1, use the pre-trained prediction sub-network to predict the delay and Doppler parameters at time t; wherein, the prediction sub-network includes: trained based on the PR - subnet network.
[0153] Specifically, PR - subnet: To make full use of the temporal connection of the estimation metrics, a PR - subnet is proposed for parameter prediction. At the input layer of the PR - subnet, the input data is obtained from the output of the first ESnet in step S100:
[0154]
[0155] Among them, represents the estimated previous time - series parameters, that is, the target delay and Doppler parameter information of all OTFS frames extracted by the target parameter extraction network in the time period τ - 1.
[0156] Figure 8 Fig. shows the structure and data flow of the PR - subnet. Among them, the input data passes through two LSTM layers, a fully - connected layer and outputs the predicted parameters. The data is passed to the two LSTM layers. By improving the recurrent neural network RNN, LSTM can solve the problems of gradient vanishing and explosion in RNN and better control longer - term temporal relationships. Each LSTM layer consists of multiple LSTM units. Among them, each unit uses gates to regulate the information flow. These gates, including the forget gate, input gate and output gate, control the update of the unit state and output, enabling the LSTM layer to process sequential data and effectively capture long - term dependencies. As Figure 8 shown, the LSTM unit obtains random initial states h0 and c0 and updates them through the input data. The unit passes the hidden state h and the cell state c to the next unit, which is regarded as memory information. In addition, the hidden state is passed to the same unit in the next layer as the output. The LSTM prediction is the output derived from the last cell of LSTM layer - 2 A fully - connected layer is used to improve the overall performance, which can extract the output features of LSTM layer - 2 and reshape the output size. Due to the normalization of the non - linear activation function, this fully - connected layer also serves as a numerical mapping to prevent scaling problems. The output of the PR - subnet can be expressed as the following formula (34):
[0157]
[0158] Among them, represents the predicted parameters at time t, that is, the predicted delay - Doppler parameters, P(·) is the total function of the PR - subnet, is the predicted delay - Doppler index value of the specific target i at time t.
[0159] Step S300: Compensate and splice the predicted delay and Doppler parameter information at time t with the OTFS frame at time t.
[0160] Step S310: Map the predicted OTFS frame delay and Doppler parameters at time t to the DD domain plane to obtain corresponding coordinate data.
[0161] Specifically, map the predicted specific index value to the DD domain plane to obtain DD domain plane data for the data fusion process. F(·) represents the mapping function, and the output can be expressed as the following formula (35):
[0162]
[0163] where can be regarded as the predicted time-delay Doppler domain signal of the input at time t based on previous data.
[0164] Step S320: Respectively fuse the real and imaginary parts of the predicted delay and Doppler parameters at time t with the imaginary and real parts of the preprocessed compressed frame at time t to obtain a fused frame.
[0165] Subsequently, fuse in the channels where the imaginary and real parts are separated to and stitch:
[0166]
[0167] where is the comprehensive compensation and observation data, that is, the data after compensating and stitching the predicted data and the current observation data, and finally fed through the output network.
[0168] The output ESnet2 shares the same architecture configuration as ESnet1, only differing in parameterization to adapt to different input dimensions.
[0169] Step S400: Input the compensated and stitched data into a pre-trained estimation network to obtain the delay and Doppler parameters of the OTFS frame at time t, where the estimation network includes: trained based on the second ESnet of the residual network.
[0170] As Figure 5 shown, input the fused frame into the estimation network to obtain the estimated delay and Doppler parameters of the OTFS frame at time t. Since the two types of information in are respectively affected by impulse noise and prediction error. ESnet2 will learn to optimize the estimation result by adjusting the weights of using prediction and observation information. Finally, the output of PAEnet can be expressed as follows:
[0171]
[0172] where and f PAERepresents the total function of the proposed PAENnet.
[0173] For the training and estimation of the PAEnet network, the offline training and network estimation processes will be introduced. The total training set can be expressed as Equation (38):
[0174]
[0175] where w ∈ {1, 2,..., n s}, is the w-th training example generated by the correlation function (27). And S (w) = [(x1, y1), (x2, y2),..., (x i , y i )] T represent the input matrix dimension and output parameters respectively.
[0176] In addition, the training of PAEnet requires a cost function. According to the MMSE criterion (Minimum Mean Square Error), the optimal MMSE estimator is obtained by minimizing the Bayesian MSE. Specifically
[0177]
[0178] where and S t represent the estimated and true parameters of the UAV. The Bayesian MSE expression is an intuitive choice for the cost function to obtain the best performance. However, in practical scenarios, the statistical Bayesian MSE cannot be accessed with a finite number of training examples, making the adoption of the empirical MSE inevitable. The empirical MSE cost function can be expressed as Equation (40) as follows:
[0179]
[0180] where F represents the Frobenius norm, considering is the i-th sensing parameter of PAEnet based on , represents the predicted parameter of the i-th sample. The empirical MSE cost function can be expressed as Equation (41) as follows:
[0181]
[0182] where F represents the Frobenius norm, which is a commonly used norm in the matrix space. This loss function is used together with the gradient descent and backpropagation algorithms to iteratively refine the weight parameters of the network, and finally a well-trained PAEnet is obtained through multiple iterations.
[0183] As Figure 9 shown, through comprehensive simulation results, among them, the PAEnet method with a compensation mechanism has the best performance, proving the effectiveness of the compensation mechanism in unreliable sensing scenarios, and at the same time demonstrating the ability of PAEnet to achieve excellent accuracy for OTFS-based UAV sensing.
[0184] As Figure 5 shown, on the other hand, an embodiment of the present disclosure proposes a UAV sensing system according to an OTFS signal. The system includes an auxiliary estimation network PAEnet and a fusion module: The auxiliary estimation network PAEnet includes: a target parameter extraction network, a prediction sub-network, and an estimation network; The system includes:
[0185] A target parameter extraction network, configured to obtain target OTFS data in real time, and extract target delay and Doppler parameter information of the current frame and previous frames. Among them, the target parameter extraction network includes a first ESnet based on a residual network as the backbone;
[0186] A prediction sub-network, configured to predict the OTFS delay and Doppler parameters of the next frame according to the extracted target delay and Doppler parameter information; Among them, the prediction sub-network includes: trained based on PR-subnet;
[0187] A compensation splicing module, configured to perform compensation splicing on the received OTFS frame of the next frame and the predicted OTFS delay and Doppler parameters of the next frame;
[0188] An estimation network, configured to estimate the data after compensation splicing to obtain the OTFS delay and Doppler parameters of the next frame. Among them, the estimation network includes a second ESnet based on a residual network trained.
[0189] On the other hand, an embodiment of the present disclosure proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described above are implemented.
[0190] On the other hand, an embodiment of the present disclosure proposes a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described above are implemented.
[0191] In some processes described in the specification, claims, and the above-mentioned drawings of the embodiments of the present disclosure, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" herein are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0192] The above are only the preferred embodiments of the embodiments of the present disclosure. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the embodiments of the present disclosure, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present disclosure.
Claims
1. A drone sensing method based on OTFS signals, characterized in that, The pre-trained Prediction-Assisted Estimation Network (PAEnet) is adopted. The Prediction-Assisted Estimation Network (PAEnet) includes: a target parameter extraction network, a prediction sub-network, and an estimation network. The method includes: Obtain target OTFS data in real time, and use the pre-constructed target parameter extraction network to extract the target delay and Doppler parameter information of all OTFS frames in the preset time period τ-1 before the current time t. Among them, the target parameter extraction network is obtained by training the first ESnet based on the residual network as the backbone; According to the extracted target delay and Doppler parameter information of the preset time period τ-1, use the pre-trained prediction sub-network to predict the delay and Doppler parameters at time t. Among them, the prediction sub-network includes: being obtained by training based on the PR-subnet network; Compensate and splice the predicted delay and Doppler parameters at time t with the OTFS frame at time t; Input the compensated and spliced data into the pre-trained estimation network to obtain the delay and Doppler parameters of the OTFS frame at time t. Among them, the estimation network includes: being obtained by training the second ESnet based on the residual network.
2. The method according to claim 1, characterized in that, After obtaining the OTFS data in real time, the method further includes: Adopt a data preprocessing estimator based on two-dimensional correlation to preprocess all OTFS frames in the preset time period τ-1 before and including the current time t to obtain corresponding compressed frames. The data preprocessing formula is as follows; Among them, V[k, l] represents the preprocessed data, and Y dd is the signal before correlation, that is, the set of received OTFS frames the data of the corresponding frames in the data, represents the data set of all OTFS frames in the time series from time t - τ + 1 to time t, where t represents the current time, τ ∈ R represents the length of the time series data set, and t - τ + 1 represents the start time; * represents the conjugate operation, and X dd is the input delayed Doppler frame, k represents the Doppler index, l represents the delay index, n and m are the time index and frequency index respectively, and the output matrix related to 2D pulse compression is expressed as is regarded as the set of compressed frames; After obtaining the compressed frames, divide all the compressed frame data into the real part and the imaginary part for neural network channel processing respectively. The expression is as follows: Among them, represents the result of signal frame compression and separation into real and imaginary parts of the input at time t.
3. The method according to claim 2, wherein The compensation and splicing of the predicted delay and Doppler parameters at time t with the OTFS frame at time t includes: Map the predicted OTFS frame delay and Doppler parameters at time t to the DD domain plane to obtain corresponding coordinate data; Fuse the real part and the imaginary part of the predicted delay and Doppler parameters at time t with the imaginary part and the real part of the preprocessed compressed frame at time t respectively to obtain a fused frame.
4. The method according to any one of claims 1 to 3, characterized in that, The first ESnet and the second ESnet respectively include: a pre-training layer, a fine-tuning layer, and a fully connected layer; The pre-training layer includes a first residual block layer1 and a second residual block layer2, and the fine-tuning layer includes a third residual block layer3 and a fourth residual block layer4; The first, second, third, and fourth residual blocks respectively include in sequence: a convolutional layer CONV, a batch normalization layer BN and an activation function ReLU, a convolutional layer CONV, and a batch normalization layer BN; The expression of the remaining connectivity is as follows: Among them, is the output of each residual block, including the original input x of the corresponding residual block and the optimized output of the corresponding residual block 5. The method according to claim 4, wherein Add an attention mechanism before the pre-training layer. The attention mechanism equation is as follows: M s (x) = σ(conv([avg(x), max(x)])) where M s (·) represents the function of the spatial attention mechanism, σ represents the sigmoid non-linear activation function, conv(·) is the convolution operation, and avg(·) and max(·) respectively take the mean and maximum values of the input; Connect the two feature sets of the mean and the maximum value, and use the sigmoid activation function to calculate the weight coefficient, scale the parameters to the range, and learn the non-linearity of the features. Among them, σ(x) represents the output value after applying the Sigmoid function, and this value is between 0 and 1, where e -x represents the negative exponential power of the base e of the natural logarithm, e -x is an exponential decay term, and the spatial attention mechanism calculates the importance weights of each position in the input data through convolution.
6. The method according to any one of claims 1 to 3, characterized in that, The PR-subnet network includes: two LSTM layers and a fully connected layer.
7. The method according to any one of claims 1 to 3, characterized in that, The training process of the PAEnet network includes: During the offline training process, the empirical MSE is used as the cost function. The empirical MSE loss function is used together with the gradient descent and backpropagation algorithms to iteratively refine the weight parameters of the network. Finally, the trained PAEnet network is obtained through multiple iterations.
8. An unmanned aerial vehicle perception system based on OTFS signals, characterized in that, The system includes a pre-trained Prediction-Assisted Estimation Network (PAEnet) and a fusion module. The Prediction-Assisted Estimation Network (PAEnet) includes: a target parameter extraction network, a prediction sub-network, and an estimation network. The target parameter extraction network is used to extract the target delay and Doppler parameter information of all OTFS frames in a preset time period τ-1 before the current time t according to the real-time acquired target OTFS data. Among them, the target parameter extraction network is trained by a first ESnet based on a residual network as the backbone. The prediction sub-network is used to predict the delay and Doppler parameters at time t according to the extracted target delay and Doppler parameter information of all in the preset time period τ-1. Among them, the prediction sub-network includes: being trained based on PR-subnet. The compensation splicing module is used to perform compensation splicing on the predicted delay and Doppler parameters at time t and the OTFS frame at time t. The estimation network is used to estimate the data after compensation splicing to obtain the delay and Doppler parameters of the OTFS frame at time t. Among them, the estimation network is trained by a second ESnet based on a residual network.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.