Unmanned aerial vehicle time-frequency image reconstruction method based on deep learning

The SelfNet model is used to reconstruct the time-frequency images of drones, which solves the problem of distortion of the time-frequency characteristics of drone echo signals in complex environments, achieves high-precision image reconstruction and classification under low signal-to-noise ratio conditions, and improves the accuracy of drone detection.

CN120779353APending Publication Date: 2025-10-14ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510802745.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

UAV echo signals are susceptible to interference in complex environments, resulting in distortion of their time-frequency characteristics. Existing methods are unable to effectively identify and analyze UAV time-frequency images under low signal-to-noise ratio conditions.

Method used

An autoencoder model (SelfNet) based on convolutional neural networks is used to reconstruct the time-frequency images of drones. The encoder extracts features and combines them with fully connected layers for feature fusion. The decoder reconstructs the image, and transfer learning is used to improve the generalization ability of the model on small sample data.

Benefits of technology

It effectively extracts and reconstructs high-quality time-frequency images, improves the accuracy of drone time-frequency images under low signal-to-noise ratio conditions, enhances the accuracy of drone detection and classification, and provides technical support for subsequent applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120779353A_ABST
    Figure CN120779353A_ABST
Patent Text Reader

Abstract

An unmanned aerial vehicle time-frequency image reconstruction method based on deep learning belongs to the technical field of radar signal processing and image reconstruction, solves the technical problem that unmanned aerial vehicle time-frequency images are difficult to identify and analyze under the condition of low signal-to-noise ratio, and comprises the following steps: S1, performing short-time Fourier transform on echo signals, and establishing an unmanned aerial vehicle radar echo model STFT; s2, designing a SelfNet model based on a convolutional neural network, and reconstructing an unmanned aerial vehicle time-frequency curve graph; s3, according to an unmanned aerial vehicle radar echo model STFT and a time-frequency curve output image, establishing an unmanned aerial vehicle time-frequency image data set, and performing training by using a SelfNet model; and S4, reconstructing the time-frequency image data set of the unmanned aerial vehicle by using the trained SelfNet model weight. The method can be used for actually measuring data, the precision of the time-frequency image of the unmanned aerial vehicle under the low signal-to-noise ratio can be improved, and follow-up multi-class classification and micro-motion parameter estimation of the unmanned aerial vehicle are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of UAV radar signal processing, and specifically relates to a UAV time-frequency image reconstruction method based on deep learning. Background Art

[0002] In recent years, drone technology has been widely used in various fields. Radar detection has gained widespread application due to its advantages in long-range, high-precision positioning, and rapid response. Furthermore, research on the micro-Doppler characteristics of drones has attracted considerable attention. However, drone echo signals are susceptible to interference in complex environments, resulting in distortion of their time-frequency characteristics. Traditional time-frequency analysis methods (such as short-time Fourier transform and wavelet transform) have limitations in addressing such problems.

[0003] Deep learning-based time-frequency image reconstruction algorithms aim to leverage the adaptive learning capabilities of deep neural networks to extract effective information from noise interference and channel distortion, reconstructing high-quality time-frequency images. Currently, research on drone time-frequency image reconstruction is still in its exploratory stages, and existing methods face numerous challenges in model structure design and training data generation. Summary of the Invention

[0004] In order to overcome the shortcomings of the existing technology and solve the technical problem that UAV time-frequency images are difficult to effectively identify and analyze under low signal-to-noise ratio conditions, the present invention provides a UAV time-frequency image reconstruction method based on deep learning. It uses an autoencoder model (SelfNet model) based on a convolutional neural network to reconstruct UAV time-frequency images, improve the accuracy of UAV time-frequency images under low signal-to-noise ratio conditions, and provide technical support for subsequent multi-category classification and micro-motion parameter estimation of UAVs.

[0005] The present invention is implemented through the following technical solutions: A method for reconstructing time-frequency images of a UAV based on deep learning, comprising the following steps:

[0006] S1. Establish the UAV radar echo model STFT, perform short-time Fourier transform on the echo signal, and obtain the initial time-frequency image;

[0007] The UAV radar echo model STFT is:

[0008]

[0009] In formula ①, S sum (t) is the original echo signal; w(t) is the window function of the short-time Fourier transform; τ is the position where the window function moves; ω is the Doppler frequency in Hertz (Hz); j is the imaginary unit, and j is specified as 2 =-1;

[0010] Original echo signal S sumThe expression of (t) is:

[0011]

[0012] In formula ②, L is the blade length, in meters (m); θ k =θ0+2kπ / N represents the initial rotation angle under different blade numbers N, in degrees (°), Φ k (t) is the phase function of the initial rotation angle:

[0013]

[0014] In formula ③, λ is the radar wavelength, in meters (m); l is the distance from any scattering point on the blade to the center of the blade, in meters (m); θ0 is the initial phase, in degrees (°); f r is the blade rotation rate, in revolutions per second (r / s); t represents the elapsed time, in seconds (s);

[0015] S2. Design a convolutional neural network-based autoencoder model (SelfNet) to reconstruct the drone’s time-frequency curve graph;

[0016] The autoencoder model includes an encoder and a decoder, and is combined with a fully connected layer to perform feature fusion; wherein:

[0017] Encoder: It consists of five convolutional blocks, each of which extracts the time-frequency curve features of the drone through convolution operations, and then gradually reduces the size of the multi-dimensional feature map through pooling operations;

[0018] Intermediate feature fusion: First, the multi-dimensional feature map processed by the encoder is flattened to compress the multi-dimensional feature map into a one-dimensional feature vector. Then, the feature fusion is performed through four fully connected layers in sequence, and a dropout layer is used to enhance the generalization ability of the model.

[0019] Decoder: First, the one-dimensional feature vector after the intermediate feature fusion is unflattened to restore the one-dimensional feature vector to a multidimensional feature map. Then, the resolution of the feature map is gradually increased through five transposed convolutional layers, and finally an output image of the drone's time-frequency curve with the same size as the input image is reconstructed.

[0020] S3, based on the drone radar echo model STFT established in step S1 and the drone time-frequency curve output image output in step S2, establish a drone time-frequency image dataset and use the autoencoder model for training;

[0021] The UAV time-frequency image dataset includes a training set and a validation set, which are divided into 6 categories according to the number of rotors N = 1, 2, 3, 4, 5, and 6. Each category includes 1,800 training set images and 200 validation set images; that is, the dataset contains a total of 12,000 images, including 10,800 training set images and 1,200 validation set images;

[0022] During the training process of the training set, set the batch size batch size = 16, the total number of training epochs = 50, and the initial learning rate learning_rate = 10 -3 , define the loss function as mean squared error loss function, define the optimizer as AdamW optimizer, define the learning rate scheduler as step learning rate scheduler, and set step size = 5, decay factor gamma = 0.9, that is, multiply the learning rate by 0.9 after every 5 epochs; on this basis, use early stopping method to prevent model overfitting during training, and output the trained autoencoder model;

[0023] S4. Use the trained autoencoder model weights to reconstruct the drone time-frequency image dataset and complete the drone time-frequency image reconstruction based on deep learning.

[0024] Furthermore, in step S2, each convolution block includes two groups of "convolution layer + batch normalization + Leaky ReLU activation function + Dropout layer" and one group of "maximum pooling layer + Dropout layer", and the two groups of convolution layers are used to extract complex features in the time-frequency curve graph; the Leaky ReLU activation function is used to alleviate the problem of neuron "death"; the batch normalization operation is used to accelerate training and improve feature expression capabilities; the Dropout layer is used to enhance the generalization ability of the model and prevent the model from overfitting; the maximum pooling layer is used to gradually reduce the size of the feature map, reduce the amount of calculation during training, and accelerate the training process.

[0025] Furthermore, in step S3, first, a small sample data set is used to further verify the performance of the trained autoencoder model. The small sample data set includes a training set and a validation set, and is divided into 6 categories according to the number of rotors N = 1, 2, 3, 4, 5, and 6. Each category includes 90 training set images and 10 validation set images; that is, the small sample data set has a total of 600 images, including 540 training set images and 60 validation set images;

[0026] Secondly, directly using the autoencoder model for training and reconstruction on small samples results in poor reconstruction results. Integrating the concept of transfer learning, we leverage the trained model weights, freeze some of the model parameters, and then transfer them to the small sample for further training and reconstruction. Specifically, we freeze the first two convolutional blocks of the encoder, the intermediate feature fusion layer, and the decoder in the trained autoencoder model, ensuring that the parameters of these layers remain unchanged during training. Unfreezing the last three convolutional blocks of the encoder allows the parameters of these layers to be updated during training, allowing transfer training of the trained autoencoder model. Transfer training results show improved reconstructed image quality.

[0027] The beneficial effects of the present invention are:

[0028] 1. This paper addresses the problem of distortion of the time-frequency characteristics of drone echo signals caused by interference in complex environments. It proposes a deep learning-based time-frequency image reconstruction method. This method effectively extracts time-frequency curve features through the SelfNet model, reconstructs high-quality time-frequency images, and improves the accuracy of drone time-frequency images under low signal-to-noise ratio conditions.

[0029] 2. The SelfNet model designed in this paper has a reasonable structure. It combines the "encoder-decoder" concept of the U-Net model, extracts features through the encoder, reconstructs the image through the decoder, and adds a fully connected layer in the middle for feature fusion.

[0030] 3. This paper verifies the generalization ability of SelfNet in data-scarce conditions through small-sample experiments and transfer learning. It can achieve good reconstruction through transfer learning on small-sample datasets, providing a new solution for data-limited practical application scenarios.

[0031] 4. This invention lays the foundation for subsequent UAV classification and micro-motion parameter estimation. The reconstructed time-frequency diagram has clear time-frequency characteristics, which helps to improve the accuracy of UAV detection and classification.

[0032] In short, this deep learning-based UAV time-frequency image reconstruction method has certain promotion and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flow chart of the present invention;

[0034] Figure 2 This is the UAV radar echo model and the schematic diagram of the relationship between the rotor blades and the radar;

[0035] Figure 3 Schematic diagram of the UAV micro-Doppler time-frequency characteristic curve when N=2;

[0036] Figure 4Schematic diagram of the time-frequency characteristic curve of UAV micro-Doppler under different signal-to-noise ratios; Figure 4 (a) The signal-to-noise ratio is SNR = -5dB, Figure 4 (b) The signal-to-noise ratio is SNR = -10dB, Figure 4 (c) The signal-to-noise ratio is SNR = -15dB;

[0037] Figure 5 This is the SelfNet model structure diagram;

[0038] Figure 6 Examples of time-frequency curves of UAVs with different numbers of rotors;

[0039] Figure 7 This is an example of the time-frequency curve of UAVs with different numbers of rotors reconstructed by SelfNet;

[0040] Figure 8 This is an example of the time-frequency curve of drones in a small sample;

[0041] Figure 9 This is the reconstruction result image obtained by directly training on a small sample;

[0042] Figure 10 This is the reconstruction result image obtained after SelfNet weight migration;

[0043] Figure 11 This is an example of the original time-frequency curve in the public dataset;

[0044] Figure 12 For Figure 11 Time-frequency diagram after adding noise;

[0045] Figure 13 The reconstructed time-frequency graph obtained by training the weights of SelfNet with a small sample size;

[0046] Figure 14 Reconstructed time-frequency map obtained by transferring the training weights for SelfNet. DETAILED DESCRIPTION

[0047] The present invention will be described in further detail below with reference to the accompanying drawings and examples.

[0048] Example 1

[0049] like Figure 1 The method for reconstructing time-frequency images of a UAV based on deep learning is shown, comprising the following steps:

[0050] S1. Perform short-time Fourier transform on the echo signal to establish the UAV radar echo model STFT:

[0051]

[0052] In formula ①, S sum (t) is the original echo signal; w(t) is the window function of the short-time Fourier transform; τ is the position where the window function moves; ω is the Doppler frequency in Hertz; j is the imaginary unit, and j is specified 2 =-1;

[0053] Original echo signal S sum The expression of (t) is:

[0054]

[0055] In formula ②, L is the blade length, in meters (m); θ k =θ0+2kπ / N represents the initial rotation angle under different blade numbers N, in degrees (°), Φ k (t) is the phase function of the initial rotation angle:

[0056]

[0057] In formula ③, λ is the radar wavelength, in meters; l is the distance from any scattering point on the blade to the center of the blade, in meters; θ0 is the initial phase, in degrees; f r is the blade rotation rate in revolutions per second; t is the elapsed time in seconds.

[0058] In this step S1, if Figure 2 As shown, assuming that the radar and the rotating blades are in the same plane, take a rotor blade as an example: let the XOY coordinate system represent the radar coordinate system, and the origin O is the position of the radar; the xoy coordinate system represents the drone coordinate system, which is the horizontal translation of the radar coordinate system, and point o is the geometric center of the drone blade. Assume that v is the radial velocity of the drone in the radar line of sight, M0(x0,y0) is a scattering point on the blade, which rotates around the center point o with a rotation rate of f r At t = 0, the distance from point o to M0 is l, the distance from the radar to the blade center is D, the distance from the radar to M0 is D0, and the initial phase is θ0. Assume that after time t, the scattering point M0 rotates to point M t (x t ,y t ), at this time the radar reaches M t The distance is D t , the phase is θ t , and θ t =θ0+2πf r Since it is assumed that the radar and the rotating blades are in the same plane, the radar reaches M t Distance D t It can be expressed as:

[0059]

[0060] Generally, the length of the drone's blades is often much smaller than the distance between the radar and the drone. Therefore, formula (1) can be simplified as shown in formula (2):

[0061]

[0062] Φ(t) is the scattering phase function, which is a function of time t. Combining formula (2), the phase function of the UAV blade can be obtained as shown in formula (3):

[0063]

[0064] At this time, point M on the leaf t The echo signal is shown in formula (4):

[0065]

[0066] In formula (4), ρ is the scattering coefficient, f d is the change in Doppler frequency caused by the UAV's own motion during flight. In order to study the micro-Doppler characteristics caused by the UAV's micro-motion, ρ = 1 is taken here, and the influence of the constant term and the UAV's speed is ignored. Then Equation (4) can be further simplified:

[0067]

[0068] Formula (5) can be understood as the echo of a single scattering point on the UAV blade. Then the echo signal of an entire blade is actually the integration of all scattering points on [0, L], where the blade length is L, that is:

[0069]

[0070] From formula (6), it can be seen that an oscillation phenomenon similar to the sinc function will occur in the echo signal of the entire blade.

[0071] Extending formula (6) to the case where the number of blades is N (i.e. formula ② mentioned above), we have:

[0072]

[0073] In formula (7), θ k =θ0+2kπ / N represents the initial rotation angle under different blade numbers N, Φ k (t) is its phase function, as shown in formula (8) (i.e. formula ③ mentioned above):

[0074]

[0075] For the echo signal S sum(t) is subjected to short-time Fourier transform, and the corresponding time-frequency spectrum can be obtained, as shown in formula (9) (i.e. formula ① mentioned above):

[0076]

[0077] In formula (9), w(t) is the window function of short-time Fourier transform, τ is the position where the window function moves, and ω is the Doppler frequency.

[0078] S2, design an autoencoder model based on convolutional neural network (SelfNet model, such as Figure 5 As shown), reconstruct the time-frequency curve of the UAV;

[0079] The autoencoder model includes an encoder and a decoder, and is combined with a fully connected layer to perform feature fusion; wherein:

[0080] Encoder: It consists of five convolution blocks. Each convolution block extracts the time-frequency curve features of the drone through convolution operations, and then gradually reduces the size of the multi-dimensional feature map through pooling operations. Each convolution block contains two sets of "convolution layer + batch normalization + LeakyReLU activation function + Dropout layer" and one set of "maximum pooling layer + Dropout layer". The two sets of convolution layers are used to extract complex features in the time-frequency curve map. The LeakyReLU activation function is used to alleviate the problem of neuron "death". The batch normalization operation is used to accelerate training and improve feature expression capabilities. The Dropout layer is used to enhance the generalization ability of the model and prevent overfitting. The maximum pooling layer is used to gradually reduce the size of the feature map, reduce the amount of calculation during training, and accelerate the training process.

[0081] Intermediate feature fusion: First, the multi-dimensional feature map processed by the encoder is flattened to compress the multi-dimensional feature map into a one-dimensional feature vector. Then, the feature fusion is performed through four fully connected layers in sequence, and a dropout layer is used to enhance the generalization ability of the model.

[0082] Decoder: First, the one-dimensional feature vector after the intermediate feature fusion is unflattened to restore the one-dimensional feature vector to a multidimensional feature map. Then, the resolution of the feature map is gradually amplified through five transposed convolutional layers, and finally an output image of the drone time-frequency curve with the same size as the input image is reconstructed.

[0083] In this embodiment 1, a time-frequency image dataset with a signal-to-noise ratio range of [-5, 0] dB is generated, and the number of rotors N = 1, 2, 3, 4, 5, 6. 2000 time-frequency images are generated for each number of rotors, 90% of which are used as training sets and 10% as validation sets. Figure 3 The time-frequency curve diagram of the UAV when the number of rotors N=2 is shown. Figure 4 yes Figure 3 Schematic diagram of the time-frequency curve of the UAV under different signal-to-noise ratios. Figure 3 and Figure 4 It can be seen that as the signal-to-noise ratio decreases, the effective information of the time-frequency diagram is gradually severely obscured, making it difficult to perform effective analysis.

[0084] S3, based on the drone radar echo model STFT established in step S1 and the drone time-frequency curve output image output in step S2, establish a drone time-frequency image dataset and use the autoencoder model for training;

[0085] The UAV time-frequency image dataset includes a training set and a validation set, which are divided into 6 categories according to the number of rotors N = 1, 2, 3, 4, 5, and 6. Each category includes 1800 training set images and 200 validation set images.

[0086] The software platform used for training is Pycharm Community Edition 2022.3.1, and the operating system is Windows 11. During the training process of the training set, the batch size is set to 16, the total number of training epochs is set to 50, and the initial learning rate is set to 10. -3 , define the loss function as the mean square error loss function, define the optimizer as the AdamW optimizer, define the learning rate scheduler as the step learning rate scheduler, and set the step size to 5, the decay factor gamma to 0.9; on this basis, use the early stopping method to prevent the model from overfitting during training, and output the trained autoencoder model; after each training cycle (epoch), compare the current validation loss with the best validation loss recorded previously. If the current validation loss is lower, it means that the performance of the model on the validation set has improved. At this time, the best validation loss is updated and the patience counter is reset to "0"; conversely, if the validation loss has not improved, the patience counter will increase by "1". When the patience counter reaches the preset patience value (patience = 5), it is considered that the model can no longer be optimized further, triggering the early stopping mechanism and stopping the training process;

[0087] S4. Use the trained autoencoder model weights to reconstruct the drone time-frequency image dataset and complete the drone time-frequency image reconstruction based on deep learning.

[0088] Figure 6 The following diagram shows the time-frequency curves of drones with different numbers of rotors. Figure 7 The time-frequency curve after reconstruction by the SelfNet model is shown. Figure 6 and Figure 7It can be seen that the SelfNet model can effectively extract the time-frequency curve features and suppress noise, thereby improving the resolution of time-frequency images.

[0089] Example 2

[0090] Based on Example 1, this Example 2 further verifies the performance of SelfNet on a small sample dataset, including the following steps:

[0091] First, a small sample dataset is used to further verify the performance of the trained autoencoder model. The small sample dataset consists of two parts: a training set and a validation set. The dataset is divided into six categories according to the number of rotors N = 1, 2, 3, 4, 5, and 6. Each category includes 90 training set images and 10 validation set images.

[0092] Secondly, the first two convolutional blocks of the encoder, the intermediate feature fusion part and the decoder in the trained autoencoder model are frozen so that the parameters of these layers remain unchanged during the training process; the last three convolutional blocks of the encoder are unfrozen, allowing the parameters of these layers to be updated during the training process, and the trained autoencoder model is transferred.

[0093] In Example 2, a small sample dataset with a signal-to-noise ratio range of [-10, -5] dB was generated, with the number of rotors N = 1, 2, 3, 4, 5, and 6. For each rotor number, 100 time-frequency graphs were generated, with 90% used as the training set and 10% used as the validation set. Directly training and reconstructing the small sample using the SelfNet model resulted in poor reconstructed image quality. However, by transferring the weights trained in Example 1 using the SelfNet model to the small sample and retraining it, the time-frequency curve features were successfully extracted, achieving good reconstruction.

[0094] Figure 8 Shows an example of a time-frequency curve in a small sample data set. Figure 9 The results of training and reconstructing small samples directly using the SelfNet model are shown. Figure 10 The results of SelfNet weight transfer, retraining and reconstruction are shown. Figures 8-10 It can be seen that the effect of directly training and reconstructing small samples is very poor. The features of the original time-frequency curve are not extracted at all, and the effective feature information is almost completely lost. However, by using the model weights trained in Example 1 and migrating them to small samples for further training and reconstruction, the SelfNet model successfully and effectively extracts the feature information of the time-frequency curve, and suppresses the noise of the original situation, achieving good reconstruction.

[0095] Example 3

[0096] Building on Example 1, Example 3 validates the performance of the SelfNet model using a publicly available dataset of field-measured radar detection drones. The raw echo data is preprocessed to generate 50 time-frequency plots, which are then expanded to 500 time-frequency plots after adding noise (with a signal-to-noise ratio range of [-5, 0] dB). 90% of these plots serve as the training set, and 10% as the validation set. The training results demonstrate that directly using weights trained on a small sample yields poor reconstruction results. However, using weights trained after transfer training successfully extracts time-frequency curve features that approximate the original curves, achieving excellent reconstruction.

[0097] Figure 11 It shows a typical time-frequency diagram obtained by preprocessing the raw echo data of the public dataset. Figure 12 Shows the Figure 11 The time-frequency diagram obtained after adding noise, Figure 13 The results of directly reconstructing the weights obtained by small sample training using the SelfNet model in Example 2 are shown. Figure 14 The results of reconstruction using the weights obtained after migration training of the SelfNet model in Example 2 are shown. Figure 13 and Figure 14 , we can see that the effect of directly using the weights trained with small samples for reconstruction is very poor, and the characteristic information of the time-frequency curve is destroyed; after reconstruction using the weights trained with transfer, the characteristic information of the time-frequency curve is successfully extracted, which is closer to Figure 11 The original time-frequency curves shown achieve good reconstruction.

[0098] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A deep learning-based UAV time-frequency image reconstruction method, characterized in that: The following steps are involved: S1. Perform short-time Fourier transform on the echo signal to establish the UAV radar echo model STFT: In formula ①, S sum (t) is the original echo signal; w(t) is the window function of short-time Fourier transform; τ is the position where the window function moves; ω is the Doppler frequency in Hertz; j is an imaginary unit, and j is specified 2 =-1; Original echo signal S sum The expression of (t) is: In formula ②, L is the blade length in meters; θ k =θ0+2kπ / N represents the initial rotation angle under different blade numbers N, in degrees, Φ k (t) is the phase function of the initial rotation angle: In formula ③, λ is the radar wavelength, in meters; l is the distance from any scattering point on the blade to the center of the blade, in meters; θ0 is the initial phase, in degrees; f r is the blade rotation rate in revolutions per second; t represents the elapsed time in seconds; S2. Design an autoencoder model based on convolutional neural network to reconstruct the time-frequency curve of the drone; The autoencoder model includes an encoder and a decoder, and is combined with a fully connected layer to perform feature fusion; wherein: Encoder: It consists of five convolutional blocks, each of which extracts the time-frequency curve features of the drone through convolution operations, and then gradually reduces the size of the multi-dimensional feature map through pooling operations; Intermediate feature fusion: First, the multi-dimensional feature map processed by the encoder is flattened to compress the multi-dimensional feature map into a one-dimensional feature vector. Then, the feature fusion is performed through four fully connected layers in sequence, and a dropout layer is used to enhance the generalization ability of the model. Decoder: First, the one-dimensional feature vector after the intermediate feature fusion is unflattened to restore the one-dimensional feature vector to a multidimensional feature map. Then, the resolution of the feature map is gradually increased through five transposed convolutional layers, and finally an output image of the drone's time-frequency curve with the same size as the input image is reconstructed. S3, based on the drone radar echo model STFT established in step S1 and the drone time-frequency curve output image output in step S2, establish a drone time-frequency image dataset and use the autoencoder model for training; The UAV time-frequency image dataset includes a training set and a validation set, which are divided into 6 categories according to the number of rotors N = 1, 2, 3, 4, 5, and 6. Each category includes 1800 training set images and 200 validation set images. During the training process of the training set, set the batch size batch size = 16, the total number of training epochs = 50, and the initial learning rate learning_rate = 10 -3 , define the loss function as mean squared error loss function, define the optimizer as AdamW optimizer, define the learning rate scheduler as step learning rate scheduler, and set step size = 5, decay factor gamma = 0.9; on this basis, use early stopping method in the training process to prevent model overfitting, and output the trained autoencoder model; S4. Use the trained autoencoder model weights to reconstruct the drone time-frequency image dataset and complete the drone time-frequency image reconstruction based on deep learning.

2. The deep learning-based UAV time-frequency image reconstruction method according to claim 1, characterized in that: In step S2, each convolution block includes two groups of "convolution layer + batch normalization + Leaky ReLU activation function + Dropout layer" and one group of "maximum pooling layer + Dropout layer". The two groups of convolution layers are used to extract complex features in the time-frequency curve graph; the Leaky ReLU activation function is used to alleviate the problem of neuron "death"; the batch normalization operation is used to accelerate training and improve feature expression capabilities; the Dropout layer is used to enhance the generalization ability of the model and prevent the model from overfitting; the maximum pooling layer is used to gradually reduce the size of the feature map, reduce the amount of calculation during training, and accelerate the training process.

3. The deep learning-based UAV time-frequency image reconstruction method according to claim 1, characterized in that: In step S3, first, a small sample data set is used to further verify the performance of the trained autoencoder model. The small sample data set includes a training set and a validation set. The data set is divided into 6 categories according to the number of rotors N = 1, 2, 3, 4, 5, and 6. Each category includes 90 training set images and 10 validation set images. Secondly, the first two convolutional blocks of the encoder, the intermediate feature fusion part and the decoder in the trained autoencoder model are frozen so that the parameters of these layers remain unchanged during the training process; the last three convolutional blocks of the encoder are unfrozen, allowing the parameters of these layers to be updated during the training process, and the trained autoencoder model is transferred.