Method and system for detecting radar targets of weak RCS unmanned aerial vehicles based on frequency-modulated continuous wave radar

By aligning and calibrating RTK information between the radar and the UAV, and combining two-dimensional Fourier transform and adaptive compression, a spatiotemporal feature decoupling enhancement network is constructed. This solves the problem of weak RCS UAV detection by FMCW radar in complex environments, and achieves efficient and accurate UAV target detection.

CN122110083APending Publication Date: 2026-05-29SOUTH CHINA UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2026-03-13
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing FMCW radars struggle to effectively detect UAVs with weak RCS in complex environments, exhibiting problems such as weak echo energy, strong clutter, severe multipath ambiguity, difficulty in labeling, large temporal redundancy, and insufficient real-time performance.

Method used

Time alignment and spatial calibration are performed using RTK information from both radar and UAV ends to generate high-precision annotation information. Two-dimensional fast Fourier transform preprocessing is then performed to adaptively compress redundant time-series data in a non-uniform manner. Spatiotemporal feature decoupling modeling and enhancement are constructed, and an encoder-decoder noise reduction and reconstruction network is used for detection.

Benefits of technology

It reduces labeling costs, improves label consistency, reduces redundancy and latency, suppresses multipath interference, and enhances detection accuracy and robustness, making it suitable for urban low-altitude and complex terrain scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122110083A_ABST
    Figure CN122110083A_ABST
Patent Text Reader

Abstract

The application provides a weak radar cross section unmanned aerial vehicle radar target detection method and system based on frequency modulation continuous wave radar, belongs to the field of radar signal processing and intelligent sensing, and aims to solve the detection problem of low-altitude weak reflection targets caused by weak echo energy, serious multipath effect and clutter interference. The method synchronously acquires and spatiotemporally calibrates the real-time dynamic differential positioning information of the unmanned aerial vehicle and the radar, automatically generates high-precision labels for supervised learning, and reduces the artificial labeling cost and error. After two-dimensional fast Fourier transform is performed on the radar echo, an adaptive non-uniform compression strategy is adopted to aggregate key timing features to reduce data redundancy. Further, a spatiotemporal feature decoupling enhancement module is used to suppress multipath ambiguity and improve feature separability. Finally, a target detection network based on an encoder-decoder noise reduction reconstruction is constructed to denoise and finely reconstruct the features, so that the weak reflection unmanned aerial vehicle target can be accurately and robustly detected and positioned in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar signal processing and intelligent sensing technology, and in particular relates to a method and system for detecting weak RCS UAV radar targets based on frequency modulated continuous wave radar. Background Technology

[0002] The rapid development of the low-altitude economy has spurred emerging application scenarios such as urban air traffic, logistics delivery, and unmanned inspection. With the rapid development of drone technology and 6G communication technology, the application of drones in urban environments is becoming increasingly widespread. However, in areas such as around airports, important facilities, and urban low-altitude zones, illegal intrusion and interference by drones, and the resulting security risks, are becoming increasingly prominent. Therefore, achieving real-time and reliable detection and identification of drone targets in complex environments has become a key technological requirement for low-altitude safety monitoring. Existing drone detection methods mainly include photoelectric / infrared, acoustic, radio spectrum detection, and radar detection. Among these, photoelectric and infrared methods are easily affected by lighting conditions, obstruction factors, and weather changes; acoustic methods have limited detection range and are easily affected by environmental noise interference; radio spectrum detection relies on the drone's communication link and is less suitable for "silent" flight or modified drones. In comparison, radar detection has advantages such as all-day, all-weather, and long-range operation; in particular, frequency-modulated continuous wave (FMCW) radar has high application value in low-altitude small target detection scenarios due to its small size, relatively controllable cost, and strong range and velocity measurement capabilities.

[0003] Despite the advantages of FMCW radar, UAVs typically exhibit characteristics such as small radar cross-section (RCS), weak echo energy, and high maneuverability. In environments with urban buildings, woodlands, and complex terrain, they are also susceptible to interference from ground clutter, multipath effects, and other moving targets. This leads to weak separability, ambiguity, or even submersion of UAV echoes in the Doppler representation domain, resulting in decreased detection performance. To improve detection accuracy, some solutions tend to introduce longer time-series data for accumulation and analysis, but this often introduces information redundancy, increases data processing latency and computational complexity, and is not conducive to real-time deployment. Furthermore, deep learning-based radar target detection methods have received widespread attention in recent years, but their performance usually depends on high-quality labeled datasets. Because radar data is difficult to visualize intuitively, and target echoes are weak and susceptible to clutter, existing labeling methods often face problems such as insufficient labeling accuracy, high labeling costs, and difficulties in spatiotemporal alignment. At the same time, in complex scenarios, the spatiotemporal features of radar are prone to coupling, and without targeted feature modeling and enhancement mechanisms, the robustness of the model to multipath ambiguity and noise disturbances remains limited. Therefore, researching an FMCW radar UAV detection method and system that is effective and robust for detecting weak RCS UAV targets in complex environments is not only of great significance for urban airspace management and security monitoring, but also one of the key technologies that urgently need to be broken through in order to promote the rapid development of the low-altitude economy.

[0004] Therefore, there is an urgent need in this field for a technical solution that can solve the above problems.

[0005] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for detecting weak RCS UAV radar targets based on frequency modulated continuous wave radar.

[0007] To achieve the above objectives, the present invention provides the following solution:

[0008] A method for detecting UAV radar targets with weak radar cross section based on frequency modulated continuous wave radar includes: The system acquires real-time dynamic differential positioning information of the target UAV and real-time dynamic differential positioning information of the frequency modulated continuous wave radar, and performs time alignment and spatial calibration on the two to obtain the relative distance and relative azimuth angle of the UAV relative to the radar. Based on the relative distance and relative azimuth, annotation information for supervised learning is generated; The target echo signal received by the frequency modulated continuous wave radar is preprocessed with two-dimensional fast Fourier transform to construct range-azimuth spectral features. The range-azimuth spectral features are adaptively non-uniformly compressed, and redundant time series data are adaptively aggregated and compressed to obtain compressed time series features. The compressed temporal features are decoupled and enhanced in the spatiotemporal domain to suppress the ambiguity of echo representation caused by multipath effects; The UAV radar target detection network is based on encoder-decoder noise reduction and reconstruction of the input features after decoupling and enhancing features in the spatiotemporal domain. The encoder-decoder structure is used for noise reduction and reconstruction, and the detection and localization results of UAV targets in the radar spectral domain are output. The annotation information is used to train the UAV radar target detection network.

[0009] When training the UAV radar target detection network, auxiliary prediction heads are set in multiple intermediate layers, and multi-scale supervision is performed on each auxiliary prediction head; The supervision loss for each prediction branch adopts the binary cross-entropy loss, and the losses of each prediction branch are weighted and summed according to preset weights to obtain the total training loss.

[0010] Optionally, the time alignment includes: By unifying the real-time dynamic differential data from the UAV and the radar to a preset sampling frequency based on radar acquisition, frame synchronization between the two is achieved by downsampling the real-time dynamic differential data from the side with the higher update rate.

[0011] Optionally, the spatial calibration includes: The distance between the radar and the UAV is calculated based on the fixed geographic coordinates of the radar and the dynamic geographic coordinates of the UAV. Based on the great circle flight path azimuth calculation method, the azimuth angle of the UAV relative to due north is calculated; Based on the pre-calibrated angle between the radar zero-degree azimuth and true north, the azimuth angle relative to true north is converted into a relative azimuth angle in the radar coordinate system.

[0012] Optionally, the annotation information is a confidence map with the same size as the input feature map of the UAV radar target detection network; the generation of the confidence map includes: The actual coordinates of the UAV are mapped to the target location point of the range-azimuth spectrum, and a two-dimensional Gaussian distribution is constructed with the target location point as the mean. The pixel values ​​of the confidence map are then assigned values.

[0013] Optionally, the adaptive non-uniform compression of the range-azimuth spectral features includes: The range-azimuth spectrum is divided into multiple range-azimuth spectrum intervals; Principal component analysis is performed on the chirped sequences within each range-azimuth spectral interval. The micro-temporal dynamic information within the corresponding interval is adaptively aggregated into a preset number of principal components, and the compressed temporal features are output. The preset number of principal components is adaptively determined based on the accuracy and time delay constraints of the target detection task, so as to prioritize the retention of key dynamic features related to UAV identification while reducing the scale of input data and computational complexity.

[0014] Optionally, the spatiotemporal feature decoupling modeling and enhancement of the compressed temporal features is performed on a frame-by-frame basis, including spatial and temporal branches. The processing of the spatial branch and the temporal branch are as follows: Spatial branching: The complex distance-azimuth spectrum of a single frame is decomposed into real and imaginary parts and stacked along the channel dimension. Local spatial texture features are extracted through 1×3×3 three-dimensional convolution. The extracted features are group normalized and GELU nonlinear activation is performed, and average pooling is performed along the chirp dimension to suppress anomalous noise. Temporal branch: The amplitude of the complex distance-azimuth spectrum of a single frame is used as input, and the dynamic distribution features within the frame are modeled by 3×3 two-dimensional convolution; and the dynamic distribution features are subjected to group normalization and GELU nonlinear activation. The output features of the spatial branch and the temporal branch are concatenated along the channel dimension, and channel fusion is performed through 1×1 two-dimensional convolution, group normalization and GELU nonlinear activation to obtain the output features of the frame.

[0015] To achieve parallel acceleration, the batch dimension and frame dimension of the input data are merged, and a one-time forward computation is performed on the merged dimension to reduce the computational overhead caused by frame-by-frame processing.

[0016] Optionally, the UAV radar target detection network based on encoder-decoder noise reduction and reconstruction adopts a U-shaped encoder-decoder structure, which includes an encoding path and a decoding path: The encoding path extracts deep semantic features through step-by-step downsampling; The decoding path restores spatiotemporal resolution through step-by-step upsampling; The encoding path and the decoding path are connected by skip connections to achieve multi-scale feature fusion, so as to perform noise reduction modeling and reconstruction of input features.

[0017] The skip connection includes a skip fusion module, which consists of 3×3×3 three-dimensional convolution, group normalization, and GELU nonlinear activation, and is used to fuse the detailed features of the encoder side and the semantic features of the decoder.

[0018] Optionally, training the UAV radar target detection network includes: Auxiliary prediction heads are set in multiple intermediate layers of the UAV radar target detection network, and multi-scale supervision is performed on each auxiliary prediction head; The supervision loss for each prediction branch adopts the binary cross-entropy loss, and the losses of each prediction branch are weighted and summed according to preset weights to obtain the total training loss.

[0019] A weak radar cross section UAV radar target detection system based on frequency modulated continuous wave radar includes: The radar-UAV spatiotemporal alignment and calibration subsystem is used to acquire the real-time dynamic differential positioning information of the target UAV and the real-time dynamic differential positioning information of the frequency modulated continuous wave radar, and to perform time alignment and spatial calibration on the two to obtain the relative distance and relative azimuth angle of the UAV relative to the radar. The label generation module is used to generate annotation information for supervised learning based on the relative distance and relative azimuth angle output by the radar-UAV spatiotemporal alignment and calibration subsystem. A two-dimensional fast Fourier transform data preprocessing module is used to perform two-dimensional fast Fourier transform preprocessing on the target echo signal received by the frequency modulated continuous wave radar to construct range-azimuth spectral features. An adaptive non-uniform compression module is used to input the range-azimuth spectrum features into the adaptive non-uniform compression module, and to perform adaptive information aggregation and compression on redundant time series data to obtain compressed time series features. The spatiotemporal feature decoupling enhancement module is used to input the compressed temporal features into the spatiotemporal feature decoupling enhancement module to decouple and enhance the spatiotemporal domain features, so as to suppress the echo characterization ambiguity caused by multipath effect; The UAV radar target detection network based on encoder-decoder denoising and reconstruction is used to input the features processed by the spatiotemporal feature decoupling enhancement module into the UAV radar target detection network, use the encoder-decoder structure to denoise and reconstruct the features, and output the detection and localization results of UAV targets in the radar spectral domain. The annotation information is used to train the UAV radar target detection network.

[0020] Optionally, the backbone network of the UAV radar target detection network based on encoder-decoder noise reduction and reconstruction is constructed based on the MetaFormer paradigm and includes multi-level ConvFormer blocks: The ConvFormer block uses 3×7×7 three-dimensional separable convolutions as feature mixing units to jointly model the dependencies between the time dimension, distance dimension, and angle dimension.

[0021] The three-dimensional separable convolution employs an asymmetric convolution kernel, setting a smaller receptive field in the time dimension to capture rapidly changing dynamic features, and a larger receptive field in the distance-azimuth spectral domain to cover the target shape and its spatial extension range.

[0022] Compared with the prior art, the present invention has the following beneficial effects: This invention proposes a weak RCS UAV radar target detection method and system based on frequency-modulated continuous wave radar. Addressing the key technical bottlenecks of weak-reflection targets in complex environments, such as "weak echo energy, strong clutter, severe multipath ambiguity, difficult labeling, large temporal redundancy, and insufficient real-time performance," this invention constructs an end-to-end technical route of "high-precision automatic labeling - weak target adaptive compression - spatiotemporal decoupling enhancement - noise reduction and reconstruction detection." The beneficial effects of this invention include at least: 1) Reduce annotation costs and improve label consistency: Utilize RTK information from radar and UAV terminals to automatically calculate the relative distance and azimuth of targets and generate confidence labels consistent with the RA spectrum, significantly reducing the workload of manual annotation and reducing errors caused by manual alignment; 2) Reduce redundancy and latency while preserving the characteristics of weak targets: By using an adaptive compression strategy, redundant time-series information is compressed while preserving the key dynamic characteristics of weak targets, thereby reducing data size and computational load, reducing inference latency, and improving real-time detection capabilities. 3) Suppressing representation ambiguity caused by multipath and spurious peaks: A spatiotemporal decoupling enhancement mechanism is adopted to weaken the interference of multipath superposition and spurious peaks on feature expression, thereby alleviating the problems of target representation ambiguity and false detection in complex scenarios; 4) Improve training stability and generalization robustness: Introduce an encoder-decoder denoising and reconstruction network and combine it with a multi-scale supervision strategy to enhance the recovery and discrimination of weak target echoes in strong clutter backgrounds, thereby improving training convergence stability, robustness and cross-scene generalization performance.

[0023] Therefore, this invention can achieve more efficient and accurate weak RCS UAV detection in various scenarios such as urban low-altitude areas, airport perimeters, and complex terrains, possessing good engineering feasibility and promotional value. In summary, this invention achieves synergistic optimization of "accuracy improvement, redundancy reduction, and robustness enhancement" in weak RCS UAV detection at the technical level; at the application level, it can reduce data annotation and computational resource costs, improve the real-time response capability and reliability of low-altitude safety monitoring, and has significant engineering application value and promotional significance. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 The overall system flowchart provided for embodiments of the present invention.

[0026] Figure 2 This is a time synchronization diagram provided for an embodiment of the present invention.

[0027] Figure 3 This is a schematic diagram of radar-UAV position spatial calibration provided in an embodiment of the present invention.

[0028] Figure 4 This is a schematic diagram of the label generation process provided in an embodiment of the present invention.

[0029] Figure 5 This is a flowchart of radar echo data processing provided in an embodiment of the present invention.

[0030] Figure 6 The flowchart illustrates the single-frame data preprocessing of the spatiotemporal feature decoupling enhancement module provided in this embodiment of the invention.

[0031] Figure 7 This is a schematic diagram of a UAV radar target detection network based on encoder-decoder noise reduction and reconstruction, provided in an embodiment of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] The purpose of this invention is to provide a method and system that can improve the real-time response capability and reliability of low-altitude safety monitoring.

[0034] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Example 1: This embodiment provides a method and system for detecting UAV targets with weak radar cross sections based on Frequency-Modulated Continuous Wave (FMCW) radar. It aims to address the problem of effectively detecting UAVs and other weakly reflective targets with a certain degree of concealment and small RCS in complex environments. Existing UAV detection schemes typically rely on the analysis of large amounts of time-series data to improve detection accuracy, but this leads to information redundancy, increasing data processing latency and computational complexity. To reduce data redundancy and improve detection efficiency and accuracy in complex scenarios, the method includes: acquiring the UAV's real-time kinematic (RTK) positioning information and performing time alignment and spatial calibration with the FMCW radar's RTK positioning information to generate annotation information for supervised learning; performing 2D FFT preprocessing on the target echo signal received by the FMCW radar; and then inputting the preprocessed data into an adaptive non-uniform compression module to extract key information from the redundant time-series data, ensuring the validity of the input data and reducing data redundancy. To further improve detection performance, a spatiotemporal feature decoupling enhancement module is designed. This module models and enhances the spatiotemporal features of the data processed by the adaptive non-uniform compression module to alleviate the ambiguity of UAV echo signals caused by multipath effects, thereby improving the clarity and detection accuracy of UAV echo representations in different scenarios. Finally, a UAV radar target detection network based on denoising modeling is constructed. This network uses an encoder-decoder structure to denoise and reconstruct the features output by the spatiotemporal feature decoupling enhancement module, thereby improving the accuracy and robustness of UAV detection. The overall system flowchart is shown below. Figure 1 As shown.

[0036] Step 1: Label creation for deep learning model training: (1) Spatiotemporal alignment and calibration subsystem of radar RTK information and UAV RTK information: This embodiment generates supervisory labels for deep learning model training by real-time recording of RTK information corresponding to the FMCW radar pointing direction and the UAV's RTK positioning information. Given that the output update rate of the RTK information acquisition device may fluctuate, and that RTK data is typically recorded with true north as the reference coordinate system, it is necessary to synchronize and align the RTK information from the radar and UAV before generating the labels, and to complete coordinate system transformation and spatial calibration. Time synchronization in this project is achieved by unifying the sensor data from both ends to the radar's reference frequency. In this implementation, the reference frequency is set to 10 Hz, and downsampling is used to complete frame synchronization. The specific process is as follows... Figure 2 As shown.

[0037] In spatial calibration, this invention uses an automated process to label radar values. The core process utilizes RTK technology to accurately acquire the global geographic coordinates of the radar sensor and the target UAV. Fixed geographic coordinates of the radar are acquired separately using RTK equipment. Dynamic geographic coordinates of the target at every moment Based on these coordinates, the distance between the two is calculated using the Haversine formula, such as... Figure 3 As shown in (a), this formula is applicable to accurate calculations over short distances, and its formula is as follows: (1) Where R represents the Earth's radius, taken as 6,371,009 meters. and These are the differences in latitude and longitude, respectively. Then, to determine the target's azimuth relative to the radar... In this embodiment, the azimuth angle of the target UAV relative to true north is calculated using the great circle flight path azimuth formula. w 2. The formula is as follows: (2) However, this angle is based on a geographic coordinate system and cannot be directly used as the radar's relative azimuth. Therefore, before inputting it into the tag generation module, the angle between the radar's zero-degree azimuth and true north must be pre-calibrated. Finally, through difference operations... This allows us to obtain the precise relative azimuth angle of the target UAV in the radar coordinate system, such as... Figure 3 As shown in (b). Ultimately, the calculated distance and relative azimuth will be used as the target UAV position in the radar frame at that moment, and will be used to generate the training labels required for the subsequent deep learning model training. Figure 3 (c) shows an example of the punctuation results, where the red pentagram precisely indicates the target UAV position calculated by the method.

[0038] (2) Tag generation module The tag generation module receives the actual UAV coordinates obtained from the aforementioned spatiotemporal alignment and spatial calibration. This process generates supervisory labels for training deep learning models. The supervisory labels are generated as confidence maps with the same size as the network input feature map. Specifically, the label generation module maps the UAV's real coordinates to the target location points on the range-angle (RA) spectrum of the radar's two-dimensional representation domain, and assigns Gaussian distribution values ​​to the pixel values ​​of the confidence map: where the mean of the Gaussian distribution corresponds to the target object's position, and the variance is associated with the target category and scale information. The calculation formula is shown below: (3) in,α Used to control the scale of the Gaussian ellipse in the distance dimension, with a value of 5; b Used to control the scale of the Gaussian ellipse in the angular dimension, with a value of 10; and The confidence level distribution's expansion range and decay rate are controlled by a value of 10. The generated confidence level plot is shown below. Figure 4 As shown.

[0039] Step 2: Radar echo data preprocessing and weak target adaptive non-uniform compression: (1) FMCW radar data acquisition subsystem: The FMCW radar data acquisition subsystem of this invention adopts a time-division multiplexing multiple-input multiple-output (TDM-MIMO) strategy. By performing time-sequential excitation and joint processing on multiple transmit and receive channels, it can acquire data containing... The equivalent mapping of a MIMO array with one transmit antenna and one receive antenna is as follows: A two-dimensional virtual array composed of individual array elements. This enables higher-dimensional spatial sampling capabilities and improves angular resolution. During the operation of the FMCW radar, the transmitting end transmits linear frequency modulated continuous wave signals. When radiation reaches the target detection area, and a target UAV is present, its echo can be considered as a backscattering response to the transmitted signal. This backscattered echo, under reasonable approximation conditions, can be expressed as: (4) Where A is the amplitude. Indicates round-trip time. Indicates a fast time. Indicates the pulse duration. ( () represents the number of pulses. For the number of chirps emitted, This is the initial tilt distance. The radial velocity of the target, At the speed of light, For wavelength, For carrier frequency, and These are the initial azimuth and elevation angles, respectively. and These are the column and row numbers of the two-dimensional virtual array, respectively. The rectangular function is shown in the following formula: (5) (2) 2D FFT data preprocessing and PCA-based adaptive non-uniform compression module: This module uses a single antenna in the array as a reference element to perform frequency domain transformation on the echo data described in equation (4) in the range direction (fast time dimension) and the angular direction (virtual array dimension), obtaining a frequency domain representation that facilitates subsequent weak target enhancement and compression characterization. Theoretically, by applying FFT to the four dimensions of range, velocity (Doppler dimension), azimuth angle, and elevation angle respectively, equation (4) can be expressed as a four-dimensional spectrum, as shown in equation (6): (6) in, , , and These represent the frequency variables of distance, Doppler, azimuth, and elevation, respectively. Represents a rectangular window function. These are coefficient terms related to system parameters.

[0040] in, , , and These represent the frequency variables of distance, Doppler, azimuth, and elevation, respectively. Represents a rectangular window function. These are coefficient terms related to system parameters.

[0041] Given that target detection in this embodiment only needs to cover a preset field of view, and that pitch dimension information contributes relatively little to target discrimination (or is limited by the effective aperture in the pitch direction, resulting in insufficient resolution), explicit frequency domain expansion and angle estimation are not performed on the pitch angle dimension. Specifically, the pitch dimension can be marginalized, retaining only the range and azimuth dimensions to complete the two-dimensional frequency domain transformation and feature construction, i.e., the RA spectrum. Meanwhile, this invention addresses the chirp dimension... An adaptive non-uniform compression strategy based on PCA is designed: Each RA spectral interval is used as the basic processing unit, and PCA is independently performed on the chirped sequences within that interval. The micro-temporal feature energies are adaptively aggregated and compressed, ultimately condensed into... Principal component representation. Interval-level independent dimensionality reduction significantly reduces the input data size and computational complexity. Furthermore, it prioritizes the preservation of key dynamic features relevant to target recognition during compression, thereby improving the robustness and accuracy of subsequent detection. Detailed processing flow is as follows: Figure 5 As shown.

[0042] Step 3: UAV spatiotemporal enhancement and noise reduction detection network based on radar echo: (1) Spatiotemporal feature decoupling enhancement module: The RA spectrum obtained after preprocessing in step 2 is susceptible to multipath reflection superposition and spatiotemporal phase mismatch and coupling effects caused by target motion and radar system in complex scenarios, leading to phenomena such as energy diffusion, false peaks, and target characterization blurring, thereby reducing the performance of weak target detection. To address this, this invention proposes a spatiotemporal feature decoupling and enhancement module for jointly modeling, decoupling, and enhancing the slow-time / cross-frame dynamic information corresponding to the RA spectrum machine, thereby mitigating the blurring effect caused by multipath interference and improving target feature clarity. The module structure is as follows: Figure 6 As shown.

[0043] The spatiotemporal feature decoupling module receives the dimensionality-reduced complex form RA spectrum sequence data and processes the input independently frame by frame. The dimension of the single-frame input data is... ,in, For batch size, The number of chirps contained in each frame. and These correspond to the distance dimension and the angle dimension, respectively. In one embodiment, the spatiotemporal feature decoupling module processes the data independently per frame: for any frame's complex RA spectrum... Spatial feature branches and temporal dynamic branches are constructed separately, and feature decoupling, enhancement, and fusion are performed within the frame to obtain the output features of that frame. .

[0044] First, input a single frame of complex numbers. Decompose into real part With the imaginary part And stacked along the channel dimension to form a tensor The number of channels, 2, corresponds to the real and imaginary parts respectively, and is used for texture modeling in subsequent spatial branches. Next, the input spatial branches are processed using a convolution kernel with a size of... The three-dimensional convolution extracts local spatial texture features; then group normalization (GN) and GELU activation are performed sequentially to achieve feature scale normalization and nonlinear mapping; further, along the chirped dimension... Average pooling is performed to enhance the perception of global trends and suppress anomalous noise interference, resulting in spatial branch output features. .

[0045] To preserve the intensity variation information between chirps, complex amplitudes are constructed as inputs for the time branch. ,in The imaginary unit, This indicates amplitude calculation. (For...) Apply Two-dimensional convolution is used to model intra-frame dynamic distribution features, and GN and GELU processing are performed sequentially to obtain temporal branch features consistent with spatial branches in terms of spatial scale. .

[0046] Finally, and It is obtained by splicing along the channel, and then through Two-dimensional convolution is used for channel fusion, and GN and GELU are applied sequentially to obtain the output features of the frame. The above steps are repeated for each frame in the sequence to obtain a high-quality feature representation of the entire sequence frame by frame. Furthermore, to achieve parallel acceleration, since the normalization calculation of GN does not depend on the statistics of the batch dimension, the batch dimension can be... With frame dimension merged into Perform a one-time forward computation to improve processing efficiency and reduce implementation overhead.

[0047] (2) UAV radar target detection network based on encoder-decoder noise reduction and reconstruction: This invention proposes a UAV radar target detection network based on encoder-decoder denoising and reconstruction. The network employs a U-shaped encoder-decoder structure, with the backbone built on the MetaFormer paradigm to achieve refined target detection for discrete units within the radar cube. The core of the backbone consists of multi-level ConvFormer block modules, using three-dimensional separable convolutions as feature mixing units to efficiently model the joint dependencies of radar signals across time, range, and angle dimensions, thereby enhancing the feature representation capability for weak echo targets and improving detection accuracy. To match the physical characteristics of radar spectral domain data, in one embodiment, the depthwise convolution uses asymmetric kernels: a smaller receptive field is set in the time dimension to capture rapidly changing dynamic features; a larger receptive field is set in the RA spectral domain to cover the target shape and its spatial extension range, thus balancing dynamic response and spatial structure modeling capabilities. The network as a whole follows a U-shaped framework: the encoding path extracts deep semantic features through progressive downsampling; the decoding path gradually restores spatiotemporal resolution through progressive upsampling to achieve feature denoising, modeling, and reconstruction. The encoder and decoder transmit multi-scale features via skip connections, and feature fusion is completed through a skip fusion module. This skip fusion module, which can be composed of 3D convolution, GN, and GELU activation functions, effectively integrates high-frequency detail information from the encoder side with deep semantic information from the decoder side, thereby preserving local structural features such as edges during upsampling. The specific structure is as follows: Figure 7 As shown.

[0048] Furthermore, to enhance feature modeling capabilities and training stability, the proposed network adopts the following structural design: GN and GELU are uniformly used in each stage to achieve stable feature normalization and nonlinear representation; the channel mapping within the network is implemented using 3D convolution, thereby constructing a fully convolutional feature transformation unit, enabling the network to support inputs of different sizes while maintaining spatiotemporal feature consistency. Moreover, a GN layer is set after each downsampling and upsampling operation that changes the feature map resolution to stabilize the inter-layer activation distribution, and a 3D convolution is configured after this GN layer to adaptively adjust the number of feature map channels.

[0049] At the input end, the present invention adopts The encoder input structure with large convolutional kernels quickly establishes a large receptive field and combines it with GN to complete the initial feature extraction. At the prediction output, GN is applied to the final feature map. Three-dimensional convolution is used to classify and regress the discrete units of the radar data cube, thereby enabling the detection and localization of UAV targets. Through the integrated network design of "spatiotemporal enhancement feature extraction - noise reduction and reconstruction - refined detection" described above, this invention can improve the clarity of echo representation and detection robustness under complex conditions such as multipath interference, noise disturbance and phase mismatch, and is suitable for UAV target detection scenarios based on radar echoes.

[0050] (3) Design of training objective function: To achieve multi-scale supervision, this invention introduces auxiliary prediction heads (corresponding to) in multiple intermediate layers of the network. Figure 6 Feature layer in and The design applies independent supervisory losses to each prediction branch. This approach provides more direct gradient feedback for feature layers of different depths, thereby improving the discriminative power of deep features, reducing the model's dependence on shallow features for prediction, and alleviating underfitting. It's important to note that each supervisory loss is calculated using Binary Cross-Entropy (BCE). This is because the task essentially involves binary classification on each discrete spectral unit in the RA spectral domain. BCE effectively characterizes the difference between predicted probabilities and true labels, exhibiting good gradient properties and numerical stability, making it particularly suitable for classification scenarios with sparse labels. Furthermore, since the density map prediction output is a continuous value ranging from 0 to 1, BCE can also be used to measure the deviation between predicted probabilities and true density labels. Therefore, based on its stable numerical behavior and good gradient characteristics, this embodiment selects BCE as the training loss function. Specifically, for the ... Intermediate feature layer Corresponding auxiliary output The monitoring loss is defined as follows: (7) in, Indicates the number of discrete units participating in the supervision; Let be the true label of the i-th discrete unit; and be the predicted probability value of the i-th branch in that discrete unit. In one embodiment, the total training loss of the network can be obtained by weighted summation of the supervision losses of each branch, to achieve joint optimization of prediction branches at different scales, i.e.: (8) in, For the first The loss weights corresponding to each prediction branch are used to balance the contributions of different branches to the overall optimization objective.

[0051] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0052] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for detecting UAV radar targets with weak radar cross section based on frequency-modulated continuous wave radar, characterized in that, include: The system acquires real-time dynamic differential positioning information of the target UAV and real-time dynamic differential positioning information of the frequency modulated continuous wave radar, and performs time alignment and spatial calibration on the two to obtain the relative distance and relative azimuth angle of the UAV relative to the radar. Based on the relative distance and relative azimuth, annotation information for supervised learning is generated; The target echo signal received by the frequency modulated continuous wave radar is preprocessed with two-dimensional fast Fourier transform to construct range-azimuth spectral features. The range-azimuth spectral features are adaptively non-uniformly compressed, and redundant time series data are adaptively aggregated and compressed to obtain compressed time series features. The compressed temporal features are decoupled and enhanced in the spatiotemporal domain to suppress the ambiguity of echo representation caused by multipath effects; The UAV radar target detection network is based on encoder-decoder noise reduction and reconstruction of the input features after decoupling and enhancing features in the spatiotemporal domain. The encoder-decoder structure is used for noise reduction and reconstruction, and the detection and localization results of UAV targets in the radar spectral domain are output. The annotation information is used to train the UAV radar target detection network.

2. The method according to claim 1, characterized in that, The time alignment includes: The real-time dynamic differential data from the UAV and the radar are unified to a preset sampling frequency based on radar acquisition. Frame synchronization between the two is achieved by downsampling the real-time dynamic differential data from the side with the higher update rate.

3. The method according to claim 1, characterized in that, The spatial calibration includes: The distance between the radar and the UAV is calculated based on the fixed geographic coordinates of the radar and the dynamic geographic coordinates of the UAV. Based on the great circle flight path azimuth calculation method, the azimuth angle of the UAV relative to due north is calculated; Based on the pre-calibrated angle between the radar zero-degree azimuth and true north, the azimuth angle relative to true north is converted into a relative azimuth angle in the radar coordinate system.

4. The method according to claim 1, characterized in that, The annotation information is a confidence map with the same size as the input feature map of the UAV radar target detection network; the generation of the confidence map includes: The actual coordinates of the UAV are mapped to the target location point of the range-azimuth spectrum, and a two-dimensional Gaussian distribution is constructed with the target location point as the mean. The pixel values ​​of the confidence map are then assigned values.

5. The method according to claim 1, characterized in that, The adaptive non-uniform compression of the range-azimuth spectral features includes: The range-azimuth spectrum is divided into multiple range-azimuth spectrum intervals; Principal component analysis is performed on the chirped sequences in each distance-azimuth spectral interval. The micro-temporal dynamic information in the corresponding interval is adaptively aggregated into a preset number of principal components, and the compressed temporal features are output.

6. The method according to claim 1, characterized in that, The compressed temporal features are subjected to spatiotemporal domain feature decoupling modeling and enhancement, processed on a frame-by-frame basis, and include spatial and temporal branches. The processing of the spatial branch and the temporal branch are as follows: Spatial branching: The complex distance-azimuth spectrum of a single frame is decomposed into real and imaginary parts and stacked along the channel dimension. Local spatial texture features are extracted through 1×3×3 three-dimensional convolution. The extracted features are group normalized and GELU nonlinear activation is performed, and average pooling is performed along the chirp dimension to suppress anomalous noise. Temporal branch: The amplitude of the complex distance-azimuth spectrum of a single frame is used as input, and the dynamic distribution features within the frame are modeled by 3×3 two-dimensional convolution; and the dynamic distribution features are subjected to group normalization and GELU nonlinear activation. The output features of the spatial branch and the temporal branch are concatenated along the channel dimension, and channel fusion is performed through 1×1 two-dimensional convolution, group normalization and GELU nonlinear activation to obtain the output features of the frame.

7. The method according to claim 1, characterized in that, The UAV radar target detection network based on encoder-decoder noise reduction and reconstruction adopts a U-shaped encoder-decoder structure, including an encoding path and a decoding path: The encoding path extracts deep semantic features through step-by-step downsampling; The decoding path restores spatiotemporal resolution through step-by-step upsampling; The encoding path and the decoding path are connected by skip connections to achieve multi-scale feature fusion, so as to perform noise reduction modeling and reconstruction of input features.

8. The method according to claim 1, characterized in that, The training of the UAV radar target detection network includes: Auxiliary prediction heads are set in multiple intermediate layers of the UAV radar target detection network, and multi-scale supervision is performed on each auxiliary prediction head; The supervision loss for each prediction branch adopts the binary cross-entropy loss, and the losses of each prediction branch are weighted and summed according to preset weights to obtain the total training loss.

9. A radar target detection system for unmanned aerial vehicles with weak radar cross section based on frequency-modulated continuous wave radar, characterized in that, include: The radar-UAV spatiotemporal alignment and calibration subsystem is used to acquire the real-time dynamic differential positioning information of the target UAV and the real-time dynamic differential positioning information of the frequency modulated continuous wave radar, and to perform time alignment and spatial calibration on the two to obtain the relative distance and relative azimuth angle of the UAV relative to the radar. The label generation module is used to generate annotation information for supervised learning based on the relative distance and relative azimuth angle output by the radar-UAV spatiotemporal alignment and calibration subsystem. A two-dimensional fast Fourier transform data preprocessing module is used to perform two-dimensional fast Fourier transform preprocessing on the target echo signal received by the frequency modulated continuous wave radar to construct range-azimuth spectral features. Furthermore, the pitch dimension information is marginalized, rather than explicitly expanding the pitch angle dimension in the frequency domain. An adaptive non-uniform compression module is used to input the range-azimuth spectrum features into the adaptive non-uniform compression module, and to perform adaptive information aggregation and compression on redundant time series data to obtain compressed time series features. The spatiotemporal feature decoupling enhancement module is used to input the compressed temporal features into the spatiotemporal feature decoupling enhancement module to decouple and enhance the spatiotemporal domain features, so as to suppress the echo characterization ambiguity caused by multipath effect; The UAV radar target detection network based on encoder-decoder denoising and reconstruction is used to input the features processed by the spatiotemporal feature decoupling enhancement module into the UAV radar target detection network, use the encoder-decoder structure to denoise and reconstruct the features, and output the detection and localization results of UAV targets in the radar spectral domain. The annotation information is used to train the UAV radar target detection network.

10. The system according to claim 9, characterized in that, The backbone network of the UAV radar target detection network based on encoder-decoder noise reduction and reconstruction is constructed based on the MetaFormer paradigm and includes multi-level ConvFormer blocks: The ConvFormer block uses 3×7×7 three-dimensional separable convolutions as feature mixing units to jointly model the dependencies between the time dimension, distance dimension, and angle dimension.