SAR (Synthetic Aperture Radar) small-scale target detection system and method based on deep learning

Through adaptive filtering and multi-scale CFAR preprocessing, content-aware feature reorganization network and multi-stage cascade detection, the accuracy and robustness problems in SAR small-scale object detection are solved, and efficient and real-time small-scale object detection is achieved, which is suitable for power inspection and disaster monitoring in complex backgrounds.

CN120495626APending Publication Date: 2025-08-15QUJING POWER SUPPLY BUREAU YUNNAN POWER GRID CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510563960.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing SAR small-scale object detection method has low detection accuracy, inaccurate positioning and poor robustness in complex backgrounds, insufficient discrimination of shallow feature, unstable bounding box matching, insufficient optimization capabilities of single-stage detectors and serious interference in complex backgrounds.

Method used

Adaptive guidance filtering and improved multi-scale local adaptive CFAR method are used for preprocessing, combined with content-aware feature recombination network and multi-head attention transformer for feature extraction, introduced the modified Papist distance loss function to optimize bounding box matching, and gradually optimize candidate target screening and regression through a multi-stage cascade detection architecture, and improved detection performance with the multi-source data fusion module.

Benefits of technology

It significantly improves the spatial positioning accuracy and signal-to-noise ratio of small-scale targets, reduces the risks of missed detection and missed detection, has good platform generalization capabilities and real-time capabilities, meets the needs of edge computing deployment, reduces system costs, and improves the automation and intelligence level of critical infrastructure monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495626A_ABST
    Figure CN120495626A_ABST
Patent Text Reader

Abstract

The invention discloses an SAR (Synthetic Aperture Radar) small-scale target detection system and method based on deep learning, and aims to solve the problems of low detection precision, inaccurate positioning and poor robustness of a small-scale target under a complex background. The system performs preprocessing through adaptive guided filtering and improved multi-scale local self-adaption to generate a high-quality candidate region; a content perception feature recombination network is combined with a lightweight multi-head attention converter and a dynamic deformable fusion unit, cross-level dynamic aggregation of shallow texture and deep semantics is achieved, and the small target perception ability is enhanced; a modified Bhattacharyya distance loss function is introduced to optimize bounding box matching, and the small target positioning precision is remarkably improved. In addition, the system adopts a multi-stage cascade detection architecture to gradually optimize candidate target screening and regression, and the recall rate is improved through weak target recovery and geometric consistency verification. The multi-source data fusion module combines polarization characteristics and optical image information to enhance the cross-modal detection capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection and recognition, and in particular to a SAR small-scale target detection system and method based on deep learning. Background Art

[0002] Synthetic Aperture Radar (SAR) is a remote sensing technology with all-weather, all-day imaging capabilities, widely used in military reconnaissance, power inspections, disaster monitoring, and other fields. However, in practical applications, the detection of small-scale targets in SAR images faces many challenges, especially in complex backgrounds. The weak response characteristics of small targets are easily overwhelmed by coherent speckle noise and background clutter, resulting in a significant performance degradation of traditional detection algorithms in weak signal and strong interference environments.

[0003] Traditional Constant False Alarm Rate (CFAR) detectors have good detection performance for large-scale targets in uniform backgrounds. However, their reliance on statistical clutter modeling makes them difficult to adapt to the non-uniform distribution of complex natural or artificial backgrounds. Furthermore, CFAR methods rely on local intensity differences for detection and are unable to effectively utilize contextual information or multi-scale semantic features, resulting in a high risk of missed detections in small-scale detection.

[0004] In recent years, with the development of deep learning technology, target detection methods based on convolutional neural networks (CNNs) have gradually become a research hotspot. However, existing deep learning methods still have the following major problems in SAR small-scale target detection:

[0005] 1. Insufficient discriminability of shallow features: Although shallow features have high spatial resolution, they lack high-level semantic information, resulting in limited ability to express small objects. Deep features, while having strong semantic information, have a significantly reduced spatial resolution due to downsampling, making small objects easily overlooked.

[0006] 2. Bounding box matching is sensitive to small objects: Mainstream object detection algorithms typically use Intersection over Union (IoU) or its variants (such as GIoU, DIoU, and CIoU) as metrics for bounding box matching and regression. However, these metrics are extremely sensitive to positional deviations in small object scenarios. Even pixel-level jitter can cause significant fluctuations in the loss function, affecting the stability and convergence speed of model training.

[0007] 3. Insufficient optimization capabilities of single-stage detectors: Existing methods are mostly single-stage or lightweight structures, lacking a staged candidate box screening and optimization mechanism. The coarse-grained preliminary regression is easily guided to the local optimum by background noise, resulting in limited overall detection performance.

[0008] 4. Complex background interference is serious: Coherent speckle noise and background clutter in SAR images will significantly reduce the signal-to-noise ratio (SNR) and contrast ratio (CNR) of small targets, making it difficult to effectively extract and identify target features.

[0009] In summary, existing SAR small-scale target detection methods have obvious deficiencies in shallow feature expression, bounding box matching robustness, model optimization capability and anti-interference ability. There is an urgent need for a new deep learning detection method that can combine multi-level feature enhancement, precise position measurement and step-by-step optimization structure to solve the above technical problems and improve the detection accuracy and stability of small-scale targets in complex backgrounds. Summary of the Invention

[0010] To solve the above problems, the present invention proposes a SAR small-scale target detection system and method based on deep learning, aiming to solve the problems of poor small-scale target detection accuracy, poor positioning stability and difficulty in model optimization in the existing technology.

[0011] In order to achieve the above object, the present invention adopts the following technical solutions:

[0012] A SAR small-scale target detection system based on deep learning includes: a data acquisition and preprocessing module, a deep learning core detection network connected to the data acquisition and preprocessing module for data transmission, a post-processing and weak target repair module connected to the deep learning core detection network for data transmission, and a multi-source data fusion module connected to the post-processing and weak target repair module for data transmission;

[0013] The data acquisition and preprocessing module obtains raw data from the hardware device and performs preliminary processing on the image to provide high-quality input for subsequent modules; the deep learning core detection network extracts multi-level features from the preprocessed image and generates a coarse-screened candidate target list; the post-processing and weak target repair module further optimizes the output of the deep learning core detection network, performs secondary evaluation and recovery on weak targets and targets with blurred boundaries, and improves the detection recall rate; the multi-source data fusion module improves the system's cross-modal generalization capability by fusing multi-source data such as polarization features and optical images.

[0014] Furthermore, the data acquisition process of the data acquisition and preprocessing module is as follows: using millimeter wave synthetic aperture radar SAR to obtain high-resolution SAR images, combining the inertial navigation unit IMU and the global navigation satellite system GNSS to record attitude and position information, and using PPS clock pulses to synchronize the time base of the radar, inertial navigation unit IMU and computing unit;

[0015] The image preprocessing process of the data acquisition and preprocessing module is as follows: using the adaptive guided coherence filtering (AGCF) method to dynamically adjust the filter template response to retain target details while suppressing non-structured noise;

[0016] The filtering equation is as follows:

[0017]

[0018] Where, I out (x, y) is the grayscale value or intensity value of the output image at the pixel position (x, y); Ω(x, y) is the local window area centered at the pixel position (x, y); w ij is the weight coefficient, which indicates the contribution of the pixel (i, j) in the neighborhood to the central pixel (x, y); I(i, j) is the grayscale value or intensity value of the input image at the pixel position (i, j);

[0019] Through the improved multi-scale local adaptive CFAR method, the suspected target area is quickly extracted by combining the gradient direction and response intensity;

[0020] The outputs of the data acquisition and preprocessing module are: preprocessed SAR images, preliminarily screened target proposal boxes, and time synchronization data.

[0021] Furthermore, the improved multi-scale local adaptive CFAR method introduces the image gradient enhancement channel as an edge prior to guide noise modeling and sliding window configuration, which includes the following steps:

[0022] First, the gradient intensity map is constructed using the grayscale gradient direction of the SAR image: G(x,y) = ‖▽I(x,y)‖, which is used to mark the area where the target edge may exist; where G(x,y) is the value of the gradient intensity map at the pixel position (x,y); ▽I(x,y) is the gradient vector of the image I(x,y) at the pixel position (x,y);

[0023] Then, in the CFAR sliding window construction, the gradient direction and response intensity are combined to construct the protection band and background band;

[0024] Finally, the high response areas are scored and the suspected target candidate areas are screened out to generate target suggestion boxes.

[0025] Furthermore, the feature extraction and reconstruction process of the deep learning core detection network is as follows: using the content-aware feature reconstruction network CAFR to extract shallow, mid-level and deep feature maps, introducing the lightweight multi-head attention transformer LMHT to enhance position and channel information, and using the dynamic deformable fusion unit DCAU to adaptively aggregate texture details and semantic features;

[0026] The bounding box optimization process of the deep learning core detection network is as follows: The modified Bhattacharyya distance loss function (MBD-Loss) is introduced to model the target box as a two-dimensional Gaussian distribution, accurately matching the predicted box with the true box.

[0027] The modified Bhattacharyya distance is defined as follows:

[0028]

[0029] Where D mbd (P, Q) is the modified Bhattacharyya distance, which is used to measure the spatial distribution difference between the two target boxes P and Q; μ P and μ Q Represents the center coordinate vectors of the target boxes P and Q, μ P =(xP,yP) and μ Q = (xQ, yQ) is the center position of the two target boxes; (μ P -μ Q ) T is the transposed vector of the target frame center position deviation; Σ- 1 is the inverse matrix of the covariance matrix; det(Σ), det(Σ P )、det(Σ Q ) is the determinant value of the covariance matrix, which reflects the shape scale information of the target box and is used to calculate the shape difference; ∈ is a small positive number introduced to avoid logarithmic divergence, which is a very small positive value to ensure that the formula will not be numerically unstable during the calculation process;

[0030] The multi-stage cascade detection process of the deep learning core detection network is as follows: in the first stage, the graph neural network EA-GNN is used to remove isolated noise; in the second stage, the high-resolution cross attention module CLSR is used to optimize the positioning accuracy; in the third stage, the residual confidence calibration network RCR is designed to improve the confidence stability;

[0031] The output of the deep learning core detection network is a coarse list of candidate objects including categories, bounding box coordinates and confidence, as well as a high-quality feature map.

[0032] Furthermore, the content-aware feature reconstruction network CAFR introduces the lightweight multi-head attention transformer LMHT and the dynamic deformable fusion unit DCAU to build a cross-level feature reconstruction mechanism;

[0033] The structure of the content-aware feature reconstruction network CAFR is as follows:

[0034] The backbone network extracts multi-scale features: extracts feature maps F1, F2, and F3 at three scales from the original input image, corresponding to the shallow, medium, and deep layers respectively;

[0035] Lightweight Multi-Head Attention Transformer (LMHT): Each layer’s features are input into the LMHT module, and the position and channel information are jointly enhanced through the multi-head self-attention mechanism (MHSA).

[0036] The joint enhancement is expressed as:

[0037] F′ i =LMHT(F i )=Concat(MHSA1,...,MHSA h )W O

[0038] Where F′ i is the feature representation after being processed by the lightweight multi-head attention transformer LMHT; LMHT(F i ) is a lightweight multi-head attention transformer for the input feature map F i The processing process; Concat(MHSA1,...,MHSA h ) is to concatenate the outputs of multiple attention heads; MHSA k represents the output of the kth attention head; W O is the output mapping matrix; h is the number of attention heads;

[0039] Dynamic Deformable Fusion Unit (DCAU): Adaptively aggregates texture details and semantic features by nonlinearly recombining features at various scales and combining them with spatial attention masks.

[0040] Adaptive aggregation is expressed as:

[0041]

[0042] Where, F agg is the feature representation after adaptive aggregation; For all input feature maps F′ i The weighted sum operation of α i (x) is the attention weight, representing the i-th feature map F′ i The importance of the current position x; DeformConv(F′ i ,Δp i ) represents the deformable convolution operation, acting on the feature map F′ i and dynamic offset Δp i ;Δp iis the dynamic offset, which indicates the spatial offset of the deformable convolution kernel of the i-th feature map; i is the feature map index, which indicates the number of the feature map currently being processed.

[0043] Furthermore, the multi-stage cascade detection process uses a three-level cascade detection architecture, with each stage focusing on different tasks. By gradually optimizing the screening, positioning, and classification of candidate targets, the detection stability and accuracy are improved;

[0044] Candidate region structure modeling stage: The input is the CAF feature map. The graph neural network module EA-GNN is used to construct a spatial similarity graph for the initially extracted candidate boxes. The nodes represent candidate targets, and the edge weights represent texture or geometric similarity. Contextual information is propagated through graph convolution, isolated noise candidates are eliminated, and target regions with contextual structure are retained.

[0045] Fine position regression stage: Based on the screening results of the first stage, a high-resolution cross-attention module CLSR is used to focus on the center of gravity of fine features to enhance the spatial positioning ability of the local response area; the position adjustment adopts a two-stage regression method to improve the fitting ability;

[0046] Semantic confidence recalibration stage: For targets with blurred boundaries and low response, a residual confidence calibration network (RCR) is constructed to correct the bias of the category distribution of detection candidates; an additional soft label smoothing mechanism is used to improve the stability of the confidence output of boundary targets and reduce the risk of false detection.

[0047] Furthermore, the weak target recovery process of the post-processing and weak target restoration module is as follows: first, record the low-confidence candidate boxes and their local texture features; then, establish a geometric graph model, and construct a candidate graph by combining texture similarity and spatial proximity; finally, introduce a shape invariant moment or gradient direction histogram based on the shape invariant moment or gradient direction histogram. Figure 1 A consistent evaluation method is used to recover valid candidate targets that were mistakenly eliminated;

[0048] The geometric consistency check process of the post-processing and weak target restoration module is as follows: if the similarity constraint is met, that is, the distance difference is less than 5 pixels and the shape invariant moment conformity is greater than 0.85, the threshold is dynamically adjusted and the target is recycled;

[0049] The non-maximum suppression (NMS) process of the post-processing and weak target restoration module is as follows: removing redundant candidate boxes and generating the final target detection results;

[0050] The outputs of the post-processing and weak target restoration module are: category, bounding box coordinates and confidence, as well as recovered weak target information.

[0051] Furthermore, the polarization feature extraction process of the multi-source data fusion module is as follows: the polarization reflectance PR and polarization coherence index PCI are calculated respectively; the calculation formula is as follows:

[0052]

[0053] In the formula, PR HH / W is the polarization reflection ratio, representing the reflection intensity ratio between the HH polarization channel and the broadband signal; S HH is the complex scattering value of the HH polarization channel, representing the scattering signal intensity when the radar wave is transmitted horizontally and received horizontally; S W is the complex scattering value of the broadband signal, representing the comprehensive scattering signal intensity of the radar wave within the entire broadband range; |S HH | 2 and |S W | 2 are the power values of the HH polarization channel and the broadband signal respectively, that is, the modulus square of the complex scattering value; PCI is the polarization coherence index, used to measure the coherence between different polarization channels; S HV is the complex scattering value of the HV polarization channel, representing the scattering signal intensity when the radar wave is transmitted horizontally and received vertically; |S HV | 2 is the power value of the HV polarization channel, that is, the modulus square of the complex scattering value, reflecting the intensity or energy of the HV polarization signal; |S HH | and |S W | are the amplitude values of the HH polarization channel and the broadband signal respectively, that is, the modulus of the complex scattering value;

[0054] The feature fusion process of the multi-source data fusion module is as follows: combining the polarization statistical features and the SAR image-guided features, and inputting them into the auxiliary channel before the deep learning network;

[0055] The post-processing regularization process of the multi-source data fusion module is as follows: introducing the edge map or foreground map in the optical image as the regularization information, and constructing the soft fusion weight to improve the final prediction accuracy;

[0056] The output of the multi-source data fusion module is: fused features and regularization weights.

[0057] A method for SAR small-scale target detection based on deep learning. This method is based on the above-mentioned SAR small-scale target detection system based on deep learning, and includes the following steps:

[0058] Step 1: Data acquisition and preprocessing: Use the millimeter-wave synthetic aperture radar SAR to obtain high-resolution SAR images, and synchronously collect the attitude and position information provided by the inertial navigation device IMU and the global navigation satellite system GNSS; perform despeckling filtering on the original SAR image, and use the adaptive guided filtering AGCF to dynamically suppress the speckle noise and retain the target details; use the improved multi-scale local adaptive CFAR to generate high-quality candidate regions by combining the gradient direction and the response intensity;

[0059] Step 2: The Content-Aware Feature Reconstruction Network (CAFR) constructs high-quality feature representations by extracting shallow, mid-level, and deep feature maps corresponding to different levels of semantic information. It introduces a lightweight multi-head attention transformer (LMHT) to jointly enhance position and channel information through a multi-head self-attention mechanism. It also uses a dynamic deformable fusion unit (DCAU) combined with a spatial attention mask to achieve adaptive aggregation of texture details and semantic features.

[0060] Step 3: Modified Bhattacharyya distance bounding box matching mechanism MBD-Loss optimizes positioning accuracy: Model the target box as a two-dimensional Gaussian distribution; define the modified Bhattacharyya distance MBD loss function to calculate the distribution difference between the predicted box and the true box; combine the classification loss and confidence loss to form an end-to-end optimization target;

[0061] Step 4: The multi-stage cascade detection structure gradually optimizes candidate target screening and regression: the graph neural network EA-GNN is used to construct a spatial similarity graph, remove isolated noise, and retain the target area with contextual structure; the high-resolution cross attention module CLSR is used to focus on the center of gravity of fine features to optimize the spatial positioning ability of the local response area; the residual confidence calibration network RCR is designed to correct the bias of the category distribution of low-confidence targets and improve the confidence output stability of targets with blurred boundaries; the position and classification results of the candidate targets are optimized step by step;

[0062] Step 5: Weak target recovery and geometric consistency check to improve recall rate: record all low-confidence candidate boxes and their local texture features, and build a geometric graph model; for weak response candidates in the vicinity of high-confidence targets, build a candidate graph based on texture similarity and spatial proximity; introduce shape invariant moments or gradient direction histograms to improve recall rate. Figure 1 A consistency evaluation method is used to determine whether the similarity constraint is met; if the condition is met, the threshold is dynamically adjusted and the target is restored;

[0063] Step 6: Multi-source data fusion enhances cross-modality versatility: extract polarization features, including polarization reflectance (PR) and polarization coherence index (PCI); fuse edge maps or foreground maps in optical images as regularization information, and combine soft fusion weights to guide target confidence correction;

[0064] Step 7: Post-processing and model deployment optimization: Perform non-maximum suppression (NMS) on the final detection results to remove redundant candidate boxes. Use TensorRT inference acceleration to optimize FP16 precision for the Jetson Xavier NX embedded AI platform, ensuring that single-frame image processing latency does not exceed 90ms. Output the final target detection results, including target category, bounding box coordinates, and confidence level, to meet online real-time detection requirements.

[0065] The present invention proposes a SAR small-scale target detection system and method based on deep learning. Compared with existing technologies, the system has significant technical advantages and practical benefits in terms of detection accuracy, robustness, real-time performance, and engineering applications. Specifically, the present invention has the following beneficial effects:

[0066] 1. Significantly improve the spatial positioning accuracy of small-scale targets: By proposing a GNN-Transformer fusion structure, the feature dependency relationship between the target area and the candidate areas is accurately modeled. This significantly improves the spatial positioning accuracy of small-scale targets under complex backgrounds and multiple interference conditions, and significantly improves the problems of blurred target boundaries and inaccurate spatial positioning. It is particularly suitable for the precise detection of foreign objects on power transmission lines, such as bird nests, plastic bags, and ice.

[0067] 2. Improving the signal-to-noise ratio and contrast of small targets: By optimizing bounding box metrics using the Content-Aware Feature Reconstruction Network (CAFR) and the Modified Bhattacharyya Distance (MBD), this method significantly improves the signal-to-noise ratio (SNR) and contrast-to-noise ratio (CNR) of small targets in SAR images. This method effectively enhances the detectability of small targets. In terms of deep detection model performance, target detection accuracy is significantly improved, significantly reducing the risk of missed and false detections, and enhancing the reliability of power facility monitoring and hazard detection.

[0068] 3. Possess good platform generalization capabilities and versatility: The system and method proposed in this invention can achieve stable target detection and classification results under the conditions of SAR data of different polarization channels or consumer-grade SAR sensors, without relying on additional high-cost auxiliary equipment. It has been verified in various scenarios, such as complex mountain terrain, interference environments in forest areas, and densely populated urban areas. The target detection error is stably controlled within a reasonable range, and the model has good robustness, making it suitable for deployment on various airborne platforms and large-scale inspection tasks. The system has cross-scenario and cross-modal universal capabilities, providing reliable technical support for target detection in complex environments.

[0069] 4. Meet real-time requirements and support edge computing deployment: Through deep model compression and TensorRT inference acceleration optimization, this system achieves efficient real-time processing capabilities on a low-power embedded computing platform. The processing latency of single-frame SAR images is stable, meeting the requirements of online real-time processing and edge computing deployment for drones, providing efficient and reliable technical support for real-time target detection tasks such as power inspections and disaster monitoring.

[0070] 5. Significantly reduce system costs and promote technology adoption: By reducing reliance on high-precision navigation equipment or auxiliary sensors, overall system costs are reduced, significantly reducing hardware investment and maintenance costs. This also significantly reduces image scrapping and task repetition rates due to missed targets, improving the first-time success rate of SAR imaging systems and aerial inspection missions, and further enhancing economic benefits.

[0071] 6. Improving the Automation and Intelligence of Critical Infrastructure Monitoring: This invention significantly enhances the automation and intelligence of critical infrastructure monitoring, such as transmission lines, laying a solid technical foundation for building efficient, all-weather, automated grid sensing and monitoring systems. By improving the efficiency, safety, and reliability of power equipment inspections, this invention has demonstrated significant socioeconomic benefits in practical engineering applications and holds broad potential for application.

[0072] This invention significantly improves the detection accuracy, recall, and robustness of small-scale targets in SAR imagery through innovative deep learning techniques and multi-module collaborative optimization, while simultaneously meeting the requirements of real-time performance and engineering deployment. Its technical advantages and practical benefits are reflected not only in the overall improvement of detection performance but also in the significant reduction of system costs, the promotion of technology promotion and industrialization, and the provision of efficient and reliable technical support for fields such as power inspection and disaster monitoring. It possesses significant economic value and social application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0074] Figure 1 This is an algorithm flow chart of the SAR small-scale target detection system based on deep learning of the present invention;

[0075] Figure 2 This is a structural diagram of the SAR motion compensation deep learning network based on the fusion of graph neural network and temporal Transformer in the present invention. DETAILED DESCRIPTION

[0076] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0077] To address the problems of low detection accuracy, inaccurate positioning, and poor robustness of small-scale targets in complex backgrounds. This embodiment provides a SAR small-scale target detection system based on deep learning. The system is designed for SAR image scenes in complex natural and artificial backgrounds, and significantly improves the detection capability of low signal-to-noise ratio and small-sized targets. The system combines a content-aware feature enhancement network, a modified Bhattacharyya distance boundary matching mechanism, and a multi-stage cascade detector structure to construct a target detection system that takes into account discrimination ability, positioning accuracy, and training stability without relying on traditional convolution stacking.

[0078] The deep learning-based SAR small-scale target detection system consists of the following core modules:

[0079] Data acquisition and preprocessing module:

[0080] like Figure 1 and Figure 2 As shown in Figure 1, the data acquisition and preprocessing module is the foundation of the entire system, responsible for acquiring raw data from hardware devices and performing preliminary image processing to provide high-quality input for the subsequent deep learning detection network. Its main functions include: Data Acquisition: It acquires high-resolution SAR images using millimeter-wave synthetic aperture radar (SAR) and simultaneously records attitude and position information; Time Base Synchronization: It uses PPS clock pulses to uniformly control the time base of the radar, inertial navigation unit (IMU), and computing unit to ensure temporal consistency of multi-source data; Image Preprocessing: It uses adaptive guided coherent filtering (AGCF) and an improved multi-scale local adaptive CFAR method to suppress coherent speckle noise, quickly extract suspected target regions, and generate high-quality candidate boxes. The module's output includes preprocessed SAR images, preliminarily selected target proposals, and time synchronization data, providing optimized input for feature extraction and target detection in the subsequent deep learning core detection network.

[0081] The data collection process of the data collection and preprocessing module is as follows:

[0082] For SAR image acquisition: The system is equipped with a Ka-band SAR imaging radar with a center frequency of 34.5GHz and a bandwidth of 200MHz, which is used to acquire high-resolution SAR images; SAR images have all-weather and all-day imaging capabilities and are suitable for small target detection tasks in complex backgrounds.

[0083] For attitude and position information measurement, the MEMSIMU inertial navigation device and the GNSS global navigation satellite system are used to record the pitch angle, yaw angle, roll angle and other attitudes of the UAV platform and the longitude, latitude, altitude and other position information; these auxiliary information provides important support for subsequent multi-source data fusion and target positioning correction.

[0084] For time base synchronization, PPS clock pulses are used to synchronize the radar, IMU and computing unit to ensure the time consistency of multi-source data streams; all raw data streams are transmitted to the Jetson Xavier NX edge computing platform through high-speed LVDS and UART interfaces, laying the foundation for real-time processing.

[0085] The image preprocessing process of the data acquisition and preprocessing module is as follows:

[0086] To suppress speckle noise, we use the adaptive guided coherence filtering (AGCF) method to enhance the original image during preprocessing. AGCF statistically analyzes the local coherence matrix of the image and dynamically adjusts the filter template response to preserve target details while suppressing unstructured noise.

[0087] The filtering equation is as follows:

[0088]

[0089] Where, I out (x, y) is the grayscale value or intensity value of the output image at the pixel position (x, y); Ω(x, y) is the local window area centered at the pixel position (x, y); w ij is the weight coefficient, which indicates the contribution of the pixel (i, j) in the neighborhood to the central pixel (x, y); I(i, j) is the grayscale value or intensity value of the input image at the pixel position (i, j).

[0090] The filter equation dynamically adjusts the weight w by performing weighted averaging on the pixels in the local window. ij , thereby achieving: retaining target edges and important features through high weights, retaining target details; weakening coherent speckle noise and other interference through low weights, suppressing unstructured noise; using neighborhood information to smooth the image while avoiding blurring the target outline and enhancing local consistency; this filtering method is particularly suitable for the SAR image preprocessing stage and can effectively improve the detection accuracy and robustness of small-scale targets.

[0091] In the specific implementation, AGCF adopts a 3×3 guided filter kernel, dynamically adjusts the filter weights in 5×5 image blocks, and maintains the integrity of the target edge through the guided graph structure.

[0092] For candidate target region generation, an improved multi-scale local adaptive CFAR method, namely A-CFAR-EP, is used to quickly extract suspected target regions by combining gradient direction and response intensity. This method includes the following steps:

[0093] First, the gradient intensity map is constructed using the grayscale gradient direction of the SAR image: G(x,y)=‖▽I(x,y)‖, which is used to mark the areas where the target edge may exist.

[0094] Where G(x,y) is the value of the gradient intensity map at the pixel position (x,y); ▽I(x,y) is the gradient vector of the image I(x,y) at the pixel position (x,y), Indicates the grayscale change rate of the image in the horizontal direction, Represents the grayscale change rate of the image in the vertical direction; ‖▽I(x,y)‖ is the modulus of the gradient vector, I(x,y) is the grayscale value or intensity value of the input SAR image at the pixel position (x,y).

[0095] Then, in the CFAR sliding window construction, the gradient direction and response intensity are combined to jointly construct the protection band and background band; finally, the high response area is scored and the suspected target candidate area is screened out to generate the target suggestion box.

[0096] A-CFAR-EP is based on an 11×11 sliding window design and significantly improves the positive sample coverage by setting a gradient direction response threshold of 20. This method can effectively reduce redundant background features and improve the computational efficiency of subsequent deep learning models.

[0097] After the aforementioned data acquisition and preprocessing process, the data acquisition and preprocessing module outputs: 1. Preprocessed SAR images: Removes speckle noise and enhances target details; 2. Initially screened target proposals: High-quality candidate regions generated using the A-CFAR-EP method; 3. Time-synchronized data: Ensures temporal consistency across multi-source data streams, supporting subsequent multi-source data fusion and positioning correction. These outputs provide high-quality input for the subsequent deep learning core detection network, significantly improving the system's detection accuracy and robustness.

[0098] Deep learning core detection network:

[0099] like Figure 1 and Figure 2As shown in the figure, the deep learning core detection network is the core module of the entire system. It connects to the data acquisition and preprocessing module and is responsible for extracting multi-level features from preprocessed SAR images and generating a coarse list of candidate targets. Its main functions include: Feature Extraction and Reconstruction: Using the Content-Aware Feature Reconstruction Network (CAFR), it constructs high-quality, multi-scale feature representations, alleviating the semantic information loss and boundary blurring of small-scale targets during deep downsampling. Bounding Box Optimization: Introducing the Modified Bhattacharyya Distance Loss (MBD-Loss), it models the target bounding box as a two-dimensional Gaussian distribution, accurately matching the predicted and true bounding boxes to improve the localization and regression accuracy of small targets. Multi-Stage Cascade Detection: Utilizing a three-stage cascade detection architecture (EA-GNN, CLSR, and RCR), it progressively optimizes the screening, localization, and classification of candidate targets, significantly improving detection stability and accuracy. This module outputs a coarse list of candidate targets with categories, bounding box coordinates, and confidence scores, as well as a high-quality feature map, which provides optimized input for subsequent post-processing and weak target repair modules.

[0100] The feature extraction and reorganization process of the deep learning core detection network is as follows:

[0101] The backbone network extracts multi-scale features: extracts feature maps F1, F2, and F3 at three scales from the original input image, corresponding to shallow, middle, and deep features respectively.

[0102] Lightweight Multi-Head Attention Transformer (LMHT): Each layer feature is input into the LMHT module, and the position and channel information are jointly enhanced through the multi-head self-attention mechanism (MHSA) to obtain a joint enhanced representation of position and channel:

[0103] F′ i =LMHT(F i )=Concat(MHSA1,...,MHSA h )W O

[0104] Where F′ i is the feature representation after being processed by the lightweight multi-head attention transformer LMHT; LMHT(F i ) is a lightweight multi-head attention transformer for the input feature map F i The processing process; Concat(MHSA1,...,MHSA h ) is to concatenate the outputs of multiple attention heads; MHSA k represents the output of the kth attention head; W O is the output mapping matrix; h is the number of attention heads. This formula is obtained by using the lightweight multi-head attention transformer LMHT to transform the input feature map F i Enhance and generate richer feature representation F′ i. Each note head MHSA k Calculate the attention weights independently to capture the information of different subspaces in the feature map; concatenate the outputs of all attention heads to form a high-dimensional feature representation; and output the mapping matrix W. O Mapping the concatenated features to the target dimension ensures that the output feature dimensions are compatible with subsequent modules. This method can effectively enhance the position and channel information of features, especially in small-scale object detection tasks, significantly improving the feature expression capability and model detection accuracy.

[0105] Dynamic Deformable Fusion Unit (DCAU): DCAU nonlinearly reorganizes features at each scale and combines them with spatial attention masks to achieve adaptive aggregation of texture details and semantic features.

[0106]

[0107] Where, F agg is the feature representation after adaptive aggregation; For all input feature maps F′ i The weighted sum operation of α i (x) is the attention weight, representing the i-th feature map F′ i The importance of the current position x; DeformConv(F′ i ,Δp i ) represents the deformable convolution operation, acting on the feature map F′ i and dynamic offset Δp i ;Δp i is the dynamic offset, which represents the spatial offset of the deformable convolution kernel of the i-th feature map; i is the feature map index, which represents the number of the feature map currently being processed. The purpose of this formula is to adaptively aggregate multi-scale features through the dynamic deformable fusion unit DCAU to generate a high-quality feature representation F agg . By the attention weight α i (x) Dynamically adjust the contribution of different feature maps to enhance the focus on small-scale targets and key areas; use dynamic offset Δp i Adjust the position of the convolution kernel so that it can flexibly adapt to the geometric deformation and local details of the target; by multi-scale feature map F′ i A weighted summation is performed to achieve an organic fusion of shallow texture details and deep semantic information. This method can significantly alleviate the problems of semantic information loss and blurred boundaries of small-scale objects during deep downsampling, while enhancing the model's ability to perceive target structure and details, providing higher-quality feature input for subsequent object detection tasks.

[0108] This mechanism significantly enhances the shallow layer's structural perception of small-scale targets, while solving the problem of uneven fusion of deep and shallow features.

[0109] The bounding box optimization process of the deep learning core detection network is as follows:

[0110] Introducing the modified Bhattacharyya distance loss function MBD-Loss: Modeling the target box as a two-dimensional Gaussian distribution Where μ = (x, y) is the center coordinate of the target frame, Σ = diag (w 2 ,h 2 ) is the target box size covariance matrix.

[0111] The modified Bhattacharyya distance is defined as follows:

[0112]

[0113] Where D mbd (P, Q) is the modified Bhattacharyya distance, which is used to measure the spatial distribution difference between the two target boxes P and Q; μ P and μ Q Represents the center coordinate vectors of the target boxes P and Q, μ P =(xP,yP) and μ Q = (xQ, yQ) is the center position of the two target boxes; (μ P -μ Q ) T is the transposed vector of the target frame center position deviation; Σ- 1 is the inverse matrix of the covariance matrix; det(Σ), det(Σ P )、det(Σ Q ) is the determinant value of the covariance matrix, which reflects the shape scale information of the target box and is used to calculate the shape difference; ∈ is a small positive number introduced to avoid logarithmic divergence, which ensures that the formula will not be numerically unstable during the calculation process.

[0114] The loss function combination is:

[0115] L=λ1·L cls +λ2·D mbd +λ3·L conf

[0116] Where L is the total loss function used to optimize the target detection model; λ1, λ2, and λ3 are weight coefficients used to balance the importance of different loss terms; L cls is the classification loss, which is used to evaluate the accuracy of target category prediction; D mbd To correct the Bhattacharyya distance loss and optimize the positioning accuracy of the target frame; L conf is the target confidence loss, which is used to evaluate the accuracy of confidence prediction of whether the target box exists or not.

[0117] Represent the target box as a two-dimensional Gaussian distribution The position and shape of the target box are described by the mean μ and covariance matrix Σ; D mbd (P,Q) combines the center position deviation and shape difference of the target box to robustly reflect the degree of match between the predicted box and the ground-truth box, making it particularly suitable for small target localization tasks. The total loss function L combines classification loss, localization loss, and confidence loss. By adjusting the weight coefficients λ1, λ2, and λ3, it achieves multi-task joint optimization, significantly improving the overall performance of the target detection model. Compared with traditional IoU-based loss functions, this mechanism is more tolerant to positional offsets and scale changes of small targets and has a more stable gradient, making it particularly suitable for small target detection tasks in complex backgrounds.

[0118] The multi-stage cascade detection process of the deep learning core detection network is as follows:

[0119] Phase 1: Candidate Region Structure Modeling (EA-GNN): The input is the CAFR feature map. The graph neural network module, EA-GNN, constructs a spatial similarity graph for the initially extracted candidate boxes. Nodes represent candidate objects, and edge weights represent texture or geometric similarity. Graph convolution propagates contextual information, eliminating isolated noisy candidates and retaining target regions with contextual structure.

[0120] Phase 2: Fine-grained Position Regression (CLSR): Based on the results of the first phase, a high-resolution cross-attention module (CLSR) is used, focusing on the center of gravity of fine features to enhance the spatial localization capability of the local response area. Position adjustment uses a two-stage regression approach, using two layers of cross attention and two layers of multi-layered layer processing (MLP) to achieve fine-grained position regression with 128 channels, improving fitting capabilities.

[0121] Phase 3: Semantic Confidence Recalibration (RCR): For objects with blurred boundaries and low response, a residual confidence calibration network (RCR) is designed to correct the bias in the category distribution of detection candidates. An additional soft label smoothing mechanism with a temperature coefficient set to 0.7 improves the stability of the confidence output for boundary objects and reduces the risk of false detection.

[0122] After the aforementioned feature extraction, bounding box optimization, and multi-stage cascade detection process, the deep learning core detection network outputs: 1. A coarse list of candidate objects, including object categories, bounding box coordinates, and confidence scores; and 2. High-quality feature maps, which are used for further optimization of the post-processing and weak object restoration modules. These outputs provide high-quality input data for subsequent modules, significantly improving the system's detection accuracy and robustness, particularly for small-scale object detection in complex backgrounds.

[0123] Post-processing and weak target repair module:

[0124] like Figure 1 and Figure 2 As shown in the figure, the post-processing and weak target repair module is an important part of the system. It is connected to the deep learning core detection network data transmission and is responsible for further optimizing the output results of the deep learning core detection network. Its main functions include: weak target recovery: by recording low-confidence candidate boxes and their local texture features, combined with the geometric consistency verification mechanism, it recovers valid candidate targets that were mistakenly eliminated; redundant box removal: using the non-maximum suppression (NMS) algorithm to remove redundant candidate boxes and generate the final target detection result; improving the recall rate: without significantly increasing the false detection rate, it significantly improves the recall rate of small-scale targets and targets with blurred boundaries, which is especially suitable for dense detection scenarios. The output of this module is the final target detection result including category, bounding box coordinates and confidence, as well as the recovered weak target information, providing more comprehensive and accurate detection results for subsequent tasks.

[0125] The weak target recovery process of the post-processing and weak target repair module is as follows:

[0126] First, record low-confidence candidate boxes and their local texture features: The system first records all candidate boxes judged as low-confidence by the deep learning core detection network and extracts their local texture features, such as edge structure and gradient direction; this information provides basic data for subsequent evaluation of the rationality of the candidate boxes.

[0127] Then, a geometric graph model is built: Based on the weak response candidate boxes in the vicinity of high-confidence targets, a geometric graph model is constructed. The nodes in the graph represent candidate targets, and the edge weights represent the texture similarity and spatial proximity between the candidate boxes.

[0128] Then, a candidate map is constructed: the relationship between candidate boxes is evaluated by texture similarity and spatial proximity, and candidate objects with potential relevance are screened. For example, a candidate box with a distance less than 5 pixels and a shape invariant moment matching greater than 0.85 is considered a possible missed target.

[0129] Finally, we introduce the shape invariant moment or gradient direction histogram Figure 1 Consistency evaluation method: A second evaluation is performed on the candidate box to determine whether it should be restored as a valid target. If the similarity constraints are met, such as the distance difference is less than 5 pixels and the shape invariant moment conformance is greater than 0.85, the confidence threshold is dynamically adjusted and the target is re-included in the detection results.

[0130] The geometric consistency check process of the post-processing and weak target repair module is as follows:

[0131] Based on the similarity constraint definition, the system sets clear geometric consistency verification conditions. These include: distance constraint: the distance between the candidate bounding box and the high-confidence target must be less than 5 pixels; shape constraint: the shape invariant moments of the candidate bounding box, such as the Hu moment, must match the high-confidence target with a degree greater than 0.85. Candidate bounding boxes that meet these conditions are considered missed targets. For dynamic threshold adjustment, the confidence threshold of candidate bounding boxes that meet the geometric consistency verification conditions is dynamically adjusted, upgrading them from low-confidence status to valid targets. For target recycling, candidate bounding boxes that pass verification are re-incorporated into the final detection results to prevent small-scale targets from being mistakenly rejected due to low confidence.

[0132] The non-maximum suppression (NMS) process of the post-processing and weak target restoration module is as follows:

[0133] In the final list of candidate boxes, the non-maximum suppression (NMS) algorithm is used to remove redundant boxes. NMS compares the confidence scores and the degree of overlap (IoU) of the candidate boxes, retaining the objects with the highest confidence and the lowest overlap with other boxes. After NMS processing, the final object detection result is generated, including the category, bounding box coordinates, and confidence score.

[0134] After the aforementioned weak target recovery, geometric consistency check, and non-maximum suppression processes, the post-processing and weak target repair module outputs: 1. Final target detection results, including target category, bounding box coordinates, and confidence; 2. Recovered weak target information, along with missed targets recovered through a geometric consistency check mechanism. These outputs significantly improve the system's detection recall while maintaining high detection accuracy, particularly for small-scale target detection in complex backgrounds. The post-processing and weak target repair module, through weak target recovery and the geometric consistency check mechanism WTR-GC, effectively addresses the problem of small-scale targets being falsely rejected due to low confidence, significantly improving detection recall. Furthermore, non-maximum suppression (NMS) removes redundant bounding boxes, ensuring the accuracy and conciseness of the final detection results. This module is particularly suitable for dense detection scenarios, significantly enhancing the system's practicality and robustness without significantly increasing the false detection rate.

[0135] Multi-source data fusion module:

[0136] The multi-source data fusion module is a crucial component of the system. It connects to the post-processing and weak target repair modules for network data transmission. By fusing multi-source data, such as SAR polarimetric features and optical imagery, it significantly enhances the system's cross-modal generalization capabilities. Its main functions include: Polarimetric Feature Extraction: Calculating the polarimetric reflectance (PR) and polarimetric coherence index (PCI) from multi-polarimetric SAR images to provide complementary statistical features; Feature Fusion: Combining polarimetric statistical features with SAR image guidance features and feeding them into an auxiliary channel before the deep learning network, enhancing the model's target perception; Post-processing Regularization: Introducing edge maps or foreground images from optical images as regularization information, constructing soft fusion weights to correct target confidence and improve final prediction accuracy. The module outputs fused features and regularization weights, which effectively enhance the system's detection performance in complex scenarios and are particularly well-suited for collaborative detection tasks involving multimodal remote sensing data.

[0137] The polarization feature extraction process of the multi-source data fusion module is as follows:

[0138] First, calculate the polarization reflectance ratio PR: The polarization reflectance ratio reflects the difference in reflection intensity between different polarization channels. Its calculation formula is as follows:

[0139]

[0140] Where PR HH / W is the polarization reflectance, which represents the reflection intensity ratio between the HH polarization channel and the broadband signal; S HH is the complex scattering value of the HH polarization channel, which represents the scattered signal intensity of the radar wave when it is horizontally transmitted and received; S W is the complex scattering value of the broadband signal, which represents the comprehensive scattering signal intensity of the radar wave in the entire broadband range; |S HH | 2 and |S W | 2 are the power values of the HH polarization channel and the broadband signal, that is, the square of the modulus of the complex scattering value.

[0141] Then, the polarization coherence index (PCI) is calculated. The PCI is used to measure the coherence between different polarization channels. Its calculation formula is as follows:

[0142]

[0143] Where PCI is the polarization coherence index, which is used to measure the coherence between different polarization channels; S HV is the complex scattering value of the HV polarization channel, which represents the scattered signal intensity of the radar wave when it is transmitted horizontally and received vertically; |S HV | 2is the power value of the HV polarization channel, which is the squared modulus of the complex scattering value and reflects the intensity or energy of the HV polarization signal; |S HH | and |S W | are the amplitude values of the HH polarization channel and the broadband signal respectively, which are the moduli of the complex scattering values.

[0144] The polarization reflectivity and polarization coherence index provide descriptions of the physical characteristics of the target, which helps to distinguish different types of targets, such as metal objects, vegetation, etc., and enhance the model's ability to identify small-scale targets.

[0145] The feature fusion process of the multi-source data fusion module is as follows:

[0146] Fuse the extracted polarization statistical features PR and PCI with the guiding features of the SAR image; the guiding features usually come from information such as the texture and gradient direction of the SAR image and are used to enhance the target's structural perception ability. Input the fused features into the auxiliary channel of the deep learning network as additional input information to enhance the model's multi-dimensional perception ability of the target. This lightweight fusion mechanism can significantly improve the model's feature expression ability without significantly increasing the computational burden.

[0147] The post-processing regularization process of the multi-source data fusion module is as follows:

[0148] In the post-processing stage, use the edge map or foreground map in the optical image as regularization information to guide the correction of the target confidence. The edge map can highlight the contour information of the target, while the foreground map can provide the semantic prior information of the target. Generate soft fusion weights through the edge map or foreground map to dynamically adjust the confidence score of the target box. The construction mechanism of the soft fusion weights can effectively alleviate the problem of unstable confidence of the boundary-blurred target and reduce the false detection rate. Through regularization correction, further optimize the confidence distribution of the detection results to ensure the high precision and high robustness of target detection.

[0149] After the aforementioned polarimetric feature extraction, feature fusion, and post-processing regularization process, the multi-source data fusion module outputs: 1. Fusion features: enhanced features generated by combining polarimetric statistical features with SAR image guidance features, which serve as auxiliary inputs to the deep learning network; and 2. Regularization weights: soft fusion weights generated based on optical imagery, which are used to correct target confidence and improve final prediction accuracy. By fusing multi-source data, including SAR polarimetric features and optical imagery, the multi-source data fusion module significantly enhances the system's cross-modal generalization and detection performance. Its core lies in utilizing polarimetric statistical features (PR) and (PCI) to enhance the target's physical characteristics and guiding target confidence correction using edge maps or foreground maps in optical imagery. This lightweight fusion mechanism not only enhances the model's multidimensional perception capabilities but also significantly improves the system's detection accuracy and robustness in complex scenarios, providing strong technical support for collaborative detection tasks involving multimodal remote sensing data.

[0150] The aforementioned modules form the deep learning-based SAR small-scale target detection system. The entire system utilizes the ROS2 platform as software middleware for module scheduling and data communication, employing a publisher / subscriber architecture for module decoupling. The system model, optimized with TensorRT, is deployed on the Jetson Xavier NX embedded platform, supporting FP16 precision inference and achieving single-frame image latency of less than 90ms, meeting the requirements for online detection and processing during flight.

[0151] In addition, based on the above-mentioned SAR small-scale target detection system based on deep learning, this embodiment also proposes a SAR small-scale target detection method based on deep learning, which includes the following steps:

[0152] Step 1: Data acquisition and preprocessing: Use millimeter-wave synthetic aperture radar (SAR) to acquire high-resolution SAR images, and simultaneously collect attitude and position information provided by the inertial navigation unit (IMU) and the global navigation satellite system (GNSS). Perform despeckling on the original SAR image, using adaptive guided filtering (AGCF) to dynamically suppress coherent speckle noise and preserve target details. Utilize an improved multi-scale local adaptive CFAR to combine gradient direction and response intensity to generate high-quality candidate regions.

[0153] Step 2: The Content-Aware Feature Reconstruction Network (CAFR) constructs high-quality feature representations by extracting shallow, mid-level, and deep feature maps corresponding to different levels of semantic information. It introduces a lightweight multi-head attention transformer (LMHT) to jointly enhance position and channel information through a multi-head self-attention mechanism. It also uses a dynamic deformable fusion unit (DCAU) combined with a spatial attention mask to achieve adaptive aggregation of texture details and semantic features.

[0154] Step 3: Modified Bhattacharyya distance bounding box matching mechanism MBD-Loss optimizes positioning accuracy: Model the target box as a two-dimensional Gaussian distribution; define the modified Bhattacharyya distance MBD loss function to calculate the distribution difference between the predicted box and the true box; combine the classification loss and confidence loss to form an end-to-end optimization target;

[0155] Step 4: The multi-stage cascade detection structure gradually optimizes candidate target screening and regression: the graph neural network EA-GNN is used to construct a spatial similarity graph, remove isolated noise, and retain the target area with contextual structure; the high-resolution cross attention module CLSR is used to focus on the center of gravity of fine features to optimize the spatial positioning ability of the local response area; the residual confidence calibration network RCR is designed to correct the bias of the category distribution of low-confidence targets and improve the confidence output stability of targets with blurred boundaries; the position and classification results of the candidate targets are optimized step by step;

[0156] Step 5: Weak target recovery and geometric consistency check to improve recall rate: record all low-confidence candidate boxes and their local texture features, and build a geometric graph model; for weak response candidates in the vicinity of high-confidence targets, build a candidate graph based on texture similarity and spatial proximity; introduce shape invariant moments or gradient direction histograms to improve recall rate. Figure 1 A consistency evaluation method is used to determine whether the similarity constraint is met; if the condition is met, the threshold is dynamically adjusted and the target is restored;

[0157] Step 6: Multi-source data fusion enhances cross-modality versatility: extract polarization features, including polarization reflectance (PR) and polarization coherence index (PCI); fuse edge maps or foreground maps in optical images as regularization information, and combine soft fusion weights to guide target confidence correction;

[0158] Step 7: Post-processing and model deployment optimization: Perform non-maximum suppression (NMS) on the final detection results to remove redundant candidate boxes. Use TensorRT inference acceleration to optimize FP16 precision for the Jetson Xavier NX embedded AI platform, ensuring that single-frame image processing latency does not exceed 90ms. Output the final target detection results, including target category, bounding box coordinates, and confidence level, to meet online real-time detection requirements.

[0159] During the field validation phase, the system and method were applied to drone inspection missions in mountainous and forested areas, as well as in suburban power transmission corridors. They were able to reliably detect foreign objects and ice-covered targets smaller than 2% of the image size, even in moderately complex backgrounds. The system achieved an average detection accuracy of 91.3% and a missed detection rate of less than 5%. Without relying on RTK-GPS or high-precision IMUs, the overall system exhibits excellent detection accuracy, robustness, and engineering deployability, making it suitable for high-reliability sensing scenarios such as power inspections, disaster monitoring, and border security.

[0160] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A SAR small-scale target detection system based on deep learning, characterized by: include: A data acquisition and preprocessing module, a deep learning core detection network connected to the data acquisition and preprocessing module for data transmission, a post-processing and weak target repair module connected to the data transmission of the deep learning core detection network, and a multi-source data fusion module connected to the data transmission of the post-processing and weak target repair module; The data acquisition and preprocessing module obtains raw data from the hardware device and performs preliminary processing on the image to provide high-quality input for subsequent modules; the deep learning core detection network extracts multi-level features from the preprocessed image and generates a coarse-screened candidate target list; the post-processing and weak target repair module further optimizes the output of the deep learning core detection network, performs secondary evaluation and recovery on weak targets and targets with blurred boundaries, and improves the detection recall rate; the multi-source data fusion module improves the system's cross-modal generalization capability by fusing multi-source data such as polarization features and optical images.

2. The SAR small-scale target detection system based on deep learning according to claim 1, characterized in that: The data acquisition and preprocessing module uses millimeter-wave synthetic aperture radar (SAR) to acquire high-resolution SAR images, combines the inertial navigation unit (IMU) and the global navigation satellite system (GNSS) to record attitude and position information, and uses PPS clock pulses to synchronize the radar, inertial navigation unit (IMU), and computing unit. The image preprocessing process of the data acquisition and preprocessing module is as follows: using the adaptive guided coherence filtering (AGCF) method to dynamically adjust the filter template response to retain target details while suppressing non-structured noise; The filtering equation is as follows: Where, I out (x, y) is the grayscale value or intensity value of the output image at the pixel position (x, y); Ω(x, y) is the local window area centered at the pixel position (x, y); w ij is the weight coefficient, which indicates the contribution of the pixel (i, j) in the neighborhood to the central pixel (x, y); I(i, j) is the grayscale value or intensity value of the input image at the pixel position (i, j); Through the improved multi-scale local adaptive CFAR method, the suspected target area is quickly extracted by combining the gradient direction and response intensity; The outputs of the data acquisition and preprocessing module are: preprocessed SAR images, preliminarily screened target proposal boxes, and time synchronization data.

3. The SAR small-scale target detection system based on deep learning according to claim 2, characterized in that: The improved multi-scale local adaptive CFAR method introduces the image gradient enhancement channel as the edge prior to guide noise modeling and sliding window configuration, which includes the following steps: First, the gradient intensity map is constructed using the grayscale gradient direction of the SAR image: G(x,y) = ‖▽I(x,y)‖, which is used to mark the area where the target edge may exist; where G(x,y) is the value of the gradient intensity map at the pixel position (x,y); ▽I(x,y) is the gradient vector of the image I(x,y) at the pixel position (x,y); Then, in the CFAR sliding window construction, the gradient direction and response intensity are combined to construct the protection band and background band; Finally, the high response areas are scored and the suspected target candidate areas are screened out to generate target suggestion boxes.

4. The SAR small-scale target detection system based on deep learning according to claim 1, characterized in that: The feature extraction and reconstruction process of the deep learning core detection network is as follows: using the content-aware feature reconstruction network (CAFR) to extract shallow, mid-level, and deep feature maps, introducing the lightweight multi-head attention transformer (LMHT) to enhance position and channel information, and using the dynamic deformable fusion unit (DCAU) to adaptively aggregate texture details and semantic features; The bounding box optimization process of the deep learning core detection network is as follows: The modified Bhattacharyya distance loss function (MBD-Loss) is introduced to model the target box as a two-dimensional Gaussian distribution, accurately matching the predicted box with the true box. The modified Bhattacharyya distance is defined as follows: Where D mbd (P, Q) is the modified Bhattacharyya distance, which is used to measure the spatial distribution difference between the two target boxes P and Q; μ P and μ Q Represents the center coordinate vectors of the target boxes P and Q, μ P =(xP,yP) and μ Q = (xQ, yQ) is the center position of the two target boxes; (μ P -μ Q ) T is the transposed vector of the target frame center position deviation; Σ- 1 is the inverse matrix of the covariance matrix; det(Σ), det(Σ P )、det(Σ Q ) is the determinant value of the covariance matrix, which reflects the shape scale information of the target box and is used to calculate the shape difference; ∈ is a small positive number introduced to avoid logarithmic divergence, which is a very small positive value to ensure that the formula will not be numerically unstable during the calculation process; The multi-stage cascade detection process of the deep learning core detection network is as follows: in the first stage, the graph neural network EA-GNN is used to remove isolated noise; In the second stage, a high-resolution cross-attention module (CLSR) is used to optimize positioning accuracy. In the third stage, a residual confidence calibration network (RCR) is designed to improve confidence stability. The output of the deep learning core detection network is a coarse list of candidate objects including categories, bounding box coordinates and confidence levels, as well as a high-quality feature map.

5. The SAR small-scale target detection system based on deep learning according to claim 4, characterized in that: The content-aware feature reconstruction network CAFR introduces the lightweight multi-head attention transformer LMHT and the dynamic deformable fusion unit DCAU to build a cross-level feature reconstruction mechanism; The structure of the content-aware feature reconstruction network CAFR is as follows: The backbone network extracts multi-scale features: extracts feature maps F1, F2, and F3 at three scales from the original input image, corresponding to the shallow, medium, and deep layers respectively; Lightweight Multi-Head Attention Transformer (LMHT): Each layer’s features are input into the LMHT module, and the position and channel information are jointly enhanced through the multi-head self-attention mechanism (MHSA). The joint enhancement is expressed as: F i ′=LMHT(F i )=Concat(MHSA1,...,MHSA h )W O Where, F i ′ is the feature representation after processing by the lightweight multi-head attention transformer LMHT; LMHT(F i ) is a lightweight multi-head attention transformer for the input feature map F i The processing process; Concat(MHSA1,…,MHSA h ) is to concatenate the outputs of multiple attention heads; MHSA k represents the output of the kth attention head; W O is the output mapping matrix; h is the number of attention heads; Dynamic Deformable Fusion Unit (DCAU): Adaptively aggregates texture details and semantic features by nonlinearly recombining features at various scales and combining them with spatial attention masks. Adaptive aggregation is expressed as: Where, F agg is the feature representation after adaptive aggregation; For all input feature maps F i ′ weighted sum operation; α i (x) is the attention weight, representing the i-th feature map F i ′ is important at the current position x; DeformConv(F i ′,Δp i ) represents the deformable convolution operation, acting on the feature map F i ′ and dynamic offset Δp i ;Δp i is the dynamic offset, which indicates the spatial offset of the deformable convolution kernel of the i-th feature map; i is the feature map index, which indicates the number of the feature map currently being processed.

6. The SAR small-scale target detection system based on deep learning according to claim 4, characterized in that: The multi-stage cascade detection process uses a three-level cascade detection architecture, with each stage focusing on different tasks. By gradually optimizing the screening, positioning, and classification of candidate targets, the detection stability and accuracy are improved. Candidate region structure modeling stage: The input is the CAF feature map. The graph neural network module EA-GNN is used to construct a spatial similarity graph for the initially extracted candidate boxes. The nodes represent candidate targets, and the edge weights represent texture or geometric similarity. Contextual information is propagated through graph convolution, isolated noise candidates are eliminated, and target regions with contextual structure are retained. Fine position regression stage: Based on the screening results of the first stage, a high-resolution cross-attention module CLSR is used to focus on the center of gravity of fine features to enhance the spatial positioning ability of the local response area; the position adjustment adopts a two-stage regression method to improve the fitting ability; Semantic confidence recalibration stage: For targets with blurred boundaries and low response, a residual confidence calibration network (RCR) is constructed to correct the bias of the category distribution of detection candidates; An additional soft label smoothing mechanism is used to improve the confidence output stability of boundary objects and reduce the risk of false detection.

7. The SAR small-scale target detection system based on deep learning according to claim 1, characterized in that: The weak target recovery process of the post-processing and weak target restoration module is as follows: first, low-confidence candidate boxes and their local texture features are recorded; then, a geometric graph model is established, and a candidate graph is constructed by combining texture similarity and spatial proximity; finally, an evaluation method based on shape invariant moments or gradient direction histogram consistency is introduced to recover valid candidate objects that were mistakenly rejected; The geometric consistency check process of the post-processing and weak target restoration module is as follows: if the similarity constraint is met, that is, the distance difference is less than 5 pixels and the shape invariant moment conformity is greater than 0.85, the threshold is dynamically adjusted and the target is recycled; The non-maximum suppression (NMS) process of the post-processing and weak target restoration module is as follows: removing redundant candidate boxes and generating the final target detection results; The outputs of the post-processing and weak target restoration module are: category, bounding box coordinates and confidence, as well as recovered weak target information.

8. The SAR small-scale target detection system based on deep learning according to claim 1, characterized in that: The polarization feature proposal process of the multi-source data fusion module is as follows: the polarization reflectance PR and polarization coherence index PCI are calculated respectively; The calculation formula is as follows: where PR HH / W is the polarization reflection ratio, representing the reflection intensity ratio between the HH polarization channel and the broadband signal; S HH is the complex scattering value of the HH polarization channel, representing the scattering signal intensity when the radar wave is transmitted horizontally and received horizontally; S W is the complex scattering value of the broadband signal, representing the comprehensive scattering signal intensity of the radar wave within the entire broadband range; |S HH | 2 and |S W | 2 are the power values of the HH polarization channel and the broadband signal respectively, that is, the modulus square of the complex scattering value; PCI is the polarization coherence index, used to measure the coherence between different polarization channels; S HV is the complex scattering value of the HV polarization channel, representing the scattering signal intensity when the radar wave is transmitted horizontally and received vertically; |S HV | 2 is the power value of the HV polarization channel, that is, the modulus square of the complex scattering value, reflecting the intensity or energy of the HV polarization signal; |S HH | and |S W | are the amplitude values of the HH polarization channel and the broadband signal respectively, that is, the modulus of the complex scattering value; The feature fusion process of the multi-source data fusion module is as follows: combining polarization statistical features with SAR image guidance features and inputting them into the auxiliary channel before the deep learning network; The post-processing regularization process of the multi-source data fusion module is as follows: the edge map or foreground map in the optical image is introduced as regularization information, and soft fusion weights are constructed to improve the final prediction accuracy; The output of the multi-source data fusion module is: Fusion features and regularization weights.

9. A SAR small-scale target detection method based on deep learning, the method being based on the SAR small-scale target detection system based on deep learning according to any one of claims 1 to 8, characterized in that: The following steps are involved: Step 1: Data acquisition and preprocessing: Use millimeter-wave synthetic aperture radar (SAR) to acquire high-resolution SAR images, and simultaneously collect attitude and position information provided by the inertial navigation unit (IMU) and the global navigation satellite system (GNSS); The original SAR image is despeckled and filtered, and adaptive guided filtering (AGCF) is used to dynamically suppress coherent speckle noise and retain target details. An improved multi-scale local adaptive CFAR is used to combine gradient direction and response intensity to generate high-quality candidate regions. Step 2: The Content-Aware Feature Reconstruction Network (CAFR) constructs high-quality feature representations by extracting shallow, mid-level, and deep feature maps corresponding to different levels of semantic information. It introduces a lightweight multi-head attention transformer (LMHT) to jointly enhance position and channel information through a multi-head self-attention mechanism. It also uses a dynamic deformable fusion unit (DCAU) combined with a spatial attention mask to achieve adaptive aggregation of texture details and semantic features. Step 3: Modified Bhattacharyya distance bounding box matching mechanism MBD-Loss optimizes positioning accuracy: Model the target box as a two-dimensional Gaussian distribution; define the modified Bhattacharyya distance MBD loss function to calculate the distribution difference between the predicted box and the true box; combine the classification loss and confidence loss to form an end-to-end optimization target; Step 4: The multi-stage cascade detection structure gradually optimizes candidate target screening and regression: the graph neural network EA-GNN is used to construct a spatial similarity graph, remove isolated noise, and retain target areas with contextual structure; Using the high-resolution cross-attention module CLSR, focusing on the center of gravity of fine features, we optimize the spatial positioning ability of the local response area. We also design the residual confidence calibration network RCR to correct the bias of the category distribution of low-confidence targets and improve the confidence output stability of targets with blurred boundaries. Optimize the location and classification results of candidate targets step by step; Step 5: Weak target recovery and geometric consistency check to improve recall rate: record all low-confidence candidate boxes and their local texture features, and build a geometric graph model; for weak response candidates in the vicinity of high-confidence targets, construct a candidate graph based on texture similarity and spatial proximity; introduce an evaluation method based on shape invariant moments or gradient direction histogram consistency to determine whether the similarity constraints are met; if the conditions are met, dynamically adjust the threshold and recover the target; Step 6: Multi-source data fusion enhances cross-modality versatility: extract polarization features, including polarization reflectance (PR) and polarization coherence index (PCI); fuse edge maps or foreground maps in optical images as regularization information, and combine soft fusion weights to guide target confidence correction; Step 7: Post-processing and model deployment optimization: Perform non-maximum suppression (NMS) on the final detection results to remove redundant candidate boxes. Use TensorRT inference acceleration to optimize FP16 precision for the Jetson Xavier NX embedded AI platform, ensuring single-frame image processing latency does not exceed 90ms. Output final object detection results, including object category, bounding box coordinates, and confidence level, to meet online real-time detection requirements.

Citation Information

Cited By

  • Power transmission channel hidden danger target detection method and device, electronic equipment and storage medium

    CN121190742A

  • Illumination estimation method and system based on image semantic driving and system training method

    CN121214076A

  • Multi-target detection method and device

    CN121259350A

  • A multi-target detection method and device

    CN121259350B

  • Multi-modal interaction target detection method and system based on soft fusion supplement, and readable storage medium

    CN121392676A