Incremental few-sample target detection method for SAR (Synthetic Aperture Radar) image

By introducing the PGCI module and adaptive knowledge distillation strategy, the problems of catastrophic forgetting and speckle noise interference in SAR image target detection are solved, and efficient and accurate target detection in dynamic environments is achieved. It is suitable for edge devices with limited computing resources.

CN120689627APending Publication Date: 2025-09-23YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510798194.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

SAR image target detection in dynamic environments faces problems such as catastrophic forgetting caused by category expansion, speckle noise interference, and pre-training bias, which affect model performance and adaptability.

Method used

A physics-guided causal intervention (PGCI) module is combined with an adaptive knowledge distillation strategy to eliminate noise interference through speckle noise correction and noise-aware backdoor adjustment, and the base class knowledge is preserved in incremental learning through an adaptive mask module (AMM).

Benefits of technology

The model's detection performance and adaptability in complex environments have been significantly improved, and it can efficiently adapt to new categories under conditions of a small number of samples and limited resources while maintaining high-precision recognition of basic categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689627A_ABST
    Figure CN120689627A_ABST
Patent Text Reader

Abstract

The invention relates to an increment few-sample target detection method for an SAR image, and belongs to the field of remote sensing information processing. The problems of disastrous forgetting, speckle noise interference and pre-training deviation caused by sample scarcity and inaccessible basic categories during dynamic environment category expansion in the prior art are solved. The method comprises the following steps: data preprocessing: eliminating speckle noise through 1%-99% normalization; in the base class training, a teacher model is constructed by utilizing Deformable DETR; a physically guided causal intervention module performs speckle noise correction and noise aware backdoor adjustment (generating mask intervention features by using equivalent looks ENL); the self-adaptive knowledge distillation strategy is combined with a self-adaptive mask module, and the base class knowledge is reserved and a new class is learned through mask weighted fusion features and KL divergence loss; and new target detection is realized through class expansion. The noise suppression capability is remarkably improved, disastrous forgetting is effectively relieved, the model adaptability and robustness under the condition of few samples are improved, deployment to an edge geographic information system is facilitated, and real-time dynamic detection is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of remote sensing information processing, artificial intelligence, machine vision, etc., and provides an incremental few-sample target detection method for SAR images. Background Art

[0002] Synthetic aperture radar (SAR) has become an indispensable remote sensing technology in geographic information systems (GIS) due to its unique all-weather, all-time, and penetrating imaging capabilities. However, in practical GIS applications, object detection in SAR imagery faces the need for category expansion in dynamic environments, requiring the continuous identification of new target categories that may not have been included in the initial training phase. This demand has led to an urgent need for efficient category expansion techniques. In non-cooperative scenarios (such as disaster-stricken areas), obtaining a sufficient number of labeled samples for training traditional deep learning models is impractical. While existing deep learning methods have made significant progress in SAR object detection, they typically assume the existence of a predefined closed set of categories. However, real-world GIS applications require systems that can dynamically expand to new categories that have never been observed or included in the original base class. Furthermore, due to operational constraints, base class data may be inaccessible during the incremental update process, further complicating traditional training methods. Therefore, achieving efficient and accurate category expansion in scenarios with scarce data, inaccessible base classes, and limited computational resources is key to fully leveraging the power of SAR data in dynamic environments.

[0003] In recent years, the few-shot learning (FSL) and incremental learning paradigms have garnered significant attention. Few-shot learning aims to generalize from a small number of labeled examples by leveraging prior knowledge learned from base classes. However, existing few-shot learning methods are mostly targeted at the task of automatic target recognition (ATR), which struggles to meet the requirements of open-world detection scenarios in practical GIS applications. Furthermore, incremental few-shot object detection (iFSOD), an emerging paradigm, allows object detectors to be trained on base classes and then register an unlimited number of new classes using only a few examples, and is capable of running on resource-constrained devices. However, applying iFSOD to SAR imagery faces two major challenges: first, catastrophic forgetting due to the limited number of new class examples, inaccessible base class data, and diverse SAR imaging conditions (e.g., different acquisition angles and modes); and second, performance degradation caused by inherent speckle noise and pre-training bias in SAR data. These issues severely impact the reliability of SAR image feature extraction, resulting in poor generalization of the model to new scenarios.

[0004] In response to the above-mentioned needs and technical difficulties of SAR image target detection, the present invention proposes an incremental few-shot target detection method for SAR images, aiming to solve problems such as catastrophic forgetting and speckle noise interference, provide an efficient and accurate solution for target detection in SAR images, improve the adaptability and robustness of the system's detection performance in complex dynamic environments, and facilitate deployment in end-side geographic information systems, providing strong technical support for dynamic target detection in geographic information systems. Summary of the Invention

[0005] The present invention aims to address the performance degradation caused by catastrophic forgetting, speckle noise, and pretraining bias in SAR image target detection in dynamic environments. This paper proposes an incremental few-shot target detection method for SAR images. By introducing a physically guided causal intervention (PGCI) module and combining it with an adaptive knowledge distillation strategy, this method effectively mitigates catastrophic forgetting and improves the model's adaptability and robustness in dynamic environments.

[0006] In order to achieve the above object, the present invention adopts the following technical means:

[0007] The present invention provides an incremental few-shot target detection method for SAR images, comprising the following steps:

[0008] Step 1: Data Preprocessing

[0009] The SAR images were normalized to 1%-99% and speckle noise interference was eliminated by truncation and gray value mapping.

[0010] Step 2: Base class training

[0011] Use the Deformable DETR architecture to train the base categories and build a teacher model to provide stable features and classification outputs;

[0012] Step 3: Physically Guided Causal Intervention PGCI Module

[0013] Eliminate speckle noise and pre-training bias in SAR images through speckle noise correction and noise-aware backdoor adjustment;

[0014] Step 4: Adaptive Knowledge Distillation Strategy

[0015] Reuse the teacher model built in step 2 and combine it with the adaptive mask module AMM to retain the base class knowledge and learn new category features during the incremental learning process;

[0016] Step 5: Class Extension

[0017] Based on the PGCI module and the adaptive knowledge distillation strategy, target detection is performed on SAR images containing new classes, achieving efficient incremental learning under the conditions of inaccessible base classes and scarce samples, where the base classes refer to the set of target classes included in the initial training stage.

[0018] In the above scheme, step 1 is as follows:

[0019] Step 1: Data preprocessing: SAR images are normalized to 1%-99% and speckle noise is eliminated by truncation and grayscale mapping.

[0020] Among them, the normalization formula is:

[0021]

[0022] Where I(i, j) represents the gray value of the original image after truncation, N(i, j) represents the gray value of the transformed image, P1 is the 1% quantile in the gray value distribution of the original image, and P 99 is the 99% quantile in the grayscale value distribution of the original image;

[0023] The truncation process is specifically as follows: pixel values ​​less than P1 are uniformly adjusted to P1, and pixel values ​​greater than P 99 The pixel values ​​are uniformly adjusted to P 99 .

[0024] In the above scheme, the speckle noise correction in step 3 includes the following steps:

[0025] Step 3.al: Extract features of SAR images using pre-trained backbone network Where B is the batch size, C is the number of channels, H′ and W′ are the height and width of the feature;

[0026] Step 3.a2: Perform a two-dimensional discrete wavelet transform on the feature X and decompose it into a low-frequency component LL, a high-frequency component HL, and a high-frequency component HH. The low-frequency component LL is used to estimate the speckle noise level.

[0027] Step 3.a3: Process the LL component after global average pooling through a lightweight multi-layer perceptron MLP to predict the channel noise coefficient α∈[0,1] C , where α c represents the noise pollution level of channel c;

[0028] Step 3.a4: Use pre-trained statistics (μ clean ,σ clean ) to normalize the features and correct them by the noise coefficient α:

[0029]

[0030] Where ⊙ represents element-wise multiplication.

[0031] In the above scheme, the speckle noise correction in step 3 includes the following steps: Noise-aware backdoor adjustment includes the following sub-steps:

[0032] Step 3.b1: Calculate the local mean of the corrected features:

[0033]

[0034] Where μ X is the mean of the feature, σ X is the variance of the feature, ∈ is the minimum value to prevent division by zero;

[0035] Step 3.b2: ENL pixel Perform global average pooling GAP to obtain the sample-level ENL score

[0036] Step 3.b3: Use a learnable threshold τ = {τ1, ..., τ K-1}Disperse the ENL scores into K noise level intervals and generate a spatial mask M∈[0, 1] for each sample B×(G·K)×H′×W ′, where G represents the number of feature channel groups;

[0037] Step 3.b4: Select the corresponding mask slice M according to the discretized ENL intervalb [k], and apply it to the grouped features

[0038]

[0039] In the above scheme, step 4 includes:

[0040] Step 4.1: Construction of teacher model: Use the teacher model constructed in step 2 as the knowledge source to provide stable features f base and classification output q base ;

[0041] Step 4.2: Student model training: AMM module generates mask through convolutional network AMM ∈[0, 1] B×c×h×w Adaptively mask the area of ​​the new category so that the student model can retain the knowledge of the basic category when learning the new category, where B is the batch size, c is the number of channels, h and w are the height and width of the feature map, and the value of the mask indicates the degree of attention of the model to the new category;

[0042] Step 4.3: Transform the teacher model’s features f base With the characteristics of the student model f novel By mask weighted fusion, the feature loss is calculated:

[0043]

[0044] Step 4.4, knowledge transfer: Transfer the knowledge of the teacher model to the student model through the KL divergence loss function to ensure that the student model maintains high-precision recognition of the basic categories during the incremental learning process, where the KL divergence loss function is defined as:

[0045]

[0046] Where, represents the KL divergence, q novel and q base are the category probability distributions of the student model and the teacher model respectively;

[0047] The AMM module consists of a feature extraction network and a mask generation network, wherein the feature extraction network extracts features through convolution operations, and the mask generation network generates masks through activation functions, and performs weighted processing on the features to achieve adaptive learning of new categories.

[0048] In the above solution, step 5 includes:

[0049] Step 5.1, classifier expansion: the classifier of the teacher model obtained in the base class training phase is scaled to adapt to the detection requirements of the new category;

[0050] Step 5.2, feature extraction and intervention:

[0051] Extract high-dimensional features of input SAR images through pre-trained backbone network

[0052] The features are input into the PGCI module to perform speckle noise correction and noise-aware backdoor adjustment to obtain the post-intervention features.

[0053] Step 5.3, feature fusion and decoding:

[0054] Fuse the intervention features with the mask information output by the Adaptive Mask Module (AMM);

[0055] The Transformer encoder-decoder processes the fused features and regresses the bounding box coordinates and category probability distribution of the target;

[0056] Step 5.4: Incremental learning and deployment:

[0057] Iteratively optimize model parameters using only a small number of labeled new class samples;

[0058] The optimized model is deployed to the geographic information system of the edge device to achieve real-time target detection under conditions where the base class is inaccessible and samples are scarce.

[0059] Because the present invention adopts the above technical means, it has the following beneficial effects:

[0060] 1) In terms of improved noise processing capabilities, the PGCI module can effectively correct speckle noise in SAR images, improving the accuracy and reliability of feature extraction. Through noise-aware backdoor adjustments, it further suppresses noise interference on target detection, significantly improving the model's performance in complex noisy environments.

[0061] 2) Regarding catastrophic forgetting, the adaptive knowledge distillation strategy combined with the AMM module significantly alleviates the catastrophic forgetting problem in incremental learning. The model effectively retains base class knowledge when learning new categories, avoiding the degradation of base class performance caused by the limited number of new category samples.

[0062] 3) In terms of efficient training and generalization, with a small number of samples and limited iterations, the model can quickly adapt to new categories while maintaining high-accuracy recognition of basic categories. This method is applicable to a variety of SAR image scenarios and has good generalization and robustness.

[0063] 4) We propose an incremental few-shot object detection method that is easy to deploy and apply. This method is lightweight and easy to deploy on edge devices with limited computing resources. This enables GIS to perform dynamic object detection in real time and efficiently, providing strong technical support for practical applications.

[0064] Through the above technical means, the present invention effectively solves the key problems in SAR image target detection, significantly improves the detection performance and adaptability of the model in complex environments, and provides strong technical support for dynamic target detection in geographic information systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is the overall technical roadmap of the present invention. The technologies involved are from left to right: base class training, knowledge distillation, new class training, and GIS system deployment.

[0066] Figure 2 This is the overall architecture diagram of the class extension network in the present invention;

[0067] Figure 3 Schematic diagram of the adaptive mask module (AMM) in the present invention;

[0068] Figure 4 The structural causal model involved in the present invention, from left to right, corresponds to the structural causal model under sufficient sample conditions, under few sample conditions, and after causal intervention;

[0069] Figure 5 This is a structural diagram of the physically guided causal intervention (PGCI) module in the present invention. DETAILED DESCRIPTION

[0070] The following is a detailed description of the embodiments of the present invention. Although the present invention will be described and illustrated in conjunction with certain specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, modifications or equivalent substitutions of the present invention are intended to fall within the scope of the claims of the present invention.

[0071] In addition, in order to better illustrate the present invention, numerous specific details are given in the following detailed description. It will be understood by those skilled in the art that the present invention can also be implemented without these specific details.

[0072] The present invention provides an incremental few-shot target detection method for SAR images, comprising the following steps:

[0073] Step 1: Data preprocessing: Perform 1%-99% normalization on the SAR image to adjust the pixel value range of the image to reduce the impact of speckle noise and speed up model training;

[0074] Considering that the collected image is a single-channel grayscale image, grayscale normalization is adopted, and the grayscale value of the pixel is distributed between 0 and 255. The 1%-99% normalization method can avoid insufficient image contrast, that is, unbalanced image pixel brightness distribution, and speed up the network training speed.

[0075] The SAR image normalization formula is:

[0076]

[0077] Where I(i, j) and N(i, j) represent the grayscale value of the original image after truncation and the grayscale value of the transformed image respectively, and P 99 and P1 represent the 1% quantile and 99% quantile in the grayscale value distribution of the original image respectively. Specifically, the truncation process is to uniformly adjust the pixel values ​​less than the 1% quantile to P1, and the pixel values ​​greater than the 99% quantile to P 99 .

[0078] Step 2: Base Class Training: Use the Deformable DETR architecture to train the base classes and build a stable teacher model. This architecture combines the advantages of Transformer and convolutional neural networks to efficiently extract target features from SAR images and provide reliable base class knowledge for subsequent incremental learning.

[0079] The core of Deformable DETR lies in its unique deformable attention mechanism, which allows the model to adaptively sample on the feature map to better capture the shape and position information of the target. Specifically, the DeformableDETR architecture consists of multiple encoder and decoder layers. The encoder layer is responsible for extracting feature representations of the input image, while the decoder layer uses these features for target detection. In the encoder, the Transformer architecture fully extracts features and fuses information on the input feature map through its multi-head self-attention (MHSA) module. This mechanism enables the model to capture long-distance dependencies in the image, thereby better understanding the contextual information of the target. In the decoder, the deformable attention mechanism further enhances the model's perception of the target shape and position. Through this mechanism, the model can adaptively adjust the position of the sampling points on the feature map to better adapt to the shape changes of the target. This flexibility makes Deformable

[0080] DETR performs well in processing complex targets in SAR images.

[0081] Step 3: Physics-Guided Causal Intervention (PGCI): The SAR image is processed using the PGCI module, which includes speckle noise correction and noise-aware backdoor adjustments. This module leverages the physical properties of SAR imaging and uses a causal intervention strategy to eliminate the interference of noise and pre-training bias on target detection.

[0082] Step 4: Knowledge distillation structure, through the Adaptive Mask Module (AMM), retains the prior knowledge of the base class during incremental learning while focusing on feature learning of the new class, effectively alleviating the catastrophic forgetting phenomenon. Specific contents include:

[0083] Step 4.1: Use the teacher model built in step 2 as a knowledge source to provide stable features and classification outputs.

[0084] Step 4.2, student model training: Through the AMM module, the area of ​​the new category is adaptively masked so that the student model can retain the knowledge of the basic category when learning the new category.

[0085] Step 4.3, knowledge transfer: Transfer the knowledge of the teacher model to the student model through loss functions such as KL divergence, ensuring that the student model maintains high-precision recognition of basic categories during the incremental learning process;

[0086] Step 5, Class Expansion: Combining the PGCI module and the adaptive knowledge distillation strategy, we perform target detection on SAR images containing new classes. During the new class training phase, we use only a small number of labeled samples for incremental learning. The AMM and PGCI modules ensure that the model retains knowledge of base classes when learning new classes and effectively suppresses the interference of speckle noise.

[0087] In the above technical solution, the specific implementation algorithm of the physically guided causal intervention is as follows:

[0088] Step a. Speckle Noise Correction

[0089] Step a1. Feature extraction: Use the pre-trained backbone network to extract the features of the SAR image , where B is the batch size, C is the number of channels, H′ and W′ are the length and width dimensions of the feature;

[0090] Step a2. Two-dimensional discrete wavelet transform (2D DWT): For the feature X in step a1, perform a two-dimensional discrete wavelet transform and decompose it into low-frequency components (LL) and high-frequency components (HL, LH, HH). The low-frequency component LL is used to estimate the speckle noise level;

[0091] Step a3. Noise level estimation: a lightweight multi-layer perceptron (MLP) processes the LL component after global average pooling and predicts the channel noise coefficient α∈[0,1] C , where α c represents the noise pollution level of channel c;

[0092] Step a4. Feature correction, using pre-trained statistics (μ clean , σ clean ) to normalize the features and correct them by the noise coefficient α:

[0093]

[0094] Where ⊙ represents element-wise multiplication.

[0095] Step b. Noise-Aware Backdoor Adjustment

[0096] Step b1. Local statistics and ENL estimation, calculate the local mean of the corrected features:

[0097]

[0098] Where μ X is the mean of the feature, σ X is the variance of the feature, ∈ is the minimum value to prevent division by zero.

[0099] Step b2. Global Average Pooling (GAP): Perform global average pooling on ENL to obtain the sample-level ENL score

[0100] Step b3. Discretization and mask generation: Use a set of learnable thresholds τ = {τ1, ..., τ K-1} Divide the ENL scores into K noise level intervals. For each sample, generate a spatial mask M∈[0, 1] B×(G·K)×H′×W′ , where G represents the number of feature channel groups.

[0101] Step b4. Feature causal intervention: Select the corresponding mask slice M according to the discretized ENL interval b [k], and apply it to the grouped features

[0102]

[0103] In the above technical solution, in step 5:

[0104] The overall architecture of the class expansion network is based on the Deformable DETR structure. Based on the model rich in prior knowledge of the base classes in step 2, the class expansion network first scales the classifier according to the new class to adapt to the unseen category. In this stage, the high-dimensional features of the SAR image after passing through the feature extractor (Backbone) are first subjected to PGCI for feature causal intervention. The features are then passed through the Transformer encoder and decoder, and finally regressed to obtain the target rectangle and category information. The training process of the new class is combined with an adaptive knowledge distillation strategy. The final model can achieve efficient detection of new classes in the case of a small number of samples, inaccessible base classes, and limited iterations, while maintaining high-precision recognition performance for the base classes. The design of this method fully considers the requirements of deploying geographic information systems on edge devices with limited computing resources. By optimizing the model structure and training strategy, it ensures that the model can run efficiently and maintain good detection performance in the case of a small number of samples, inaccessible base classes, and limited iterations.

[0105] In the above technical solution, the Physics-Guided Causal Intervention (PGCI) module proposes a causal intervention strategy based on physical properties to address the problems of speckle noise and pre-training bias in SAR images. Through speckle noise correction and noise-aware backdoor adjustment, the interference of noise and pre-training bias on target detection is effectively eliminated. In the specific implementation, a two-dimensional discrete wavelet transform (2D DWT) is used to estimate the speckle noise level, and a lightweight multi-layer perceptron (MLP) is used to predict the noise coefficient, thereby correcting the features. At the same time, the equivalent number of views (ENL) is used to discretize the noise level and generate a spatial mask for feature causal intervention.

[0106] In this technical solution, an adaptive knowledge distillation strategy introduces a teacher-student model structure, combined with an adaptive masking module (AMM). This allows for incremental learning while retaining base class knowledge while focusing on learning features for new categories. The AMM module adaptively masks regions of new categories to avoid catastrophic forgetting. Using loss functions such as KL divergence, the knowledge of the teacher model is efficiently transferred to the student model, ensuring the student model's high-precision recognition of base classes during incremental learning.

[0107] Example 1

[0108] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0109] The hardware environment used for algorithm testing is Intel(R) AMD Ryzen 5975WX*2 CPU and NVIDIA GeForce RTX 4090 GPU, and the operating environment is Python 3.8 and the related extension package Pytorch.

[0110] The core algorithms of the model include Transformer target detection network, knowledge distillation, causal intervention, and incremental few-shot learning.

[0111] In particular, for the Transformer object detection network, experiments have shown that the core advantage of Deformable DETR lies in its combination of the Transformer architecture and the deformable attention mechanism, which can efficiently process image features and perform object detection. The Transformer architecture uses its Multi-Head Self-Attention (MHSA) module to fully extract features and fuse information from the input feature map. The deformable attention mechanism allows the model to adaptively sample on the feature map, thereby better capturing the shape and position information of the target. Since targets in SAR images are often affected by speckle noise and imaging conditions, the shape and position information may not be clear enough. The deformable attention mechanism is particularly suitable for processing complex targets in SAR images. Through the deformable attention mechanism, the model can more flexibly locate the target and generate accurate bounding boxes. Therefore, in this method, Deformable DETR is used as the basic architecture in the base class training stage and the class expansion stage.

[0112] like Figure 1 As shown in the figure, the basic process of the present invention begins with base class training, which uses a large amount of labeled base class data to train the basic model and build a stable teacher model. Then, combined with the physically guided causal intervention (PGCI) module, incremental learning is performed on new class data. Through an adaptive knowledge distillation strategy, the knowledge of the teacher model is transferred to the student model, ensuring that the model does not forget the base class knowledge when learning new categories. Finally, the model is deployed in a geographic information system (GIS) for real-time SAR image target detection. It includes data preprocessing, base class training, knowledge distillation, and causal intervention.

[0113] The specific steps are as follows:

[0114] The original satellite or airborne SAR image is cropped to a grayscale image of size 1024*1024. The pixel values ​​of the data are normalized to the range of 0 and 1 to accelerate model training.

[0115] The formula for grayscale transformation normalization is:

[0116]

[0117] Where I(i, j) and N(i, j) represent the grayscale value of the original image after truncation and the grayscale value of the transformed image respectively, and P 99and P1 represent the 1% quantile and 99% quantile in the grayscale value distribution of the original image respectively. Specifically, the truncation process is to uniformly adjust the pixel values ​​less than the 1% quantile (P1) to P1, and to adjust the pixel values ​​greater than the 99% quantile (P 99 ) is uniformly adjusted to P 99 .

[0118] The Deformable DETR architecture is used to train the base categories and build a stable teacher model.

[0119] refer to Figure 2 The pre-trained base category model is used as the teacher model to provide stable features and classification outputs. Through the adaptive knowledge distillation strategy and the adaptive mask module (AMM), the area of ​​the new category is adaptively masked, so that the student model can retain the knowledge of the base category when learning new categories.

[0120] refer to Figure 3 ,The implementation of the AMM module is divided into mask generation and feature loss calculation. First, the convolutional network is used to generate the mask AMM ∈[0, 1] B×c×h×w , the value of the mask indicates the degree of attention the model pays to the new category. The features fbase of the teacher model and the features f novel Calculate feature loss by mask Let the model fit the base class area at the feature level, and learn new feature expressions in the new class area:

[0121]

[0122] Where w, h, and c represent the length, width, and number of channels of the feature, respectively.

[0123] The knowledge distillation strategy uses loss functions such as KL divergence to transfer the knowledge of the teacher model to the student model, ensuring that the student model maintains high-precision recognition of the basic categories during the incremental learning process. for:

[0124]

[0125] Where, represents the KL divergence, q novel and q base are the category probability distributions of the student model and the teacher model, respectively.

[0126] In order to deeply understand the pre-training bias and the noise interference of SAR images in incremental learning of new data in few-shot learning, this method constructs a structural causal graph (SCM). Figure 4This figure graphically illustrates the causal relationships between different variables. In the SAR image target detection task, the following key variables and their causal relationships are defined: N: speckle noise, determined by the physical properties of SAR imaging; D: pre-trained knowledge, derived from a pre-trained model of a natural image dataset; X: features extracted from the SAR image, influenced by both speckle noise N and pre-trained knowledge D; C: latent representation, determined by both feature X and pre-trained knowledge D, which further influences the classification result; Y: classification label, the final output of target detection. The causal relationship between these variables is reflected in the fact that both D and N affect X, which in turn affects Y; X affects C, which in turn directly affects Y. Analysis of SCM reveals that in few-shot learning, the limited number of samples easily introduces bias, leading to inaccurate causal inference. Speckle noise N and pre-training bias D introduce additional interference paths, further exacerbating the problem.

[0127] In order to solve the above problems, the present invention uses a physically guided causal intervention (PGCI) module which is divided into speckle noise correction (Speckle Noise Correction) and noise-aware backdoor adjustment (Noise-Aware Backdoor Adjustment). Figure 5 , the algorithm implemented in combination with the physically guided causal intervention (PGCI) model body is as follows:

[0128] 1. Feature extraction, using a pre-trained backbone network to extract features of SAR images Where B is the batch size, C is the number of channels, H′ and W′ are the length and width dimensions of the feature;

[0129] 2. Two-dimensional discrete wavelet transform (2D DWT): Perform a two-dimensional discrete wavelet transform on feature X and decompose it into low-frequency components (LL) and high-frequency components (HL, LH, HH). The low-frequency component LL is used to estimate the speckle noise level.

[0130] 3. Noise level estimation, through a lightweight multi-layer perceptron (MLP) processing the LL component after global average pooling, predicting the channel noise coefficient α∈[0,1] C , where α c Indicates the noise pollution level of channel c.

[0131] 4. Feature correction, using pre-trained statistics (μ clean , σ clean ) to normalize the features and correct them by the noise coefficient α:

[0132]

[0133] Where ⊙ represents element-wise multiplication.

[0134] 5. Local statistics and ENL estimation, calculate the local mean of the corrected features:

[0135]

[0136] Where μ X is the mean of the feature, σ X is the variance of the feature, ∈ is the minimum value to prevent division by zero.

[0137] 6. Global average pooling (GAP), perform global average pooling on ENL to obtain the sample-level ENL score

[0138] 7. Discretization and mask generation, using a set of learnable thresholds τ = {τ1, ..., τ K-1} Divide the ENL scores into K noise level intervals. For each sample, generate a spatial mask M∈[0, 1] B×(G·K)×H′×W′ , where G represents the number of feature channel groups.

[0139] 8. Feature causal intervention, select the corresponding mask slice M according to the discretized ENL interval b [k], and apply it to the grouped features

[0140]

[0141] After training, the model is deployed to a GIS system on edge devices for real-time SAR imagery target detection. The model is able to run efficiently and maintain good detection performance with a small number of samples, inaccessible base classes, and limited iterations.

[0142] like Figure 2 As shown in the figure, the overall architecture of the class extension network is based on Deformable DETR. It extracts features from the input SAR image through the feature extractor (Backbone), then performs feature intervention through the PGCI module, and then processes it through the Transformer encoder and decoder. Finally, the target rectangle and category information are regressed to complete the SAR image target detection.

[0143] like Figure 3 As shown in Figure 3, the adaptive mask module structure diagram reveals the internal structure of AMM.

[0144] like Figure 4As shown in the figure, the structural causal models (from left to right) correspond to the sufficient sample condition, the few-sample condition, and the structural causal model after causal intervention. Under the sufficient sample condition, the model can learn accurate causal relationships; under the few-sample condition, speckle noise and pre-training bias introduce interference. After causal intervention, the model can eliminate interference and restore unbiased predictions through a physics-guided causal intervention strategy.

[0145] like Figure 5 The physics-guided causal intervention (PGCI) module structure diagram shows the detailed architecture of the PGCI module. The PGCI module effectively addresses the issues of speckle noise and pre-training bias in SAR images through speckle noise correction and noise-aware backdoor adjustment.

[0146] In order to facilitate those skilled in the art to better understand the technical concept of the present invention and the contribution of the present invention compared with the prior art, the relationship between the technical problem solved by the present invention, the technical solution, and the technical effects is further explained:

[0147] This paper introduces a Physics-Guided Causal Intervention (PGCI) module and an adaptive knowledge distillation strategy to effectively address issues such as catastrophic forgetting, speckle noise, and pre-training bias in incremental few-shot target detection in SAR images. This significantly improves the model's detection performance and adaptability in dynamic environments. The following provides a detailed analysis of the proposed method from three perspectives: technical approach, technical issues addressed, and beneficial effects achieved.

[0148] 1. By introducing the physically guided causal intervention (PGCI) module, the present invention effectively alleviates the problems of speckle noise interference and pre-training bias in SAR images, and improves the reliability of feature extraction and detection accuracy.

[0149] Technical means:

[0150] The present invention introduces a PGCI module in step 3, which includes two key submodules: speckle noise correction and noise-aware backdoor adjustment.

[0151] In speckle noise correction, the features are decomposed using a two-dimensional discrete wavelet transform (2D DWT), and the low-frequency component LL is used to estimate the noise level. A lightweight MLP is then used to predict the channel-level noise coefficient α. The features are then normalized and corrected to eliminate the effects of speckle noise.

[0152] In the noise-aware backdoor adjustment, the noise level is evaluated by calculating the local mean and the equivalent number of views (ENL), and a learnable threshold is used to divide the samples into different noise intervals to generate a spatial mask for causal intervention on the features.

[0153] Technical issues solved:

[0154] Speckle noise is prevalent in SAR images, and due to the diversity of imaging conditions (such as different angles and modes), image features are susceptible to noise interference, affecting the accuracy of target detection. Furthermore, pre-trained models are often based on natural image datasets (such as ImageNet), which deviate from the imaging characteristics of SAR images, further exacerbating the performance degradation of models in SAR scenarios.

[0155] Beneficial effects achieved:

[0156] Improving feature extraction reliability: The PGCI module performs noise correction and causal intervention on SAR images, effectively suppressing the interference of speckle noise and pre-training bias on feature representation, enabling the model to more accurately extract the key features of the target.

[0157] Enhanced detection accuracy: In SAR images with severe noise interference, the PGCI module significantly improves the model's ability to recognize targets, especially maintaining high detection accuracy under low signal-to-noise ratio conditions.

[0158] Enhance model robustness: Through a physics-guided causal intervention strategy, the model has stronger environmental adaptability and can cope with the variable imaging conditions and noise distribution in SAR images.

[0159] 2. This paper effectively alleviates the catastrophic forgetting problem in incremental learning by combining an adaptive knowledge distillation strategy with an adaptive mask module (AMM), ensuring that the model retains base class knowledge during the learning process of new categories.

[0160] Technical means:

[0161] The present invention introduces an adaptive knowledge distillation strategy in step 4, and combines it with the AMM module to achieve the retention of base class knowledge and the learning of new category features.

[0162] The AMM module generates adaptive masks through a convolutional network, which is used to dynamically focus on the areas of new categories during training to avoid excessive interference with the base category features.

[0163] The knowledge of the teacher model (trained based on the base class) is transferred to the student model through the KL divergence loss function, ensuring that the base class can still be recognized with high accuracy when learning new categories.

[0164] Technical issues solved:

[0165] During incremental learning, due to the limited number of new category samples and the inaccessibility of base category data, the model is prone to catastrophic forgetting, that is, when learning new categories, the model's ability to recognize base categories decreases significantly. This severely limits the model's application in dynamic environments.

[0166] Beneficial effects achieved:

[0167] Mitigating catastrophic forgetting: The AMM module adaptively masks regions of new categories, preventing the model from over-focusing on new categories and neglecting base category features during learning. Furthermore, the knowledge distillation strategy transfers prior knowledge from the teacher model to the student model, further strengthening base category recognition capabilities.

[0168] Improving incremental learning efficiency: With a small number of samples, the combination of the AMM module and the knowledge distillation strategy enables the model to quickly adapt to new categories while maintaining high-precision recognition of the base categories.

[0169] Enhanced model generalization capability: By retaining base class knowledge, the model can still maintain a certain level of generalization capability when facing unseen categories, improving the system's adaptability in dynamic environments.

[0170] 3. By combining the PGCI module with the adaptive knowledge distillation strategy, this paper achieves efficient incremental learning under conditions where the base class is inaccessible and samples are scarce, thereby improving the deployment flexibility and applicability of the model.

[0171] Technical means:

[0172] In step 5, the present invention proposes an incremental few-sample target detection method, which combines the PGCI module with the adaptive knowledge distillation strategy to achieve efficient detection of new categories under the conditions that the base class is inaccessible and samples are scarce.

[0173] In the class expansion stage, only a small number of labeled new category samples are used for training, and the PGCI module is used to perform noise correction and causal intervention on the input image to improve feature quality.

[0174] Through the AMM module and knowledge distillation strategy, the base class knowledge is retained during the incremental learning process, ensuring that the model does not forget old knowledge when training new categories.

[0175] Technical issues solved:

[0176] In practical applications, SAR image target detection often faces the following challenges:

[0177] Inaccessible base class data: During the incremental update process, the base class data from the original training phase may not be available.

[0178] Sample scarcity: In non-cooperative scenarios (such as disaster-stricken areas), it is difficult to obtain a sufficient number of labeled samples for model training.

[0179] Limited computing resources: In edge devices (such as GIS systems), models must be lightweight and have efficient reasoning capabilities.

[0180] Beneficial effects achieved:

[0181] Adaptability to dynamic environment requirements: The present invention can achieve efficient detection of new categories under the conditions where base classes are inaccessible and samples are scarce, thus meeting the real-time target detection requirements of GIS systems in dynamic environments.

[0182] Improved deployment flexibility: Through lightweight design and efficient training strategies, the model can be deployed to GIS systems in edge devices, achieving low-power, high-efficiency real-time detection.

[0183] Enhanced model practicality: This method is not only applicable to SAR image target detection, but can also be extended to other remote sensing and machine vision tasks with similar challenges, and has good application prospects.

Claims

1. An incremental few-shot target detection method for SAR images, characterized in that: The following steps are involved: Step 1: Data Preprocessing The SAR images were normalized to 1%-99% and speckle noise interference was eliminated by truncation and gray value mapping. Step 2: Base class training Use the Deformable DETR architecture to train the base categories and build a teacher model to provide stable features and classification outputs; Step 3: Physically Guided Causal Intervention PGCI Module Eliminate speckle noise and pre-training bias in SAR images through speckle noise correction and noise-aware backdoor adjustment; Step 4: Adaptive Knowledge Distillation Strategy Reuse the teacher model built in step 2 and combine it with the adaptive mask module AMM to retain the base class knowledge and learn new category features during the incremental learning process; Step 5: Class Extension Based on the PGCI module and the adaptive knowledge distillation strategy, target detection is performed on SAR images containing new classes, achieving efficient incremental learning under the conditions of inaccessible base classes and scarce samples, where the base classes refer to the set of target classes included in the initial training stage.

2. The method according to claim 1, characterized in that Step 1: Step 1: Data preprocessing: SAR images are normalized to 1%-99% and speckle noise is eliminated by truncation and grayscale mapping. Among them, the normalization formula is: Where I(i, j) represents the gray value of the original image after truncation, N(i, j) represents the gray value of the transformed image, P1 is the 1% quantile in the gray value distribution of the original image, and P 99 is the 99% quantile in the grayscale value distribution of the original image; The truncation process is specifically as follows: pixel values ​​less than P1 are uniformly adjusted to P1, and pixel values ​​greater than P 99 The pixel values ​​are uniformly adjusted to P 99 .

3. The method according to claim 1, characterized in that The speckle noise correction in step 3 includes the following steps: Step 3.a1: Extract features from SAR images using a pre-trained backbone network Where B is the batch size, C is the number of channels, H′ and W′ are the height and width of the feature; Step 3.a2: Perform a two-dimensional discrete wavelet transform on the feature X and decompose it into a low-frequency component LL, a high-frequency component HL, and a high-frequency component HH. The low-frequency component LL is used to estimate the speckle noise level. Step 3.a3: Process the LL component after global average pooling through a lightweight multi-layer perceptron MLP to predict the channel noise coefficient α∈[0,1] C , where α c represents the noise pollution level of channel c; Step 3.a4: Use pre-trained statistics (μ clean ,σ clean ) to normalize the features and correct them by the noise coefficient α: Where o represents element-by-element multiplication.

4. The method according to claim 3, characterized in that The speckle noise correction described in step 3 includes the following steps: Noise-aware backdoor adjustment includes the following sub-steps: Step 3.b1: Calculate the local mean of the corrected features: Where μ X is the mean of the feature, σ X is the variance of the feature, ∈ is the minimum value to prevent division by zero; Step 3.b2: ENL pixel Perform global average pooling GAP to obtain the sample-level ENL score Step 3.b3: Use a learnable threshold τ = {τ1, ..., τ K-1 }Disperse the ENL scores into K noise level intervals and generate a spatial mask M∈[0, 1] for each sample B×(G·K)×H×W , where G represents the number of feature channel groups; Step 3.b4: Select the corresponding mask slice M according to the discretized ENL interval b [k], and apply it to the grouped features 5. The method according to claim 1, wherein Step 4 includes: Step 4.1: Construction of teacher model: Use the teacher model constructed in step 2 as the knowledge source to provide stable features r base and classification output q base ; Step 4.2: Student model training: AMM module generates mask through convolutional network AMM ∈[0, 1] B×c×h×w Adaptively mask the area of ​​the new category so that the student model can retain the knowledge of the basic category when learning the new category, where B is the batch size, c is the number of channels, h and w are the height and width of the feature map, and the value of the mask indicates the degree of attention of the model to the new category; Step 4.3: Transform the teacher model’s features f base With the characteristics of the student model f novel By mask weighted fusion, the feature loss is calculated: Step 4.4, knowledge transfer: Transfer the knowledge of the teacher model to the student model through the KL divergence loss function to ensure that the student model maintains high-precision recognition of the basic categories during the incremental learning process, where the KL divergence loss function is defined as: Where, represents the KL divergence, q novel and q base are the category probability distributions of the student model and the teacher model respectively; The AMM module consists of a feature extraction network and a mask generation network, wherein the feature extraction network extracts features through convolution operations, and the mask generation network generates masks through activation functions, and performs weighted processing on the features to achieve adaptive learning of new categories.

6. The method according to claim 1, characterized in that Step 5 includes: Step 5.1, classifier expansion: the classifier of the teacher model obtained in the base class training phase is scaled to adapt to the detection requirements of the new category; Step 5.2, feature extraction and intervention: Extract high-dimensional features of input SAR images through pre-trained backbone network The features are input into the PGCI module to perform speckle noise correction and noise-aware backdoor adjustment to obtain the post-intervention features. Step 5.3, feature fusion and decoding: Fuse the intervention features with the mask information output by the Adaptive Mask Module (AMM); The Transformer encoder-decoder processes the fused features and regresses the bounding box coordinates and category probability distribution of the target; Step 5.4: Incremental learning and deployment: Iteratively optimize model parameters using only a small number of labeled new class samples; The optimized model is deployed to the geographic information system of the edge device to achieve real-time target detection under conditions where the base class is inaccessible and samples are scarce.