Generalized SAR (Synthetic Aperture Radar) target detection method for cross-source scene unified framework
Through dynamic statistical matching and feature decoupling modules, the SAR image features are optimized, which solves the problem of insufficient generalization ability of SAR image object detection in cross-domain scenarios, and achieves higher detection accuracy and adaptability.
Patent Information
- Application Number
- CN202510575221.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-05
AI Technical Summary
The existing SAR image object detection methods lack generalization capabilities in cross-domain scenarios, which are affected by imaging conditions and scarcity of labeled samples, resulting in a decline in domain offset and detection performance.
By defining the scattering characteristics of SAR images, dynamically match the characteristics and designing the domain-invariant features and domain-specific features decoupling module, eliminating the differences between domains, optimizing the feature distribution, and using the Faster RCNN baseline network for detection.
It enhances the adaptability of the model to unknown domains, effectively eliminates domain offsets, and improves the generalization ability and detection accuracy of SAR target detection.
Smart Images

Figure CN120431320A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing target detection, and in particular to a generalized SAR target detection method with a unified framework for cross-source scenarios. Background Art
[0002] Target detection in Synthetic Aperture Radar (SAR) images holds significant application value in fields such as military reconnaissance and disaster monitoring. In recent years, with breakthroughs in deep learning technology, target detection methods based on Convolutional Neural Networks (CNNs) have made significant progress in the SAR field. Common frameworks such as Faster RCNN and You Only Look Once (YOLO) have significantly improved detection accuracy through end-to-end feature learning. However, the inherent imaging characteristics of SAR images, such as speckle noise, viewpoint sensitivity, and complex target scattering, as well as the diversity of practical application scenarios, pose significant challenges to the generalization ability of existing methods in cross-domain scenarios.
[0003] By integrating multi-scale features, designing attention mechanisms, and implementing data augmentation strategies, researchers have gradually improved the model's robustness to scale variations and noise in SAR images. Furthermore, the construction of public datasets, such as Moving and Stationary Target Acquisition and Recognition (MSTAR) and OpenSAR, provides critical support for algorithm training and validation. However, existing techniques still have significant limitations. First, SAR data is significantly affected by imaging conditions such as resolution, angle of incidence, polarization mode, and geographic environment, resulting in a significant domain shift between training and test sets, which severely degrades model performance across domains. Second, SAR image annotation relies heavily on specialized prior knowledge. Due to data confidentiality and acquisition costs, the scarcity of annotated samples and the lack of domain diversity severely limit the model's generalization ability. For example, a detection model trained for near-shore vessels may fail in open ocean scenarios with complex sea conditions due to differences in background clutter, while a detection algorithm for urban buildings may generate false positives due to interference from mountainous terrain.
[0004] The essence of these challenges can be summarized as the domain generalization (DG) problem, which focuses on how to enable a model to maintain stable detection performance in an unknown target domain when exposed to only limited source domain data. In recent years, DG technology has become a key direction for addressing this challenge by synergizing the extraction of domain-invariant features and the modeling of domain-specific features. Traditional domain adaptation methods rely on target domain data, while DG requires cross-domain generalization when the target domain is unknown. It is difficult to achieve good results through feature alignment alone, which places higher demands on feature decoupling capabilities. Summary of the Invention
[0005] The purpose of this invention is to solve the problem of higher feature decoupling capability requirements and enhance the adaptability of the model to unknown domain distributions, thereby providing a new solution for SAR target detection in complex scenarios. A generalized SAR target detection method with a unified framework for cross-source scenarios is proposed, which includes the following steps:
[0006] Step 1: define the scattering characteristics of the SAR image, and expand the domain feature space by dynamically statistically matching the SAR image scattering characteristics;
[0007] Step 2: Design a domain-invariant feature and domain-specific feature decoupling module for the feature space, and remove redundant domain discrimination information to achieve final feature optimization.
[0008] Furthermore, the step 1 is specifically as follows:
[0009] The source domain SAR image is defined as
[0010] in, It is in N s The i-th image in the fully annotated SAR images, express The corresponding category label, In this source domain there are N s images, respectively, with x i and y i Represents the i-th image and its label information; at the same time, is the jth sample in the i-th image, where and are the category and bounding box of the sample respectively;
[0011] The target domain for testing is denoted as Among them, x i is the i-th image, there are N t images.
[0012] Define a non-empty input space and any output space The goal of the algorithm is to train a model using labeled source domain data S and ensure that the model is in the unknown distribution Target domain Good performance on.
[0013] Furthermore, the step 1 is specifically as follows:
[0014] The dynamic statistical matching SAR image scattering characteristics are specifically:
[0015] In the encoder, the domain distribution of the SAR image is expanded by changing the statistical characteristics of the SAR image to simulate other domains, and the initial information of the SAR image in the domain is complete;
[0016] By aligning statistical parameters and adjusting feature space, we can eliminate the differences between domains. We can set the statistical characteristics of the mean and variance of SAR images in a specific domain. Reflects the backscattering intensity, variance Reflecting the intensity of texture or speckle noise, the statistical parameter alignment and feature space adjustment are expressed as:
[0017]
[0018] Where μ represents the mean of feature f, σ represents the variance of feature f, B, C, H and W represent the sizes of batch, channel, height and width respectively, and b, c, h and w represent the values of batch, channel, height and width respectively;
[0019] Convolution is performed through convolutional layers. Assuming that after some convolution blocks, the corresponding feature map is expressed as
[0020] Considering the distribution differences between multiple domains, a linear transformation is used to map each source domain data into the statistical space of other domains.
[0021]
[0022] Among them, σ s′ is the mean value mapped to another domain, μ s′ is the variance mapped to another domain;
[0023] Subsequently, an adaptive batch normalization (AdaBN) layer is inserted into the model. In the AdaBN network model, the batch normalization parameters of the source domain are replaced by the mean and variance of other source domains to dynamically adapt to the attribute dispersion characteristics of each source domain.
[0024] Furthermore, the step 2 is specifically as follows:
[0025] For the matching SAR image features, a feature decoupling method is designed to separate domain-invariant features from domain-specific features. The feature decoupling is internally divided into two branches: a domain-invariant feature branch consists of a channel attention module, which uses a convolution block consisting of a 1×1 convolution layer, a BN layer, and a ReLU activation function. The channel attention module focuses on the global feature representation of the SAR image. The operation of one branch is as follows:
[0026]
[0027] Among them, δ represents the sigmoid function, F con , F avg and F max are convolutional layers, representing a convolutional layer, a channel average pooling layer, and a channel maximum pooling layer. represents the function combination operator, ⊕ represents the channel dimension concatenation, and ⊙ is the Hadamard product;
[0028] The other branch is used to extract domain-specific features, including a spatial attention module that focuses on local salient representations related to the detected objects through the following equation:
[0029]
[0030] in, It is a tensor product operator. The domain-invariant features and domain-specific features obtained by the above method have the same dimensions as their inputs.
[0031] Generally, the larger the cosine distance between domain-specific features and domain-invariant features, the more non-category information the domain-invariant features contain; therefore, and Both are mapped into the embedding space by the visual projector Vp, whose weights are pre-trained and initialized from CLIP, and then two corresponding embedding vectors can be obtained. and Where R is the dimension of the region of interest and Ne is the embedding dimension; the following equation is designed to optimize the mutual information between these two vectors:
[0032]
[0033] Specifically, the above equation indirectly encourages and The cosine similarity between them decreases;
[0034] In addition, in order to impose regularization constraints on the feature decoupling module, which can stably separate domain-specific features and domain-invariant features in different scenarios, the above mutual information optimization objective function is further constrained to the following form:
[0035]
[0036] Here, γ is a trade-off hyperparameter that controls the learning strategy.
[0037] Furthermore, the step 2 is specifically as follows: the γ is set to 0.7.
[0038] Furthermore, the step 2 is specifically as follows:
[0039] Remove domain discriminative information from the above decoupled domain invariant features and train the domain decoupler through domain classification loss To learn domain discriminative features:
[0040]
[0041] in, is the cross entropy loss function;
[0042] The detection loss function is defined by Faster RCNN, which includes the classification loss and bounding box regression loss generated by the Region Proposal Network (RPN), as well as the loss generated by the Region Proposal Classifier (RPC). The detection loss function is:
[0043]
[0044] Based on the baseline network architecture of Faster RCNN, the total loss function is:
[0045]
[0046] Among them, α and β are hyperparameters used to weigh the corresponding loss functions.
[0047] Furthermore, the step 2 is specifically as follows: the α is set to 0.7, and the β is set to 0.4.
[0048] The beneficial effect of the invention is that this paper proposes a unified target detection framework based on the decoupling and optimization of domain-invariant features and domain-specific features, aiming to enhance the model's adaptability to unknown domain distributions, thereby providing a novel solution for SAR target detection in complex scenarios. This optimized feature distribution profile is sufficient to eliminate domain shift and complete downstream detection tasks in the detection head of the baseline network. Experimental results show that the proposed model framework runs on the FasterRCNN baseline network and obtains target detection results through its downstream detection head. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram of the overall unified framework object detection of DG heterogeneous SAR images of the present invention;
[0050] Figure 2 are four ship images in the corresponding dataset, where Figure 2 (a) is AIR-SARShip, Figure 2 (b) is HRSID, Figure 2 (c) Ships, Figure 2 (d) is SMCDD.
[0051] Figure 3 is the DG object detection result of the HRSID dataset as the target domain, where Figure 3 (a)- Figure 3 In (e), green, red, and blue represent correct detection, false detection, and missed detection, respectively. Figure 3 (a) is DGOD, Figure 3 (b) is GEFS, Figure 3 (c) is ADGU, Figure 3 (d) is CLIP, Figure 3 (e) for our Figure 3 (f) is the ground truth.
[0052] Figure 4 The DG object detection results with the SMCDD dataset as the target domain are shown in Figure 4 (a)- Figure 4 In (e), green, red, and blue represent correct detection, false detection, and missed detection, respectively; Figure 4 (a) is DGOD, Figure 4 (b) is GEFS, Figure 4 (c) is ADGU, Figure 4 (d) is CLIP, Figure 4 (e) for our Figure 4 (f) is the ground truth.
[0053] Figure 5In order to use t-SNE to visualize the features of the four dataset domains, Figure 5 (a) and Figure 5 In (b), red, blue, green, and yellow represent Ships, AIR SARShip, SMCDD, and HRSID, respectively. Figure 5 (a) is the initial feature distribution, Figure 5 (b) is the feature distribution after DG. DETAILED DESCRIPTION
[0054] The present invention will be further described below with reference to the accompanying drawings and examples of the present invention.
[0055] To address the problem of higher feature decoupling capability requirements and enhance the model's adaptability to unknown domain distributions, thereby providing a new solution for SAR target detection in complex scenarios, the present invention provides a generalized SAR target detection method with a unified framework for cross-source scenarios, comprising the following steps:
[0056] Step 1: define the scattering characteristics of the SAR image, and expand the domain feature space by dynamically statistically matching the SAR image scattering characteristics;
[0057] Step 2: Design a domain-invariant feature and domain-specific feature decoupling module for the feature space, and remove redundant domain discrimination information to achieve final feature optimization.
[0058] The step 1 is specifically as follows:
[0059] The source domain SAR image is defined as
[0060] in, It is in N s The i-th image in the fully annotated SAR images, express The corresponding category label, In this source domain there are N s images, respectively, with x i and y i Represents the i-th image and its label information; at the same time, is the jth sample in the i-th image, where and are the category and bounding box of the sample respectively;
[0061] The target domain for testing is denoted as Among them, x i is the i-th image, there are N t images.
[0062] Define a non-empty input space and any output space The goal of the algorithm is to train a model using labeled source domain data S and ensure that the model is in the unknown distribution Target domain At the same time, it should be noted that the unknown distribution Not equal to the distribution of any source domain S The proposed generalized SAR target detection framework is as follows Figure 1 As shown, the training phase involves using three datasets as source domains together as input, and when the iterative optimization of the network model parameters is completed, another dataset is used as the target domain for testing.
[0063] The step 1 is specifically as follows:
[0064] The dynamic statistical matching SAR image scattering characteristics are specifically:
[0065] Imaging results from different devices have different domain distributions, which are related to the scattering properties that characterize different SAR images. These distributions can be characterized by statistical metrics in the image feature maps. There is a large domain gap between different domains, which is caused by the different distributions. Therefore, in the encoder, the domain distribution of the SAR image is expanded by changing its statistical features to simulate other domains. At the same time, the feature decoupling strategy used in the next step of network processing ensures that the initial SAR image information in the domain is complete when processing the SAR image's statistical features in this step.
[0066] In the encoder, the domain distribution of the SAR image is expanded by changing the statistical characteristics of the SAR image to simulate other domains, and the initial information of the SAR image in the domain is complete; the differences between domains are eliminated by statistical parameter alignment and feature space adjustment; the statistical characteristics of the mean and variance of the SAR image in a specific domain are set, where the mean Reflects the backscattering intensity, variance Reflecting the intensity of texture or speckle noise, the statistical parameter alignment and feature space adjustment are expressed as:
[0067]
[0068] Where μ represents the mean of feature f, σ represents the variance of feature f, B, C, H and W represent the sizes of batch, channel, height and width respectively, and b, c, h and w represent the specific parameters of the sizes of batch, channel, height and width respectively;
[0069] Convolution is performed through convolutional layers. Assuming that after some convolution blocks, the corresponding feature map is expressed as
[0070] Considering the distribution differences between multiple domains, a linear transformation is used to map each source domain data into the statistical space of other domains.
[0071]
[0072] Among them, σ s′ is the mean value mapped to another domain, μ s′ is the variance mapped to another domain;
[0073] Subsequently, an adaptive batch normalization (AdaBN) layer is inserted into the model. In the AdaBN network model, the batch normalization parameters of the source domain are replaced by the mean and variance of other source domains to dynamically adapt to the attribute dispersion characteristics of each source domain.
[0074] The step 2 is specifically as follows:
[0075] For the matching SAR image features, a feature decoupling method is designed to separate domain-invariant features from domain-specific features. The feature decoupling is internally divided into two branches: a domain-invariant feature branch consists of a channel attention module, which uses a convolution block consisting of a 1×1 convolution layer, a BN layer, and a ReLU activation function. The channel attention module focuses on the global feature representation of the SAR image. The operation of one branch is as follows:
[0076]
[0077] Among them, δ represents the sigmoid function, F con ,F avg and F max are convolutional layers, representing a convolutional layer, a channel average pooling layer, and a channel maximum pooling layer. represents the function combination operator, ⊕ represents the channel dimension concatenation, and ⊙ is the Hadamard product;
[0078] The other branch is used to extract domain-specific features, including a spatial attention module that focuses on local salient representations related to the detected objects through the following equation:
[0079]
[0080] in, It is a tensor product operator. The domain-invariant features and domain-specific features obtained by the above method have the same dimensions as their inputs.
[0081] Generally, the larger the cosine distance between domain-specific features and domain-invariant features, the more non-category information the domain-invariant features contain; therefore, and Both are mapped into the embedding space through the visual projector Vp, whose weights are pre-trained and initialized from CLIP (Contrastive Language-Image Pre-training, Learning Transferable Visual Models From Natural Language Supervision), and then two corresponding embedding vectors can be obtained. and Where R is the dimension of the region of interest and Ne is the embedding dimension; if these two vectors are orthogonal, it means that they are semantically unrelated. The following equation is designed to optimize the mutual information between these two vectors:
[0082]
[0083] Specifically, the above equation indirectly encourages and The cosine similarity between them decreases;
[0084] In addition, in order to impose regularization constraints on the feature decoupling module, which can stably separate domain-specific features and domain-invariant features in different scenarios, the above mutual information optimization objective function is further constrained to the following form:
[0085]
[0086] Among them, γ is a trade-off hyperparameter that controls the learning strategy, and the γ is set to 0.7.
[0087] The step 2 is specifically as follows:
[0088] Remove domain discriminative information, such as domain shift, from the above decoupled domain invariant features and train the domain decoupler through domain classification loss To learn domain discriminative features:
[0089]
[0090] in, is the cross entropy loss function;
[0091] The detection loss function is defined by Faster RCNN, which includes the classification loss and bounding box regression loss generated by the Region Proposal Network (RPN), as well as the loss generated by the Region Proposal Classifier (RPC). The detection loss function is:
[0092]
[0093] Based on the baseline network architecture of Faster RCNN, the total loss function is:
[0094]
[0095] Among them, α and β are hyperparameters used to weigh the corresponding loss functions. The α is set to 0.7 and the β is set to 0.4.
[0096] Experimental results and analysis
[0097] a. Data and Experimental Setup
[0098] Four different remote sensing image ship datasets, namely one optical remote sensing dataset and three SAR datasets, are used to verify the proposed algorithm: HRSID, AIR-SARShip, SMCDD, and ship. These four datasets are the imaging results of sea surface ships obtained by different imaging platforms in different scenarios. Figure 2 Four ship images from the corresponding datasets are shown in Figure 2, showing significant domain differences between them. We conducted experiments on these four datasets. Since different remote sensing datasets come from different platforms, three of the four datasets were used as source domains for the experiment. The results of exploring how each dataset, when used as a target domain, can be generalized using the other three datasets as source domains.
[0099] Faster R-CNN was used as the baseline network, while ResNet-101 was used as the backbone network for the entire model. During training, stochastic gradient descent with a momentum of 0.9 was used as the optimization method, and the initial learning rate was set to 0.001. The trade-off hyperparameters α and β in both equations were empirically chosen to be 0.7 and 0.4, respectively. A batch size of 4 was used for training, with a total of 50 epochs. Similar to many existing SAR target detection methods, all experiments used a uniform intersection-over-union (IoU) threshold of 0.5 to evaluate performance. To evaluate detector performance, mean average precision (mAP), precision (PR), recall (RE), and F1 score were used to measure detector quality. All experiments were conducted using PyTorch 2.0 on a workstation equipped with a 14-vCPU Intel(R) Xeon(R) Gold 6348 CPU and an Nvidia Tesla A800 Turbo GPU.
[0100] b. Comparative experimental results analysis
[0101] We validate the object detection performance of the proposed method from both quantitative and qualitative perspectives. First, Table 1 compares the object detection performance of the proposed method with the state-of-the-art (SOTA) methods in the generalization experiments described above, showing the object detection performance when the four datasets are used as target domains. The compared SOTA methods include domain generalization object detection methods in remote sensing and vision, such as Domain Generalized Object Detection (DGOD), Generalization-Enhanced Few-Shot (GEFS), Achieving Domain Generalization for Underwater (ADGU), and CLIP (Contrastive Language-Image Pre-training).
[0102] Table 1 compares the target detection performance of the proposed method and the SOTA method in the generalization experiment above.
[0103] TABLE I
[0104] COMPARATIVE EXPERIMENTAL RESULTS OF DOMAIN GENERALIZED OBJECTDETECTION(%)
[0105]
[0106] The experimental results presented in Table 1 show that the proposed method surpasses state-of-the-art methods in object detection across all four DG experiments. Because we use three source domains, the model is able to capture more comprehensive information about image object features. Domain-invariant features focus on interesting foregrounds, while domain-specific features are related to backgrounds. This effective decoupling allows these two functions to function better and be effective in heterogeneous image object detection tasks. Despite significant data variability, including and excluding optical remote sensing datasets, the proposed method maintains a detection rate advantage of over 4% over the state-of-the-art method.
[0107] Therefore, when all three source domains are SAR data and the experiments are conducted on the target domain of optical remote sensing data, the detection performance will drop significantly, but the proposed method is always the best performer.
[0108] In addition to quantitatively comparing the detection results of each method, we also selected two typical scenarios to demonstrate Figure 3 and Figure 4 Detection results in Figure 2. For the sake of generality, the object detection results of the comparison methods are shown here, using the target domains HRSID and SMCDD as examples. It can be seen that the proposed method achieves the closest correct detections to the ground truth and has the lowest false detection rate. Furthermore, especially for objects like ships, extracting features such as contours and textures from the foreground and background, focusing on domain-invariant and domain-specific features, respectively, is crucial for correct object detection.
[0109] C. Stripping Studies and Related Analysis
[0110] 1) Dynamic Statistics of Image Features: Table 2 shows the results of network models targeting AIR-SARship and SMCDD, respectively. Compared to the baseline model and the full approach, the decoupling of domain-invariant and domain-specific features without expanding the feature space using dynamic statistics is achieved. As can be seen, compared to the baseline network, mAP still improves by over 7% and the F1 score by over 6%, demonstrating that generalization performance can be improved simply by partitioning features into task-dependent and task-independent components. Furthermore, the experiment shows that performing this step improves object detection mAP by 4.6% and 5.4% in the two tasks, respectively.
[0111] Table 2 Results of network models with AIR-SARship and SMCDD as target domains
[0112] TABLE II
[0113] OBJECT DETECTION RESULTS(%)OF DYNAMIC STATISTICS
[0114]
[0115] 2) Backbone network: In all experiments, ResNet-101 has been the backbone network model. However, in fact, we all know that in addition to the 101 layer, other parameter layer backbone networks of ResNet can also be used. Therefore, an ablation experiment is conducted here to verify whether the ResNet-101 used is the most suitable backbone network model. Table 3 shows the target detection results of different backbone networks (taking two domain generalization tasks as an example, the target domains are HRSID and Ships respectively). It can be clearly seen that the ResNet-101 backbone network can achieve the best target detection performance, which is consistent with our understanding, because when used for Faster R-CNN, ResNet's more parameters can optimize the network parameters to achieve the best performance.
[0116] Table 3. Object detection results of different backbone networks
[0117] TABLE III
[0118] OBJECT DETECTION RESULTs(%)OF DIFFERENT BACKBONE NETWORKS
[0119]
[0120] 3) Feature Visualization: SAR images captured by different devices exhibit different stylistic and texture features due to differences in radar parameters, imaging modes, viewing angles, and scenes. The styles of different images can be characterized by statistical metrics of image feature maps. Here, we visualize the four datasets selected in the experiment as object features on a two-dimensional plane using t-distributed stochastic neighbor embedding (t-SNE). Figure 5 The feature visualization results for the initial feature distributions of the four data domains and the reduction of the domain gap after domain generalization (the target domain is HRSID) are shown. It can be seen that initially, there is a large gap in the feature distributions of the four datasets. However, after domain generalization, the domain shift between them is significantly reduced, demonstrating the effectiveness of the proposed method in enhancing generalization performance on heterogeneous data.
[0121] This paper proposes a unified object detection framework based on the decoupled optimization of domain-invariant and domain-specific features. This framework aims to enhance the model's adaptability to unknown domain distributions, thereby providing a novel solution for SAR target detection in complex scenarios. This optimized feature distribution profile is sufficient to eliminate domain shift and complete downstream detection tasks within the baseline network's detection head. Experimental results demonstrate that the proposed model framework runs on the Faster R-CNN baseline network and achieves excellent object detection results through its downstream detection head.
Claims
1. A generalized SAR target detection method with a unified framework for cross-source scenarios, comprising the following steps: Step 1: define the scattering characteristics of the SAR image, and expand the domain feature space by dynamically statistically matching the SAR image scattering characteristics; Step 2: Design a domain-invariant feature and domain-specific feature decoupling module for the feature space, and remove redundant domain discrimination information to achieve final feature optimization. Furthermore, the step 1 is specifically as follows: The source domain SAR image is defined as in, It is in N s The i-th image in the fully annotated SAR images, express The corresponding category label, In this source domain there are N s images, respectively, with x i and y i Represents the i-th image and its label information; at the same time, is the jth sample in the i-th image, where and are the category and bounding box of the sample respectively; The target domain for testing is denoted as Among them, x i is the i-th image, there are N t images. Define a non-empty input space and any output space The goal of the algorithm is to train a model using labeled source domain data S and ensure that the model is in the unknown distribution Target domain Good performance on.
2. The generalized SAR target detection method with a unified framework for cross-source scenarios according to claim 1, wherein step 1 is specifically: The dynamic statistical matching SAR image scattering characteristics are specifically: In the encoder, the domain distribution of the SAR image is expanded by changing the statistical characteristics of the SAR image to simulate other domains, and the initial information of the SAR image in the domain is complete; By aligning statistical parameters and adjusting feature space, we can eliminate the differences between domains. We set the statistical characteristics of the mean and variance of SAR images in a specific domain, where: mean Reflects the backscattering intensity, variance Reflecting the intensity of texture or speckle noise, the statistical parameter alignment and feature space adjustment are expressed as: Where μ represents the mean of feature f, σ represents the variance of feature f, B, C, H and W represent the size of batch, channel, height and width respectively, b, c, h and w represent the numerical value of the size of batch, channel, height and width respectively; Convolution is performed through convolutional layers. Assuming that after some convolution blocks, the corresponding feature map is expressed as Considering the distribution differences between multiple domains, a linear transformation is used to map each source domain data into the statistical space of other domains. Among them, σ s′ is the mean value mapped to another domain, μ s′ is the variance mapped to another domain; Subsequently, the AdaBN layer is inserted into the model. In the AdaBN network model, the batch normalization parameters of the source domain are replaced by the mean and variance of other source domains to dynamically adapt to the attribute dispersion characteristics of each source domain.
3. In the generalized SAR target detection method with a unified framework for cross-source scenarios according to claim 1, step 2 specifically comprises: For the matching SAR image features, a feature decoupling method is designed to separate domain-invariant features from domain-specific features. The feature decoupling is internally divided into two branches: a domain-invariant feature branch consists of a channel attention module, which uses a convolution block consisting of a 1×1 convolution layer, a BN layer, and a ReLU activation function. The channel attention module focuses on the global feature representation of the SAR image. The operation of one branch is as follows: in, δ represents the sigmoid function, F con ,F avg and F max is a convolutional layer, representing a convolutional layer, a channel average pooling layer and a channel maximum pooling layer, o represents a function combination operator, ⊕ represents channel size cascade, and ⊙ is the Hadamard product; The other branch is used to extract domain-specific features, including a spatial attention module that focuses on local salient representations related to the detected objects through the following equation: in, It is a tensor product operator. The domain-invariant features and domain-specific features obtained by the above method have the same dimensions as their inputs. Generally, the larger the cosine distance between domain-specific features and domain-invariant features, the more non-category information the domain-invariant features contain; therefore, and Both are mapped into the embedding space by the visual projector Vp, whose weights are pre-trained and initialized from CLIP, and then two corresponding embedding vectors can be obtained. and Where R is the dimension of the region of interest and Ne is the embedding dimension; the following equation is designed to optimize the mutual information between these two vectors: Specifically, the above equation indirectly encourages and The cosine similarity between them decreases; In addition, in order to impose regularization constraints on the feature decoupling module, which can stably separate domain-specific features and domain-invariant features in different scenarios, the above mutual information optimization objective function is further constrained to the following form: Here, γ is a trade-off hyperparameter that controls the learning strategy.
4. The generalized SAR target detection method with a unified framework for cross-source scenarios as claimed in claim 3, wherein the step 2 specifically comprises: setting γ to 0.
7.
5. The generalized SAR target detection method with a unified framework for cross-source scenarios according to claim 2, wherein the second step is specifically: Remove domain discriminative information from the above decoupled domain invariant features and train the domain decoupler through domain classification loss To learn domain discriminative features: in, is the cross entropy loss function; The detection loss function is defined by Faster RCNN, which includes the classification loss and bounding box regression loss generated by the region proposal network, as well as the loss generated by the region proposal classifier. The detection loss function is: Based on the baseline network architecture of Faster RCNN, the total loss function is: Among them, α and β are hyperparameters used to weigh the corresponding loss functions.
6. The generalized SAR target detection method with a unified framework for cross-source scenarios according to claim 4, wherein the step 2 specifically comprises: setting α to 0.7 and setting β to 0.4.
Citation Information
Cited By
An aircraft target domain adaptation detection method based on scene feature decoupling
CN122510880A