SAR image change detection method based on guide fusion and multi-scale feature aggregation network

By using guided fusion and multi-scale feature aggregation network methods, the problems of insufficient difference map generation quality and feature expression in SAR image change detection are solved, high-precision change area identification and noise suppression are achieved, and the detection effect is improved.

CN120635561APending Publication Date: 2025-09-12TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510740250.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-31
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing SAR image change detection methods have shortcomings in difference map generation quality, feature expression and multi-scale feature loss, making it difficult to achieve high-precision change area identification in complex application scenarios.

Method used

A method based on guided fusion and multi-scale feature aggregation network is adopted. The guided fusion difference map is generated by multiplicative fusion of ER, R and LR difference maps, and the MSFANet model is constructed. The MSFAF module and CII module are combined to enhance feature expression and multi-scale aggregation capabilities. The FCM algorithm is used for pre-classification, and finally classification is performed through the MSFANet model.

Benefits of technology

The quality of the difference map is improved, the feature expression and multi-scale aggregation capabilities are enhanced, the recognition accuracy and contrast of the changed area are improved, the coherent speckle noise is suppressed, and the positioning accuracy of the changed area under complex background is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635561A_ABST
    Figure CN120635561A_ABST
Patent Text Reader

Abstract

The invention provides an SAR (Synthetic Aperture Radar) image change detection method based on guide fusion and a multi-scale feature aggregation network. The specific implementation steps are as follows: (1) generating an ER (Error Rate), an R (Rate) and an LR (Local Rate) difference chart; (2) carrying out multiplication fusion on the ER difference chart and the R difference chart to obtain a guide difference chart; (3) guiding the LR difference chart by using the guiding difference chart to generate a final guiding fusion difference chart; (4) constructing an MSFANet model; (5) pre-classifying the guide fusion difference chart; (6) selecting a training sample and a test sample; and (7) performing final classification on the test samples by using the trained MSFANet model. According to the method, through guide fusion and multi-scale feature aggregation, the structure and edge features of the target area can be improved, and through learning rich hierarchical feature representation, the recognition of the change area in the image and the filtering of irrelevant information in the background are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of SAR image change detection, specifically a SAR image change detection method based on guided fusion and multi-scale feature aggregation networks. This invention can be applied to fields such as computer vision, remote sensing image processing, and pattern recognition for change detection and target recognition. Background Art

[0002] Synthetic Aperture Radar (SAR) image change detection technology provides critical data support for disaster assessment, environmental monitoring, and other fields by analyzing differences in surface scattering properties in multi-temporal SAR images. However, due to the limitations of SAR imaging mechanisms and complex application scenarios, existing methods still face challenges in difference map generation, feature expression, and classification accuracy, including insufficient difference map quality, limited feature expression capabilities, and missing multi-scale features.

[0003] Currently, research on change detection in SAR images focuses on two main areas: designing effective algorithms to suppress speckle noise and improve the quality of difference maps; and building lightweight, high-precision models to accurately and efficiently identify changed regions. As outlined in this review, developing a noise-robust, multi-scale feature-rich, lightweight, and efficient SAR image change detection method is a key requirement for improving remote sensing monitoring accuracy. Summary of the Invention

[0004] To overcome the limitations of existing technologies, this paper proposes a SAR image change detection method based on guided fusion and multi-scale feature aggregation networks. This method not only improves the quality of difference maps but also enhances feature expression and multi-scale aggregation capabilities.

[0005] The specific steps of implementing the present invention include the following:

[0006] (1) Generate ER, R and LR difference maps:

[0007] (1a) Input two dual-phase SAR images;

[0008] (1b) Using the energy ratio (ER) operator, the ratio (R) operator, and the log-ratio (LR) operator to obtain a single difference map;

[0009] (2) Multiply and fuse the ER and R difference maps to obtain the guided difference map:

[0010] The ER difference map and the R difference map are fused by multiplication to obtain the guided difference map;

[0011] (3) Use the guided difference map to guide the LR difference map to generate the final guided fusion difference map:

[0012] Use the guided difference map to guide the LR difference map to perform local linear transformation fusion to obtain the final guided fusion difference map;

[0013] (4) Constructing the MSFANet model:

[0014] (4a) Based on the lightweight architecture of MobileNetV3, SE attention is replaced by coordinate attention (CA) to improve the spatial position modeling capability;

[0015] (4b) Design a multi-scale feature adaptive fusion (MSFAF) module to extract multi-scale features through convolution with different kernel sizes;

[0016] (4c) Designing a contextual information interaction (CII) module that combines cross-covariance attention to enhance the interaction between global and local features;

[0017] (5) Pre-classify the guided fusion difference map:

[0018] Fuzzy C-Means (FCM) algorithm is used to pre-classify the guided fusion difference map to obtain the initial classification result {Ω c ,Ω i ,Ω u};

[0019] (6) Obtain training samples and test samples:

[0020] (6a) After obtaining the pre-classification result {Ω c ,Ω i ,Ω u}, select the c and Ω u The neighborhood around the class pixel is used as a training sample;

[0021] (6b) After obtaining the pre-classification result {Ω c ,Ω i ,Ω u}, select the i The neighborhood around the class pixel is used as a test sample;

[0022] (7) Use the trained MSFANet model to perform final classification on the test samples:

[0023] The MSFANet model is trained using training samples. After the MSFANet model training is completed, the test samples are input for final classification.

[0024] Compared with the prior art, the present invention has the following advantages:

[0025] First, a two-stage guided fusion strategy is proposed. The contrast of the changing area is enhanced by the multiplicative fusion of ER and R operators. The LR difference map is optimized by guided filtering to suppress coherent speckle noise while preserving edge details.

[0026] Second, we designed the MSFAF module to extract multi-scale features through convolutions of different kernel sizes and adaptively weighted fusion to solve the problem of detail loss caused by the fixed receptive field of traditional convolutional networks.

[0027] Third, we design the CII module and use cross-covariance attention to model global dependencies.

[0028] Fourth, CA replaces the traditional SE module, and captures spatial position information through horizontal and vertical pooling, thereby improving the positioning accuracy of changing areas under complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is an implementation flow chart of the present invention;

[0030] Figure 2 This is the MobileNetV3 network framework diagram of the present invention;

[0031] Figure 3 It is the Bneck module diagram of the present invention;

[0032] Figure 4 This is the MSFAF module diagram of the present invention

[0033] Figure 5 is the CII module of the present invention;

[0034] Figure 6 It is the overall structure diagram of MSFANet of the present invention;

[0035] Figure 7 This is the Bern dataset in an embodiment of the present invention, where (a) and (b) are two SAR images of the same area in Bern at different times, (c) is the Bern real change reference map, and (d) is the change detection effect map of the present invention;

[0036] Figure 8 is the Ottawa dataset in an embodiment of the present invention, where (a) and (b) are two SAR images of the same area of ​​Ottawa at different times, (c) is a real change reference map of Ottawa, and (d) is a change detection effect map of the present invention;

[0037] Figure 9 This is the inland river dataset in the embodiment of the present invention, where (a) and (b) are two SAR images of the same area of ​​the inland river at different times, (c) is a reference map of the actual changes in the inland river, and (d) is a change detection effect map of the present invention. DETAILED DESCRIPTION

[0038] The present invention will be further described below with reference to the accompanying drawings.

[0039] Reference Figure 1 , the specific steps of the present invention are as follows:

[0040] Step 1: Generate ER, R and LR difference maps.

[0041] Step 1: Input two SAR images I1 and I2 of the same area at different times, with a size of M × N. First, perform PPB filtering on I1 and I2 to obtain denoised images X1 and X2;

[0042] Step 2: Replace the mean filter in the traditional mean-ratio (MR) operator with a local energy algorithm to enhance the structural and local detail features of the changing regions in the difference map, thereby generating an ER difference map. The R operator is normalized to remap the difference values ​​to a uniform relative scale, unifying the numerical distribution of the SAR image. The LR operator can convert multiplicative speckle noise into additive speckle noise, effectively resisting the influence of noise.

[0043] The specific algorithms for generating difference maps using the three operators are shown in formulas (11), (12), (13), and (14):

[0044]

[0045] DI LR (i,j)=|log(X1(i,j)+1)-log(X2(i,j)+1)| (14)

[0046] Where η = (r-1) / 2, the local energy window size is set to r×r, and E(i, j) represents the sum of the squared grayscale values ​​of all pixels in the r×r neighborhood. E1(i, j) and E2(i, j) represent the local energy values ​​of images X1 and X2, respectively.

[0047] Step 2: Multiply and fuse the ER and R difference maps to obtain the guided difference map.

[0048] The difference maps generated by different algorithms have different advantages. In order to further optimize the local details and structural features of the difference map, the ER difference map and the R difference map are fused by multiplication to obtain the guided difference map.

[0049] Step 3: Use the guided difference map to guide the LR difference map to generate the final guided fusion difference map.

[0050] Use the guided difference map to guide the LR difference map to perform local linear transformation fusion to obtain the final guided fusion difference map;

[0051] Step 1: DI ER and DI R Fusion is performed by multiplication to generate DI that combines local intensity changes and regional feature information ERR , the specific algorithm is shown in formula 15;

[0052] DI ERR (i, j) = DI ER (i, j) DI R (i, j) (15)

[0053] Step 2: DI ERR Use the difference map as input to guide DI LR Perform local linear transformation fusion. The specific algorithm is shown in Formula 16.

[0054] DI GF (i,j)=guided_fusion(DI ERR , DI LR ) (16)

[0055] Fusion DI GF Solve DI to some extent LR The problem of blurred edge features and aggregated DI ERR The structural features and local detail features in the image can enhance the contrast of the changing area while suppressing the coherent speckle noise and smoothing the useless information in the background.

[0056] Step 4: Build the MSFANet model.

[0057] Step 1: Based on the MobileNetV3 network framework, SE attention is replaced by coordinate attention (CA) in the Bneck module. First, the input feature map is globally averaged pooled in the horizontal and vertical directions to generate a direction-sensitive feature vector V h and V v ; Secondly, V h and V v After concatenation, the spatial attention weights are generated by the shared MLP; then, the weights are decomposed into horizontal components W h and the vertical component W v , and multiply it element-by-element with the original feature map to enhance the feature response that is sensitive to spatial position;

[0058] Step 2: Design the Multi-Scale Feature Adaptive Fusion (MSFAF) module. First, four parallel convolution branches are used to extract multi-scale features using depthwise separable convolutions with different kernel sizes of 1×1, 3×3, 5×5, and 7×7. Second, the multi-scale features F1, F2, F3, and F4 are element-wise added to generate aggregated features, as shown in Equation 17.

[0059] F=F1+F2+F3+F4 (17)

[0060] Then, the aggregated multi-scale feature map F is globally averaged pooled in the spatial dimension to extract the global information in the image, as shown in Formula 18;

[0061]

[0062] Among them, F s is the output after global average pooling, H ms and W ms is the height and width of the multi-scale feature map;

[0063] Finally, four parallel convolutions are used to improve the feature vector F′ s The channel dimension is calculated, and the SoftMax function is used to adaptively calculate the weight parameters s1, s2, s3 and s4 corresponding to the features of different scales. The obtained weight parameters are recalibrated for the feature maps of different scales, and all the feature maps are aggregated to obtain the final multi-scale features. The aggregation formula is shown in Formula 19:

[0064] F ms =F1·s1+F2·s2+F3·s3+F4·s4 (19)

[0065] Among them, F ms It is the multi-scale feature output extracted by the MSFAF module;

[0066] Step 3: Design the contextual information interaction (CII) module. First, use cross-covariance attention to calculate attention from the channel dimension of the feature map to enhance the interaction between the local information and the global information of the input features, as shown in Formula 20:

[0067]

[0068] Where W q , W k and W v is the weight matrix.

[0069] Will After point-by-point convolution and GELU activation function, GAP and GMP are used in the first branch to obtain global information respectively; then, 1×1 convolution is used on the second and third branches to generate the linear transformation results of the feature map and simplify the product of Query and Key; then, the first and third branches are matrix multiplied with the second branch respectively, and the two branches obtained represent cross-channel and cross-space context information respectively; finally, the final output of the CII module is obtained by applying the broadcast Hadamard product on these two branches and adding it to the original input.

[0070] Step 5: Pre-classify the guided fusion difference map.

[0071] Fuzzy C-Means (FCM) algorithm is used to pre-classify the guided fusion difference map to obtain the initial classification result {Ω c ,Ω i ,Ω u}.

[0072] Step 6: Obtain training samples and test samples.

[0073] Step 1: Get the pre-classification result {Ω c ,Ω i ,Ω u}, select the c and Ω u The neighborhood around the class pixel is used as a training sample, and a sample unification strategy is used to deal with the imbalance of positive and negative samples, selecting the same number of positive and negative samples for training the network;

[0074] Step 2: After obtaining the pre-classification result {Ω c ,Ω i ,Ω u}, select the i The neighborhood around the class pixel is used as the test sample.

[0075] Step 7: Use the trained MSFANet model to perform final classification on the test samples.

[0076] The trained MSFANet model is used to classify the test samples and obtain the final SAR image change detection results.

[0077] The simulation effect of the present invention is further described below in conjunction with simulation experiments:

[0078] 1. Simulation environment

[0079] The hardware test platform of the embodiment of the present invention is: the processor is an Intel i5-9300H CPU with a main frequency of 2.40 GHz, the memory is 16 GB, and the software platform is: Windows 10 system, Pycharm platform.

[0080] 2. Simulation content

[0081] Three real SAR image datasets are used to verify the superiority of the present invention. At the same time, in order to more objectively illustrate the effect of change detection, the present invention provides five quantitative evaluation criteria for the detection results: number of missed detections (FN), number of false detections (FP), number of overall errors (OE), percentage of correct classifications (PCC) and Kappa coefficient, and compares them with three existing methods DDNet, FFITN and LANTNet on real datasets.

[0082] The first data set selected in this embodiment is the Berne data set, such as Figure 7 The two SAR images shown in (a) and (b) were taken by ERS-2 in April 1999 and May 1999, with an image size of 301 × 301 pixels. Figure 7 (c) is the reference image of the dataset; the second dataset is the Ottawa dataset, which is the flood-affected area of ​​Ottawa, Canada, taken by Radarsat-1 in May and August 1997. The image size is 290×350 pixels. Figure 8 (a) and 8(b) depict two images taken in May and August 1997, respectively. Figure 8 (c) is the reference image of the dataset; the third dataset is the inland river dataset, which was taken by the Radarsat-2 remote sensing satellite sensor at the Yellow River estuary in China in June 2008 and June 2009. The image size is 291×444 pixels. Figure 9 (a) and Figure 9 As shown in (b), Figure 9 (c) is the reference image of this dataset.

[0083] 3. Simulation results and analysis

[0084] Table 1 shows the comparison results of change detection on the Bern dataset. Compared with the other three methods, the proposed method has the lowest FN and OE. In addition, PCC and Kappa are the highest.

[0085] Table 2 shows the comparison results of change detection on the Ottawa dataset. Compared with the other three methods, the proposed method has the lowest FP and OE. In addition, it has the highest PCC and Kappa.

[0086] Table 3 shows the comparison results of change detection on the inland river dataset. Compared with the other three methods, the method proposed in this paper has the lowest OE. In addition, the PCC and Kappa values ​​are the highest.

[0087] Table 1 Comparison results of change detection on the Bern dataset

[0088] method FP FN OE PCC (%) Kappa DDNet 81 252 333 99.63 0.8425 FFITN 216 148 364 99.60 0.8449 LANTNet 71 274 345 99.62 0.8344 The present invention 183 127 310 99.66 0.8672

[0089] Table 2 Comparison results of change detection on the Ottawa dataset

[0090] method FP FN OE PCC (%) Kappa DDNet 839 851 1690 98.34 0.9374 FFITN 1328 685 2013 98.02 0.9267 LANTNet 1307 909 2216 97.82 0.9188 The present invention 672 783 1455 98.57 0.9460

[0091] Table 3. Comparison results of change detection on inland river dataset

[0092] method FP FN OE PCC (%) Kappa DDNet 1340 587 1927 98.51 0.7843 FFITN 1065 781 1846 98.57 0.7842 LANTNet 688 934 1622 98.74 0.7972 The present invention 788 737 1525 98.82 0.8158

[0093] The data in the three tables above demonstrate that our proposed SAR image change detection method, based on guided fusion and a multiscale feature aggregation network, outperforms the other three methods on the three public datasets. This method leverages the multiscale spatial features of the changed regions and contextual differences in the image, thereby enhancing the network's feature representation capabilities and achieving excellent performance in SAR image change detection.

Claims

1. A SAR image change detection method based on guided fusion and multi-scale feature aggregation network, comprising the following steps: (1) Generate ER, R and LR difference maps: (1a) Input two dual-phase SAR images; (1b) Using the energy ratio (ER) operator, the ratio (R) operator, and the log-ratio (LR) operator to obtain a single difference map; (2) Multiply and fuse the ER and R difference maps to obtain the guided difference map: The difference maps generated by different algorithms have different advantages. In order to further optimize the local details and structural features of the difference map, the ER difference map and the R difference map are fused by multiplication to obtain the guided difference map. (3) Use the guided difference map to guide the LR difference map to generate the final guided fusion difference map: Use the guided difference map to guide the LR difference map to perform local linear transformation fusion to obtain the final guided fusion difference map; (4) Constructing the MSFANet model: (4a) Based on the lightweight architecture of MobileNetV3, SE attention is replaced by coordinate attention (CA) to improve the spatial position modeling capability; (4b) Design a multi-scale feature adaptive fusion (MSFAF) module to extract multi-scale features through convolution with different kernel sizes; (4c) Designing a contextual information interaction (CII) module that combines cross-covariance attention to enhance the interaction between global and local features; (5) Pre-classify the guided fusion difference map: Fuzzy C-Means (FCM) algorithm is used to pre-classify the guided fusion difference map to obtain the initial classification result {Ω c ,Ω i ,Ω u }; (6) Obtain training samples and test samples: (5a) After obtaining the pre-classification result {Ω c ,Ω i ,Ω u }, select the c and Ω u The neighborhood around the class pixel is used as a training sample; (5b) After obtaining the pre-classification result {Ω c ,Ω i ,Ω u }, select the i The neighborhood around the class pixel is used as a test sample; (7) Use the trained MSFANet model to perform final classification on the test samples: The MSFANet model is trained using training samples. After the MSFANet model training is completed, the test samples are input for final classification.

2. The SAR image change detection method based on guided fusion and multi-scale feature aggregation network according to claim 1 is characterized in that: In step (1b), the ER operator, R operator, and LR operator are used to obtain a single difference map. The specific steps are as follows: Step 1: Perform PPB filtering on two original SAR images I1 = {I1(i, j), 1≤i≤M, 1≤j≤N} and I2 = {I2(i, j), 1≤i≤M, 1≤j≤N} obtained at different times in the same geographical area to obtain the corresponding denoised images X1 and X2; Step 2: Replace the mean filter in the traditional mean-ratio (MR) operator with a local energy algorithm to enhance the structural and local detail features of the changing regions in the difference map, thereby generating an ER difference map. The R operator is normalized to remap the difference values ​​to a uniform relative scale, unifying the numerical distribution of the SAR image. The LR operator can convert multiplicative speckle noise into additive speckle noise, effectively resisting the influence of noise. The specific algorithms for generating difference maps using the three operators are shown in formulas (1), (2), (3) and (4): DI LR (i,j)=|log(X1(i,j)+1)-log(X2(i,j)+1)| (4) Where η = (r-1) / 2, the local energy window size is set to r×r, and E(i, j) represents the sum of the squared grayscale values ​​of all pixels in the r×r neighborhood. E1(i, j) and E2(i, j) represent the local energy values ​​of images X1 and X2, respectively.

3. The SAR image change detection method based on guided fusion and multi-scale feature aggregation network according to claim 1 is characterized in that: The method described in step (3) of using the guided difference map to guide the LR difference map to generate the final guided fusion difference map is: Step 1: DI ER and DI R Fusion is performed by multiplication to generate DI that combines local intensity changes and regional feature information ERR , the specific algorithm is shown in Formula 5; IN ERR (i,j)=DI ER (i, j)·DI R (i,j) (5) Step 2: Considering that the LR operator is robust to speckle noise, but the edge features of the changing area are easily blurred, the DI LR To overcome this weakness, guided filtering is used as a fusion method to ERR Use the difference map as input to guide DI LR Perform local linear transformation fusion. The specific algorithm is shown in Formula 6. THE GF (i,j)=guided_fusion(DI ERR ,THE LR ) (6) Fusion DI GF Solve DI to some extent LR The problem of blurred edge features and aggregated DI ERR The structural features and local detail features in the image can enhance the contrast of the changing area while suppressing the coherent speckle noise and smoothing the useless information in the background.

4. The SAR image change detection method based on guided fusion and multi-scale feature aggregation network according to claim 1 is characterized in that: The specific steps for constructing the MSFANet model described in step (4) are as follows: Step 1: Based on the MobileNetV3 network framework, SE attention is replaced by coordinate attention (CA) in the Bneck module. First, the input feature map is globally averaged pooled in the horizontal and vertical directions to generate a direction-sensitive feature vector V h and V v ; Secondly, V h and V v After concatenation, the spatial attention weights are generated by the shared MLP; then, the weights are decomposed into horizontal components W h and the vertical component W v , and multiply it element-by-element with the original feature map to enhance the feature response that is sensitive to spatial position. The MobileNetV3 network framework is shown in Figure 2, and the Bneck module is shown in Figure 3; Step 2: Design the Multi-Scale Feature Adaptive Fusion (MSFAF) module. The MSFAF module is shown in Figure 4. First, four parallel convolution branches are used to extract multi-scale features using depthwise separable convolutions with different kernel sizes of 1×1, 3×3, 5×5, and 7×7. Second, the multi-scale features F1, F2, F3, and F4 are added element-by-element to generate aggregated features, as shown in Formula 7. F=F1+F2+F3+F4 (7) Then, the aggregated multi-scale feature map F is globally averaged pooled in the spatial dimension to extract the global information in the image, as shown in Formula 8; Among them, F s is the output after global average pooling, H ms and W ms is the height and width of the multi-scale feature map; Finally, four parallel convolutions are used to improve the feature vector F′ s The channel dimension is calculated, and the SoftMax function is used to adaptively calculate the weight parameters s1, s2, s3 and s4 corresponding to the features of different scales. The obtained weight parameters are recalibrated for the feature maps of different scales, and all feature maps are aggregated to obtain the final multi-scale features. The aggregation formula is shown in Formula 9: <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> ms <h2 style=";text-align:left;direction:ltr"> =F1·s1+F2·s2+F3·s3+F4·s4 (9) Among them, F ms It is the multi-scale feature output extracted by the MSFAF module; Step 3: Design the contextual information interaction (CII) module. First, use cross-covariance attention to calculate attention from the channel dimension of the feature map to enhance the interaction between the local information and the global information of the input features, as shown in Formula 10: Where W q , W k and W v is the weight matrix. Will After point-by-point convolution and GELU activation function, GAP and GMP are used in the first branch to obtain global information respectively; then, 1×1 convolution is used in the second and third branches to generate the linear transformation results of the feature map and simplify the product of Query and Key; then, the first and third branches are matrix multiplied with the second branch respectively, and the two branches obtained represent cross-channel and cross-space context information respectively; finally, by applying broadcast Hadamard product on these two branches and adding them to the original input, the final output of the CII module is obtained. The CII module is shown in Figure 5, and the overall structure of MSFANet is shown in Figure 6.

5. The SAR image change detection method based on guided fusion and multi-scale feature aggregation network according to claim 1 is characterized in that: The step (6a) described in obtaining the pre-classification result {Ω c ,Ω i ,Ω u }, select the c and Ω u The neighborhood around the class pixel is used as the training sample, and the sample unification strategy is adopted to deal with the imbalance problem of positive and negative samples. The same number of positive and negative samples are selected to train the network.